Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 10:00 PM ET on Thursday, July 16 until 12:00 AM ET on Friday 17 due to maintenance. We apologize for the inconvenience.


Search for: All records

Creators/Authors contains: "An, S"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. The objective of this study was to evaluate the performance of an XGBoost model trained with behavioral, physiological, performance, environmental, and cow feature data for classifying cow health status (HS). The model predicted HS based on physical activity, resting, reticulo-rumen temperature, rumination and eating behavior, milk yield, conductivity and components, temperature and humidity index, parity, calving features, and stocking density. Daily at 5 a.m., the model generated a HS prediction [0 = no health disorder (HD); 1 = health disorder]. At 7 a.m., technicians blind to the prediction conducted clinical exams on cows from 3 to 11 DIM to classify cows (n = 625) as affected (HD = 1) or not (HD = 0) by metritis, mastitis, ketosis, indigestion, displaced abomasum, and pneumonia. Using each day a cow presented clinical signs of HD as a positive case (i.e., HD = 1), metrics of performance (%; 95% CI) were: sensitivity (Se) = 57 [52, 62], specificity = 81 [80, 82]; positive predictive value (PPV) = 20 [18, 22], negative predictive value = 96 [95, 96], accuracy = 79 [78, 80], balanced accuracy = 69 [66, 72], F-1 Score = 29 [26, 32]. Sensitivity was also evaluated using fixed time intervals around clinical diagnosis of disease as a positive case (Table 1). Our findings suggest that the ability of an XGBoost algorithm trained on diverse sensor and nonsensor data to identify cows with HD was moderate when only days when cows presented clinical signs of disease were considered a positive case. Sensitivity and PPV can be improved substantially when all days within fixed intervals before and after clinical diagnosis are used as positive cases. Table 1 (Abstr. 2614). Sensitivity and PPV for an XGBoost algorithm trained to predict cow health status using fixed intervals before and after clinical diagnosis as positive cases Day relative to CD Se (%) 95% CI PPV (%) 95% CI −5 to 0 58 49, 67 21 16, 25 −3 to 0 55 46, 64 19 15, 24 −5 to 1 69 61, 78 24 20, 29 −5 to 3 81 73, 88 28 23, 33 −5 to 5 86 80, 92 30 25, 34 −3 to 1 67 58, 75 23 18, 27 −3 to 3 78 70, 86 27 22, 31 −3 to 5 83 76, 90 28 24, 33 0 to 3 75 68, 83 24 20, 29 0 to 5 81 73, 88 26 21, 31 −1 to 0 54 44, 63 18 14, 22 0 to 1 63 54, 72 20 16, 25 −1 to 1 66 57, 75 21 17, 26 
    more » « less
  2. Labeling data via rules-of-thumb and minimal label supervision is central to Weak Supervision, a paradigm subsuming subareas of machine learning such as crowdsourced learning and semi-supervised ensemble learning. By using this labeled data to train modern machine learning methods, the cost of acquiring large amounts of hand labeled data can be ameliorated. Approaches to combining the rules-of-thumb falls into two camps, reflecting different ideologies of statistical estimation. The most common approach, exemplified by the Dawid-Skene model, is based on probabilistic modeling. The other, developed in the work of Balsubramani-Freund and others, is adversarial and game-theoretic. We provide a variety of statistical results for the adversarial approach under log-loss: we characterize the form of the solution, relate it to logistic regression, demonstrate consistency, and give rates of convergence. On the other hand, we find that probabilistic approaches for the same model class can fail to be consistent. Experimental results are provided to corroborate the theoretical results. 
    more » « less
  3. Abstract The recharge oscillator (RO) is a simple mathematical model of the El Niño Southern Oscillation (ENSO). In its original form, it is based on two ordinary differential equations that describe the evolution of equatorial Pacific sea surface temperature and oceanic heat content. These equations make use of physical principles that operate in nature: (a) the air‐sea interaction loop known as the Bjerknes feedback, (b) a delayed oceanic feedback arising from the slow oceanic response to winds within the equatorial band, (c) state‐dependent stochastic forcing from fast wind variations known as westerly wind bursts (WWBs), and (d) nonlinearities such as those related to deep atmospheric convection and oceanic advection. These elements can be combined at different levels of RO complexity. The RO reproduces ENSO key properties in observations and climate models: its amplitude, dominant timescale, seasonality, and warm/cold phases amplitude asymmetry. We discuss the RO in the context of timely research questions. First, the RO can be extended to account for ENSO pattern diversity (with events that either peak in the central or eastern Pacific). Second, the core RO hypothesis that ENSO is governed by tropical Pacific dynamics is discussed from the perspective of influences from other basins. Finally, we discuss the RO relevance for studying ENSO response to climate change, and underline that accounting for ENSO diversity, nonlinearities, and better links of RO parameters to the long term mean state are important research avenues. We end by proposing important RO‐based research problems. 
    more » « less
  4. We present a study that examines the effects of guidance on learning about addressing ill-defined problems in undergraduate bi- ology education. Two groups of college students used an online labo- ratory named VERA to learn about ill-defined ecological phenomena. While one group received guidance, such as giving the learners a specific problem and instruction on problem-solving methods, the other group re- ceived minimal guidance. The results indicate that, while performance in a problem-solving task was not different between groups receiving more vs. minimal guidance, the group that received minimal guidance adopted a more exploratory strategy and generated more interesting models of the given phenomena in a problem-solving task. 
    more » « less
  5. Virtual laboratories that enable novice scientists to construct, evaluate and revise models of complex systems heavily involve parameter estimation tasks. We seek to understand novice strategies for parameter estimation in model exploration to design better cognitive supports for them. We conducted a study of 50 college students for a parameter estimation task in exploring an ecological model. We identified three types of behavioral patterns and their underlying cognitive strategies. Specifically, the students used systematic search, problem decomposition and reduction, and global search followed by local search as their cognitive strategies 
    more » « less
  6. Modeling is an important aspect of scientific problem-solving. How- ever, modeling is a difficult cognitive process for novice learners in part due to the high dimensionality of the parameter search space. This work investigates 50 college students’ parameter search behaviors in the context of ecological modeling. The study revealed important differences in behaviors of successful and unsuccessful students in navigating the parameter space. These differences suggest opportunities for future development of adaptive cognitive scaffolds to support different classes of learners 
    more » « less
  7. Citizen scientists have the potential to expand scientific research. The virtual research assistant called VERA empowers citizen scientists to engage in environmental science in two ways. First, it automatically generates simulations based on the conceptual models of ecological phenomena for repeated testing and feedback. Second, it leverages the Encyclopedia of Life biodiversity knowledgebase to support the process of model construction and revision. 
    more » « less
  8. Free, publicly-accessible full text available June 1, 2027
  9. Free, publicly-accessible full text available June 1, 2027