Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 5:00 PM ET until 8:00 PM ET on Friday, September 11 due to maintenance. We apologize for the inconvenience.


Title: Assessing Statistical Consultations and Collaborations
As practitioners and teachers of statistical consulting and collaboration, how do we assess the effectiveness of ours and our students’ engagement on projects with domain experts? We propose that assessments of the effectiveness of statistical collaborations should be based on the four areas of attitude, skills, performance, and improvement. In this brief paper, we describe several ways for conducting assessments in these four areas and conclude with a call for the statistics and data science education community to build upon these ideas.  more » « less
Award ID(s):
1955109 2022138
PAR ID:
10227759
Author(s) / Creator(s):
; ;
Date Published:
Journal Name:
JSM Proceedings, Statistical Consulting Section
Format(s):
Medium: X
Sponsoring Org:
National Science Foundation
More Like this
  1. Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it remains unclear whether large language models (LLMs) can handle rigorous and trustworthy statistical causal inference. Current benchmarks usually involve simplified tasks. For example, these tasks might only ask LLMs to identify semantic causal relationships or draw conclusions directly from raw data. As a result, models may overlook important statistical pitfalls, such as Simpson’s paradox or selection bias. This oversight limits the applicability of LLMs in the real world. To address these limitations, we propose CausalPitfalls, a comprehensive benchmark designed to rigorously evaluate the capability of LLMs in overcoming common causal inference pitfalls. Our benchmark features structured challenges across multiple difficulty levels, each paired with grading rubrics. This approach allows us to quantitatively measure both causal reasoning capabilities and the reliability of LLMs’ responses. We evaluate models using two protocols: (1) direct prompting, which assesses intrinsic causal reasoning, and (2) code-assisted prompting, where models generate executable code for statistical analysis. Additionally, we validate the effectiveness of this judge by comparing its scoring with assessments from human experts. Our results reveal significant limitations in current LLMs when performing statistical causal inference. The CausalPitfalls benchmark provides essential guidance and quantitative metrics to advance the development of trustworthy causal reasoning systems. Our code is publicly available at CausalPitfalls. 
    more » « less
  2. Marine protected areas (MPAs) are among the most widely used strategy to protect marine ecosystems and are typically designed to protect specific habitats rather than a single and/or multiple species. To inform the con- servation of species of conservation concern there is the need to assess whether existing and proposed MPA designs provide protection to these species. For this, information on species spatial distribution and exposure to threats is necessary. However, this information if often lacking, particularly for mobile migratory species, such as marine turtles. To highlight the importance of this information when designing MPAs and for assessments of their effectiveness, we identified high use areas of post-nesting hawksbill turtles (Eretmochelys imbricata) in Brazil as a case study and assessed the effectiveness of Brazilian MPAs to protect important habitat for this group based on exposure to threats. Most (88%) of high use areas were found to be exposed to threats (78% to artisanal fishery and 76.7% to marine traffic), where 88.1% were not protected by MPAs, for which 86% are exposed to threats. This mismatch is driven by a lack of explicit conservation goals and targets for turtles in MPA management plans, limited spatial information on species' distribution and threats, and a mismatch in the scale of conservation initiatives. To inform future assessments and design of MPAs for species of conservation concern we suggest that managers: clearly state and make their goals and targets tangible, consider ecological scales instead of political boundaries, and use adaptative management as new information become available. 
    more » « less
  3. Protected Areas (PAs) form the basis of biodiversity conservation and remain the primary strategy to mitigate the global biodiversity crisis. Accelerating population declines and extinctions at a global scale require the establishment of conservation priorities, often guided by the identification of biodiversity hotspots, with endemism serving as a key indicator of environmental uniqueness. Reptiles have often been overlooked in such assessments despite their vulnerability and contribution to biodiversity metrics. Angola, a highly biodiverse yet understudied African country, has been highlighted for its reptile diversity and endemism, although no conservation actions have been implemented towards this fauna. Here we present a first systematic analysis of spatial patterns of reptile endemism in Angola. Using two endemism metrics and spatial statistics, we identify hotspots and conservation priority areas for endemic reptiles and evaluate the effectiveness of existing PAs through a gap analysis. We update the endemism rate of Angolan reptiles from 12% to 28%, identifying significant hotspots in western Angola. Southwestern Angola emerges as the country’s main center of endemism, characterized by narrow ranged local endemics. Our results reveal substantial conservation gaps, with less than 10% of priority hotspots currently protected and over half of Angola’s endemic reptiles lacking formal conservation assessments and representation in the PA network. These findings underscore the need for targeted conservation planning and the establishment of new protected areas in southwestern Angola to safeguard the country’s unique reptile fauna. 
    more » « less
  4. The process of regionalization involves clustering a set of spatial areas into spatially contiguous regions. Given the NP-hard nature of regionalization problems, all existing algorithms yield approximate solutions. To ascertain the quality of these approximations, it is crucial for domain experts to obtain statistically significant evidence on optimizing the objective function, in comparison to a random reference distribution derived from all potential sample solutions. In this paper, we propose a novel spatial regionalization problem, denoted as SISR (Statistical Inference for Spatial Regionalization), which generates random sample solutions with a predetermined region cardinality. The driving motivation behind SISR is to conduct statistical inference on any given regionalization scheme. To address SISR, we present a parallel technique named PRRP (P-Regionalization through Recursive Partitioning). PRRP operates over three phases: the region-growing phase constructs initial regions with a predetermined region cardinality, while the region merging and region-splitting phases ensure the spatial contiguity of unassigned areas, allowing for the growth of subsequent regions with predetermined cardinalities. An extensive evaluation shows the effectiveness of PRRP using various real datasets. 
    more » « less
  5. null (Ed.)
    In the spring of 2020, universities across America, and the world, abruptly transitioned to online learning. The online transition required faculty to find novel ways to administer assessments and in some cases, for students to utilize novel ways of cheating in their classes. The purpose of this paper is to provide a retrospective on cheating during online exams in the spring of 2020. It specifically looks at honor code violations in a sophomore level engineering course that enrolled more than 200 students. In this particular course, four pre-COVID assessments were given in class and six mid-COVID assessments were given online. This paper examines the increasing rate of cheating on these assessments and the profiles of the students who were engaged in cheating. It compares students who were engaged in violations of the honor code by uploading exam questions vs. those who those who looked at solutions to uploaded questions. This paper also looks at the abuse of Chegg during exams and the responsiveness of Chegg’s honor code team. It discusses the effectiveness of Chegg’s user account data in pursuing academic integrity cases. Information is also provided on the question response times for Chegg tutors in answering exam questions and the actual efficacy of cheating in this fashion. 
    more » « less