Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
Free, publicly-accessible full text available June 2, 2027
-
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal InferenceReliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it remains unclear whether large language models (LLMs) can handle rigorous and trustworthy statistical causal inference. Current benchmarks usually involve simplified tasks. For example, these tasks might only ask LLMs to identify semantic causal relationships or draw conclusions directly from raw data. As a result, models may overlook important statistical pitfalls, such as Simpson’s paradox or selection bias. This oversight limits the applicability of LLMs in the real world. To address these limitations, we propose CausalPitfalls, a comprehensive benchmark designed to rigorously evaluate the capability of LLMs in overcoming common causal inference pitfalls. Our benchmark features structured challenges across multiple difficulty levels, each paired with grading rubrics. This approach allows us to quantitatively measure both causal reasoning capabilities and the reliability of LLMs’ responses. We evaluate models using two protocols: (1) direct prompting, which assesses intrinsic causal reasoning, and (2) code-assisted prompting, where models generate executable code for statistical analysis. Additionally, we validate the effectiveness of this judge by comparing its scoring with assessments from human experts. Our results reveal significant limitations in current LLMs when performing statistical causal inference. The CausalPitfalls benchmark provides essential guidance and quantitative metrics to advance the development of trustworthy causal reasoning systems. Our code is publicly available at CausalPitfalls.more » « lessFree, publicly-accessible full text available April 23, 2027
-
Free, publicly-accessible full text available May 27, 2027
-
Free, publicly-accessible full text available December 6, 2026
-
Free, publicly-accessible full text available January 1, 2027
-
This paper examines the design and evaluation of Large Language Model (LLM) tutors for Python programming, focusing on personalization that accommodates diverse student backgrounds. It highlights the challenges faced by socioeconomically disadvantaged students in computing courses and proposes LLM tutors as a solution to provide inclusive educational support. The study explores two LLM tutors, Khanmigo and CS50.ai, assessing their ability to offer personalized learning experiences. By employing a focus group methodology at a public minority-serving institution, the research evaluates how these tutors meet varied educational goals and adapt to students’ diverse needs. The findings underscore the importance of advanced techniques to tailor interactions and integrate programming tools based on students' progress. This research contributes to the understanding of educational technologies in computing education and provides insights into the design and implementation of LLM tutors that effectively support equitable student success.more » « less
-
Collisionless plasma systems are often studied using fully kinetic simulations, where protons and electrons are treated as particles. Due to their computational expense, it is necessary to reduce the ion-to-electron mass ratio or the ratio between plasma and cyclotron frequencies in simulations of large systems. In this Letter we show that when electron-scale waves are present in larger-scale systems, numerical parameters affect their amplitudes and effects on the larger system. Using lower-hybrid drift waves during magnetic reconnection as an example, we find that the ratio between the wave electric field and the reconnection electric field scales as , while the phase relationship is also affected. The combination of these effects means that the anomalous drag that contributes to momentum balance in the reconnection region can be underestimated by an order of magnitude. The results are relevant to the coupling of electron-scale waves to ion-scale reconnection regions, and other systems such as collisionless shocks. Published by the American Physical Society2024more » « less
An official website of the United States government

Full Text Available