Attention:The NSF Public Access Repository (PAR) system and access will be unavailable from 10:00 PM ET on Thursday, July 16 until 12:00 AM ET on Friday 17 due to maintenance. We apologize for the inconvenience.


Search for: All records

Creators/Authors contains: "Luo, A"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. The rapid adoption of generative AI chatbots has intensified discussions around trustworthy AI. Understanding how lay users perceive and evaluate chatbot trustworthiness is essential to keeping these discussions inclusive, making formal assessments user-centered, and revealing instances of misplaced user trust. To this end, we conducted an interactive online study with 254 U.S.-based participants who were asked to investigate whether a generative AI chatbot produced problematic, questionable, or unfair responses. Using a researcher-supplied probing tool, participants could freely interact with the chatbot and flag any issues. Participants engaged in 551 open-ended conversations and primarily sought to probe for issues and topics related to their everyday use. However, participants frequently failed to uncover the issues they expected to find. Overall, trustworthiness perceptions increased after the probing intervention, regardless of whether issues were detected, and participants with higher initial trustworthiness were less likely to flag problems in general. Our findings suggest that lay users may have the right instincts about concerns with generative AI, but are not well equipped to surface these issues on their own. We also find that trustworthiness perceptions can increase rapidly, even among initially skeptical users, likely driven by users' tendency to treat outputs as deliberate and reasoned rather than probabilistic, and by their reliance on surface cues—particularly performance and utility—when judging trustworthiness. Our work complements prior work by illuminating how everyday users assess trustworthiness in generative AI, broadening debates on trustworthy AI, and highlighting the need for stronger guidance to help lay users accurately assess and calibrate trustworthiness perceptions. 
    more » « less
    Free, publicly-accessible full text available June 1, 2027
  2. Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it remains unclear whether large language models (LLMs) can handle rigorous and trustworthy statistical causal inference. Current benchmarks usually involve simplified tasks. For example, these tasks might only ask LLMs to identify semantic causal relationships or draw conclusions directly from raw data. As a result, models may overlook important statistical pitfalls, such as Simpson’s paradox or selection bias. This oversight limits the applicability of LLMs in the real world. To address these limitations, we propose CausalPitfalls, a comprehensive benchmark designed to rigorously evaluate the capability of LLMs in overcoming common causal inference pitfalls. Our benchmark features structured challenges across multiple difficulty levels, each paired with grading rubrics. This approach allows us to quantitatively measure both causal reasoning capabilities and the reliability of LLMs’ responses. We evaluate models using two protocols: (1) direct prompting, which assesses intrinsic causal reasoning, and (2) code-assisted prompting, where models generate executable code for statistical analysis. Additionally, we validate the effectiveness of this judge by comparing its scoring with assessments from human experts. Our results reveal significant limitations in current LLMs when performing statistical causal inference. The CausalPitfalls benchmark provides essential guidance and quantitative metrics to advance the development of trustworthy causal reasoning systems. Our code is publicly available at CausalPitfalls. 
    more » « less
    Free, publicly-accessible full text available April 23, 2027
  3. Free, publicly-accessible full text available August 11, 2026
  4. Abstract A hallmark of many unconventional superconductors is the presence of many-body interactions that give rise to broken-symmetry states intertwined with superconductivity. Recent resonant soft X-ray scattering experiments report commensurate 3a0charge density wave order in infinite-layer nickelates, which has important implications regarding the universal interplay between charge order and superconductivity in both cuprates and nickelates. Here we present X-ray scattering and spectroscopy measurements on a series of NdNiO2+xsamples, which reveal that the signatures of charge density wave order are absent in fully reduced, single-phase NdNiO2. The 3a0superlattice peak instead originates from a partially reduced impurity phase where excess apical oxygens form ordered rows with three-unit-cell periodicity. The absence of any observable charge density wave order in NdNiO2highlights a crucial difference between the phase diagrams of cuprate and nickelate superconductors. 
    more » « less