Search for: All records

Creators/Authors contains: "Xu, Xuhai"

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

  1. Free, publicly-accessible full text available April 13, 2027
  2. Abstract Psychological stress is a key driver of short-term blood pressure (BP) elevations and cardiovascular risk, yet its moment-to-moment impact in daily life remains difficult to predict. In this longitudinal observational study, we collected multimodal data from 20 adults with self-reported hypertension, including continuous wearable-derived heart rate and activity, ecological momentary assessment (EMA) stress ratings, and ambulatory BP measurements in free-living conditions. The dataset comprised 3694 EMA responses and 3812 BP measurements collected over approximately four weeks per participant (mean 24.1 ± 8.5 days). We evaluated whether participant-specific (“personalized”) models outperform a single pooled population model. Two prediction tasks were examined: (i) prediction of near-term BP elevations from wearable signals and stress EMA responses and (ii) prediction of self-reported stress from wearable signals and BP. Across both tasks, personalized models consistently improved predictive performance. For BP prediction, personalized models achieved a mean AUROC of 0.803, exceeding the population model by 0.235, while for stress prediction they achieved a mean AUROC of 0.849, exceeding the population model by 0.208. These findings suggest that personalized wearable-based models can capture individual patterns of stress and BP dynamics, with direct implications for precision mental health assessment and just-in-time adaptive intervention design in future work. 
    more » « less
    Free, publicly-accessible full text available May 8, 2027
  3. Free, publicly-accessible full text available November 25, 2026
  4. Free, publicly-accessible full text available September 27, 2026
  5. Abstract Individuals are increasingly utilizing large language model (LLM)-based tools for mental health guidance and crisis support in place of human experts. While AI technology has great potential to improve health outcomes, insufficient empirical evidence exists to suggest that AI technology can be deployed as a clinical replacement; thus, there is an urgent need to assess and regulate such tools. Regulatory efforts have been made and multiple evaluation frameworks have been proposed, however,field-wide assessment metrics have yet to be formally integrated. In this paper, we introduce a comprehensive online platform that aggregates evaluation approaches and serves as a dynamic online resource to simplify LLM and LLM-based tool assessment:MindBench.ai. At its core,MindBench.aiis designed to provide easily accessible/interpretable information for diverse stakeholders (patients, clinicians, developers, regulators, etc.). To createMindBench.ai, we built off our work developing MINDapps.org to support informed decision-making around smartphone app use for mental health, and expanded the technical MINDapps.org framework to encompass novel large language model (LLM) functionalities through benchmarking approaches. TheMindBench.aiplatform is designed as a partnership with the National Alliance on Mental Illness (NAMI) to provide assessment tools that systematically evaluate LLMs and LLM-based tools with objective and transparent criteria from a healthcare standpoint, assessing both profile (i.e. technical features, privacy protections, and conversational style) and performance characteristics (i.e. clinical reasoning skills). With infrastructure designed to scale through community and expert contributions, along with adapting to technological advances, this platform establishes a critical foundation for the dynamic, empirical evaluation of LLM-based mental health tools—transforming assessment into a living, continuously evolving resource rather than a static snapshot. 
    more » « less
    Free, publicly-accessible full text available December 1, 2026
  6. Advances in large language models (LLMs) have empowered a variety of applications. However, there is still a significant gap in research when it comes to understanding and enhancing the capabilities of LLMs in the field of mental health. In this work, we present a comprehensive evaluation of multiple LLMs on various mental health prediction tasks via online text data, including Alpaca, Alpaca-LoRA, FLAN-T5, GPT-3.5, and GPT-4. We conduct a broad range of experiments, covering zero-shot prompting, few-shot prompting, and instruction fine-tuning. The results indicate a promising yet limited performance of LLMs with zero-shot and few-shot prompt designs for mental health tasks. More importantly, our experiments show that instruction finetuning can significantly boost the performance of LLMs for all tasks simultaneously. Our best-finetuned models, Mental-Alpaca and Mental-FLAN-T5, outperform the best prompt design of GPT-3.5 (25 and 15 times bigger) by 10.9% on balanced accuracy and the best of GPT-4 (250 and 150 times bigger) by 4.8%. They further perform on par with the state-of-the-art task-specific language model. We also conduct an exploratory case study on LLMs' capability on mental health reasoning tasks, illustrating the promising capability of certain models such as GPT-4. We summarize our findings into a set of action guidelines for potential methods to enhance LLMs' capability for mental health tasks. Meanwhile, we also emphasize the important limitations before achieving deployability in real-world mental health settings, such as known racial and gender bias. We highlight the important ethical risks accompanying this line of research. 
    more » « less