Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher.
Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?
Some links on this page may take you to non-federal websites. Their policies may differ from this site.
-
BackgroundIn the modern economy, shift work is prevalent in numerous occupations. However, it often disrupts workers’ circadian rhythms and can result in shift work sleep disorder. Proper management of shift work sleep disorder involves comprehensive and patient-specific strategies, some of which are similar to cognitive behavioral therapy for insomnia. ObjectiveOur goal was to develop and evaluate machine learning algorithms that predict physicians’ sleep advice using wearable and survey data. We developed a web- and app-based system to provide individualized sleep and behavior advice based on cognitive behavioral therapy for insomnia for shift workers. MethodsData were collected for 5 weeks from shift workers (N=61) in the intensive care unit at 2 hospitals in Japan. The data comprised 3 modalities: Fitbit data, survey data, and sleep advice. After the first week of enrollment, physicians reviewed Fitbit and survey data to provide sleep advice and selected 1 to 5 messages from a list of 23 options. We handcrafted physiological and behavioral features from the raw data and identified clusters of participants with similar characteristics using hierarchical clustering. We explored 3 models (random forest, light gradient-boosting machine, and CatBoost) and 3 data-balancing approaches (no balancing, random oversampling, and synthetic minority oversampling technique) to predict selections for the 7 most frequent advice messages related to bedroom brightness, smartphone use, and nap and sleep duration. We tested our predictions under participant-dependent and participant-independent settings and analyzed the most important features for prediction using permutation importance and Shapley additive explanations. ResultsWe found that the clusters were distinguished by work shifts and behavioral patterns. For example, one cluster had days with low sleep duration and the lowest sleep quality when there was a day shift on the day before and a midnight shift on the current day. Our advice prediction models achieved a higher area under the precision-recall curve than the baseline in all settings. The performance differences were statistically significant (P<.001 for 13 tests and P=.003 for 1 test). Sensitivity ranged from 0.50 to 1.00, and specificity varied between 0.44 and 0.93 across all advice messages and dataset split settings. Feature importance analysis of our models found several important features that matched the corresponding advice messages sent. For instance, for message 7 (darken the bedroom when you go to bed), the models primarily examined the average brightness of the sleep environment to make predictions. ConclusionsAlthough our current system requires physician input, an accurate machine learning algorithm shows promise for automatic advice without compromising the trustworthiness of the selected recommendations. Despite its decent performance, the algorithm is currently limited to the 7 most popular messages. Further studies are needed to enable predictions for less frequent advice labels.more » « less
-
Abstract Individuals are increasingly utilizing large language model (LLM)-based tools for mental health guidance and crisis support in place of human experts. While AI technology has great potential to improve health outcomes, insufficient empirical evidence exists to suggest that AI technology can be deployed as a clinical replacement; thus, there is an urgent need to assess and regulate such tools. Regulatory efforts have been made and multiple evaluation frameworks have been proposed, however,field-wide assessment metrics have yet to be formally integrated. In this paper, we introduce a comprehensive online platform that aggregates evaluation approaches and serves as a dynamic online resource to simplify LLM and LLM-based tool assessment:MindBench.ai. At its core,MindBench.aiis designed to provide easily accessible/interpretable information for diverse stakeholders (patients, clinicians, developers, regulators, etc.). To createMindBench.ai, we built off our work developing MINDapps.org to support informed decision-making around smartphone app use for mental health, and expanded the technical MINDapps.org framework to encompass novel large language model (LLM) functionalities through benchmarking approaches. TheMindBench.aiplatform is designed as a partnership with the National Alliance on Mental Illness (NAMI) to provide assessment tools that systematically evaluate LLMs and LLM-based tools with objective and transparent criteria from a healthcare standpoint, assessing both profile (i.e. technical features, privacy protections, and conversational style) and performance characteristics (i.e. clinical reasoning skills). With infrastructure designed to scale through community and expert contributions, along with adapting to technological advances, this platform establishes a critical foundation for the dynamic, empirical evaluation of LLM-based mental health tools—transforming assessment into a living, continuously evolving resource rather than a static snapshot.more » « lessFree, publicly-accessible full text available December 1, 2026
-
BackgroundAs mobile health (mHealth) studies become increasingly productive owing to the advancements in wearable and mobile sensor technology, our ability to monitor and model human behavior will be constrained by participant receptivity. Many health constructs are dependent on subjective responses, and without such responses, researchers are left with little to no ground truth to accompany our ever-growing biobehavioral data. This issue can significantly impact the quality of a study, particularly for populations known to exhibit lower compliance rates. To address this challenge, researchers have proposed innovative approaches that use machine learning (ML) and sensor data to modify the timing and delivery of surveys. However, an overarching concern is the potential introduction of biases or unintended influences on participants’ responses when implementing new survey delivery methods. ObjectiveThis study aims to demonstrate the potential impact of an ML-based ecological momentary assessment (EMA) delivery system (using receptivity as the predictor variable) on the participants’ reported emotional state. We examine the factors that affect participants’ receptivity to EMAs in a 10-day wearable and EMA–based emotional state–sensing mHealth study. We study the physiological relationships indicative of receptivity and affect while also analyzing the interaction between the 2 constructs. MethodsWe collected data from 45 healthy participants wearing 2 devices measuring electrodermal activity, accelerometer, electrocardiography, and skin temperature while answering 10 EMAs daily, containing questions about perceived mood. Owing to the nature of our constructs, we can only obtain ground truth measures for both affect and receptivity during responses. Therefore, we used unsupervised and supervised ML methods to infer affect when a participant did not respond. Our unsupervised method used k-means clustering to determine the relationship between physiology and receptivity and then inferred the emotional state during nonresponses. For the supervised learning method, we primarily used random forest and neural networks to predict the affect of unlabeled data points as well as receptivity. ResultsOur findings showed that using a receptivity model to trigger EMAs decreased the reported negative affect by >3 points or 0.29 SDs in our self-reported affect measure, scored between 13 and 91. The findings also showed a bimodal distribution of our predicted affect during nonresponses. This indicates that this system initiates EMAs more commonly during states of higher positive emotions. ConclusionsOur results showed a clear relationship between affect and receptivity. This relationship can affect the efficacy of an mHealth study, particularly those that use an ML algorithm to trigger EMAs. Therefore, we propose that future work should focus on a smart trigger that promotes EMA receptivity without influencing affect during sampled time points.more » « less
-
Abstract Objectives: This study aimed to evaluate, using wearable sensors, the impact of transitioning from an 8-hour to a 12-hour shift schedule on sleep patterns and well-being in intensive care unit (ICU) nurses with pre-existing sleep disturbances. We also examined differences in outcome based on chronotype. Methods: We conducted an observational study at a university hospital ICU between November 2020 and October 2023, before and after a hospital-wide shift schedule change. Nurses wore wearable sensors and completed daily surveys over 5 weeks under each shift system. Rotating-shift ICU nurses with a Pittsburgh Sleep Quality Index score >5 were eligible. Sleep metrics and subjective well-being were compared using linear mixed models, adjusting for age. Sleep episodes were categorized relative to shift timing, and chronotype-stratified subgroup analyses were performed. Results: Eighty nurses completed the study (12-hour shift: 37; 8-hour shift: 43). The interval between shifts was greater for the 12-hour shift group (36.12 vs 26.78 hours). Total sleep duration did not significantly differ between groups (12-hour shift: 418.5 minutes; 8-hour shift: 398 minutes); however, the 12-hour shift group had less fragmented sleep, higher subjective well-being scores, and lower reported stress and fatigue. Evening chronotypes appeared to benefit more from 12-hour shifts, with longer sleep duration and higher well-being scores, though these differences were not statistically significant. Conclusions: Transitioning to a 12-hour shift schedule was associated with reduced sleep fragmentation and improved well-being, particularly among evening chronotypes. These findings suggest that shift schedule structure and individual chronotype may influence adaptation to shift work in ICU settings.more » « less
-
Free, publicly-accessible full text available January 19, 2027
-
Free, publicly-accessible full text available January 1, 2027
-
Background/Objectives: Nurses are at high risk for burnout. Identification of biomarkers associated with early manifestations of distress is essential to support effective intervention efforts. Methods: Fifty nurses from a large hospital system participated in a 30-day study of biopsychosocial factors that may contribute to burnout. Nurses wore an Oura ring that collected behavioral data and they completed a self-report burnout questionnaire at baseline and the end of the study period. Machine learning models were developed to evaluate whether objective measures could predict burnout states and changes at the end of the study period. Analyses were exploratory and hypothesis-generating for future work. Results: Data for 45 participants were included in the analyses. Participants with burnout had significantly higher sleep variability. Sleep measures provided 75.75% accuracy in ability to discriminate between burnout states. Heart rate-based measures better modeled changes in symptomatic components of burnout (Emotional Exhaustion, Depersonalization) over time. Heart rate-based measures provided a R-squared value of 0.13 (p < 0.05) (RMSE of 7.41) in a regression model of changes in Emotional Exhaustion evaluated in a leave-one-participant-out cross-validation. Conclusions: Sleep measures’ association with a state of burnout may reflect the longer-term manifestations of chronic exposure to workplace stress. Short-term changes in burnout symptoms are associated with disturbances in heart rate measures. Wearable technology may support monitoring/early identification of those at risk for burnout.more » « lessFree, publicly-accessible full text available January 1, 2027
-
Free, publicly-accessible full text available November 26, 2026
-
Free, publicly-accessible full text available November 15, 2026
-
Free, publicly-accessible full text available October 11, 2026
An official website of the United States government
