Improving medical machine learning models with generative balancing for equity and excellence

Theodorou, Brandon; Danek, Benjamin; Tummala, Venkat; Kumar, Shivam_Pankaj; Malin, Bradley; Sun, Jimeng

doi:10.1038/s41746-025-01438-z

Citation Details

Improving medical machine learning models with generative balancing for equity and excellence

Abstract Applying machine learning to clinical outcome prediction is challenging due to imbalanced datasets and sensitive tasks that contain rare yet critical outcomes and where equitable treatment across diverse patient groups is essential. Despite attempts, biases in predictions persist, driven by disparities in representation and exacerbated by the scarcity of positive labels, perpetuating health inequities. This paper introduces , a synthetic data generation approach leveraging large language models, to address these issues. enhances algorithmic performance and reduces bias by creating realistic, anonymous synthetic patient data that improves representation and augments dataset patterns while preserving privacy. Through experiments on multiple datasets, we demonstrate that boosts mortality prediction performance across diverse subgroups, achieving up to a 21% improvement in F1 Score without requiring additional data or altering downstream training pipelines. Furthermore, consistently reduces subgroup performance gaps, as shown by universal improvements in performance and fairness metrics across four experimental setups. more »

Award ID(s):: 2205289

PAR ID:: 10571913

Author(s) / Creator(s):: Theodorou, Brandon; Danek, Benjamin; Tummala, Venkat; Kumar, Shivam_Pankaj; Malin, Bradley; Sun, Jimeng

Publisher / Repository:: Nature Publishing Group

Date Published:: 2025-02-14

Journal Name:: npj Digital Medicine

Volume:: 8

Issue:: 1

ISSN:: 2398-6352

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Journal Article:
https://doi.org/10.1038/s41746-025-01438-z

More Like this