NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval

Jiang, Pengcheng; Xiao, Cao; Jiang, Minhao; Bhatia, Parminder; Kass-Hout, Taha; Sun, Jimeng; Han, Jiawei (May 2025, ICLR/OpenReview)

Free, publicly-accessible full text available May 18, 2026
Bi-level Contrastive Learning for Knowledge-Enhanced Molecule Representations

https://doi.org/10.1609/aaai.v39i1.32013

Jiang, Pengcheng; Xiao, Cao; Fu, Tianfan; Bhatia, Parminder; Kass-Hout, Taha; Sun, Jimeng; Han, Jiawei (April 2025, Proceedings of the AAAI Conference on Artificial Intelligence)

Molecular representation learning is vital for various downstream applications, including the analysis and prediction of molecular properties and side effects. While Graph Neural Networks (GNNs) have been a popular framework for modeling molecular data, they often struggle to capture the full complexity of molecular representations. In this paper, we introduce a novel method called Gode, which accounts for the dual-level structure inherent in molecules. Molecules possess an intrinsic graph structure and simultaneously function as nodes within a broader molecular knowledge graph. Gode integrates individual molecular graph representations with multi-domain biochemical data from knowledge graphs. By pre-training two GNNs on different graph structures and employing contrastive learning, Gode effectively fuses molecular structures with their corresponding knowledge graph substructures. This fusion yields a more robust and informative representation, enhancing molecular property predictions by leveraging both chemical and biological information. When fine-tuned across 11 chemical property tasks, our model significantly outperforms existing benchmarks, achieving an average ROC-AUC improvement of 12.7% for classification tasks and an average RMSE/MAE improvement of 34.4% for regression tasks. Notably, Gode surpasses the current leading model in property prediction, with advancements of 2.2% in classification and 7.2% in regression tasks.
more » « less
Free, publicly-accessible full text available April 11, 2026
Certifiably Byzantine-Robust Federated Conformal Prediction

Kang, Mintong; Lin, Zhen; Sun, Jimeng; Xiao, Cao; Li, Bo (July 2024, International Conference on Machine Learning (ICML 2024))

Full Text Available
Synthesizing Multimodal Electronic Health Records via Predictive Diffusion Models

https://doi.org/10.1145/3637528.3671836

Zhong, Yuan; Wang, Xiaochen; Wang, Jiaqi; Zhang, Xiaokun; Wang, Yaqing; Huai, Mengdi; Xiao, Cao; Ma, Fenglong (August 2024, ACM)

Full Text Available
Recent Advances in Predictive Modeling with Electronic Health Records

https://doi.org/10.24963/ijcai.2024/914

Wang, Jiaqi; Luo, Junyu; Ye, Muchao; Wang, Xiaochen; Zhong, Yuan; Chang, Aofei; Huang, Guanjie; Yin, Ziyi; Xiao, Cao; Sun, Jimeng; et al (August 2024, International Joint Conferences on Artificial Intelligence Organization)

The development of electronic health records (EHR) systems has enabled the collection of a vast amount of digitized patient data. However, utilizing EHR data for predictive modeling presents several challenges due to its unique characteristics. With the advancements in machine learning techniques, deep learning has demonstrated its superiority in various applications, including healthcare. This survey systematically reviews recent advances in deep learning-based predictive models using EHR data. Specifically, we introduce the background of EHR data and provide a mathematical definition of the predictive modeling task. We then categorize and summarize predictive deep models from multiple perspectives. Furthermore, we present benchmarks and toolkits relevant to predictive modeling in healthcare. Finally, we conclude this survey by discussing open challenges and suggesting promising directions for future research.
more » « less
Full Text Available
TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale

https://doi.org/10.18653/V1/2024.NAACL-LONG.154

Jiang, Pengcheng; Xiao, Cao; Wang, Zifeng; Bhatia, Parminder; Sun, Jimeng; Han, Jiawei (January 2024, Association for Computational Linguistics)
Duh, Kevin; G'omez-Adorno, Helena; Bethard, Steven (Ed.)
The advent of large language models (LLMs) has significantly advanced natural language processing tasks like text summarization. However, their large size and computational demands, coupled with privacy concerns in data transmission, limit their use in resourceconstrained and privacy-centric settings. To overcome this, we introduce TriSum, a framework for distilling LLMs’ text summarization abilities into a compact, local model. Initially, LLMs extract a set of aspect-triple rationales and summaries, which are refined using a dualscoring method for quality. Next, a smaller local model is trained with these tasks, employing a curriculum learning strategy that evolves from simple to complex tasks. Our method enhances local model performance on various benchmarks (CNN/DailyMail, XSum, and ClinicalTrial), outperforming baselines by 4.5%, 8.5%, and 7.4%, respectively. It also improves interpretability by providing insights into the summarization rationale.
more » « less
Full Text Available
BIPEFT: Budget-Guided Iterative Search for Parameter Efficient Fine-Tuning of Large Pretrained Language Models

https://doi.org/10.18653/v1/2024.findings-emnlp.437

Chang, Aofei; Wang, Jiaqi; Liu, Han; Bhatia, Parminder; Xiao, Cao; Wang, Ting; Ma, Fenglong (January 2024, Association for Computational Linguistics)

Full Text Available
ClinicalRisk: A New Therapy-related Clinical Trial Dataset for Predicting Trial Status and Failure Reasons

https://doi.org/10.1145/3583780.3615113

Luo, Junyu; Qiao, Zhi; Glass, Lucas; Xiao, Cao; Ma, Fenglong (October 2023, ACM)
Unity in Diversity: Collaborative Pre-training Across Multimodal Medical Sources

https://doi.org/10.18653/v1/2024.acl-long.199

Wang, Xiaochen; Luo, Junyu; Wang, Jiaqi; Zhong, Yuan; Zhang, Xiaokun; Wang, Yaqing; Bhatia, Parminder; Xiao, Cao; Ma, Fenglong (January 2024, Association for Computational Linguistics)

Full Text Available
Artificial intelligence foundation for therapeutic science

https://doi.org/10.1038/s41589-022-01131-2

Huang, Kexin; Fu, Tianfan; Gao, Wenhao; Zhao, Yue; Roohani, Yusuf; Leskovec, Jure; Coley, Connor W.; Xiao, Cao; Sun, Jimeng; Zitnik, Marinka (October 2022, Nature Chemical Biology)

Full Text Available

« Prev Next »

Search for: All records