NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval

Jiang, Pengcheng; Xiao, Cao; Jiang, Minhao; Bhatia, Parminder; Kass-Hout, Taha; Sun, Jimeng; Han, Jiawei (May 2025, ICLR/OpenReview)

Free, publicly-accessible full text available May 18, 2026
Bi-level Contrastive Learning for Knowledge-Enhanced Molecule Representations

https://doi.org/10.1609/aaai.v39i1.32013

Jiang, Pengcheng; Xiao, Cao; Fu, Tianfan; Bhatia, Parminder; Kass-Hout, Taha; Sun, Jimeng; Han, Jiawei (April 2025, Proceedings of the AAAI Conference on Artificial Intelligence)

Molecular representation learning is vital for various downstream applications, including the analysis and prediction of molecular properties and side effects. While Graph Neural Networks (GNNs) have been a popular framework for modeling molecular data, they often struggle to capture the full complexity of molecular representations. In this paper, we introduce a novel method called Gode, which accounts for the dual-level structure inherent in molecules. Molecules possess an intrinsic graph structure and simultaneously function as nodes within a broader molecular knowledge graph. Gode integrates individual molecular graph representations with multi-domain biochemical data from knowledge graphs. By pre-training two GNNs on different graph structures and employing contrastive learning, Gode effectively fuses molecular structures with their corresponding knowledge graph substructures. This fusion yields a more robust and informative representation, enhancing molecular property predictions by leveraging both chemical and biological information. When fine-tuned across 11 chemical property tasks, our model significantly outperforms existing benchmarks, achieving an average ROC-AUC improvement of 12.7% for classification tasks and an average RMSE/MAE improvement of 34.4% for regression tasks. Notably, Gode surpasses the current leading model in property prediction, with advancements of 2.2% in classification and 7.2% in regression tasks.
more » « less
Free, publicly-accessible full text available April 11, 2026
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models

https://doi.org/10.18653/V1/2024.NAACL-LONG.155

Jiang, Pengcheng; Lin, Jiacheng; Wang, Zifeng; Sun, Jimeng; Han, Jiawei (January 2024, Association for Computational Linguistics)
Duh, Kevin; G'omez-Adorno, Helena; Bethard, Steven (Ed.)
The field of relation extraction (RE) is experiencing a notable shift towards generative relation extraction (GRE), leveraging the capabilities of large language models (LLMs). However, we discovered that traditional relation extraction (RE) metrics like precision and recall fall short in evaluating GRE methods. This shortfall arises because these metrics rely on exact matching with human-annotated reference relations, while GRE methods often produce diverse and semantically accurate relations that differ from the references. To fill this gap, we introduce GENRES for a multidimensional assessment in terms of the topic similarity, uniqueness, granularity, factualness, and completeness of the GRE results. With GENRES, we empirically identified that (1) precision/recall fails to justify the performance of GRE methods; (2) human-annotated referential relations can be incomplete; (3) prompting LLMs with a fixed set of relations or entities can cause hallucinations. Next, we conducted a human evaluation of GRE methods that shows GENRES is consistent with human preferences for RE quality. Last, we made a comprehensive evaluation of fourteen leading LLMs using GENRES across document, bag, and sentence level RE datasets, respectively, to set the benchmark for future research in GRE.
more » « less
Full Text Available
TriSum: Learning Summarization Ability from Large Language Models with Structured Rationale

https://doi.org/10.18653/V1/2024.NAACL-LONG.154

Jiang, Pengcheng; Xiao, Cao; Wang, Zifeng; Bhatia, Parminder; Sun, Jimeng; Han, Jiawei (January 2024, Association for Computational Linguistics)
Duh, Kevin; G'omez-Adorno, Helena; Bethard, Steven (Ed.)
The advent of large language models (LLMs) has significantly advanced natural language processing tasks like text summarization. However, their large size and computational demands, coupled with privacy concerns in data transmission, limit their use in resourceconstrained and privacy-centric settings. To overcome this, we introduce TriSum, a framework for distilling LLMs’ text summarization abilities into a compact, local model. Initially, LLMs extract a set of aspect-triple rationales and summaries, which are refined using a dualscoring method for quality. Next, a smaller local model is trained with these tasks, employing a curriculum learning strategy that evolves from simple to complex tasks. Our method enhances local model performance on various benchmarks (CNN/DailyMail, XSum, and ClinicalTrial), outperforming baselines by 4.5%, 8.5%, and 7.4%, respectively. It also improves interpretability by providing insights into the summarization rationale.
more » « less
Full Text Available
Taxonomy-guided Semantic Indexing for Academic Paper Search

https://doi.org/10.18653/v1/2024.emnlp-main.407

Kang, SeongKu; Zhang, Yunyi; Jiang, Pengcheng; Lee, Dongha; Han, Jiawei; Yu, Hwanjo (January 2024, Association for Computational Linguistics)

Full Text Available
Text Augmented Open Knowledge Graph Completion via Pre-Trained Language Models

https://doi.org/10.18653/v1/2023.findings-acl.709

Jiang, Pengcheng; Agarwal, Shivam; Jin, Bowen; Wang, Xuan; Sun, Jimeng; Han, Jiawei (July 2023, Association for Computational Linguistics)
Rogers, Anna; Boyd-Graber, Jordan; Okazaki, Naoaki (Ed.)
The mission of open knowledge graph (KG) completion is to draw new findings from known facts. Existing works that augment KG completion require either (1) factual triples to enlarge the graph reasoning space or (2) manually designed prompts to extract knowledge from a pre-trained language model (PLM), exhibiting limited performance and requiring expensive efforts from experts. To this end, we propose TagReal that automatically generates quality query prompts and retrieves support information from large text corpora to probe knowledge from PLM for KG completion. The results show that TagReal achieves state-of-the-art performance on two benchmark datasets. We find that TagReal has superb performance even with limited training data, outperforming existing embedding-based, graph-based, and PLM-based methods.
more » « less
Full Text Available

Search for: All records