NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Retrieval and Structuring Augmented Generation with LLMs for Web Applications

https://doi.org/10.1145/3701716.3715870

Jiao, Yizhu; Ouyang, Siru; Zhong, Ming; Zhang, Yunyi; Ding, Linyi; Zhou, Sizhe; Han, Jiawei (May 2025, ACM)

Free, publicly-accessible full text available May 8, 2026
Automated Mining of Structured Knowledge from Text in the Era of Large Language Models

https://doi.org/10.1145/3637528.3671469

Zhang, Yunyi; Zhong, Ming; Ouyang, Siru; Jiao, Yizhu; Zhou, Sizhe; Ding, Linyi; Han, Jiawei (August 2024, ACM)
Baeza-Yates, Ricardo; Bonchi, Francesco (Ed.)
Massive amount of unstructured text data are generated daily, ranging from news articles to scientific papers. How to mine structured knowledge from the text data remains a crucial research question. Recently, large language models (LLMs) have shed light on the text mining field with their superior text understanding and instructionfollowing ability. There are typically two ways of utilizing LLMs: fine-tune the LLMs with human-annotated training data, which is labor intensive and hard to scale; prompt the LLMs in a zero-shot or few-shot way, which cannot take advantage of the useful information in the massive text data. Therefore, it remains a challenge on automated mining of structured knowledge from massive text data in the era of large language models. In this tutorial, we cover the recent advancements in mining structured knowledge using language models with very weak supervision. We will introduce the following topics in this tutorial: (1) introduction to large language models, which serves as the foundation for recent text mining tasks, (2) ontology construction, which automatically enriches an ontology from a massive corpus, (3) weakly-supervised text classification in flat and hierarchical label space, (4) weakly-supervised information extraction, which extracts entity and relation structures.
more » « less
Full Text Available
Grasping the Essentials: Tailoring Large Language Models for Zero-Shot Relation Extraction

https://doi.org/10.18653/v1/2024.emnlp-main.747

Zhou, Sizhe; Meng, Yu; Jin, Bowen; Han, Jiawei (January 2024, Association for Computational Linguistics)

Full Text Available
Text2DB: Integration-Aware Information Extraction with Large Language Model Agents

https://doi.org/10.18653/v1/2024.findings-acl.12

Jiao, Yizhu; Li, Sha; Zhou, Sizhe; Ji, Heng; Han, Jiawei (January 2024, Association for Computational Linguistics)

Full Text Available
Topic-Oriented Open Relation Extraction with A Priori Seed Generation

https://doi.org/10.18653/v1/2024.emnlp-main.766

Ding, Linyi; Xiao, Jinfeng; Zhou, Sizhe; Yang, Chaoqi; Han, Jiawei (January 2024, Association for Computational Linguistics)

Full Text Available
Geospatial Knowledge Hypercube

https://doi.org/10.1145/3589132.3625629

Wang, Zhaonan; Jin, Bowen; Hu, Wei; Jiang, Minhao; Kang, Seungyeon; Li, Zhiyuan; Zhou, Sizhe; Han, Jiawei; Wang, Shaowen (November 2023, ACM)

Today a tremendous amount of geospatial knowledge is hidden in massive volumes of text data. To facilitate flexible and powerful geospatial analysis and applications, we introduce a new architecture: geospatial knowledge hypercube, a multi-scale, multidimensional knowledge structure that integrates information from geospatial dimensions, thematic themes and diverse application semantics, extracted and computed from spatial-related text data. To construct such a knowledge hypercube, weakly supervised language models are leveraged for automatic, dynamic and incremental extraction of heterogeneous geospatial data, thematic themes, latent connections and relationships, and application semantics, through combining a variety of information from unstructured text, structured tables, and maps. The hypercube lays a foundation for many knowledge discovery and in-depth spatial analysis, and other advanced applications. We have deployed a prototype web application of proposed geospatial knowledge hypercube for public access at: https://hcwebapp.cigi.illinois.edu/.
more » « less
Full Text Available
Corpus-Based Relation Extraction by Identifying and Refining Relation Patterns

https://doi.org/10.1007/978-3-031-43421-1_2

Zhou, Sizhe; Ge, Suyu; Shen, Jiaming; Han, Jiawei (January 2023, Springer Nature Switzerland)

Full Text Available

Search for: All records