Tackling Structured Knowledge Extraction from Polymer Nanocomposite Literature as an NER/RE Task with seq2seq

Hu, Bingyin; Lin, Anqi; Brinson, L Catherine

doi:10.1007/s40192-024-00363-5

Citation Details

Tackling Structured Knowledge Extraction from Polymer Nanocomposite Literature as an NER/RE Task with seq2seq

There is an urgent need for ready access to published data for advances in materials design, and natural language processing (NLP) techniques offer a promising solution for extracting relevant information from scientific publications. In this paper, we present a domain-specific approach utilizing a Transformer-based model, T5, to automate the generation of sample lists in the field of polymer nanocomposites (PNCs). Leveraging large-scale corpora, we employ advanced NLP techniques including named entity recognition and relation extraction to accurately extract sample codes, compositions, group references, and properties from PNC papers. The T5 model demonstrates competitive performance in relation extraction using a TANL framework and an EM-style input sequence. Furthermore, we explore multi-task learning and joint-entity-relation extraction to enhance efficiency and address deployment concerns. Our proposed methodology, from corpora generation to model training, showcases the potential of structured knowledge extraction from publications in PNC research and beyond. more »

Award ID(s):: 2022040 1835677 1835648

PAR ID:: 10539882

Author(s) / Creator(s):: Hu, Bingyin; Lin, Anqi; Brinson, L Catherine

Publisher / Repository:: Springer

Date Published:: 2024-09-01

Journal Name:: Integrating Materials and Manufacturing Innovation

Volume:: 13

Issue:: 3

ISSN:: 2193-9764

Page Range / eLocation ID:: 656 to 668

Format(s):: Medium: X

Sponsoring Org:: National Science Foundation

Free Publicly Accessible Full Text
Accepted Manuscript1.0
Journal Article:
https://doi.org/10.1007/s40192-024-00363-5

More Like this