scieee AI-readable full text Open interactive document viewer

Polymer Data Extraction and KG Population Using LLMs

Anmol Saini; Ethier, Jeffrey; Shimizu, Cogan

Abstract

Computer-aided research techniques for accelerating scientific discovery in polymer (and materials) science has continued to grow in both utilization and access. There remain limitations, however, especially in the curation of data. Currently, data is primarily extracted and compiled from publications and other forms of text manually. This process can be time-consuming and many existing forms of representation are rigid, unable to account for the evolution of data. Knowledge graphs – and ontology – provide a representation that allows for the complex nature of polymer data, but still needs to be populated with data from literature. Given the recent successes of large language models in interpreting massive natural language corpora, we propose a pipeline for populating a modular knowledge graph that captures state-of-the-art polymer characterizations in combination with experimental metadata and methodology.

Full text

Polymer Data Extraction and KG Population Using LLMs Anmol Saini, Jeffrey Ethier, Cogan Shimizu Wright State University US Air Force Research Laboratory 2 A KG-Powered Research Assistant for Polymer Science Context - Conversational agent augmented with knowledge graphs to help with polymer experiments - Expediting discovery of polymers - Reducing rote aspects of polymer scientists’ workload - Advancement of self-driving labs in the domain of polymer science 3 KGWRAPS: Foundations Experiment Design - Source of Truth (SoT) - Graph Analytics - Knowledge Gap Identification Conversational Capabilities - Conversational Model - Integrate Conversational Model with LLM - RAG, Fine-tuning, Prompt Engineering 4 KGWRAPS: Foundations Experiment Design - Source of Truth (SoT) - Graph Analytics - Knowledge Gap Identification Conversational Capabilities - Conversational Model - Integrate Conversational Model with LLM - RAG, Fine-tuning, Prompt Engineering 5 Introduction Problems - Process: Manual data curation - Storage: Disparate collections of data Solutions - Process: Automatic data extraction - Current experiments - Previous work - Storage: Knowledge graphs (KGs) - Rich semantics and contextual relevance - Precise data modeling and limited ambiguity - Flexible and extensible 6 Approach Enslaved.Org Hub Ontology - Text obtained from Wikipedia - Ground truth from Enslaved.org KG - Text Summarization - Retrieval-Augmented Generation (RAG) - Ontology module population by schema relationships for each module and attributes described for that relation in the text files 7 Approach Extended to Polymer Science - Simpler initial version of Enslaved.org work - No chunking of text - Ontology module focusing on metadata of publication - Validation by inspection with XMLs instead of KG 8 Inputs Prompt Ontology Module 9 Llama 3.2 - Limited understanding of task and prompt - Subopmtial results