scieee AI-readable full text Open interactive document viewer

LLM-Assisted Expansion of Patent and Scholarly Literature Knowledge Graphs

Rattinger, Andre; Gütl, Christian

Abstract

Appeared in: Open Search Symposium 2025, 8-10 October 2025, CSC IT Center for Science, Helsinki, Finland.

Full text

LLM-ASSISTED EXPANSION OF PATENT AND SCHOLARLY LITERATURE KNOWLEDGE GRAPHS André Rattinger ∗, ISDS, Graz University of Technology, Graz, Austria Christian Gütl, ISDS, Graz University of Technology, Graz, Austria Abstract Patents and scholarly articles represent two deeply connected yet distinct sources of technical knowledge, each using specialized terminologies and referencing structures. Standard methods for connecting these sources—such as citation-based retrieval or classification overlaps—often miss nuanced or implicit relationships. To address this, we propose an LLM-assisted knowledge graph expansion pipeline, combining semantic embeddings and topological structures, validated through domain-specific constraints. Demonstrated initially within battery technology (CPC “H01M”), this pipeline generalizes effectively to other technical fields, enhancing knowledge discovery, prior art search, and strategic innovation analysis. INTRODUCTION Technological innovations documented in patents and scientific knowledge captured in scholarly articles represent two complementary yet distinct knowledge bases. Despite their interconnected nature, the integration of patents with academic literature is often minimal due to variations in language, structure, and referencing practices. Previous work has demonstrated the benefits of both semantic and topological graphs for patent retrieval and analysis [1,2], yet implicit, deeper connections between these domains frequently remain undiscovered. We propose to leverage Large Language Models (LLMs) to systematically uncover and validate latent semantic overlaps, connecting patent and publication knowledge graphs into a unified and enriched network. Patents tend to emphasize legal and commercial aspects—claims, novelty, and scope of protection—while academic publications focus on rigor, reproducibility, and theoretical grounding. This disparity leads to variations in language (highly specialized or obfuscated legalese vs. structured academic prose), structure (claims vs. hypotheses), and referencing practices (formal classification codes vs. standard bibliographic citations). Consequently, direct links between these two corpora are frequently underexploited, hampering comprehensive knowledge discovery and prior art analysis [3]. Efforts to unify patents with academic literature typically rely on classification (e.g., CPC codes). While these methods excel at connecting documents in well-traversed or semantically obvious paths, they do not always capture deeper relationships—such as an unreferenced publication describing a method that closely matches a patented process [4]. Moreover, large-scale knowledge graphs, though promising, ∗[email protected] can be limited by the static nature of their input data; novel or implicit links remain hidden unless explicitly recorded [5]. Recent work in semantic patent graphs [1, 6] has shown that text-based embeddings (e.g., doc2vec or BERT) can reveal relationships overlooked by purely topological approaches. Simultaneously, topological graphs that rely on CPC co-classification or patent citations provide a robust, expert-assigned backbone [1, 7]. Despite these advances, the resulting graphs still tend to miss subtle overlaps across different classification categories or cross-disciplinary leaps between specialized research articles and patents. In parallel, the rapid development of Large Language Models (LLMs) provides an opportunity to infer novel connections from natural language descriptions. LLMs can distill textual snippets—such as patent claims and scientific abstracts—into conceptual links, potentially labeling relationships with statements like “these two documents describe the same doping method” or “this publication could serve as potential prior art for that patent” [8]. However, LLMs are prone to overreach or “hallucinate,” meaning any pipeline that leverages them must integrate domain-aware safeguards (e.g., checking chemical doping terms or verifying consistent references) to ensure correctness [4]. Given these convergent trends, we propose a unified pipeline that fuses semantic embeddings, topological structures, and LLM-based edge inference. Our pipeline begins by constructing a baseline knowledge graph from known links (citations, classifications, textual similarity), then solicits an LLM to propose new relationships where only moderate textual overlap is present. Potentially spurious suggestions are filtered by domain heuristics and textual consistency checks. The result is an expanded, domain-agnostic knowledge graph that better reflects the true scope of innovation and scientific discovery. Though illustrated using battery technology patents (CPC “H01M”) due to its active and well-documented R&D environment [7], the approach naturally extends to other domains, such as AI (CPC “G06N”) or semiconductor devices (CPC “H01L”). RELATED WORK Early attempts to integrate patents and academic literature leaned primarily on citation extraction and classification overlaps, often treating each corpus separately or relying on manual heuristics. These methods excel at connecting well-documented prior art but often fail to expose deeper semantic relationships [4]. More recent approaches harness semantic embeddings to represent patents and publications in a shared vector space [6], enabling automated identification of potentially https://doi.org/10.5281/zenodo.17238499 related documents that lack direct citations. Parallel efforts focus on topological graphs built from co-classification or citation networks [1,7], leveraging expert labeling to anchor patent knowledge bases. However, these purely topological approaches may overlook subtle or emerging links, especially across interdisciplinary boundaries [3]. A promising trend is LLM-assisted knowledge graph completion, where large language models infer new edges in graphs by understanding textual descriptions [5, 8]. While this has been demonstrated in general knowledge graphs and biomedical contexts [3], it remains less explored in patent–publication integration pipelines. Incorporating LLMs offers the possibility to bridge semantic gaps by interpreting specialized legal or scientific language and inferring relationships not explicitly stated [4]. The main challenge lies in mitigating LLM “hallucinations”, underscoring the importance of domain-aware verification—such as chemical formula matching or consistency checks against domain ontologies. Our work leverages these advances by introducing an LLM-driven layer to a base knowledge graph of patent–publication pairs. Through a combination of semantic embeddings, co-classification, citation data, and LLMbased link suggestions, we address both the missing explicit links and the subtler overlaps across conceptually adjacent technologies and research fronts. This approach is particularly apt for high-innovation domains like battery technology (CPC “H01M”), where rapid advancements outpace the coverage of static classification and citation systems [2,7]. REFERENCES [1] A. Rattinger et al., “Semantic and topological patent graphs,” in SNAMS 2018: International Conference on Social Networks Analysis, Management and Security, 2018, https://example. org/semantic-topological-patent-graphs. [2] A. Rattinger, J.-M. Le Goff, and C. Guetl, “Semantic and topological graphs for patent retrieval,” in 2019 Sixth International Conference on Social Networks Analysis, Management and Security (SNAMS). IEEE, 2019, pp. 175–180. [3] J. Xu, C. Yu, J. Xu, Y. Ding, V. I. Torvik, J. Kang, M. Sung, and M. Song, “Pubmed knowledge graph 2.0: Connecting papers, patents, and clinical trials in biomedical science,” arXiv preprint arXiv:2410.07969, 2024. [4] H. Aras, R. Dessi, F. Saad, and L. Zhang, “Bridging the innovation gap: Leveraging patent information for scientists by constructing a patent-centric knowledge graph,” in CEUR Workshop Proceedings, vol. 3697, 2024, pp. 61–67. [5] L. Yao, J. Peng, C. Mao, and Y. Luo, “Exploring large language models for knowledge graph completion,” in ICASSP 20252025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5. [6] L. Siddharth, G. Li, and J. Luo, “Enhancing patent retrieval using text and knowledge graph embeddings: a technical note,” Journal of Engineering Design, vol. 33, no. 8-9, pp. 670–683, 2022. [7] H. Pohl and M. Marklund, “Battery research and innovation—a study of patents and papers,” World Electric Vehicle Journal, vol. 15, no. 5, p. 193, 2024. [8] B. Mo, K. Yu, J. Kazdan, P. Mpala, L. Yu, C. Cundy, C. Kanatsoulis, and S. Koyejo, “Kggen: Extracting knowledge graphs from plain text with language models,” arXiv preprint arXiv:2502.09956, 2025. https://doi.org/10.5281/zenodo.17238499