Putting the pieces together - Integrating clinical data with domain knowledge for advanced knowledge generation and prediction.
Abstract
Poster and abstract for “Putting the pieces together - Integrating clinical data with domain knowledge for advanced knowledge generation and prediction.” by Gehrmann, presented at the AIKD-SD 2025 Summer School co-located with the NFDI4DS Conference 2025.
Full text
AIKG-SD 2025 Summer School co-located with the NFDI4DS Conference 2025 November 25-26, 2025, Berlin, Germany 1 Putting the pieces together - Integrating clinical data with domain knowledge for advanced knowledge generation and prediction. Julia Gehrmann (ORCiD 0000-0002-4101-5458) Institute for Biomedical Informatics, University of Cologne, Faculty of Medicine and University Hospital Cologne Abstract Data capturing software like REDCap enables clinicians to document patient-level clinical data in a structured and well-annotated way [1]. However, the data export from REDCap is tabular, not covering potential semantic connections between features. These connections can be relevant for downstream tasks generating new knowledge or predictions from the REDCap data [2]. Therefore, we aim at developing an automated and reusable pipeline transforming REDCap exports into patient-level vector representations, enriched with domain knowledge from clinical terminology SNOMED CT and insights from scholarly literature via Knowledge Graph (KG) modeling. The enhanced vector representations can eventually serve as input for downstream tasks such as classification and clustering. The transformation process begins with the mapping of variable names to SNOMED CT concepts, leveraging the clinical terminology to establish a semantic foundation for the KG. Based on the resulting mappings and the REDCap metadata, we construct a KG in which nodes represent patients, clinical events, and individual measurements. Each measurement is connected to the respective patient and clinical event via a labeled edge. Moreover, we introduce labeled edges between measurement nodes connected to the same patient and event if their associated SNOMED CT concepts have a defined semantic relationship. This step embeds medical hierarchies and domain knowledge into the graph structure [3]. To further enrich the graph, we incorporate scholarly data by using a large language model (LLM) fine-tuned on manually selected, domain-specific literature. The fine-tuned LLM assesses the contextual strength of relationships between clinical concepts based on literaturederived relevance and assigns respective weights to edges. The final KG consists of nodes whose attributes comprise the original measurement values from the REDCap export and semantic labels from SNOMED CT, while edge weights reflect both ontology-defined and literature-informed concept relationships. In a second transformation process, we use a graph autoencoder to learn low-dimensional embeddings of this enriched KG. By jointly considering the graph structure, edge weights, and node attributes, the autoencoder generates vector representations that retain both the clinical data from REDCap and the qualitative context derived from ontologies and scholarly resources. Compared to the original REDCap export, the resulting embeddings offer several key advantages: (1) They incorporate hierarchical relationships between clinical variables as defined in SNOMED CT. (2) They comprise current biomedical knowledge through literature-informed edge weighting. (3) They provide a compact, semantically aware
AIKG-SD 2025 Summer School co-located with the NFDI4DS Conference 2025 November 25-26, 2025, Berlin, Germany 2 representation suitable for machine learning tasks in medical research. Furthermore, our KGbased approach fully supports heterogeneity by allowing for further data modalities to be integrated if available. The automated nature of our proposed pipeline, moreover, supports typical heterogeneity between medical centers such as different REDCap database structures. Above all, however, our approach aims to bridge the gap between structured clinical data and established biomedical knowledge, enabling a more informed and context-aware application of machine learning in healthcare. References [1] Harris, P. A., Taylor, R., Thielke, R., Payne, J., Gonzalez, N., & Conde, J. G. (2009). Research electronic data capture (REDCap)—a metadata-driven methodology and workflow process for providing translational research informatics support. Journal of biomedical informatics, 42(2), 377-381. [2] Oh, S. (2019). Feature interaction in terms of prediction performance. Applied Sciences, 9(23), 5191. [3] Chang, E., & Mostafa, J. (2021). The use of SNOMED CT, 2013-2020: a literature review. Journal of the American Medical Informatics Association, 28(9), 2017-2026.