Full text
HetERogeneous sEmantic Data integratIon for the guT-bRain interplaY Deliverable 5.5 Querying Components Version 1.10, 13.08.2025
EXECUTIVE SUMMARY HetERogeneous sEmantic Data integratIon for the guT-bRain interplaY (HEREDITARY) develops novel approaches for analysis and exploration of multi-modal biomedlical data. Multi-modality poses a challenge to data analysis, due to the divergent data types, scales, or levels of detail at which these may be given. Current approaches can represent multi-modal data in unified data structures, such as Knowledge Graphs (KGs), and tabular data. Due to the complex nature of the data, and also, expert-level interfaces like SQL and SPARQL to such data sets, users often have problems to effectively search, find, and discover insights in such data sets with ease. In this HEREDITARY deliverable, we present novel approaches for effective user access to complex data, including retrieval, explanation, and comparison of data. Specifically, we leverage the potential of state of the art Large Language Models. We show how using natural language, users from novice to expert level, can express their information need in a natural language statement. The system applies these statements to create answers for the user, and can involve users in an information-seeking dialogue. Our approaches operate on KG data (e.g., as obtained from information extraction algorithms of lage publication data), and on tabular data (e.g., structured complex patient information from clinical research). Preliminary evaluation shows the large potential of these approaches for retrieval and exploration in complex and multi-modal data. Building on this, in followup work we will refine, integrate and evaluate these querying components as part of requirement engineering, and scientific dissemination. DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 3 | 50
DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 4 | 50
DOCUMENT INFORMATION Deliverable ID 5.5 Deliverable Title Querying Components Work Package WP 5 Lead Partner TU Graz Due Date 31.08.2025 Date of submission 13.08.2025 Deliverable Type R – Report Dissemination level PU – Public AUTHORS Name Organization Tobias Schreck TUGRAZ Svetla Boytcheva ONTO Stefan Lengauer TUGRAZ Peter Waldert TUGRAZ Benedikt Kantz TUGRAZ Aleksis Datseris ONTO Tsvetelina Koleva ONTO Todor Primov ONTO Ivelina Nikolova-Koleva ONTO Juan Manuel Rodriguez (Internal reviewer) AAU Gianmaria Silvello (Contributor) UNIPD REVISION HISTORY Version Date Author Description 0.10 2025-02-28 Peter Waldert Structure for deliverable 0.20 2025-07-18 All Authors Complete draft including all components 0.30 2025-07-21 All Authors Added summary and finalized for internal review 1.00 2025-08-01 All Authors Revised and finalized with internal reviewer comments 1.10 2025-08-05 Gianmaria Silvello Final check, layout fixes and minor changes Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union. Neither the European Union nor the granting authority can be held responsible for them. DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 5 | 50
Contents 1 Introduction 10 2 OnSET Query Editor 11 2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 11 2.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.2.1 NLP Query Approaches . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.2.2 Visual Query Approaches . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 12 2.2.3 Graph Difference Views . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 13 2.3 Methodology............................................. 13 2.3.1 User Guidance . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 2.3.2 Graph Definition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.3.3 Difference Graphs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 2.3.4 Automatic Result Set Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.3.5 Integration of a Natural Language Query System . . . . . . . . . . . . . . . . . . . . . 16 2.3.6 User Interface . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.4 Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 2.5 Casestudy.............................................. 18 2.6 Conclusion.............................................. 19 2.7 FutureWork............................................. 20 3 Graph Queries from Natural Language using Constrained Language Models & Evaluation 21 3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 3.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 21 3.3 Methodology............................................. 22 3.3.1 Graph Extraction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.3.2 Synthetic Evaluation Methodology . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 23 3.3.3 Constraining the LM . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24 3.3.4 User Interface Integration . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 25 3.4 Results................................................ 25 3.5 Conclusion.............................................. 27 3.6 FutureWork............................................. 27 4 Talk to Your Graph: Natural Language Querying of Neuro Degenerative Diseases 28 4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 4.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 28 4.3 Data ................................................. 28 4.4 Methodology............................................. 30 4.4.1 Natural Language Querying . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 4.4.2 Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 4.5 Experiments and Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 35 4.6 Furtherwork............................................. 35 DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 6 | 50
5 Neurodegen-Vis: LLM-based, Privacy-Preserving Visual Exploration of High-Dimensional ALS, PD, MS Patient Data 39 5.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 5.1.1 Dataset............................................ 40 5.1.2 Anonymisation using Differential Privacy . . . . . . . . . . . . . . . . . . . . . . . . . . 40 5.2 Related Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41 5.3 Visualisation Approach & Methods . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 5.3.1 Natural Language Interface . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 42 5.4 Implementation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 44 5.5 UseCase .............................................. 44 5.6 Conclusion & Future Work . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 45 6 Summary and Further Development 45 References 46 List of Figures 1 Ontology and Sparse data Exploration Tool (OnSET) user flow. The user can select topics of interest and retrieve possible start links. These links are then expanded & constrained within an editor, which finally retrieves different instances of the searched graph. . . . . . . . . . . . 13 2 User guidance within OnSET. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 3 Query builder interface within OnSET. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 4 Difference graph examples for the DBpedia ontology. . . . . . . . . . . . . . . . . . . . . . . . 18 5 The Brainteaser Ontology (BTO) KG queried using the difference view to explore the relationship between different semantic attributes. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 6 Prototype Graph extraction process from query. . . . . . . . . . . . . . . . . . . . . . . . . . . 22 7 Prototype graph extraction results over different query complexities and ontologies. . . . . . . 26 8 A data inventory organized as a map of data domains . . . . . . . . . . . . . . . . . . . . . . 29 9 BioGraphTalk KG . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 29 10 BioGraphTalk Class Relationship . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 30 11 BioGraphTalk Class Hierarchy . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31 12 The High-Level Architecture of how GraphDBs NLQ works. . . . . . . . . . . . . . . . . . . . 31 13 GraphDB Talk To Your Graph setup . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 32 14 TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - the answer summary in natural language . . . . . . . . . . . . . . . . . . . . . . . . . 36 15 TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - the generated SPARQL query . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 37 16 TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - the generated SPARQL query run in GraphDB over BioGraphTalk . . . . . . . . . . . 38 17 TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - visual exploration of the result in GraphDB over BioGraphTalk . . . . . . . . . . . . . 38 18 Amyotrophic Lateral Sclerosis (ALS), Parkinson’s Disease (PD), Multiple Sclerosis (MS) visualisation dashboard . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 39 19 The feature languge z-component in the histogram view. Colour coded is the categorical feature z-diagnosis. ............................................. 42 DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 7 | 50
List of Tables 1 Comparison of query generation methods for Hermes 3B on DBpedia with one shot prompts. . 26 Acronyms ALS Amyotrophic Lateral Sclerosis. 7, 9 BGP Basic Graph Pattern. 9, 14, 15 BRAINTEASER Bringing Artificial Intelligence home for a better care of amyotrophic lateral sclerosis and multiple sclerosis. 9 BTO Brainteaser Ontology. 7, 9, 14, 18–20, 25, 26 CEST-MRI Chemical Exchange Saturation Transfer Magnetic Resonance Imaging. 9 DP Differential Privacy. 9 ETL Extraction, Transform and Load. 9 GBNF GGML Backus-Naur Form. 9, 23 GED Graph Edit Distance. 9, 24, 25 GMM Gaussian Mixture Models. 9 GPT Generative Pre-trained Transformer. 9, 42 GUI Graphical User Interface. 9 HEREDITARY HetERogeneous sEmantic Data integratIon for the guT-bRain interplaY. 3, 9, 10, 35, 45 HERO HEREDITARY Ontology. 9–11 HMM Hidden Markov Model. 9 iDPP Intelligent Disease Progression Prediction. 9 IR Information Retrieval. 9, 11, 12, 19, 21 KG Knowledge Graph. 3, 7, 9–12, 15–23, 25, 27–29, 35, 45 LLM Large Language Model. 9, 28, 43 LM Language Model. 9–12, 14, 16, 20–25, 27, 40–45 MS Multiple Sclerosis. 7, 9 MWEM Multiplicative Weights Exponential Mechanism. 9, 41 NLIs Natural Language Interface. 9, 41 NLP Natural Language Processing. 9–12, 21, 45 OnSET Ontology and Sparse data Exploration Tool. 7, 9–11, 13–16, 18–21, 25, 45 OQL Ontology Query Language. 9 DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 8 | 50
PCA Principal Component Analysis. 9, 39, 42–44 PD Parkinson’s Disease. 7, 9, 40, 44 PGO Performance-guided optimisation. 9 RAG Retrieval Augmented Generation. 9, 22 SUS System Usability Score. 9 TTYG Talk To Your Graph. 9, 10, 28, 45 UMAP Uniform Manifold Approximation and Projection. 9 VA Visual Analytics. 9, 39, 40, 42, 44 WASM WebAssembly. 9 DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 9 | 50
2.3.4 Automatic Result Set Overview To enable a more in-depth result overview, we provide value-wise fetching of node attributes using the subquery elements mentioned above. The user may be searching for the ages of patients who have a diagnosis of a specific disease and wants to know the distribution of these dates, as there might be hundreds of patients. The result set overview enables a quick visualization of the attributes of the nodes within the graph using standard techniques for common data types. We employ a heuristic to select the most suitable visualization type based on the current setting. If the user chooses only one value, we show the feature distribution. If the feature is continuous we group it into a set number of buckets, if it is discrete we show the most common discrete types, as shown in Figure 4a. Furthermore, if the user chooses two value attributes, we either display a heatmap if there are too many data points or use a scatter plot to show the distribution of the results. For the patients example, the system would choose the histogram plot. We integrate these result set charts with our difference view by comparing the different result sets of the two prototype graphs, as long as the values selected to be retrieved still match. These could be the birthdate distributions of a specific disease compared to the overall distribution for our example. We also allow users to download difference sets to apply their own statistical and visual evaluations outside our system, provided they are satisfied with the information from our queries. 2.3.5 Integration of a Natural Language Query System We utilize the difference views to manage the integration of a LM-based query feature, allowing users to view the possible changes the LM would make based on a natural language query before accepting or rejecting them. This is especially helpful if the LM makes a mistake by suggesting the wrong links or hallucination (Ji et al., 2023). Other failure cases include an information need by the user that the KG cannot satisfy, but the user specified in their query. The assistance system may still return some changes that are loosely related to the user request, but the user can still modify them before accepting the changes. The query parsing to a structured change set follows the same architecture which we outline in Section 3. The change set is then applied to the prototype graph Gp, and the difference view is automatically activated, including a first initial difference calculation according to the schema outlined above. 2.3.6 User Interface The resulting prototype graph ((a)in Figure 3) is then utilized in three ways, similar to the levels of García et al. (2022), but integrated into the core retrieval process. First, the graph is used on top of a 3D circle-packing visualization of the ontology to illustrate how the classes are distributed over the ontology with respect to its class hierarchy (c). Second, the prototype graph is used to generate a SPARQL query to retrieve the intended instance set. This instance set populates the third use case of the graph, where we show small instances of the initial prototype graph (b). These instances, or result sets, can then be inspected, and the properties of the retrieved instances can be explored individually. The presentation of smaller, visually similar instances also differentiates our system from existing approaches, as we visually relate the result set and prototype graph. 2.4 Implementation The outlined concepts are integrated into our system, OnSET, focusing on fast user responses even on larger ontologies. We furthermore base our entire stack on open-source systems, allowing institutions or users to DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 16 | 50
(a) (b)(c) Figure 3: The query can be built using a straightforward interactive process. The user can add any allowed link within the ontology, with a visual indication of its prevalence within the KG shown by the link width. Users can also add constraints to the nodes. (a) The tool immediately provides visual results (b) and provides a visual indicator of the explored classes and links using a small three-dimensional circle packing “minimap” in the bottom center (c). start and tweak their local systems. We show the implementation of our user guidance in Figure 2, the query builder in Figure 3, and the difference views in Figures 4 and 5 To achieve our first goal of fast user responses, we only compute the topic modeling and embeddings of links and classes on the first startup with a specific ontology. We store our resulting hierarchical topic map and embeddings in PostgreSQL3with the help of the pgvector extension 4, which enables fast retrieval given a query embedding. To generate these embeddings quickly, even on commodity hardware, we use the stella_en_400M_v55model, which is, at the time of submission, the best-performing smaller model w.r.t. the massive text embedding benchmark (Muennighoff et al., 2022). The precomputed topics are generated using the same embedding model, and additional topic labels are generated using the Hermes Llama 3.2 8B model (Teknium et al., 2024). While these efforts improve the query time for building the prototype graph, live updates of the result set require a similarly fast system. To achieve those fast responses, even on more complex queries and on commodity hardware, we use qlever (Bast & Buchhold, 2017), a SPARQL query engine that outperforms most existing engines in both speed and system requirements. This speedup enables our system to serve and display updates as the user builds their query, aiding the user in retrieving non-empty sets and showing intermediate results to guide the search even further. Our user interface builds on Vue.js6, in combination with three.js7and D3.js (Bostock et al., 2011). All the used database systems and models are open-source and open-weight, providing state-of-the-art performance 3https://www.postgresql.org/ 4https://github.com/pgvector/pgvector 5https://huggingface.co/NovaSearch/stella_en_400M_v5 6https://vuejs.org/ 7https://threejs.org/ DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 17 | 50
(a) Simple difference query comparing the distribution between two different queries related to musical works. (b) Differences in instance results for Olympic sports event with an additional filter. Figure 4: Difference graph examples for the DBpedia ontology. in their respective fields while still being able to run on commodity hardware. 2.5 Case study We present OnSET in the scope of two case studies for explorative search over two different ontologies. First, we show how a novice user might start their exploration of DBpedia (Lehmann et al., 2015) to discover interesting facts and relations. Our second use case covers the BTO (Faggioli, Menotti, et al., 2024) and how experts within a field might approach more specialized ontologies. Exploring DBpedia DBpedia (Lehmann et al., 2015) is a curated ontology and KG derived from Wikipedia, providing general linked information. Querying and exploring this knowledge base typically requires the use of the SPARQL language. Our visual exploration toolkit, OnSET, enables the novice user to start exploring the knowledge base immediately through the presented topics. An example of such an exploration flow could be the initial selection of the topic “Broadcasting system information” and “Athlete rankings and achievements”, as seen in Figure 2a. This selection queries our system for links similar to the selected topic and displays start link suggestions to the user. The user can, therefore, start exploring DBpedia without prior knowledge of the classes and links contained in it, hinting the user at possible links within the ontology. Next, the user selects a link from the suggested list, initiating the process of building the prototype graph. The user then adds links and constraints to the prototype graph, drilling down towards a specific result set of interest—in this case, persons who trained an athlete and presented a television show, and where they were educated. OnSET guides the user towards non-empty sets by indicating prevalence in the KG through link strength. The user can assume that the result set is not empty due to the strong links between all classes on the prototype graph, which can be immediately verified as the result set is updated directly. While building the graph, the “minimap” in the bottom of Figure 3 is dynamically updated to indicate which regions within the class hierarchy are covered. In this case, two links span the class hierarchy while one link, the athlete’s trainer, covers only the local hierarchy within the person region. The difference views of our querying system offer additional avenues in exploratory search, which we show in further examples. In Figure 4a, the user first searches for opening themes of television shows that are recorded in any popular place. They select the runtime in seconds to be plotted. The users, however, decide to contrast this initial information without the opening theme constraint and add the constraint "York" to the label of the populated place. Our interface displays this shift, first, by marking the added nodes in green and the deleted parts in red, with additional indications in the borders of the nodes. The second change is the difference view in the result distribution, showing how the runtime compares between the queries. The users DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 18 | 50
Figure 5: The BTO KG queried using the difference view to explore the relationship between different semantic attributes. could observe that there are many more results for the second query, and that the opening themes have, on average, a shorter runtime. We also consider smaller result sets, where the query might be more specific and the result relations are of interest. For this use, we provide different views in the instance result view, showing individual instances of the result graph to the user. We again calculate the difference across all results and highlight added and removed instances. In this case, the user queries for persons who are both authors of a work and gold medalists in a sports event, as shown in Figure 4b. There are, naturally, only a few instances of this particular relation within the graph, and the user may be interested in the specific works by Olympians who participated in the games of the 2000s, adding a label constraint. We show the user the difference in the result set, emphasizing how the result set might have reduced in size due to the additional constraint. Exploring BTO Our second example outlines the exploratory search over the BTO (Faggioli, Menotti, et al., 2024), an ontology and KG that contains semantic knowledge about patients, caregivers, and their diagnoses. The explorative process for an interested party starts similarly to the use case above, but is presented with different topics. The user might be interested in diseases associated with certain demographics, so they start their search with the topic “Patient demographics and health status” and “Diseases and Disorders”, starting with the node “Patient” as shown in Figure 2a. This single node is too general, so they search for “test”, as seen in Figure 2b. Using the semantic search capabilities, they find the relation “enrolledIn” within the ontology and add them to their prototype graph. Nevertheless, this relatively small prototype graph allows the user to compare and search over these links of interest in the result set. 2.6 Conclusion The presented IR system, OnSET, allows novice users without any prior information about the knowledge base to build queries and explore it in an integrated manner. We first present related systems that share a similar aim of enabling non-expert users to construct SPARQL queries and explore KGs. Our approach, however, differs in that we utilize initial guidance approaches to lower the barrier of entry even further by providing semantic search for both the initial link search and the prototype graph expansion, and ultimately, immediate result feedback is built right into the interface. We also provide an overview of the ontology to help relate the current query to the entire ontology. DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 19 | 50
We emphasize the use of specialized open database systems to provide fast response times, enabling interactive and immediate result exploration, even for minor changes to the prototype graph. We finally demonstrate two possible explorative user flows using OnSET to inspire further use cases. The system, furthermore, includes a novel difference view for graph query building that allows users to explore the evolution of their graph query in an integrated view with the prototype graph, constraints, and result sets. We also integrate a LM interface to facilitate exploration with little prior knowledge of the graph and the user interface, leveraging semantic retrieval and fuzzy matching to the strict ontology structure. We allow, through the difference implementation, reversal of the changes to reduce the effect of invalid modifications by the LM, and more experimental searches by the users. The resulting tools enable non-expert users to explore relations of interest within the KG, including the dependencies of attributes within the queried sub-graph. Our difference views, specifically, foster the understanding of differences of sub-graphs and their attributes and distributions over the whole result set. This effect in distributional change is especially evident for the change in attribute filters, which we demonstrate through several case studies on different ontologies and KGs, the DBpedia and BTO. While the DBpedia use cases demonstrate how our system informs users in exploratory use cases, the BTO illustrates the applicability of our tool to experts in their respective fields. They can build example graphs and query for distributional changes in medical disease attributes, exploring these differences within our tool. 2.7 Future Work OnSET provides a low barrier of entry for novice users to KG exploration. We intend to expand the breadth of queries that users can express through our system in the future. A current limitation of our system is the inability to specify complete graphs, as the user can, for the time being, only build tree-like graphs, while closed graphs might be of interest for more advanced or intricate use cases. We intend to refine the constraint application process as the filtering strength of the properties of an instance is not yet clearly visually defined. This refinement could aid the user in exploring and retrieving data more attuned to their need and assist in non-empty result retrieval. Another interesting avenue is the extension towards optional parts of the prototype graph, both in the form of links and constraints, to facilitate more fuzzy retrieval. Another missing aspect of the constraint-building process is the consideration of missing properties, where a search over multimodal data types, such as spatial or image data, could be interesting. DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 20 | 50
3 Graph Queries from Natural Language using Constrained Language Models & Evaluation We extend the OnSET system with another aspect on the intersection of IR and NLP, extracting prototype graphs from natural language queries and integrating these into our query editor. 3.1 Introduction Ontologies are usually queried using specialized query languages, such as SPARQL (Seaborne et al., 2024), which can pose an additional barrier of entry for users seeking to retrieve structured knowledge from such systems. This chapter, therefore, proposes a novel querying strategy, which maps relational queries in natural language to prototype graph representations, which can then be further adjusted by the user to desired retrieval criteria. The system is tuned for natural language queries formulated as relations, differing from traditional opaque question-answering approaches, which typically do not visualize the generated queries or their processing, nor allow edits to the query. Our system extracts the prototype graph structure from the query using the capabilities of LMs to extract information from natural language input – even in cases where there is no exact match for the required information within the ontology. We achieve a valid prototype graph by constraining the LM output using a dynamically created grammar to adhere to the classes and links present within the ontology. This novel addition of the grammar to the query generation by the LM enables our system, in turn, to always return prototype graphs that are valid within the context of the used ontology. Our system, furthermore, allows the user to refine and adjust the prototype graph in a visual editor, enabling the correction of mistakes the LM might make. Within this editor, the users can improve their queries or extend them towards more complex queries, enabling users to edit their queries without modifying the SPARQL query or any other intermediate representation. We evaluate the prototype graph extraction performance of our approach using a synthetic benchmark, consisting of sampled sub-graphs and corresponding LM-generated natural language queries. This synthetic benchmark is tested for alignment with human query examples using our synthetic query generation framework. Our retrieval approach, furthermore, does not require any metadata about the ontologies or query examples, unlike other systems (Emonet et al., 2024). 3.2 Related Work Previous efforts to map queries presented in natural language towards complex, constrained query languages (Emonet et al., 2024; Ferré, 2016; Lei et al., 2018) focused on retrieving and generating SPARQL queries directly, resulting in either very complex systems or retrieval with many iterative refinement steps. (Lei et al., 2018) map a natural language query onto a particular, tailored, Ontology Query Language, used as an intermediate representation, and then converted to SQL. They use the ontology to find relevant terms in the natural language query and map them to elements of the ontology. This approach limits the output to specific instances, hindering both extensibility and exploration. SPARKLIS (Ferré, 2016) approaches the mapping from natural language to a SPARQL query through a mixture of highly constrained natural language and visual exploration. The constraints placed upon the queries help to adhere to the graph schema but could hinder more straightforward exploration. More typically, KGs are queried using SPARQL (Seaborne et al., 2024), which can be generated with LMs (Emonet et al., 2024; S. Liu et al., 2024; Meyer et al., 2024). (Meyer et al., 2024) have shown that the out-of-the-box performance of multiple proprietary LMs is lacking for direct generation of SPARQL queries DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 21 | 50
Query LM extracts G′′ p(a) person child college Search for similar classes C′& links L′(b) õ Constrain graph G′ p(c) person person university RAG-assisted graph extraction person person university country Prototype Graph Gp(d) P1 C1 U1 C1 P2 C2 U2 C2 P3 C3 U3 C3 Queried instances GI,h (e) user refinements Figure 6: Our query extraction process using LMs, using the query example “a person and the child of a person have the alma mater of the same university”. We transform the natural language query into a prototype graph using a constrained LM. The graph is first approximated using a LM (a), where the generated classes and links might not match the ontology yet. This initial guess of the LM is used to search for semantically similar relations (b). With the subset of all possible links and nodes, the graph is extracted again and corrected for possible errors (c), resulting in a graph that adheres to the ontology. The resulting graph can be edited (d), e.g. an additional constraint for a country can be added. The resulting prototype graph can be used to perform queries over a KG to retrieve instances (e). from natural language queries. They demonstrate that adding the ontology, in textual representation, to the query enhances the generation. The generation is further improved if only relevant classes and relations are provided. Similarly, Emonet et al. (Emonet et al., 2024) improve upon these results by targeting large-scale federated KGs. They develop a Retrieval Augmented Generation (RAG) system that automatically augments the LM input with relevant ontology classes and manually created example queries. Finally, Liu et al. (S. Liu et al., 2024) introduce SPINACH, a question-answering LM agent for Wikidata, that iteratively performs actions: searching Wikidata for entities, properties, or example queries, and executing SPARQL queries on demand. Effectively, these methods highlight that incorporating the ontology is essential to improve generated SPARQL queries. Nevertheless, these methods require the LMs to generate a syntactically and semantically correct SPARQL query, but—due to syntax errors or hallucination of properties—can require feedback error messages to fix the query iteratively. In comparison to these works, our method utilizes the ontology to constrain the LM output, thereby creating a prototype graph that can be directly mapped to a valid SPARQL query. This intermediate graph eliminates syntax errors and hallucinations while enabling the usage of smaller models (≤8B parameters) and removing the need for feedback error messages. The system is also robust against semantic mismatches between the query and schema, allowing users to formulate their queries more freely. We also seamlessly integrate this natural language querying system with a visual interface, enabling users to adjust and refine their query after the mapping phase has been completed. 3.3 Methodology Our KG retrieval approach builds on the notion of graph extraction and graph instance retrieval. To realize this notion, we require the graph extraction from natural language as outlined in Figure 6. The graph extraction performance is evaluated using a synthetic dataset, generated with our query generation pipeline. The synthetic dataset is shown to be representative of real results through a comparison of using both synthetic and human-written queries on a subset of queries. DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 22 | 50
3.3.1 Graph Extraction Our KG retrieval system builds upon the notion of a prototype graph Gp:= (Np, Ep)from Section 2.3.2. Our graph extraction pipeline from the natural language returns this prototype graph using a multistep approach illustrated in Figure 6. This approach first extracts an unconstrained graph G′′ pfrom the natural language prompt using the structured output of a LM (M. X. Liu et al., 2024), which serves as the basis for retrieval of semantically similar and existing classes Cand links L. These are then used to perform another round of structured generation to get G′ p, which is refined to the final prototype graph Gpthat can be used to retrieve instance graphs GI,h. Graph from Natural Language This first unconstrained graph generation step is required as the basis for further querying of possible classes Cand links L. The LM output is nevertheless restrained to return only a specific JSON schema through GGML Backus-Naur Form (GBNF) (M. X. Liu et al., 2024; llama.cpp authors, 2025). This measure enforces a consistent output that can be parsed and used for further processing. These constraints allow us to construct the intermediate graph G′′ p, which has no constraints regarding the ontology, i.e., Land Care open. While we could constrain the graph types to the whole ontology at this step, we have no way to enforce the structural correctness of the graph over the outgoing link types, increasing the probability of an invalid graph. Constraints from Graph The unconstrained graph is used to retrieve candidates for the next generation step. This retrieval is performed for each node ni∈Npand edge ej∈Epusing sentence embeddings of the node and link description. These embeddings are used to retrieve the top kmost similar results in terms of cosine distance from all classes Cand links L. The LM generation is then further constrained to only include these results. This retrieval of semantically similar links results in the subset of classes C′⊆ C and links L′⊆ L. Constrained Graph from Natural Language Finally, the prototype graph Gpis generated by providing the LM with the same instruction as in the first step, but with additional constraints placed upon the output generation through GBNF (llama.cpp authors, 2025). These use the additional information of the possible candidate classes C′and links L′from the previous step. This limited set of only semantically similar classes enables the LM still to express any graph from the natural language query while adhering to the relevant query constraints. The LM-generated graph G′ pmay contain invalid or flipped links as we cannot enforce a valid graph structure on the output. This problem is mitigated by the previous step of only using the ontology subsets and by cleaning the graph using two rules. The first one exchanges the direction of the edge, essentially swapping the nodes (nt, lj, nh)7→ (nh, lj, nt), if the types are flipped, i.e., the link ljmay not go from ntto nh, but from nhto nt. Our second rule discards any invalid links from the graph if they violate the type constraints. 3.3.2 Synthetic Evaluation Methodology The described graph retrieval system is evaluated using synthetically generated queries from a sampled prototype graph Gp,s. This graph is used to generate queries in natural language using either a LM that is prompted with the graph as input structure or a template-based query generation. We additionally validate our query generation methodology against queries generated by humans by comparing the resulting evaluation metrics. The natural language prompt is, in turn, used to extract the prototype graph Gpusing the method from above. Finally, the sampled and extracted graphs are compared and evaluated for similarity. This evaluation DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 23 | 50
system does not incorporate user refinements for adjusting the output graph, as the methodology evaluates only the directly returned prototype graph Gp. Graph Sampling The first step in our synthetic evaluation pipeline is the sampling of prototype graphs Gp,s from the ontology. The sampling is based on probabilities derived from the instance counts of the links. We additionally select classes that are lower in the class hierarchy using a similar probabilistic approach. Finally, further links are added based on a random choice of node and a similar probabilistic selection. Generation of the Query Next, the natural language query is generated synthetically from the previously sampled graph Gp,s using a LM, more specifically the Hermes 3 Llama 3.2 3B model (Teknium et al., 2024). The LM is presented with a textual representation of the types of classes and links within the graph and prompted, using a one-shot approach, to generate the queries. We also employ a template-based query generation approach to prevent potential information leaks and evaluate simpler queries (Sannigrahi et al., 2024). Additionally, we create 12 human-written queries from the sampled graphs to validate our evaluation methodologies for queries used in practice. Scoring the Graphs Finally, the prototype graph Gpcan be extracted from the synthetic query using our graph extraction methodology and compared to the ground truth sampled graph Gp,s. This evaluation employed two scoring methodologies: one for the retrieval performance of the nodes Npand relations Ep, and another for the graph layout. First, the similarity between the retrieved nodes Npand sampled nodes Np,s is compared using the F1-score over the sets. This set-based F1,node-score uses the true positives TP =|Np∩Np,s|, false positives F P =|Np−Np,s|and false negatives FN =|Np,s −Np|rates from set intersections and differences. The F1,rel of the relations is computed using the same scheme using Epand Ep,s. Second, the graph similarity is computed using the Graph Edit Distance (GED) (Abu-Aisheh et al., 2015), which employs a stricter notion of similarity, requiring not only isomorphism between the graphs but also the same node and link types for the graphs. To achieve a comparable score to the F1scores and between different graph sizes, we weigh the distance by the number of nodes and links and invert it, giving the normalized GED score GEDs= 1 −GED max{|Np|,|Np,s|} + max{|Ep|,|Ep,s|} . Both evaluation methodologies provide scores for a single query example. We therefore reduce them to single values by averaging the scores over the different queries and evaluation settings. 3.3.3 Constraining the LM The LMs are constrained to a schema specified to a grammar whenever we retrieve any graph using language generation. To this end, we use llama-cpp-agent8, which can constrain the LM to only output a specific model using a predefined grammar. The first prototype graph G′′ puses a static schema and grammar. In the second generation step, the grammar is adapted to the specific, similar relations found utilizing the output of the first round. The updated grammar is then used to constrain the output only to contain valid types and links, which can be used to generate the correct prototype graph Gp. The graph Gpcan be converted into a SPARQL query if the user requires it for further use. 8https://llama-cpp-agent.readthedocs.io/ DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 24 | 50
Evaluated Models and Ontologies The evaluation of this work relies heavily on the use of LMs for graph retrieval, semantic retrieval, and constrained retrieval. We use the Hermes models (Teknium et al., 2024) (3B, 8B, and 70B parameter sizes) for generative retrieval tasks due to their fine-tuned capabilities for structured output, comparing them to two Qwen2.5 (32B parameters) models (Instruct and Code finetuned) (Qwen et al., 2025). A sentence-transformers model (Reimers & Gurevych, 2019) is used for all semantic retrieval tasks9. We evaluate our approach on two ontologies, specifically • DBpedia (Lehmann et al., 2015), an ontology and KG containing mapped information from Wikidata with 768 classes and 4,233,000 instances; • BTO (Faggioli, Marchesin, et al., 2024), a smaller ontology and KG focusing on brain-related diseases; and 3.3.4 User Interface Integration Additionally, we allow the users to refine their initial queries using our node-based editor OnSET, where links and nodes can be added to or removed from the graph, all within the constrained link set L, and class set C. This refinement can be helpful if our pipeline fails to find the correct graph or if the users want to refine their search without an additional text prompt. Furthermore, we transparently display how the queries are built using the LM through a similar flow to that shown in Figure 6. This visualization should hold our system accountable for any mistakes and errors that occur during processing and may guide the user towards specifying more precise queries or exploring other querying avenues. 3.4 Results We demonstrate our system’s capabilities to retrieve the correct graph from the query using our synthetic graph generation and evaluation pipeline. We evaluate the system at three different node sample counts, k∈ {2,3,5,7}, to model varying degrees of complexity in the queries. Each node sample count kis sampled for 128 synthetic queries. Additionally, we use four different open-weight models to test the dependence on model size and type. Our evaluation in Figure 7 shows that we can faithfully recover the prototype graph from the natural language query. While we cannot achieve a perfect recovery of all nodes and relations in all settings, we accomplish a F1score on the node retrieval on DBpedia across all models of approximately 0.7throughout the different levels of complexity. The correct relations are retrieved at an even higher F1 score of roughly 0.8for the various ontologies and settings. Similarly, the GED score GEDsshows that the graph structure can be recovered quite well for most of the smaller graph sizes and drops only for the larger, more complex queries, demonstrating that our approach is most useful for smaller graphs. Our evaluation, furthermore, shows that using a larger model is only slightly beneficial for our use case. The system performs similarly across all model sizes and types, suggesting that even smaller models can reconstruct the prototype graph quite well due to the constraints we place on the output. Our evaluation, furthermore, shows that the constraint and alignment step to the ontology, including our corrections, is essential to retrieve the correct graph as almost all combinations of query complexity, model, and ontology have a higher score after applying the constraint compared to the direct output of the model without the constraints of the ontology (reflected in the “raw” results). Finally, we compare our different query generation methods in Table 1 by evaluating two query complexities k={3,5}and comparing the scores. The LM-based generation method performs similarly to the 9We use the Stella 400M model (Zhang et al., 2025) for a great balance of model size and performance. DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 25 | 50
search the web if needed. The most important tools are the ones that provide functionality to the LLM to interact with GraphDB and the data inside it. The tool allows the LLM to interact with the data using SPARQL queries, similarity indexes, and other sophisticated methods that allow the LLM to easily find and aggregate the needed information to answer the user’s question and even provide its methodology if the user wishes to replicate or validate it. Figure 13: GraphDB Talk To Your Graph setup Zero-shot implementation we set up a Talk To Your Graph agent in GraphDB with ChatGPT 4.1 with temperature of 0(Figure 13), in ordert to make the solution as deterministic as possible and always to return the most probable response, i.e. the one that is presented in the KG. The base instructions are: 1"You are Quadro ,an assistant developed by Ontotext and you can answer questions about data stored in GraphDB .When answering a question,you generally take a few steps . 2First,you determine whether the question is relevant to the ontology and let the user know if it seems to be out of scope of the dataset you are working with . 3Second,you identify the relevant entities in the question and which would be objects in the \gls{kg }. 4Third,you try to map each object to a URI in the database using an appropriate tool such as autocomplete . 5Fourth,you decide whether you need to get more results out of the database using a SPARQL query or search method . 6Finally,you use the collected information to answer the question ." The Additional instructions are: DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 32 | 50
1"Your job is to help the user understand the molecular mechanisms of certain diseases and to identify explicit and hidden relationships between key research topics (like identifying which are the susceptibility genes for a particular disease / phenotype )and to provide provenance source for the extracted information based on the ontology provided .The NCBI gene ontology uses geneSymbol property for labels . 2 3Always add the @en language tag at the end string literals for example VALUES ?diseaseLabel \{"Multiple Sclerosis "@en\}. 4Always first use the full text search tool to find the IRIs . 5The data is very large ,so use VALUES instead of FILTER ." Multi-shot implementation For the Multi-shot implementation, we used the same base instructions. The prompt used for the additional instruction is: 1Your job is to help the user understand the molecular mechanisms of certain diseases and to identify explicit and hidden relationships between key research topics (like identify which are the susceptibility genes for a particular disease /phenotype )and to provide provenance source for the extracted information based on the ontology provided.NCBI genes ontology uses geneSymbol property for labels. 2 3Always add the @en language tag at the end string literals for example VALUES ?diseaseLabel {" Multiple Sclerosis "@en }. 4Always first use the full text search tool to find the IRIs . 5The data is very large ,so use VALUES instead of FILTER . 6 7 8Here are a few examples : 9 10 Which are the genes associated with Multiple Sclerosis ? 11 12 The query to find the answer is: 13 14 PREFIX ncbigenes : <https ://linkedlifedata.com/resource/ncbigenes/> 15 PREFIX rdf: <http:// www.w3. org/1999/02/22-rdf-syntax -ns#> 16 PREFIX skos : <http:// www .w3. org/2004/02/ skos/core#> 17 select distinct ?gene ?geneSymbol where { 18 VALUES ?diseaseLabel {" Multiple Sclerosis "@en} 19 ?gene rdf:type ncbigenes :Gene; 20 ncbigenes :geneSymbol ?geneSymbol ; 21 ?associatedWith ?disease. 22 ?disease ?label ?diseaseLabel . 23 } 24 DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 33 | 50
25 Are there any associated pathways to these disorders ,in which the identified genes participate ? 26 27 query : 28 29 PREFIX rdfs : <http:// www .w3. org/2000/01/ rdf-schema#> 30 PREFIX td: <https :// linkedlifedata.com/resource/targetdiscovery/ontology #> 31 PREFIX ncbigenes : <https ://linkedlifedata.com/resource/ncbigenes/> 32 PREFIX rdf: <http:// www.w3. org/1999/02/22-rdf-syntax -ns#> 33 PREFIX skos : <http:// www .w3. org/2004/02/ skos/core#> 34 #select distinct ?gene ?geneSymbol 35 select ?gene ?geneSymbol ?pathway ?descriptions ?labels 36 where { 37 VALUES ?diseaseLabel {" Multiple Sclerosis "@en} 38 ?gene rdf:type ncbigenes :Gene; 39 ncbigenes :geneSymbol ?geneSymbol ; 40 td:encodes ?protein; 41 ?associatedWith ?disease. 42 ?disease ?label ?diseaseLabel . 43 ?protein td:isPartOf ?pathway; 44 skos:prefLabel ?prefLabel 45 OPTIONAL {? pathway skos :definition ?descriptions} 46 OPTIONAL {? pathway skos :prefLabel ?labels} 47 } 48 49 Find the hierarchy of Alzheimer ? 50 51 query : 52 53 PREFIX skos : <http:// www .w3. org/2004/02/ skos/core#> 54 select *where { 55 <https :// linkedlifedata.com /resource/umls/id /C0002395> skos:broader+ ?o. 56 ?o skos:prefLabel ?label 57 } 58 59 Which proteins are encoded by APP gene ? 60 61 query : 62 63 PREFIX ncbigenes : <https ://linkedlifedata.com/resource/ncbigenes/> 64 PREFIX rdf: <http:// www.w3. org/1999/02/22-rdf-syntax -ns#> 65 select *where { 66 ?s?p"APP " ; 67 rdf :type ncbigenes :Gene 68 } DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 34 | 50
4.5 Experiments and Results For experiments we define competency questions that are either specific for some disease, gene etc., or explore more complex interrelations between multiple relations. Here are sample questions for both types: • (simple) Look for genes or proteins annotated with both gut and brain-related processes • (simple) Which are the genes associated with Amyotrophic Lateral Sclerosis? • (simple) Which proteins are encoded by APP gene? • (complex) What are protein protein interactions between ALS and Parkinson disease? • (complex) Which are genes related to ALS and MS? • (complex) How does intestinal permeability (leaky gut) contribute to neuroinflammation? • (complex) Can changes in diet or probiotics modulate the gut-brain axis and improve mental health outcomes? • (complex) What molecular pathways are involved in gut-brain communication? For these competence questions were done experiments both with zero-shot and multi-shot agents. For all simple questions, both agents successfully generated correct answers. However, the multi-shot agent was generally able to answer with fewer queries to GraphDB and also managed to answer more complex questions. Multi-Shot experiment For the question "What are protein protein interactions between ALS and Parkinson disease?" with the multi-shot agent we receive a result that contains the very comprehensive summary of the answer in natural language (Figure 14) as well as generated SPARQL query (Figure 15). We can further copy the generated SPARQL query and run it directly in GraphDB over the BioGraphTalk to explore further and deep dive into specific complex relations between entities in the result to investigate the actual relations between them (Figure 16 and 17). The manual validation by domain expert on the answer presents that the achieved results are sound and correct. The human validation of the result of SPARQL query presents that the extract. 4.6 Further work The next steps include design of multiple agents related to different HEREDITARY use cases. Another direction in further work is related to implementation of demonstrator with user interface that will provide end-to-end functionalities of the TTYG without exposing to the end user the GraphDB triple store and will fully automate the process from natural language question to visual exploration of the KG relations. In addition, further work includes experiments with complex questions that can bring new incites in the gut-brain interplay. To ensure compliance with data protection - the current version of LLM - GPT will be also possible to be substitute with open LLMs using local model storage. DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 35 | 50
Figure 14: TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - the answer summary in natural language DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 36 | 50
Figure 15: TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - the generated SPARQL query DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 37 | 50
Figure 16: TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - the generated SPARQL query run in GraphDB over BioGraphTalk Figure 17: TTYG Experiment 1 - What are protein protein interactions between ALS and Parkinson disease? - visual exploration of the result in GraphDB over BioGraphTalk DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 38 | 50
5 Neurodegen-Vis: LLM-based, Privacy-Preserving Visual Exploration of High-Dimensional ALS, PD, MS Patient Data 5.1 Introduction Medical datasets can be highly complex, ranging from measurements to social demographics. Visual Analytics (VA) approaches help users to get an understanding of the data, clusters, and relationships among variables. Parkinson's disease analysis Show Chat assistant. Pearson Correlation for selected features insnpsi_age npsid_ddur_v ins_npsi_sex npsid_yearsed overall_domain_sum npsid_rep_moca_c npsid_rep_mmse_c attent_z_comp exec_z_comp visuosp_z_comp memory_z_comp language_z_comp st_ter_daed st_ter_leed updrs_3_on rc_score_done sdmt_done flu_a_done phon_flu_done pc1 pc2 −1.0 −0.5 0.0 0.5 1.0 correlation insnpsi_age npsid_ddur_v npsid_yearsed ins_npsi_sex overall_domain_sum npsid_rep_moca_c npsid_rep_mmse_c attent_z_comp exec_z_comp visuosp_z_comp memory_z_comp language_z_comp st_ter_daed st_ter_leed updrs_3_on pc1 pc2 insnpsi_age npsid_ddur_v npsid_yearsed ins_npsi_sex overall_domain_sum npsid_rep_moca_c npsid_rep_mmse_c attent_z_comp exec_z_comp visuosp_z_comp memory_z_comp language_z_comp st_ter_daed st_ter_leed updrs_3_on pc1 pc2 1.00 0.01 -0.19 0.03 0.29 -0.38 -0.31 -0.14 -0.42 -0.37 -0.31 -0.37 -0.16 0.15 0.31 0.58 -0.02 0.01 1.00 0.16 0.09 -0.04 0.08 -0.01 -0.15 -0.06 0.08 0.21 0.09 -0.20 -0.18 -0.00 -0.09 0.88 -0.19 0.16 1.00 0.17 -0.21 0.25 0.32 0.20 0.33 0.35 0.25 0.17 0.04 -0.26 -0.08 -0.48 0.22 0.03 0.09 0.17 1.00 -0.06 -0.04 0.05 0.13 0.02 0.04 0.13 -0.00 -0.08 0.06 0.09 -0.07 0.07 0.29 -0.04 -0.21 -0.06 1.00 -0.47 -0.31 -0.36 -0.33 -0.50 -0.37 -0.39 -0.14 -0.03 -0.05 0.67 0.01 -0.38 0.08 0.25 -0.04 -0.47 1.00 0.42 0.25 0.39 0.50 0.39 0.45 0.04 -0.24 -0.21 -0.72 0.10 -0.31 -0.01 0.32 0.05 -0.31 0.42 1.00 0.33 0.40 0.53 0.23 0.35 -0.15 -0.39 -0.06 -0.66 -0.16 -0.14 -0.15 0.20 0.13 -0.36 0.25 0.33 1.00 0.38 0.47 0.23 0.26 0.12 -0.09 0.06 -0.55 -0.46 -0.42 -0.06 0.33 0.02 -0.33 0.39 0.40 0.38 1.00 0.45 0.36 0.35 0.13 -0.15 -0.16 -0.69 -0.18 -0.37 0.08 0.35 0.04 -0.50 0.50 0.53 0.47 0.45 1.00 0.40 0.50 -0.11 -0.28 -0.16 -0.81 -0.04 -0.31 0.21 0.25 0.13 -0.37 0.39 0.23 0.23 0.36 0.40 1.00 0.31 -0.07 -0.09 -0.08 -0.60 0.35 -0.37 0.09 0.17 -0.00 -0.39 0.45 0.35 0.26 0.35 0.50 0.31 1.00 0.17 -0.08 -0.10 -0.66 0.07 -0.16 -0.20 0.04 -0.08 -0.14 0.04 -0.15 0.12 0.13 -0.11 -0.07 0.17 1.00 0.16 -0.06 -0.06 -0.18 0.15 -0.18 -0.26 0.06 -0.03 -0.24 -0.39 -0.09 -0.15 -0.28 -0.09 -0.08 0.16 1.00 0.15 0.27 -0.11 0.31 -0.00 -0.08 0.09 -0.05 -0.21 -0.06 0.06 -0.16 -0.16 -0.08 -0.10 -0.06 0.15 1.00 0.17 -0.05 0.58 -0.09 -0.48 -0.07 0.67 -0.72 -0.66 -0.55 -0.69 -0.81 -0.60 -0.66 -0.06 0.27 0.17 1.00 0.00 -0.02 0.88 0.22 0.07 0.01 0.10 -0.16 -0.46 -0.18 -0.04 0.35 0.07 -0.18 -0.11 -0.05 0.00 1.00 overall_domain_sum vs insnpsi_age k_mean_cluster −1 0 1 2 3 4 5 6 7 8 9 ↑ overall_domain_sum 45 50 55 60 65 70 75 80 85 90 insnpsi_age → PCA Analysis Number of Clusters (k): 2 Run −2.5 −2.0 −1.5 −1.0 −0.5 0.0 0.5 1.0 1.5 2.0 2.5 ↑ Principal Component 2 −4−3−2−1 0 1 2 3 4 Principal Component 1 → npsid_ddur_v overall_domain_sum npsid_rep_mmse_c npsid_rep_moca_c Chatbot Clear History Suggest Features Or ask your questions. Enter your prompt here Submit Suggest Features! Analyzing insnpsi_age (age of the patient at the time of the neuropsychological assessment) and overall_domain_sum (a summary of cognitive assessments across various domains) is interesting because: 1. Cognitive Decline with Age: Age can significantly impact cognitive function. Investigating the relationship between age and overall cognitive performance can reveal patterns of cognitive decline, which is particularly relevant for conditions like Parkinson's disease. 2. Correlation Insights: Understanding how age correlates with overall cognitive performance can help identify risk factors for cognitive impairment. For instance, if older age is associated with lower scores in cognitive assessments, it may indicate a need for targeted interventions in older patients. These insights can ultimately aid in customizing treatment plans and improving patient care strategies. npsid_ddur_v, overall_d.. Figure 18: The Dashboard, Correlation Heatmap on the left, Scatterplot and biplot on the right. In the biplot a Principal Component Analysis (PCA) can be started, and the number of clusters can be selected dynamically. In the biplot and the scatterplot the PCA clusters are colordoded. On the very left is the chatbot. The suggest feature function was used for exploration. They are highlighted in the correlation heatmap (sixth row on the very left) and additionally shown in the Scatterplot. In the chatbot window, it is explained why the features are interesting to analyze. However, users face various challenges when interaction with a dashboard about complex data. Firstly, the user may have questions about the dataset itself, because of different conventions about feature names, or because they are not a domain expert. Secondly, the user may have problems understanding the dashboard itself. Perhaps they do not understand the individual visualisations, how to select the views of interest, and how the different parts of the dashboard interact with each other. Thirdly, during data exploration, it can be hard to identify interesting aspects, or where to start with data exploration. Motivated by these challenges, we designed and implemented an interactive dashboard to identify correlations, dependencies, and groups within a dataset about Parkinson’s disease patients via VA, cf. Figure 18 for a screenshot of the primary view after data anonymisation and configuration. As a key functionality, we DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 39 | 50
equip the tool with a seamless LM integration, aiming to assist the user. The LM is provided information about the dashboard and the dataset. Combined with the general world-knowlwedge of LM’s it can help the user with exploration and explanation (see Example 1). The user has a classic, familiar chatbot interface to interact with. Additionally, the LM is used to highlight and explain different views of the dashboard (cf. Examples 2 and 3). Especially for critical data domains such as healthcare, the interpretation and conclusions should still be made by the user dependent on the data, not the LM. Therefore, our focus is on explaining and exploring the dashboard and the data, rather than giving interpretations, and results only if the user can verify them on the dashboard. For these applications VA and LM’s very much suit each other, as the user can always verify everything the LM’s says, and the LM can enhance the dashboard by helping the user in understanding and exploration. What is the correlation between age and language z test and what does this suggest? The correlation between age (insnpsi_age) and the language z test (language_z_comp) is -0.413. This negative correlation suggests that as age increases, scores on the language z test tend to decrease. In other words, older patients may generally perform worse on language-related cognitive tasks compared to younger patients. Example 1: Chatbot integration: the user can ask questions to interpret correlations in the real data, without the LM ever seeing any of the data per se, only differentially private analysis results from it. In this chapter we analyze the LM implementation for a dashboard to analyze health care data of patients with Parkinson disease. In the dashboard, the focus lies on identifying correlations, dependencies, and groups within a dataset. 5.1.1 Dataset The dataset contains information on patients with PD. It comprises approximately 100 features for 50 patients. It includes demographic data such as patient age and duration of PD diagnosis, alongside results from various cognitive tests and treatment details. Several columns contain redundant information, with some measurements represented both numerically and categorically. Additionally, cognitive tests were conducted using multiple methodologies, leading to potential domain shifts. Binary indicators in the dataset denote the use of specific test methods. After accounting for duplicate information and test method markers, the dataset includes approximately 15 unique measurements. 5.1.2 Anonymisation using Differential Privacy In order to protect patients’ private data, our tool anonymises the incoming dataset before displaying it to a (potentially unauthorised) user. In a first step, names and directly personally identifiable information is stripped away, which only requires users to specify which columns are affected. But of course, this still leaves highly personal direct measurements of the patients’ health characteristic, as deemed necessary in conducting the recording of the dataset. Specifically, we generate (an arbitrary amount of) synthetic tabular data based on the distribution of features across the input dataset. As such tabular data is often a mix of ordinal, categorical and continuous DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 40 | 50
features, we have to treat them separately in the anonymisation phase. The categorical data is synthetically generated based on the Multiplicative Weights Exponential Mechanism (MWEM) (Hardt et al., 2012), with an implementation in the third-party package smartnoise-sdk, developed by OpenDP19 (and based on their formally verified opendp library). Next up, the continuous columns are then only sampled from their respective input distributions. For simplicity of this demonstrator, we assume a normal distribution in all numerical columns (also taking bounds and, for example, integer constraints into account). This still faces our generation procedure with the challenge of preserving correlations. Correlation can be preserved using randomized order, correlation-integrated sampling from the respective feature distributions. One starts by obtaining a (Pearson-) correlation matrix C∈Rn×nof the respective input features (in our case, a roughly 60x60 matrix) and passing it through the MWEM implementation to skew it with noise and making it differentially private with respect to the input data. During the sampling procedure, in random order, the columns are sampled from their respective (normal) distributions parametrised earlier. Each sample value is then skewed according to the correlation between all previously sampled values and the feature correlation contained in C. The output dataset then approximately follows the input distribution and keeps correlations consistent. The dashboard and corresponding LM interaction (and, of course, the figures included in this work) now only rely on the synthetically generated data. 5.2 Related Work Visualization and analysis of high-dimensional data is a long-standing challenge in data science, and to date many approaches for interactive exploration of this data have been proposed. The Rank by Feature framework (Seo & Shneiderman, 2006) was among the first approaches to identify and rank data features by their correlation, hence supporting the user in selecting interesting and relevant data. It used a heatmap to show the feature importance. In general, tabular data can be interactively visualized by general tools like spreadsheet software, and specifalized development suites like Tableau or Microsoft PowerBI. Lineup (Gratzl et al., 2013) is a tabular data visualization specifically for comparing and ranking the rows of a multivariate table, having the user interactively find appropriate weights for the different features. Scatter Plot matrices and scatter plots are widely-used tools to show an overview of pairwise correlations, amenable for overviewing and searching (Chegini et al., 2018). Recently, Language Model technology has drastically advanced the capability of Natural Language Processing, and existing implementations like Llama or ChatGPT allow to integrate LM-based Natural Language Interfaces (NLIs) into interactive visualization applications. There are many roles where LMs can help users to navigate and understand visualizations and data, as presented in the framework in (Zhao et al., 2025). In our work, we design a dashboard for exploring tabular data of numeric and categorical values. We make use of existing visualizations like correlation heatmaps, scatter plots, and principal components plots. While these techniques are not novel, in many case studies they have shown to be very understandable and an expert tool for data exploration. Our main contribution is the intergration of an LLM interface to assist the user, and to apply it to a anonymized real-world health data set. Our use case shows the dashboard design and LLM intergration are helpful and hence can allow for explaining and guiding users. 19https://opendp.org/ DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 41 | 50
Hardt, M., Ligett, K., & McSherry, F. (2012, December). A simple and practical algorithm for differentially private data release. In Guide Proceedings (pp. 2339–2347, Vol. 2). Curran Associates Inc. https://doi.org/10.5555/2999325.2999396 Hogan, A., Blomqvist, E., Cochez, M., D’amato, C., Melo, G. D., Gutierrez, C., Kirrane, S., Gayo, J. E. L., Navigli, R., Neumaier, S., Ngomo, A.-C. N., Polleres, A., Rashid, S. M., Rula, A., Schmelzeisen, L., Sequeda, J., Staab, S., & Zimmermann, A. (2021). Knowledge graphs. ACM Comput. Surv.,54(4). https://doi.org/10.1145/3447772 Hutchinson, M., Jianu, R., Slingsby, A., & Madhyastha, P. (2024). Llm-assisted visual analytics: Opportunities and challenges. arXiv preprint arXiv:2409.02691. Jadhav, S., Perumal, S., Tadavi, Y., Dash, B., & Parthiban, S. (2025). Leveraging large language models for biomedical knowledge graph construction and querying: An advanced nlp approach. Companion Proceedings of the ACM on Web Conference 2025, 2560–2566. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Comput. Surv.,55(12). https://doi.org/10.1145/3571730 Jia, R., Zhang, B., Méndez, S. J. R., & Omran, P. G. (2024). Leveraging large language models for semantic query processing in a scholarly knowledge graph. arXiv preprint arXiv:2405.15374. Kantz, B., Innerebner, K., Waldert, P., Lengauer, S., Lex, E., & Schreck, T. (2025). Onset: Ontology and semantic exploration toolkit. Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 3980–3984. https://doi.org/10. 1145/3726302.3730148 Krause, J., Perer, A., & Stavropoulos, H. (2016). Supporting iterative cohort construction with visual temporal queries. IEEE Transactions on Visualization and Computer Graphics,22(1), 91– 100. https://doi.org/10.1109/TVCG.2015.2467622 Kusupati, A., Bhatt, G., Rege, A., Wallingford, M., Sinha, A., Ramanujan, V., Howard-Snyder, W., Chen, K., Kakade, S. M., Jain, P., & Farhadi, A. (2022). Matryoshka representation learning. In A. H. Oh, A. Agarwal, D. Belgrave, & K. Cho (Eds.), Advances in neural information processing systems.https://openreview.net/forum?id=9njZa1fm35 Lehmann, J., Gattogi, P., Bhandiwad, D., Ferré, S., & Vahdati, S. (2023). Language models as controlled natural language semantic parsers for knowledge graph question answering. In Ecai 2023 (pp. 1348–1356). IOS Press. Lehmann, J., Isele, R., Jakob, M., Jentzsch, A., Kontokostas, D., Mendes, P. N., Hellmann, S., Morsey, M., van Kleef, P., Auer, S., & Bizer, C. (2015). Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia. Semantic Web,6, 167–195. https://api. semanticscholar.org/CorpusID:1181640 Lei, C., Özcan, F., Quamar, A., Mittal, A. R., Sen, J., Saha, D., & Sankaranarayanan, K. (2018). Ontology-based natural language query interfaces for data exploration. IEEE Data Eng. Bull., 41(3), 52–63. Lissandrini, M., Mottin, D., Palpanas, T., & Velegrakis, Y. (2020). Graph-query suggestions for knowledge graph exploration. Proceedings of The Web Conference 2020, 2549–2555. https://doi. org/10.1145/3366423.3380005 DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 48 | 50
Liu, M. X., Liu, F., Fiannaca, A. J., Koo, T., Dixon, L., Terry, M., & Cai, C. J. (2024). “we need structured output”: Towards user-centered constraints on large language model output. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, 1–9. https://doi.org/10.1145/3613905.3650756 Liu, P., Wang, X., Fu, Q., Yang, Y., Li, Y.-F., & Zhang, Q. (2022). Kgvql: A knowledge graph visual query language with bidirectional transformations. Knowledge-Based Systems,250, 108870. https://doi.org/https://doi.org/10.1016/j.knosys.2022.108870 Liu, S., Semnani, S. J., Triedman, H., Xu, J., Zhao, I. D., & Lam, M. S. (2024). SPINACH: sparqlbased information navigation for challenging real-world questions. In Y. Al-Onaizan, M. Bansal, & Y. Chen (Eds.), Findings of the association for computational linguistics: EMNLP 2024, miami, florida, usa, november 12-16, 2024 (pp. 15977–16001). Association for Computational Linguistics. https://aclanthology.org/2024.findings-emnlp.938 llama.cpp authors. (2025, February). Gbnf guide [[Online; accessed 2025-02-11]]. https://github. com/ggerganov/llama.cpp/blob/b9ab0a4d0b2ed19effec130921d05fb5c30b68c5/grammars/ README.md Menotti, L., & Silvello, G. (2024). The hereditary ontology for genomics data [Accessed: 2024-10-01]. https://hereditary.dei.unipd.it/ontology/genomics/ Meyer, L.-P., Frey, J., Brei, F., & Arndt, N. (2024). Assessing sparql capabilities of large language models. https://arxiv.org/abs/2409.05925 Muennighoff, N., Tazi, N., Magne, L., & Reimers, N. (2022). Mteb: Massive text embedding benchmark. arXiv preprint arXiv:2210.07316.https://doi.org/10.48550/ARXIV.2210.07316 Ongris, J. G., Tjitrahardja, E., Darari, F., & Ekaputra, F. J. (2024). Towards an open nli llm-based system for kgs: A case study of wikidata. 2024 7th International Seminar on Research of Information Technology and Intelligent Systems (ISRITI), 44–49. Peng, C., Xia, F., Naseriparsa, M., & Osborne, F. (2023). Knowledge graphs: Opportunities and challenges. Artificial Intelligence Review,56(11), 13071–13102. https://doi.org/10.1007/ s10462-023-10465-9 Pérez-Messina, I., Ceneda, D., & Miksch, S. (2023). A Methodology for Task-Driven Guidance Design. In M. Angelini & M. El-Assady (Eds.), Eurovis workshop on visual analytics (eurova). The Eurographics Association. https://doi.org/10.2312/eurova.20231094 Pienta, R., Tamersoy, A., Endert, A., Navathe, S., Tong, H., & Chau, D. H. (2016). Visage: Interactive visual graph querying. Proceedings of the International Working Conference on Advanced Visual Interfaces, 272–279. https://doi.org/10.1145/2909132.2909246 Qwen, : Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., Lin, H., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., . .. Qiu, Z. (2025). Qwen2.5 technical report. https://arxiv.org/abs/2412.15115 Reimers, N., & Gurevych, I. (2019). Sentence-bert: Sentence embeddings using siamese bertnetworks. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing.https://arxiv.org/abs/1908.10084 Sannigrahi, S., Fraga-Silva, T., Oualil, Y., & Van Gysel, C. (2024). Synthetic query generation using large language models for virtual assistants. Proceedings of the 47th International ACM SIDELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 49 | 50
GIR Conference on Research and Development in Information Retrieval, 2837–2841. https: //doi.org/10.1145/3626772.3661355 Seaborne, A., Taelman, R., Williams, G., Hartig, O., & Tanon, T. P. (2024, December). SPARQL 1.2 query language (W3C Working Draft) (https://www.w3.org/TR/2024/WD-sparql12-query20241227/). W3C. Seo, J., & Shneiderman, B. (2006). Knowledge discovery in high-dimensional data: Case studies and a user survey for the rank-by-feature framework. IEEE Trans. Vis. Comput. Graph.,12(3), 311–322. https://doi.org/10.1109/TVCG.2006.50 Teknium, R., Quesnelle, J., & Guang, C. (2024). Hermes 3 technical report. https://arxiv.org/abs/ 2408.11857 Urchade, Z., Tomeh, N., Holat, P., & Charnois, T. (2024). An autoregressive text-to-graph framework for joint entity and relation extraction. Vargas, H., Buil-Aranda, C., Hogan, A., & López, C. (2019). Rdf explorer: A visual sparql query builder. In C. Ghidini, O. Hartig, M. Maleshkova, V. Svátek, I. Cruz, A. Hogan, J. Song, M. Lefrançois, & F. Gandon (Eds.), The semantic web – iswc 2019 (pp. 647–663). Springer International Publishing. White, R. W., & Roth, R. A. (2009). Exploratory search: Beyond the query—response paradigm. Springer International Publishing. https://doi.org/10.1007/978-3-031-02260-9 Zhang, D., Li, J., Zeng, Z., & Wang, F. (2025). Jasper and stella: Distillation of sota embedding models. https://arxiv.org/abs/2412.19048 Zhao, Y., Zhang, Y., Zhang, Y., Zhao, X., Wang, J., Shao, Z., Turkay, C., & Chen, S. (2025). Leva: Using large language models to enhance visual analytics. IEEE Transactions on Visualization and Computer Graphics,31(3), 1830–1847. https://doi.org/10.1109/tvcg.2024.3368060 DELIVERABLE 5.5 13.08.2025, Ver. 1.10 GA 101137074 50 | 50