TILDE - Trustworthy Access to Knowledge from the Indexed Web
Full text
Final Report (M12) TILDE – Trustworthy Access to Knowledge from the Indexed Web Version 1.1 OpenWebSearch.EU “Piloting a Cooperative Open Web Search Infrastructure to Support Europe’s Digital Sovereignty” The Project is funded by the EC under GA 101070014
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 1 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Table of Contents 1 Period M1 – M6 4 1.1 Module: NLP 4 1.2 Module: Trustworthiness 5 1.3 Module: Visual Web Interface 7 2 Period M6 – M12 8 2.1 Module: NLP 8 2.2 Module: Trustworthiness 10 2.3 Module: Visual Web Interface 14 3 Data Availability 21 4 Table of Figures 22 5 Bibliography 23
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 2 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Preliminaries i. Project Info Project number 101070014 Project acronym ows.eu Project name OpenWebSearch.eu – Piloting a Cooperative Open Web Search Infrastructure to Support Europe's Digital Sovereignty Call HORIZON-CL4-2021-HUMAN-01 Topic HORIZON-CL4-2021-HUMAN-01-05 Type of action HORIZON-RIA Responsible unit DG CNECT Project starting date / Duration 01/09/2022 Project reporting period 2 Project Coordinator Prof. Dr. Michael Granitzer, University of Passau ii. Project Partners Acronym Partner KNOW Know Center Research GmbH iii. Deliverable Info Due Date / Delivery Date 08/09/2025 Deliverable Lead Michael Jantscher Deliverable type Report Dissemination level SEN Document Status / Version V1 Work-package / Lead Partner NN
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 3 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE iv. Deliverable Summary This document describes the “2024FSTPC2PN35” TILDE OpenWebSearch.eu project funded by the EC under the GA 101070014 within a Horizon Europe Framework programme. To facilitate access to the Open Web Index (OWI), we propose an AI-driven Open Web Search (OWS)- based component for data exploration, analysis, and aggregation - prototypically demonstrated by a use case in the health domain. TILDE thereby contributes to increasing the accuracy and trustworthiness of search results and aligns with the overall goal to foster a European ecosystem for web search infrastructure, emphasizing transparency, trustworthiness, and user empowerment. The respective milestones in this reporting period are: • Milestone 01 [Infrastructure setup] (M2) • Milestone 02 [Algorithm selection process finished] (M3) • Milestone 03 [First version of algorithms and application UI design available] (M6) • Milestone 04 [Methods to evaluate algorithms available] (M9) • Milestone 05 [Final set of algorithms available] (M12) • Milestone 06 [Online demonstrator application available] (M12) In agreement with the consortium (specifically with Shahab Khormali), a cost-neutral extension of the project was confirmed. The project has been extended until the end of October. The final demonstrator will be made available at this point.
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 4 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE 1 Period M1 – M6 Our visual web-platform comprises three modules: 1.) The NLP module semantically enriches the OWI by extracting health-related concepts (such as symptoms, diseases as well as expressions of wellbeing level and mood, etc.) and relations connecting them (knowledge graph (KG)). The combination of open-source LLMs together with retrieval-augmented generation (RAG) techniques enables trustworthy interaction with OWI by content summarization and question answering 2.) The Trustworthiness module investigates "unfair" bias to promote fairness and accuracy. Within the component's RAG architecture, the focus is on prompting techniques across LLMs, and on applying benchmarks and metrics to measure bias through retrieval list comparison. This refinement of our prompting strategies will have a positive effect on detecting and reducing (information) inequalities. 3.) The Visual Web Interface module supports browsing evidence along the KG information and delivers fact-based answers to predefined question-patterns (e.g., "Which treatments are available for a disease?", “What are common symptoms of a disease?”, “Which diagnostic methods are used to diagnose a disease?”) using visualization methods and LLM-based summarizations. 1.1 Module: NLP We analyzed health related content from the OWI by first scraping websites with respect to their curlie labels 1 (focusing on the “/en/Health/*” category): from the ~200.000 collected websites, 70.000 contained microdata, a structured data format supported by “Schema.org” that helps retrieving relevant parts of a website. Using the GliNER (Zaratiana, 2023) library, we extracted health-related concepts such as “disease", "symptoms", "medical procedure" and "drugs". An example overview of COVID-19 related, extracted concepts is visualized in Figure 1. Figure 1 Explorative statistics of COVID-19 related websites from the OWI In addition, we started to generate a medical knowledge graph, relating websites with each other as well as extracted concepts and structured information from the metadata section of the websites 1 Curlie labels: https://www.curlie.org/ (Accessed on: 06.11.2025)
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 5 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE (Figure 2). To standardize health-related mentions of websites and to include expert knowledge, we further linked these extracted entities to the UMLS ontology (Medicine, 2025), a comprehensive and precise clinical healthcare terminology system. KG’s structure in combination with encoded expert knowledge provides us with a solid basis to minimize the risk of hallucination in the upcoming RAG system. Figure 2 Health Knowledge Graph generated from a websites (structured) metadata and extracted and normalized entities from the website body. 1.2 Module: Trustworthiness We concentrated on establishing a core set of metrics for evaluating bias through retrieval list comparison. Following the work of (Melchiorre, 2021) we selected key indicators such as document neutrality, normalized fairness of retrieval results (NFaiRR), and ranker-agnostic fairness of document sets (SetFaiRR). This provides the foundational capability to quantify bias within the document lists retrieved for user queries, which is crucial for informing the re-ranking and display mechanisms within our planned RAG architecture which illustration is presented below:
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 6 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Figure 3 Illustration of RAG architecture and overview of all included components. When evaluating and re-ranking an initial set of documents retrieved by a RAG architecture for a given user query, our approach centers on integrating document neutrality and the Normalized Fairness of Retrieval Results (NFaiRR) metric. First, we assess the neutrality of each document with respect to predefined protected attributes (e.g., gender), classifying it as neutral if it shows no indication of a protected attribute or if it presents a balanced representation of such attributes. This "document neutrality" score is crucial as it forms the basis for the subsequent fairness calculations. In the re-ranking step, we then compute the NFaiRR for the initial set of documents. This metric allows me to quantify the overall fairness of the retrieved list, considering both the neutrality of individual documents and their position in the ranked list. The NFaiRR score, normalized against an ideal fairness score, provides a clear measure of how well-balanced the results are, guiding the re-ranking process to mitigate biases and ensure a more equitable representation across protected attributes, without significant loss of utility. During the fairness evaluation process, we specifically utilized the SetFaiRR metric because of its model-agnostic nature. Unlike NFaiRR, which assesses fairness for a given ranked list and thus inherently reflects a particular ranking model's output, SetFaiRR quantifies the inherent fairness of a document set independent of any specific ranking permutation. This was a deliberate choice. Our primary interest was not just to measure the fairness of a final ranked list, but rather to understand how our re-ranking mechanism—the "model" in a RAG architecture—influences the general desired outcome of fairness. By first establishing the baseline fairness of the document set itself using SetFaiRR, we could then more clearly discern the incremental impact and improvements achieved by our re-ranking strategies on the overall fairness of the results presented to the user.
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 7 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Furthermore, our research has identified a supplementary set of benchmarks such as StereoSet (Nadeem, 2020) and Bias Benchmark for QA (Parrish, 2021). These benchmarks helped us to navigate more precisely towards the desired level of fairness in the final re-ranking step of initially retrieved documents relevant for user query. 1.3 Module: Visual Web Interface To enable users to explore fact-based, trustworthy data, extracted website facts (concepts and relations) can be used to filter the data to only provide relevant information. Furthermore, concepts in the data can be highlighted, enabling users to gain evidence on their questions and easily identify relevant content (Figure 4). As access to OWI was delayed, other data sources were used for this mock-up as well as for the UI design in Figure 5. Figure 4: Mock-up highlighting facts extracted from a document. Figure 5 illustrates a first version of our application UI design providing an ‘aggregated’ and ‘visualized’ answer to the question: “Which diseases are related to the symptom respiratory failure?”. Users are provided with a summary of 44 results, highlighted facts extracted from the search results as well as a heat map to visually explore the relation strength of instances from the two categories “Diseases” and “Symptoms”. Figure 5: Application UI-Design showing highlighted text passages, extracted concepts as well as an aggregated summary to a sample question on the left side and a heat map visualization between instances of the two concepts “Diseases” and “Symptoms” on the right side.
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 8 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE 2 Period M6 – M12 2.1 Module: NLP Building on the work completed in the previous project phase, this module focuses on the development of a vertical, hybrid Retrieval Augmented Generation (RAG) search engine tailored for the healthcare domain. The goal of this module was to design and implement a robust information retrieval pipeline that leverages both structured entity-level search and dense semantic retrieval. Figure 6 visualizes the whole technology stack of this system. Figure 6 Technology stack for a vertical, hybrid search engine in the healthcare domain. Data Indexing and Preprocessing The scrapped health-related websites from the previous period are indexed into an Apache Solr database (Solr, 2025). During indexing, the following preprocessing steps were applied: • Entity extraction and normalization. The extracted and normalized entities from the websites create a structured layer of metadata to support entity-based retrieval. • Chunking. Website content was chunked into overlapping chunks of 500 tokens, facilitating more granular document representation and retrieval. • Embeddings. The title of each webpage and each content chunk were embedded using a Sentence Transformer model, enabling semantic similarity search. For embeddings, the sentence transformer model all-MiniLM-L6-v2 (Reimers, 2024) is utilized. These steps then build the backbone of the hybrid RAG system.
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 15 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Search Result List The search results are displayed below the answer in the lower left area. Users can either investigate the website content (Content View, see Figure 10) where extracted concepts are highlighted or show the list of extracted concepts (Concept View, see Figure 11). The icon below the title enables users to easily switch between the two views. Users can narrow down the result set by clicking on one of the extracted concepts in the “Concept View”. Figure 10 “Content View” showing Website content by highlighting extracted concepts. Figure 11 “Concept View” showing list of extracted concepts from the Website.
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 16 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Visualization Methods The visual web interface provides several visualization methods to analyze and narrow down the result set using the extracted concepts. The facet view, see Figure 12, enables users to identify most frequent concepts and easily narrow down the result set. The tag cloud, see Figure 13, and the bar chart, see Figure 14, show the most frequent concepts for one type at a time and enables users to filter the results. The matrix visualization enables users to analyze and filter most frequent co-occurrences of two selected concepts. Figure 15 provides an example for the concepts “drug” and “anatomy” for the query “covid-19 symptom treatment”. Figure 16 shows the knowledge graph of the most relevant Websites for the query “covid-19 symptom treatment”. Initially, only a maximum of six extracted concepts are shown. However, intelligent exploration methods, see Figure 17, enable users to explore the graph. Figure 12 Facet View for query “covid-19 symptom treatment”.
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 17 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Figure 13 Tag cloud of most frequent drugs for query "covid-19 symptom treatment".
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 18 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Figure 14 Bar chart of most frequent symptoms for the query "covid-19 symptom treatment".
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 19 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Figure 15 Most frequent co-occurrences of concepts “drug” and “anatomy” for query "covid-19 symptom treatment”.
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 20 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE Figure 16: Knowledge graph showing most relevant documents and extracted concepts for query "covid-19 symptom treatment". Figure 17 Knowledge graph node types (left) and graph exploration methods (center and right).
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 21 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE 3 Data Availability The processed website data for the Apache Solr import is available at: https://zenodo.org/records/17512328. The source code for the RAG system is available at: https://github.com/mijantscher/tilde-rag/
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 22 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE 4 Table of Figures Figure 1 Explorative statistics of COVID-19 related websites from the OWI ................................................. 4 Figure 2 Health Knowledge Graph generated from a websites (structured) metadata and extracted and normalized entities from the website body. ................................................................................................ 5 Figure 3 Illustration of RAG architecture and overview of all included components. ............................... 6 Figure 4: Mock-up highlighting facts extracted from a document. ................................................................. 7 Figure 5: Application UI-Design showing highlighted text passages, extracted concepts as well as an aggregated summary to a sample question on the left side and a heat map visualization between instances of the two concepts “Diseases” and “Symptoms” on the right side. ........................................... 7 Figure 6 Technology stack for a vertical, hybrid search engine in the healthcare domain. ....................8 Figure 7 Hybrid RAG Architecture with (i) the database indexing phase and (ii) the retrieval step ....... 9 Figure 8: Visual Web Interface for retrieving and analyzing health-related Websites. ........................... 14 Figure 9 Answer generated for the query "covid-19 symptom treatment". ................................................ 14 Figure 10 “Content View” showing Website content by highlighting extracted concepts. ...................... 15 Figure 11 “Concept View” showing list of extracted concepts from the Website. ..................................... 15 Figure 12 Facet View for query “covid-19 symptom treatment”. .................................................................... 16 Figure 13 Tag cloud of most frequent drugs for query "covid-19 symptom treatment". ......................... 17 Figure 14 Bar chart of most frequent symptoms for the query "covid-19 symptom treatment". ......... 18 Figure 15 Most frequent co-occurrences of concepts “drug” and “anatomy” for query "covid-19 symptom treatment”. ................................................................................................................................................ 19 Figure 16: Knowledge graph showing most relevant documents and extracted concepts for query "covid-19 symptom treatment". ............................................................................................................................ 20 Figure 17 Knowledge graph node types (left) and graph exploration methods (center and right). .... 20
Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 – 2024FSTPC2PN35 – TILDE 23 | OpenWebSearch.EU – Deliverable WP7 Project: 101070014 — Openwebsearch.eu — HORIZON-CL4-2021-HUMAN-01 - 2024FSTPC2PN35 – TILDE 5 Bibliography Chase, H. (2025). LangChain (Version 0.3.66) [Computer software]. Retrieved from LangChain: https://github.com/langchain-ai/langchain Cormack, G. V. (2009). Reciprocal rank fusion outperforms condorcet and individual rank learning methods. Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, 758-789. Khattab, O. a. (2022). Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP. arXiv preprint arXiv:2212.14024. Medicine, N. L. (2025). Unified Medical Language System (UMLS). U.S. Department of Health and Human Services. Retrieved from https://www.nlm.nih.gov/research/umls Melchiorre, A. B.-C. (2021). Investigating gender fairness of recommendation algorithms in the music domain. Information Processing & Management. Nadeem, M. B. (2020). StereoSet: Measuring stereotypical bias in pretrained language models. arXiv preprint arXiv:2004.09456. OpenAI. (2025). GPT-4o Mini [Large-language-model]. Retrieved from https://platform.openai.com/docs/models/gpt-4o-mini Parrish, A. C. (2021). BBQ: A hand-built bias benchmark for question answering. arXiv preprint arXiv:2110.08193. Reimers, N. &. (2024). all-MiniLM-L6-v2 [Machine learning model]. Retrieved from https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2 Solr, A. (2025). Apache Solr (Version 9.8) [Computer software]. Retrieved from https://solr.apache.org/ Zaratiana, U. T. (2023). Gliner: Generalist model for named entity recognition using bidirectional transformer. arXiv preprint arXiv:2311.08526.