scieee AI-readable full text Open interactive document viewer

Scientific Question Answering using Hybrid Retrieval Augmented Generation

el Baff, Roxanne; Hecking, Tobias

Abstract

Appeared in: Open Search Symposium 2025, 8-10 October 2025, CSC IT Center for Science, Helsinki, Finland.

Full text

SCIENTIFIC QUESTION ANSWERING USING HYBRID RETRIEVAL AUGMENTED GENERATION Roxanne el Baff, Tobias Hecking German Aerospace Center (DLR), Institute of Software Technology, Cologne, Germany Abstract Large language models have strong capabilities for different purposes, such as searching and question-answering [1]. However, they hallucinate on domain-specific tasks. Recent research shows compound systems outperform standalone LLMs [2, 3]. A scientific domain such as Earth observation (EO), used by different fields such as oceanography and environmental science, requires a a tailored approach to deal with hallucinations to ensure in-depth answers [4]. This presentation introduces a system for EO that integrates LLM-based conversational search and question-answering by focusing on two components: a) data curation for a Retrieval-Augmented Generation (RAG)-based model and b) LLM-based evaluation. For (a), our approach combines multiple data structures connecting different text genres (e.g., scientific and web data). For (b), we test LLM-based evaluation and its alignment with human evaluation. Lucene Index Scientific Abstracts Scientific Artifacts OpenAlex Geoservice Knowledge Graph OWILIX Web data curlie=science, earth language=English Filtered Data download web-data Using NASA Taxonomy Tagger filter create index Scientific Abstracts Scientific Artifacts Using NASA Taxonomy Tagger create keywords With keywords With keywords Data Pipelines create knowledge Graph Figure 1: The two data pipelines: From Data Acquisition and Preprocessing to Knowledge-Graph creation (Top), and index creation (Bottom). APPROACH This section outlines our three-stage approach: 1) Data Pipelines, 2) RAG-based LLM, and 3) Evaluation. 1. Data Pipelines. We create an exhaustive dataset of earth observation, including three text genres [5] for two data pipelines. The first pipeline, KG-pipeline, (Figure1-top) creates a knowledge graph connecting scientific abstracts to scientific artifacts via keywords. The second pipeline, INDEX-pipeline (Figure1-bottom), uses crawled data from the Web, processes it, and generates a searchable index. To filter and tag the data for the EO domain, we develop an EO Tagger, the TaxoTagger, based on the EO NASA taxonomy [6, 7] (e.g., earth storable), given a text, 𝑛 keywords are returned with scores between zero and one. Below, we detail each pipeline: • KG-pipeline: we download publication abstracts from OpenAlex [8], an open index of scholarly works tagged with topics across all scientific domains. We fetch abstracts with topics relevant to EO (e.g., Cosmic Evolution). Also, we download remote sensing and EO data from the DLR geoservice portal 1 . Then, we tag both genres using the TaxoTagger. We build a knowledge graph connecting scientific abstracts to artifacts (here, Geoservice data) via the keywords from the tagger. • INDEX-pipeline: We download a search index shard from OpenWebIndex [9] using owilix [10, 11], restricted to English data tagged with science and earth (Curlie tags 2 ). We filter the data with TaxoTagger, excluding those containing keywords below a specified threshold. Finally, we build a Lucene web index using MOSAIC [12]. 2. RAG-Based LLM. The RAG-based LLM relies on the two pipelines described. When a user queries the LLM, the knowledge graph is queried for the top 𝑘 results, where each result contains the top hits (nodes), along with their neighboring keywords and nodes. Simultaneously, the Lucene index is also queried for top 𝑘 hits. These hits are incorporated within the LLM prompt context along with the query. 3. Evaluation For our evaluation, we first curate a set of prompts from domain experts. Due to the absence of reference data, we use Generative Pseudo-Labeling (GPL) [13] to create a weakly labeled dataset for training a response quality classifier. We conduct an automatic evaluation to show the impact of each data pipeline with an ablation study by comparing zero-shot prompting using one of the pipelines, both, or none. Subsequently, domain experts assess the bestperforming model on criteria such as faithfulness, relevance, and completeness. Lastly, we compare human evaluations to LLM-based assessments to analyze their alignment. CONTRIBUTIONS Our contributions are the following: •A reservoir of heterogeneous data sources for EO. • An approach to evaluate the fusion between traditional retrieval (index-based) search with knowledge graph for a scientific domain with a RAG-based model. • A comparative evaluation approach aligning LLMand human-based evaluation for a scientific domain. 1https://geoservice.dlr.de accessed 25.03.2025 2https://curlie.org/ https://doi.org/10.5281/zenodo.17238529 REFERENCES [1] C. Zhai, “Large language models and future of information retrieval: Opportunities and challenges,” in Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024, pp. 481– 490. [2] M. Zaharia et al.,The shift from models to compound ai systems, https://bair.berkeley.edu/blog/2024/ 02/18/compound-ai-systems/, 2024. [3] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,” Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020. [4] L. Huang et al., “A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions,” ACM Transactions on Information Systems, vol. 43, no. 2, pp. 1–55, 2025. [5] J. Honeder, R. El Baff, T. Hecking, A. Nussbaumer, and C. Guetl, “A geo-contextualized multi-genre scientific search engine: A novel conceptual design and prototype evaluation,” in 8th International Conference on Geoinformatics and Data Analysis, Springer, 2025. [6] J. Dutra and J. Busch, “Nasa technical white paper-enabling knowledge and discovery: Taxonomy development for nasa,” Retrieved January, vol. 15, p. 2003, 2003. [7] D. Miranda, “2020 nasa technology taxonomy,” NASA, Tech. Rep., 2020. [8] J. Priem, H. Piwowar, and R. Orr, “Openalex: A fully-open index of scholarly works, authors, venues, institutions, and concepts,” arXiv preprint arXiv:2205.01833, 2022. [9] M. Granitzer et al., “Impact and development of an open web index for open web search,” Journal of the Association for Information Science and Technology, vol. 75, no. 5, pp. 512– 520, 2024. [10] M. Granitzer et al., “OpenWebSearch.eu - building an open web index on eurohpc ju infrastructures,” Procedia Computer Science, vol. 255, pp. 43–52, 2025. [11] M. Granitzer, OWILIX - Open Web Index Client, 2024. doi:10.5281/zenodo.13833664 [12] S. Gürtl and A. Nussbaumer, MOSAIC Search Engine Framework, version 0.1.0, 2024. doi:10.5281/zenodo.13790237 [13] K. Wang, N. Thakur, N. Reimers, and I. Gurevych, “Gpl: Generative pseudo labeling for unsupervised domain adaptation of dense retrieval,” arXiv preprint arXiv:2112.07577, 2021. https://doi.org/10.5281/zenodo.17238529