Automating Systematic Reviews: API-Powered Bibliographic Data Retrieval Module for NeutrinoReview
Abstract
Appeared in: Open Search Symposium 2025, 8-10 October 2025, CSC IT Center for Science, Helsinki, Finland.
Full text
AUTOMATING SYSTEMATIC REVIEWS: API-POWERED BIBLIOGRAPHIC DATA RETRIEVAL MODULE FOR NEUTRINOREVIEW E. Sandner∗1,3, I. Ilicic3, U. Sharma1, 4, I. Jakovljevic1, A. Simniceanu2, L. Fontana2, A. Henriques1, A. Wagner1, C. Gütl3 1CERN, 1211 Geneva, Switzerland 2WHO, 1211 Geneva, Switzerland 3Graz University of Technology, 8010 Graz, Austria 4University of Delhi, Delhi, India Abstract Systematic reviews are the gold standard for synthesizing research evidence but are highly timeand resourceintensive, often requiring months or even years to complete. While existing review management systems provide support for the screening phase, early steps such as literature retrieval typically require external execution, creating inefficiencies and potential for error. This paper presents a proof-of-concept implementation of an automated data retrieval and deduplication module integrated into the NeutrinoReview platform. The module supports multi-database querying, metadata normalization, and duplicate resolution through a similarity-based algorithm. To assess its performance, a controlled user study compared task completion times with Rayyan, a widely used review tool. Seven novice participants performed retrieval and deduplication workflows for four medical reviews using both systems. Results showed that NeutrinoReview reduced completion time by an average of 75%. These findings highlight the potential of automation to significantly reduce human workload in the early stages of systematic reviews. While NeutrinoReview serves as a proof of concept, the demonstrated efficiency gains underscore the value of integrating robust retrieval and deduplication modules into current and next-generation review management tools to enhance timeliness, consistency, and reliability in evidence synthesis. INTRODUCTION A systematic literature review is a method for identifying, evaluating, and synthesizing all research relevant to a specific question, topic, or phenomenon of interest. By integrating findings from all potentially relevant studies on a given question, a systematic review (SR) provides the most reliable methodology for drawing evidence-based conclusions [1]. Consequently, SRs hold a central role in medical research and practice, where they inform evidence-based decisionmaking and the development of clinical guidelines [2]. SRs are equally important in the context of primary research. Conducting a systematic review of the existing evidence prior to initiating a new study is critical for ensuring its quality and relevance [3]. Comprehensive knowledge of prior studies helps identify research gaps and formulate meaningful questions that warrant further investigation. Moreover, ∗[email protected] insights from earlier work support the optimal design of new studies [4,5]. However, the rigor of the process makes SRs highly timeand resource-intensive. Completing a single SR typically takes several months and, in some cases, even years [6–8]. Because SRs are both timeand resource-intensive, their lengthy process often fails to meet the needs of decisionmakers, particularly in contexts where rapid evidence syntheses are required to inform urgent decisions or where research resources are limited. Tools to reduce the human workload in systematic reviews are available. Most review management systems primarily support the study selection process by streamlining and (semi-)automating the screening tasks [9]. However, the initial search typically has to be conducted externally, after which candidate studies are imported into the tool. Deduplication is generally well supported, although not all tools provide simple one-click functionality and instead rely on more complex options. The extent to which further streamlining of the search and deduplication phases could translate into measurable time savings for human reviewers has not yet been systematically investigated. Therefore, this paper presents a proof-of-concept implementation of a data retrieval and deduplication module for review management systems. In an experimental study, the time required to perform these two steps using the proposed module was compared with the traditional workflow in an established review tool. The results demonstrate that the proposed module reduces the human workload for these tasks by 75%. The findings underscore the importance of developing robust data retrieval systems and integrating them into existing or next-generation review management tools. BACKGROUND AND RELATED WORK Conducting a systematic review typically begins with the development of a project protocol, which outlines the research objectives and provides a detailed roadmap for executing the review. Central to this protocol is the search strategy, which defines both the literature sources to be included and the search strings that will be applied to retrieve potentially relevant studies. After the protocol has been finalized and, where appropriate, published in a registry such as PROSPERO 1 , the literature search is conducted separately 1https://www.crd.york.ac.uk/prospero/ https://doi.org/10.5281/zenodo.17233862
for each database. Typically, this process involves accessing each database’s search engine via a web browser. The search string defined in the protocol is entered into the search interface, the search is executed, and the results are then exported and downloaded. Once all searches have been completed, the resulting files must be merged and deduplicated in preparation for the subsequent screening phase. During screening, the eligibility of each study is first assessed based on the title and abstract, with those deemed eligible then undergoing a full-text evaluation. After eligible studies have been identified, relevant information must be extracted, the risk of bias assessed through critical appraisal, and the findings synthesized into a manuscript. While study selection, data extraction, and critical appraisal have been identified as highly time-intensive tasks, literature search has been found to require comparatively less time [10]. However, since executing the search and transferring files between tools are tasks that do not require human judgment and follow standardized procedures, automating them can save time, reduce the risk of manual errors, and ensure consistency across the review process. While several tools exist to support researchers in conducting systematic reviews, they typically focus on the more time-consuming phases of the process. EPPI-Reviewer [11] and DestillerSR [12], two tools primarily designed to support users during the literature screening phase, also include integrated search functionalities. These search capabilities are limited to PubMed, allowing users to directly retrieve studies from this database, while studies from other sources require manual upload. Other widely used screening tools, such as Rayyan [13] and Covidence [14], rely exclusively on manual uploads. Each of these four tools provides a deduplication feature for the uploaded bibliographic data. In the broader context of the research project within which this study was conducted, a prototype for a new systematic review tool called NeutrinoReview is developed. Its conceptual architecture and vision are described in [15], and the 5-tier algorithm [16] as well as the Cal-X algorithm [17] have been integrated as LLM-based screening mechanisms. Without a data retrieval and deduplication module, searches must be performed via the web interfaces of the selected libraries, deduplication must be carried out separately, and the merged, deduplicated records must then be uploaded to NeutrinoReview. It is hypothesized that minimizing tool fragmentation through an integrated solution streamlines the workflow and reduces the time and effort spent adapting to different environments. However, to the best of our knowledge, the extent of time savings provided by an automated search feature has not yet been investigated. METHODOLOGY This paper introduces a open-source data retrieval and deduplication module for review management systems, engineered to reduce the procedure of multi-database retrieval and deduplication into a streamlined operation. This functionality is integrated into NeutrinoReview 2 using a clientserver model, where the server is a REST API implemented in FastAPI 3 and the client is a web application built with React 4 . The overall system architecture is depicted in Fig. 1. Figure 1: Data Retrieval and Deduplication Design Data Retrieval The data retrieval module streamlines an early step in systematic reviews: collecting literature from multiple databases. The module implements an automated pipeline that queries multiple databases in parallel, consolidates the retrieved records, and prepares the dataset for deduplication. The current implementation supports four databases selected for their public, keyless APIs: PubMed, Europe PMC, Medline, and ArXiv. The module supports databases central to different scientific domains: PubMed, Europe PMC, and Medline for medicine, and ArXiv for disciplines across natural science and computer science. To ensure flexibility, sources that lack a supported API, the system provides a custom import function. This feature allows users to upload a CSV file of citations, which is then processed using the same pipeline. The retrieval logic for the four native databases and the CSV importer is implemented in five distinct classes, each inheriting from a common BaseRetriever class. The data retrieval process is initiated with a user-provided search string and a selection of target databases. For each selected database, the search method of the corresponding retriever class executes the query. Subsequently, the raw 2https://gitlab.cern.ch/caimira/caimira-wp4/ neutrinoreview 3https://fastapi.tiangolo.com/ 4https://react.dev/ https://doi.org/10.5281/zenodo.17233862
data is parsed and normalized into a standardized format that abstracts away the heterogeneity of the various sources. By leveraging asynchronous programming, the system fetches and normalizes records concurrently, yielding a consolidated dataset ready for deduplication. Deduplication Figure 2: Deduplication Algorithm Design This module implements a deduplication algorithm to resolve duplicate entries inherent in data aggregated from heterogeneous sources. The implemented process is based on a robust approach inspired by the Deduklick algorithm [18] and illustrated in Fig. 2. The process begins with metadata normalization, followed by a primary grouping heuristic based on DOI. If a group contains a single article, it is marked as unique. Records in multi-member DOI groups, and all records without a DOI, undergo a pairwise similarity analysis. The algorithm applies a 95% similarity threshold for the former case and a 98% threshold for the latter. The core of the algorithm is the similarity calculation, which computes a composite score from weighted metadata fields. For articles with an abstract, the weighting is as follows: Title (40%), Authors (20%), and Abstract (40%). For those without an abstract, the weighting is: Title (60%) and Authors (40%). As the specific Deduklick weights are not public, these values were determined empirically through rigorous testing. The similarity for each field is calculated using the Levenshtein distance, normalized by the maximum string length [19]. Finally, a database priority system resolves duplicate sets. When a duplicate is identified, the algorithm retains the record according to a predefined database ranking: PubMed is preferred over Medline, which has precedence over Europe PMC, followed by arXiv, and finally custom databases. This hierarchy was established based on consultation with domain experts. After the detection process is complete, the database is updated with the deduplicated records. Evaluation To evaluate NeutrinoReview’s efficiency and usability, a user study was conducted comparing it against a widely used conventional tool called Rayyan 5 . The study involved seven participants, all of whom were novices with no prior experience using either system. Each participant carried out the complete data retrieval and deduplication workflow for four distinct medical systematic reviews using both tools. To ensure a standardized comparison, participants were provided with predefined search strings for PubMed, Medline, and EuropePMC, taken from original published reviews to closely mimic real-world scenarios. The primary performance metric was task completion time, measured from the creation of a new review project to the completion of deduplication in each system. Each participant followed a standardized protocol to ensure procedural consistency. The search strings together with the reference to the corresponding published systematic reviews, as wel as the experiment protocol are provided in the supplementary material6. Initially, participants were given a written instruction detailing the steps to follow and a demonstration of the systematic review workflow in both NeutrinoReview and Rayyan to ensure a consistent baseline of understanding for all users. During the subsequent task-performance phase, a facilitator was present to observe and, upon request, provide clarification or assistance to prevent impasses. This ‘assisted completion’ protocol was designed to ensure that the recorded task times primarily reflect the tool’s efficiency. For a valid comparison, the deduplication process in Rayyan was configured to emulate NeutrinoReview’s automated algorithm. Using Rayyan’s ’AutoResolver’ feature, a multi-step logic was established that mirrored the process in NeutrinoReview. The database priority (PubMed > Medline > EuropePMC) was replicated, and the conflict resolution rules were set to check for duplicates sequentially by DOI, normalized title and author, and a 95% similarity threshold. This methodological alignment ensured an equitable basis for benchmarking the deduplication performance. RESULTS Fig. 3 presents the time required by participants to perform data retrieval and deduplication using NeutrinoReview and Rayyan across four test reviews. The boxplots demonstrate that all participants completed the tasks more quickly with NeutrinoReview than with Rayyan. Notably, even the slowest participant using NeutrinoReview outperformed the fastest participant using Rayyan in 3 out of 4 test reviews. On average, task completion with NeutrinoReview required about 1.5 minutes, whereas the same tasks took more than 6 minutes with Rayyan. This corresponds to a 4-fold increase in speed and an overall time reduction of 75%. These findings indicate that the automated workflow implemented 5https://www.rayyan.ai/ 6https://zenodo.org/records/17075731 https://doi.org/10.5281/zenodo.17233862
Figure 3: Comparison between NeutrinoReview and Rayyan in the presented module substantially reduces the effort required for data retrieval and deduplication in systematic reviews. Integrating such a module into review management tools can streamline the process by minimizing the number of tools and user interfaces involved and by accelerating the initial stages of study selection. It is important to emphasize that these results should not be interpreted as evidence that NeutrinoReview is the superior tool overall. NeutrinoReview is a proof of concept, while Rayyan is a well-established and robust review management platform. Rayyan provides considerably greater flexibility in deduplication settings and may have deliberately refrained from incorporating automated data retrieval due to robustness concerns associated with external dependencies or other considerations. The findings presented here should therefore be understood as a comparison between two fundamentally different approaches: one that emphasizes automation through integrated data retrieval and oneclick deduplication, and another that prioritizes user control through customizable deduplication following manual data import. Which approach is more appropriate in practice will depend on factors beyond time efficiency alone, including robustness, flexibility, and the specific requirements of each review. LIMITATIONS AND FUTURE WORK This study was designed to evaluate the extent of workload reduction achieved through the implementation of automated data retrieval and deduplication modules. As such, other important aspects of system performance were not examined. In particular, the robustness of retrieval and deduplication processes—such as error rates, handling of incomplete or inconsistent metadata, and resilience to changes in source interfaces—remains unassessed. In addition, although the developed module supports automated data retrieval from arXiv, this functionality was not included in the experimental evaluation. The omission was due to technical constraints, specifically the absence of a bulk-export feature in the arXiv web search, which limited the feasibility of a systematic performance comparison. While this study demonstrated that integrating a data retrieval and deduplication module can further streamline the review process, future work should expand the range of supported data sources and systematically evaluate robustness metrics to provide a more comprehensive assessment of automated retrieval and deduplication workflows. CONCLUSION This study introduced and evaluated a proof-of-concept data retrieval and deduplication module designed to streamline the initial phases of systematic reviews. Integrated into the NeutrinoReview platform, the module automates multidatabase retrieval, normalizes metadata, and applies a deduplication algorithm. In a controlled user study, the module reduced the time required for data retrieval and deduplication by more 75% compared with an established review tool. These findings demonstrate the potential of automation to reduce human workload and accelerate the review process. The results underscore the importance of integrating robust retrieval and deduplication capabilities into review manhttps://doi.org/10.5281/zenodo.17233862
agement systems. By minimizing tool fragmentation and providing one-click functionality for otherwise repetitive tasks, such modules can help optimize the efficiency of systematic reviews and support timely evidence synthesis. While NeutrinoReview serves as a proof of concept rather than a production-ready system, the demonstrated workload reduction provides a strong argument for incorporating similar functionalities into existing or next-generation review management tools. Future research should extend the evaluation to cover robustness, scalability, and integration with a broader range of bibliographic databases. Ultimately, advancing such automation has the potential to not only accelerate systematic reviews but also to enhance their consistency, reliability, and impact on evidence-based research. ACKNOWLEDGEMENTS The joint CERN and WHO ARIA 7 project is funding the PhD project, in the context of which this paper was written. Furthermore, we greatly thank OpenWebSearch.EU 8 project and the members for their help and support with this publication. REFERENCES [1] Shekelle PG, Maglione MA, Luoto J, et al. Global Health Evidence Evaluation Framework. Rockville, MD: Agency for Healthcare Research and Quality (US); 2013. Available from: https://www.ncbi.nlm.nih.gov/books/ NBK121300/table/appb.t21/ . Table B.9, NHMRC Evidence Hierarchy: designations of ‘levels of evidence’ according to type of research question (including explanatory notes). [2] Cook DJ, Greengold NL, Ellrodt AG, Weingarten SR. The relation between systematic reviews and practice guidelines. Annals of Internal Medicine, 1997;127(3):210–216. [3] Clarke M, Hopewell S, Chalmers I. Clinical trials should begin and end with systematic reviews of relevant evidence: 12 years and waiting. The Lancet, 2010;376(9734):20–21. [4] Robinson KA, Brunnhuber K, Ciliska D, Juhl CB, Christensen R, Lund H. Evidence-based research series–paper 1: what evidence-based research is and why it is important? Journal of Clinical Epidemiology, 2021;129:151–157. [5] Lund H, Juhl CB, Nørgaard B, Draborg E, Henriksen M, Andreasen J, et al. Evidence-based research series–paper 2: using an evidence-based research approach before a new study is conducted to ensure value. Journal of Clinical Epidemiology, 2021;129:158–166. [6] Beller EM, Chen JK, Wang UL, Glasziou PP. Are systematic reviews up-to-date at the time of publication? Systematic Reviews, 2013;2:1–6. [7] Demetres MR, Wright DN, Hickner A, Jedlicka C, Delgado D. A decade of systematic reviews: an assessment of Weill Cornell Medicine’s systematic review service. Journal of the Medical Library Association (JMLA), 2023;111(3):728. 7https://partnersplatform.who.int/tools/aria 8https://openwebsearch.eu [8] Borah R, Brown AW, Capers PL, Kaiser KA. Analysis of the time and workers needed to conduct systematic reviews of medical interventions using data from the PROSPERO registry. BMJ Open, 2017;7(2):e012545. [9] S. Van der Mierden, K. Tsaioun, A. Bleich, C. H. C. Leenaars et al., “Software tools for literature screening in systematic reviews in biomedical research,” ALTEX, vol. 36, no. 3, pp. 508– 517, 2019. [10] B. Nussbaumer-Streit, M. Ellen, I. Klerings, R. Sfetcu, N. Riva, M. Mahmić-Kaknjo, G. Poulentzas, P. Martinez, E. Baladia, L. E. Ziganshina, et al., "Resource use during systematic review production varies widely: a scoping review", Journal of Clinical Epidemiology, vol. 139, pp. 287–296, 2021, Elsevier. [11] J. Thomas, S. Graziosi, J. Brunton, Z. Ghouze, P. O’Driscoll, M. Bond, A. Koryakina, "EPPI-Reviewer: advanced software for systematic reviews, maps and evidence synthesis", EPPI Centre, UCL Social Research Institute, University College London, 2023, Accessed in July 2025. Available: https: //eppi.ioe.ac.uk/cms/Default.aspx?tabid=2914 [12] DistillerSR Inc., "DistillerSR. Version 2.35", 2023, Accessed in July 2025. Available: https://www.distillersr. com/ [13] M. Ouzzani, H. Hammady, Z. Fedorowicz, A. Elmagarmid, "Rayyan—a web and mobile app for systematic reviews", Systematic Reviews, vol. 5, no. 1, p. 210, 2016, Springer. [14] Veritas Health Innovation, "Covidence systematic review software", Melbourne, Australia, 2025, Accessed July 2025. Available: https://www.covidence.org/ [15] E. Sandner, I. Jakovljevic, A. Simiceanu, L. Fontana, A. Henriques, A. Wagner, C. Gütl, "NeutrinoReview: CONCEPT PROPOSAL FOR AN OPEN SOURCE REVIEW MANAGEMENT TOOL", 6th International Open Search Symposium #ossym2024, p. 53, 2024. [16] E. Sandner, B. Hu, A. Simiceanu, L. Fontana, I. Jakovljevic, A. Henriques, A. Wagner, C. Gütl, "Screening automation for systematic reviews: a 5-tier prompting approach meeting Cochrane’s sensitivity requirement", 2024 2nd International Conference on Foundation and Large Language Models (FLLM), pp. 150–159, 2024, IEEE. [17] E. Sandner, M. Negovetić, K. Kothari, I. Taj, L. Fontana, A. Henriques, I. Jakovljević, A. Simniceanu, A. Wagner, C. Gütl, "Cal-X: Enhancing Systematic Review Screening with LLMs and Next-Token Likelihood Calibration", 2025 3rd International Conference on Foundation and Large Language Models (FLLM), IEEE, 2025, in press. [18] N. Borissov, Q. Haas, B. Minder, D. Kopp-Heim, M. von Gernler, H. Janka, D. Teodoro, P. Amini, "Reducing systematic review burden using Deduklick: a novel, automated, reliable, and explainable deduplication algorithm to foster medical research", Systematic Reviews, vol. 11, no. 1, p. 172, 2022, Springer. [19] Y. Yujian and L. Bo, “A normalized Levenshtein distance metric,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 29, no. 6, pp. 1091–1095, 2007. https://doi.org/10.5281/zenodo.17233862