scieee AI-readable full text Open interactive document viewer

RADAR: Enhancing FAIR Research Data Management with AI Support

Hofmann, Stefan; Soltau, Kerstin; Bach, Felix

Abstract

RADAR, developed and operated by FIZ Karlsruhe, is a well-established research data repository supporting secure archiving, publication, and long-term preservation of data across disciplines. Since its launch in 2017, RADAR has continuously evolved to meet the growing demands of open science. It offers comprehensive metadata support, persistent identifiers, semantic enrichment (e. g. Schema.org, FAIR Signposting), discipline-specific terminologies via TS4NFDI, and integration with platforms such as GitHub, GitLab and WebDAV. Flexible deployment options (RADAR Cloud, RADAR Local) and tailored services (e.g. RADAR4Chem, RADAR4Culture, RADAR4Memory) ensure broad usability and community alignment. As part of our ongoing innovation efforts, we are currently exploring AI-driven enhancements that further support FAIR data practices. These include: AI-assisted metadata enrichment, enabling e. g. the automatic extraction of metadata like relevant keywords AI-assisted FAIRness checks, offering feedback and suggestions to improve the FAIRness of datasets. These developments aim to help researchers meet growing expectations for quality metadata and data stewardship while reducing manual effort.

Full text

ai.radar-service.eu FIZ provides a self hosted AI infrastructure containing Fair-Way Service und RADAR Keyword Service. •20 GB GPU •Focus on SLM and domain-specific AI models (e.g. for services for NFDI communities, i.e. RADAR4Chem, RADAR4Culture, RADAR4Memory) •Data privacy •Calculable costs FAIRness Assessment Fair-Way Service •SLM/LLM-powered FAIR assessment •Open Source (GitHub) •Support for JSON, XML, HTML, … •Scalable •REST API / Python FastAPI •Docker Compose •Ollama for model loading © 2024 Copyright: Anmol Sharma, Databases and Information Systems (i5), RWTH Aachen Authors [email protected] [email protected] [email protected] Licensed under CC-BY 4.0 | https://creativecommons.org/licenses/by/4.0 (September 2025) Enhancing FAIR Research Data Management with AI Support RADAR is an established interdisciplinary repository for the archiving, publication, and longterm preservation of research data. It is developed and operated by FIZ Karlsruhe. We are currently exploring how AI can enhance metadata quality through features such as automated metadata extraction and FAIRness evaluation. TS4NFDI Terminology Services Input / Inference data Metadata Title(s), Description(s), Author(s), Contributor(s), Funder(s), Related Identifier(s), Location … File Contents … … ai.radar-service.eu Generic LLM like ChatGPT Chat AI (GWDG) Etc. AI RADAR Keyword Services Fair-Way Service Dataset Listing RADAR Keyword Service •SLM/LLM-powered keyword search •Based on KeyBERT API / KeyLLM: REST API / Python FastAPI •Can be connected with either ChatGPT or domain-specific SLM •Semantic Support via TS4NFDI © 2024 Copyright: FIZ Karlsruhe Easy to implement; flexible use of LLM (ChatGPT, Mistral). High costs; non-reproducible results; misleading scores (e.g. nearly empty vs. extended metadata); privacy concerns with external services. Open Source; runs on own infrastructure; standardized metrics; high quality reports; compatible also with any domain-specific metadata; promising test results We focus on Method 2 due to promising results, reduced privacy concerns, and the Fair-Way report, ensuring faster implementation. Not yet all FAIRsFAIR checks implemented; AI may introduce errors; high CPU load on low-end hardware; requires dedicated service. LLM & Prompt Engineering. LLM calculates a FAIRness score (0–100%) based on given metadata. Method 1 Fair-Way Service. Returns FAIRness scores with detailed feedback (FAIRsFAIR Data Object Assessment Metrics v0.5) Method 2 Aim: Motivate users to enhance a dataset with FAIR and significant (semantic) metadata. Conclusion Metadata Enhancement Easy to implement; flexible (generic or domain-specific models); analyzes both metadata and file content. Quality depends on model and prompt design, non-reproducible results, high costs; privacy concerns; poor XML metadata output. Domain-specific enrichment with semantic precision; standardized terminology improves interoperability; KeyBERT is OpenSource and thus transparent how it works. We focus on Method 2 due to promising results, plannable costs and privacy concerns. Metadata enhancement is limited to keywords; depends on coverage and quality of terminologies, requires a dedicated service. LLM & Prompt Engineering. A LLM like ChatGPT enhances RADAR metadata. To get better results, several hints and XML examples are provided in the prompt. Method 1 RADAR Keyword Service. Keywords are extracted from metadata and file content via KeyBERT. These keywords are mapped to semantic terms if possible. Method 2 Aim: Support users in enriching datasets/files with significant metadata based on metadata, file content and external resources. Conclusion More Information [email protected] www.radar-service.eu …