ai.radar-service.eu FIZ provides a self hosted AI infrastructure containing Fair-Way Service und RADAR Keyword Service. •20 GB GPU •Focus on SLM and domain-specific AI models (e.g. for services for NFDI communities, i.e. RADAR4Chem, RADAR4Culture, RADAR4Memory) •Data privacy •Calculable costs FAIRness Assessment Fair-Way Service •SLM/LLM-powered FAIR assessment •Open Source (GitHub) •Support for JSON, XML, HTML, … •Scalable •REST API / Python FastAPI •Docker Compose •Ollama for model loading © 2024 Copyright: Anmol Sharma, Databases and Information Systems (i5), RWTH Aachen Authors
[email protected] [email protected] [email protected] Licensed under CC-BY 4.0 | https://creativecommons.org/licenses/by/4.0 (September 2025) Enhancing FAIR Research Data Management with AI Support RADAR is an established interdisciplinary repository for the archiving, publication, and longterm preservation of research data. It is developed and operated by FIZ Karlsruhe. We are currently exploring how AI can enhance metadata quality through features such as automated metadata extraction and FAIRness evaluation. TS4NFDI Terminology Services Input / Inference data Metadata Title(s), Description(s), Author(s), Contributor(s), Funder(s), Related Identifier(s), Location … File Contents … … ai.radar-service.eu Generic LLM like ChatGPT Chat AI (GWDG) Etc. AI RADAR Keyword Services Fair-Way Service Dataset Listing RADAR Keyword Service •SLM/LLM-powered keyword search •Based on KeyBERT API / KeyLLM: REST API / Python FastAPI •Can be connected with either ChatGPT or domain-specific SLM •Semantic Support via TS4NFDI © 2024 Copyright: FIZ Karlsruhe Easy to implement; flexible use of LLM (ChatGPT, Mistral). High costs; non-reproducible results; misleading scores (e.g. nearly empty vs. extended metadata); privacy concerns with external services. Open Source; runs on own infrastructure; standardized metrics; high quality reports; compatible also with any domain-specific metadata; promising test results We focus on Method 2 due to promising results, reduced privacy concerns, and the Fair-Way report, ensuring faster implementation. Not yet all FAIRsFAIR checks implemented; AI may introduce errors; high CPU load on low-end hardware; requires dedicated service. LLM & Prompt Engineering. LLM calculates a FAIRness score (0–100%) based on given metadata. Method 1 Fair-Way Service. Returns FAIRness scores with detailed feedback (FAIRsFAIR Data Object Assessment Metrics v0.5) Method 2 Aim: Motivate users to enhance a dataset with FAIR and significant (semantic) metadata. Conclusion Metadata Enhancement Easy to implement; flexible (generic or domain-specific models); analyzes both metadata and file content. Quality depends on model and prompt design, non-reproducible results, high costs; privacy concerns; poor XML metadata output. Domain-specific enrichment with semantic precision; standardized terminology improves interoperability; KeyBERT is OpenSource and thus transparent how it works. We focus on Method 2 due to promising results, plannable costs and privacy concerns. Metadata enhancement is limited to keywords; depends on coverage and quality of terminologies, requires a dedicated service. LLM & Prompt Engineering. A LLM like ChatGPT enhances RADAR metadata. To get better results, several hints and XML examples are provided in the prompt. Method 1 RADAR Keyword Service. Keywords are extracted from metadata and file content via KeyBERT. These keywords are mapped to semantic terms if possible. Method 2 Aim: Support users in enriching datasets/files with significant metadata based on metadata, file content and external resources. Conclusion More Information
[email protected] www.radar-service.eu …