scieee AI-readable full text Open interactive document viewer

Kann LLM generierter Code FAIR sein? Herausforderungen für das Forschungsdatenmanagement.

Schiller-Stoff, Sebastian David; Münzer, Leona Elisabeth

Abstract

Folien zur Präsentation "Kann LLM generierter Code FAIR sein? Herausforderungen für das Forschungsdatenmanagement." auf der FORGE 2025 in Rostock, 24. - 26.09.2025

Full text

Kann LLM generierter Code FAIR sein? Herausforderungen für das Forschungsdatenmanagement EN: Can LLM generated CODE be FAIR? Challenges for research data management FORGE 2025, Universität Rostock Sebastian David Schiller-Stoff, Leona Elisabeth Münzer Universität Graz, Institut für Digitale Geisteswissenschaften Core question / thesis - Can/should we trust LLM generated Code in our context? -Short answer: NO - Longer answer: Still no but dependents on the context (e.g. architecture) - What does trust mean? - FAIR principles - What is our context? Methodology 1. Use FAIR principles as (established) analysis framework for research data in the DH 2. Two perspectives: a. Code as data (e.g. LLM-code managed by fedora to produce different views of a digital object) b. Code for data (e.g. data produced by LLM code ) Structure - What is our plan? 1. Our research background a. AI-Pair-Programmers b. Vibe coding 2. FAIR-Analysis of LLM-code a. Perspectives: code as data + code for data 3. Result a. Summary b. Discussion State of research 1. LLMs performance in software engineering / coding benchmarks (e.g. SWE-Bench) increased drastically in the last two years. a. Consequence: Claude 4 Sonnet / Opus is maybe the best software engineer in the world b. But: Still mixed opinions in the literature – especially when studying complex engineering results (design of software systems) or developer experiences (based on polls ). 2. Costs for LLM-code (output tokens AND model performance) fall rapidly: “Cheaper to generate good code” a. Consequence: AI code is everywhere – increased integration of AI tools in existing tools (IDEs, text editors etc.) b. But: Real costs are hidden and are going to rise in the future. “Creation of dependencies and market share”. Sebastian David Schiller-Stoff ○Full Stack-Engineer at the DDH: Department of Digital Humanities | Uni Graz ○https://orcid.org/0000-0001-6941-113X ○Background: Digital Humanities, GAMS, DERLA-Project, Leona Elisabeth Münzer ○Student Assistant at the DDH: Department of Digital Humanities | Uni Graz ○https://orcid.org/0009-0002-7170-8340 ○Background: Archaeology, Art-History, Digital Humanities, DERLA-Project (DDH), Cross-Cultural Heritage Mapping/Modelling Department of Digital Humanities Graz, Austria Challenges and Problems for FAIR (Research-)Data Generated Code as Research Data Lowering the barriers to entry through AI-Pair-Programmers makes it possible even for people without in-depth programming knowledge to generate code. This simplifies access to software development and simultaneously significantly increases productivity (GitHub 2023): More code is created in less time. This AI-generated code is increasingly being incorporated into research software in the Digital Humanities and must therefore be considered (research) data. More research about FAIR data for AI – but what about the data generated by AI? While research focuses heavily on how data can be prepared to be FAIR for AI, the fairness of the data generated by AI, specifically the code, is largely ignored. This is precisely where the question of whether and under what conditions AI-generated code can actually be findable, accessible, interoperable, and reusable, arises. Challenges for FINDABLE Rapid increase in code ●AI-Pair-Programmers significantly increase the speed of software development (e.g., +55.8% according to a GitHub study in 2023) ●This increases the amount of code generated (GitClear 2024), also in the context of research data Metadata Requirements ●To ensure that generated code remains findable, it requires machine-readable metadata and clear descriptions (e.g., regarding function, context, dependencies) ●Without such metadata, code remains fragmented and difficult to reuse Example from Practice: (Meta)data and PIDs GAMS/Fedora3 deliver metadata via OAI-PMH. To make the records discoverable in repositories and search engines, they must be converted into standards-compliant DC fields. ●dc:heading & dc:author sound plausible and often appear in web examples – but they don't exist in the official Dublin Core schema ●dc:uuid is an invention ●uuid:generate-id() creates local IDs that are different for each export. ●XML validation ●Increased validation requirements for infrastructures? Challenges for FINDABLE Persistent Identifiers ●To ensure citability and retrieval, PIDs (Persistent Identifiers) are necessary – comparable to DOIs for publications ●Currently, such standards are often lacking in AI-generated code (Increasing) Infrastructure Requirements ●The volume of new code increases the need for stable repositories and reference systems ●Existing platforms such as GitHub or Zenodo offer approaches, but are not consistently geared toward AI-generated code Challenges for ACCESSIBLE Licensing uncertainties ●It is often unclear whether and how AI-generated code is protected by copyright ●LLMs can reproduce protected code fragments from training data, resulting in legal risks Unclear Origins ●Training data for many models is unknown or undisclosed ●This leaves it unclear who "owns" the generated code and whether it is freely available Challenges for ACCESSIBLE Risk of license conflicts ●If LLM code is integrated into projects without review, existing software licenses may be violated ●Subsequent changes or restrictions to the license terms are possible Proprietary dependencies ●AI code often relies on external APIs, libraries, or platform services ●This can complicate long-term access and limit openness Example from Practice: Dependencies ●ext:normalize-license is a fictional function (proprietary extension, implemented locally only) ●Transformation may work on a server, but not in GAMS or Fedora3 ●License metadata is not delivered in machine-readable form, making it inaccessible Challenges for REUSABLE Declining code quality ●AI assistance systems tend to generate faulty or redundant code ●Code smells, structural weaknesses, and a lack of cleanup make reuse difficult Licensing ●Without a clear legal framework, trust in reusability is impossible ●Uncertainty about whether generated code may be freely used, modified, or redistributed Challenges for REUSABLE Code Smells ●A general term for structural problems in code that are not immediate errors but indicate poor maintainability or design weaknesses ●Typical patterns: Duplicated code, methods that are too long, classes that are too large, unclear responsibilities ●In the AI context: LLMs often generate redundant or unnecessarily complex code, which becomes difficult to read, difficult to maintain, and less reusable Example from Practice: Licensing ●LLM simply writes "Open License" in <dc:rights> ●Formally correct DC field → validated ●But: "Open License" is not clear (which one? CC0? CC-BY? GPL?). ●Other projects cannot decide whether and how the data may be reused → not reusable Challenges for REUSABLE Lack of standardization ●Different coding conventions and inconsistent structures prevent reliable reuse ●DH in particular lack established standards for the reuse of AI-generated code Technical debt ●Untested code adopted creates long-term maintenance and refactoring problems ●This makes sustainable integration into research infrastructures more difficult Conclusion & Outlook Challenges and Problems for FAIR (Research-)Data ●Findable: Increasing demands for visibility and traceability of generated code ●Accessible: Open questions regarding accessibility and legal frameworks ●Interoperable: Hurdles to exchange and portability between systems ●Reusable: Uncertainties regarding quality, trust, and sustainability In summary: The most common errors in LLMs are not only syntax errors, but small semantic inaccuracies that only become apparent much later in the workflow – and then are particularly expensive → technical debt, data loss, legal problems If “LLM generated Code can be FAIR” depends not solely on the (AI) systems used, but on human control (by specialists): FAIR is only partially achievable ●Without post-processing, code often lacks metadata, licensing clarity, and maintainability. Human-in-the-loop is essential ●Qualified experts must handle refactoring, documentation, and quality checks. ●AI tools support the process but do not replace human oversight. Need for legal & technical standards ●Binding guidelines for licensing, quality assurance, and reuse. ●Safeguards against insecure dependencies & hidden risks. Technical Debt -Inclusioncloud Digital Engineering. “AI Is Changing How We Code. But Is Technical Debt the Price Tag?” Blog, 30. April 2025. Zuletzt zugegriffen am 29. Juli 2025, https://inclusioncloud.com/insights/blog/ai-generated-code-technical-debt/. -Doyle, Evan. “AI Makes Tech Debt More Expensive - Gauge - Solving the Monolith/Microservices Dilemma”. 11. Februar 2025. Zuletzt zugegriffen am 24. Juli 2025, https://gauge.sh/blog/ai-makes-tech-debt-more-expensive. Zenodo Community & AGKI-DH -Zenodo: https://zenodo.org/communities/aiassistingdh -AGKI: https://agki-dh.github.io/index.html News July, August & September 2025 - GitHub Copilot RCE Vulnerability via Prompt Injection Leads to Full System Compromise: https://cybersecuritynews.com/github-copilot-rce-vulnerability/ - GitHub launches Copilot agents panel on GitHub.com: https://www.infoworld.com/article/4043343/github-launches-copilot-agents-panel-on-github-com.html - KI als Cybercrime-Copilot: https://www.csoonline.com/article/4048337/ki-greift-erstmals-autonom-an.html - Das machen Menschen besser: Studie zeigt Grenzen von Coding-Agenten: https://www.heise.de/news/Das-machen-Menschen-besser-Studie-zeigt-Grenzen-von-Coding-Agenten-10628483.html - ChatGPT Agent: Reasoning und Action-Modell kombiniert: https://www.heise.de/news/ChatGPT-Agent-Reasoning-und-Action-Modell-kombiniert-10490635.html - Wenn der KI-Agent im Fakeshop kauft: https://www.computerwoche.de/article/4044551/wenn-der-ki-agent-im-fakeshop-kauft.html - 9 Wege, mit Vibe Coding zu scheitern: https://www.computerwoche.de/article/4034385/9-wege-mit-vibe-coding-zu-scheitern.html - So verwundbar sind KI-Agenten: https://www.csoonline.com/article/4037404/so-verwundbar-sind-ki-agenten.html - “Scamlexity” We Put Agentic AI Browsers to the Test - They Clicked, They Paid, They Failed: https://guard.io/labs/scamlexity-we-put-agentic-ai-browsers-to-the-test-they-clicked-they-paid-they-failed - Coinbase setzt auf KI-Coding – neue Sicherheitslücke alarmiert Experten: https://www.btc-echo.de/schlagzeilen/coinbase-experten-warnen-sicherheitsrisiko-ki-coding-215017/ - Wenn Künstliche Intelligenz Programmierer bremst: https://www.it-daily.net/shortnews/wenn-ki-programmierer-bremst