scieee AI-readable full text Open interactive document viewer

Smart Slide Generation: Re-Using Training Materials Through the Power of LLMs

Kabjesz, Lea; Gihlein, Lea; Lampert, Mara Harriet; Haase, Robert

Abstract

This record contains the poster "Smart Slide Generation: Re-Using Training Materials Through the Power of LLMs" presented by Lea Kabjesz and Lea Gihlein at the All Hands Meeting of the German AI Centers, held on 5th - 6th November 2025 at DFKI Saarbrücken, Germany. We acknowledge funding by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under the National Research Data Infrastructure – NFDI 46/1 – 501864659.

Full text

CONTACT FUNDED BY IN COOPERATION WITH https://scads.ai/ Training is everywhere – especially in the age of constantly evolving new computational tools and methods. Within the field of bio-image analysis, both biologists and computational scientists need to be trained on how to apply new tools and manage their data, all in the context of FAIR principles (Findable, Accessible, Interoperable, Reusable). The NFDI4BIOIMAGE, a German national consortium of image analysis and data management experts, addresses this by fostering knowledge exchange, collecting training materials, and offering both inperson and online training schools. Lea Kabjesz [email protected] SMART SLIDE GENERATION: RE-USING TRAINING MATERIALS THROUGH THE POWER OF LLMs Lea Kabjesz1,2, Lea Gihlein1,2, Mara Lampert1,3, Robert Haase1,2 1 ScaDS.AI Dresden/Leipzig, 2 Leipzig University, ³ TU Dresden TRAINING MATERIALS ARE JUST DATA HUGGING FACE DATASET Our team in Task Area 5 (Training and Community Integration) focuses on aligning user needs with available resources and develops methods to make training materials more accessible and useful. While we create new materials ourselves through blog posts, lectures, workshops, and more, we also recognize the power that existing materials hold - as they can be treated just like any other data and fed into neural networks. Recent developments in the field of Large Language Models (LLMs) allow us to cluster training materials according to their content in an embedding space. Dimensionality reduction techniques facilitate visualization of relationships between training materials, e.g. to find related topics, even if the user does not know the terms yet he needs to learn to master bio-image data science. This representation can be used to find slides that cluster together or to detect topics that are rather sparsely covered in our dataset. SMART SLIDE GENERATOR We created a Hugging Face dataset - a collection of presentation slides along with their corresponding cached embeddings. Currently, it includes: •2,617 slides •70 presentations •Image embeddings •Text embeddings •Mixed embeddings https://hf.co/datasets/ScaDS-AI/SlideInsight_Cache Our goal is to build a model that enables the sustainable re-use of highquality training materials. Users will be able to explore slides interactively – ask questions, summarize content, etc.- and beyond this: request a fully tailored slide deck that fits their chosen topic, audience, style, and knowledge level. The model will generate slides from existing content, while ensuring proper credit is given to the original author(s) in compliance with FAIR principles. This approach will make content generation scalable, adaptive, personalized – and most importantly, FAIR. EMBEDDINGS We acknowledge funding by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under the National Research Data Infrastructure – NFDI 46/1 – 501864659 The slides come from contributors at NFDI4BIOIMAGE and cover diverse topics ranging from research data management to image analysis. We are continuously updating and expanding the dataset, and welcome contributions! You can explore it here: CENTER FOR SCALABLE DATA ANALYTICS AND ARTIFICIAL INTELLIGENCE NATIONAL RESEARCH DATA MANAGEMENT INFRASTRUCTURE FOR MICROSCOPY AND BIOIMAGE ANALYSIS [email protected] Lea Gihlein UMAP of Mixed Slide Embeddings UMAP-1 UMAP-2 TRAINING MATERIAL LIFE CYCLE Currently, we are working on a system that can: •Ingest existing training materials (e.g. slide decks) •Learn their structure, tone, and topic •Detect the re-use of slides or images to demonstrate how open sharing supports the “FAIR-ification” of data •Create new, customized materials for trainers & trainees https://github.com/NFDI 4BIOIMAGE/SlideInsight