Full text
Except where otherwise noted, content on these slides is licensed under a Creative Commons Attribution 4.0 International license (CC BY 4.0) Annoter les données scientifiques par des approches d'intelligence artificielle Essmay Touami, Giulia Di Gennaro, Matthieu Chavent, Pierre Poulain [email protected] Université Paris Cité & CNRS Open Source Experience, 2025-12-11
P. Poulain | CC BY 2 What is molecular dynamics (MD)? Senac et al, Langmuir, 2017. Movie by Patrick Fuchs Gagelin et al, Nature communications, 2023. water + detergent
P. Poulain | CC BY 3 Molecular dynamics simulations require (a lot of) resources expertise computer power Source: Photothèque CNRS/Cyril Frésillon (droits réservés)
P. Poulain | CC BY 4 MD data catalogue (MDverse) → LUMEN Linked User-driven Multidisciplinary Exploration Network Funded by EOSC (2025-2027)
P. Poulain | CC BY 5 Meta(data) Source: Zenodo
P. Poulain | CC BY 6 Meta(data)
P. Poulain | CC BY 7 Extract structured metadata from raw text Named Entity Recognition (NER) = find words belonging to certain categories in texts Source: Wikipedia Jim bought 300 shares of Acme Corp. in 2006. [Jim]Person0bought 300 shares of [Acme Corp.]Organization0in [2006]Time.
P. Poulain | CC BY 8 Extract structured metadata from raw text Named Entity Recognition (NER) = find words belonging to certain categories in texts Source: Figshare 🤔
P. Poulain | CC BY 9 Extract structured metadata from raw text
P. Poulain | CC BY 16 Extraction framework 🧰 Instructor Python – MIT license 12k - V1.13.0 - 06/11/2025 ⭐ OpenAI, Anthropic, Google, local, OpenRouter LamaIndex Python – MIT license https://github.com/run-llama/llama_index 45.7k - V0.14.10 - 04/12/2025 ⭐ OpenAI, Anthropic, Google, local, OpenRouter Pydantic AI Python – MIT license https://github.com/pydantic/pydantic-ai 13.7k - V1.29.0 - 09/12/2025 ⭐ OpenAI, Anthropic, Google, local, OpenRouter
P. Poulain | CC BY 17 Annotation pipeline Unstructured text 📝 📰 Extraction framework 🧰 Prompt 🗨️ Structured output model 🖼️ LLMs 🧠 Structured output 🎁 Evaluation 🎯
P. Poulain | CC BY 18 Evaluation 🎯 No hallucination {"entities": [ {"label": "SOFTNAME", "text": "GROMACS"}, {"label": "MOL", "text": "TIP3P"} ] } ❌ {"entities": [ {"label": "SOFTNAME", "text": "NAMD"}, {"label": "MOL", "text": "TIP3P"} ] } ✅ Correct output format (JSON) Extracted entities: - SOFTNAME: GROMACS - MOL: TIP3P ❌ {"entities": [ {"label": "SOFTNAME", "text": "GROMACS"}, {"label": "MOL", "text": "TIP3P"} ] } ✅ Correct entities detected (compared to ground truth) {"entities": [ {"label": "SOFTNAME", "text": "NAMD"}, {"label": "MOL", "text": "TIP3P"} ] } ❌ {"entities": [ {"label": "MOL", "text": "Ubiquitin"}, {"label": "SOFTNAME", "text": "NAMD"}, {"label": "FFM", "text": "TIP3P"} ] } ✅
P. Poulain | CC BY 19 Results: extraction frameworks 🧰 Models No framework Instructor LlamaIndex Pydantic AI Correct format (%) No hall. (%) Correct format (%) No hall. (%) Correct format (%) No hall. (%) Correct format (%) No hall. (%) gpt-4o 100 100 100 98 100 100 100 95 llama-4-maverick 100 100 100 94 60 55 100 100 qwen-2.5-72b-instruct 60 50 100 92 70 65 100 85 deepseek-chat-v3-0324 65 65 100 100 75 75 100 100
P. Poulain | CC BY 20 Results: LLM 🧠 Models Instructor Correct format (%) No hallucinations (%) Correct extraction (%) gemini-3-pro 100 100 84 gpt-5.1 100 100 79 deepseek-chat-v3-0324 100 100 74 llama-4-maverick 100 94 73 gpt-4o 100 98 69 qwen-2.5-72b-instruct 100 92 69
P. Poulain | CC BY 21 Results: LLM 🧠 Models Instructor Time Correct format (%) No hallucinations (%) Correct extraction (%) gemini-3-pro 100 100 84 5 h gpt-5.1 100 100 79 3.5 h deepseek-chat-v3-0324 100 100 74 3 h llama-4-maverick 100 94 73 8 min gpt-4o 100 98 69 1 h qwen-2.5-72b-instruct 100 92 69 4 min
P. Poulain | CC BY 22 NER with LLMs: Making it better? Strengthen LLMs outputs: ●Consensus / majority vote (send the same prompt to different LLMs / temp.) ●LLM-as-a-judge (one LLM evaluates another LLM’s responses)
P. Poulain | CC BY 23 Conclusion: NER + LLMs = ❤️ ●no training set (only test set) ●Pydantic + Instructor = 🚀 ●gemini-3-pro: 🥇 🐌 ●strengthen the results –reproducibility –performance https://github.com/MDverse/mdner_llm 10.5281/zenodo.17886764
P. Poulain | CC BY 24 Thanks IJM, Paris, France Lisa Bouarroudj Mohamed Oussaren IPBS, Toulouse, France Magdalena Szczuka Matthieu Chavent LBT, Paris, France Marc Baaden Karine Duong Giulia Di Gennaro Benoist Laurent Essmay Touami INRAE, Jouy-en-Josas, France Arnaud Ferré Amsterdam, Netherlands Steven Garcia Univ. Copenhagen, Denmark Johanna K. S. Tiemann Kresten Lindorff-Larsen KTH Royal Inst. Tech. Stockholm, Sweden Lucie Delemotte Erik Lindahl Stockholm Univ. , Sweden Rebecca J. Howard