scieee AI-readable full text Open interactive document viewer

Agentic Coding for STEM Research

Cardoen, Ben

Abstract

Invited talk at University of Birmingham lecture series on Artifical Intelligence: 'COSIMO-IDAI tutorial: Agentic Coding for Research.'The presentation covers high level workflow advice to maximize productivity using agents in STEM research.

Full text

Agentic coding for STEM research Ben Cardoen School of Mathematics, University of Birmingham 2025-11-20 Slides licensed under CC BY 4.0 – Ben Cardoen, 2025 1 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Introduction Disclaimer: This is not a pro or against AI presentation. I am documenting my findings on how I make it work for me, in my uses cases. That may or may not work for you I am not claiming the workflow tips capture the latest research, purely my experience There can be errors, so use wisely 2 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Setting the Scene My Research: I focus on designing Background What we’ll cover novel scalable algorithms that capture what cannot be solved or quantified revealing mechanism. PhD/Msc/BSc in Comp Sci. C++, MIPS, Julia, Python, Prolog, Java, Haskell, … Parallel computing, high performance computing, biomedical imaging, causal mechanisms, signal processing on graphs Use AI tools in everything except email and messages, if the use case warrants it. Final writeup is still my own, persuasion, intent, accuracy, style and semantics matter. Patterns, not instructions I can’t show the really powerful examples (publication), the trivial examples don’t have learning value We will explore how to unlock potential This slidedeck is part of my AI-generated workflow, so it’s the perfect example (the template, not the content). Coding ~ Papers, slides, proofs, graphs, geometry, …, not just ‘code’. 3 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Productivity gains These are my experiences, I don’t claim these transfer. ~ 2022-2023 2-3% Copilot for undocumented code Fall 2023 3.5% Interdisciplinary science translation, updating student assignments/quizzes (versioning questions) Spring 2024 4% LaTeX + semantic rephrasing (Writefull) speeds up writing/review. Summer 2024 5% Python generation & plotting mostly solved for solved problems Winter 2024 6% Related work searches reliable enough to verify with sampling, not redoing it from scratch. Spring 2025 10% Parsing funder requirements, journal scope/selection, Google search replaced by Perplexity for 99%. Brainstorming/reasoning Summer 2025 15% Linear Algebra/Complexity reasoning/proofs reliable enough for sketches. Julia generation works 99.9% of the time (even with 2-week old API) Fall 2025 20% Claude system administration (supervised) + Julia CA. Computational geometry becomes possible for non-research level tasks Perceived productivity != Actual productivity. Measure, log, quantify, reflect. 4 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Wisdom versus Knowledge Brooks distinguishes essential from accidental complexity ( ). Essential: Accidental: Brooks 1987 What is the best time and space complexity of recomputing first k eigenvalues of a sparse adjacency matrix after modifying 1 edge? Will pre-order or post-order converge faster in random walk of trees? Is my new distance measure a metric? How can I detect extreme value causality? What mechanism explains misfolded protein accumulation in neurons? Is ageing related to complexity? How do I secure funding? LaTeX fails to compile because 1 missing } at page 225 of your thesis, doesn’t tell you what page. Dependency nightmares, cluster/Cloud failures, untested code Figures (and figure layout), making slides Programming Language limitations, Interdisciplinary jargon Variable aliasing, use after free, data races, broken APIs, vendor lockin, … Wisdom is what you need to solve future essential problems. Knowledge is what, but not how, you learned from past problems. Subtle difference. 5 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References What is productivity? Adaptive Decision-Making Through Information-Theoretic Learning The cost model: Every AI attempt: (generation + verification) c = + c g c v Success probability at attempt : , where n = − ( − ) p n p ∞ p ∞ p 1 r n −1 = 1 − q n p n reflects your upfront investment: query design, model selection, context p 1 High requires excellent foundations p 1 Expected cost: E [ C ( k )] = c n ⋅ + ( ck + ) ∑ n =1 k p n ∏ i =1 n −1 q i C man ∏ i =1 k q i 6 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Sequential adaptation: The naive approach After task : Stochastic update based on task outcome: j ✓ Success → increase trust → = k j +1 K max ✗ Failure → decrease trust → = max(1, − 1) k j +1 k j j = { k j +1 , K max max(1, − 1), k j if task j succeeds (prob. S ( ) = 1 − ) k j ∏ k j i =1 q i if task j fails (prob. F ( ) = ) k j ∏ k j i =1 q i But is this the right way to learn? 7 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References The information-theoretic view: Failure is the strongest signal Let = entropy of your mental model (uncertainty about what works) Information gained from outcomes: Controlled failure = deliberate exploration: Reframe adaptation: Failure → high information → improve strategy** (not abandon) H Success: ≈ 0 I success Confirms existing model, low information Risk: model overfits to your query style Failure: ≫ I failure I success Reveals boundaries, constraints, failure modes Maximum entropy reduction: Δ H ∝ −log p failure Strategy improvement ∝ I (outcome;model capabilities) max tasks Push models to failure → align mental models - failure near boundary → strongest update signal In reality the previous expectation model in intractable, it’s a sequential formulate of the Byzantine general’s problem. 8 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References The Mental Model: Intent-driven autonomy. Wisdom needs a mental model. Ask yourself Black box model approach is not efficient But you don’t need to know Transformer architecture How KL divergence works 2-way communication of `intent’ Understand how models ‘see the world’ Effective interaction boils down to entropy minimization by both Two-way calibration: You build a mental model of AI capabilities, but the AI also models your standards within conversations. Consistent expectations create stable co-adaptation. Where do models ‘fail’? What, really, is a ‘hallucination’? When do I trigger it? Which documents stay in memory? How long? GPT5 will automatically reload your source files if you use git connector? If you never cared about accuracy in code snippets, do you expect that to affect next day’s queries? Across chats? Across models? Adding TOC helps, yes/no. What about 2-column format? If you ask for algorithm X, do you get the optimal, the easy, or the outdated version, and why? How do models encode very large sets of information? Does it matter? 9 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Plan your exits Complex research level task: Simulate high fidelity HIV1 capsid coat with nm accuracy Coding agent (Claude) and Research agent (GPT-5) agreed on plan + logs Both warned that task 4 would have 60% chance of success, as it was identified as research class. Completed 3/5 tasks, then stuck Hit subtle issue in combination of de-novo implementation of two algorithms from papers 4-5 debugging attempts at high level (using logs) Consensus to abort, confidence to resolve ~ 20% Time exploring dead ends is also time saved You want to have the ability to answer fast : Can X be done? Right hand side shows the extract minimum reproducer, showing mesh corruption (edge case selected) 16 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Lost in Translation Translation can work, but you need to guide the model and understand where it can go wrong. Not low-level, high level. Python (C/row-major family) Source ❌ Direct Translation to column major family, order of mangitude performance regression 1 result = 0.02 for i in range(rows):3 for j in range(cols):4 result += matrix[i][j] * 2.05 Mathematical notation of a simple, stable function f ( p , q ) = , p , q ∈ R ,( p − q ) ≠ −1 p 1 + ( p − q ) result = 0.01 for i in 1:rows2 for j in 1:cols3 result += matrix[i, j] * 2.04 end5 end6 ✅ Cache-Optimized (note GPT-5/Claude get this right the first time) result = 0.01 for j in 1:cols2 for i in 1:rows3 result += matrix[i, j] * 2.04 end5 end6 Python implementation of formula can be numerically very unstable def f(p, q): #AI generated, if p ~ q this can have catastrophic cancellation1 return p / (1 + (p - q))2 def f_stable(p, q): #AI generated + told to watch out for stability3 denominator = (1.0 - q) + p4 return p / denominator5 Related translation pitfalls to guide models on: Indexing conventions (0-based vs 1-based) Integer division semantics (floor vs truncate), float/Inf Pass-by-value vs pass-by-reference (mutation), threading model. Short-circuit evaluation order Type promotion and coercion rules Remember the tests you needed? 17 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References You get what you push for At heart, still optimization solvers. Lack native algorithmic backtracking (no explicit search trees), but can backtrack when scaffolded Give explicit capability: “if stuck, ask me/revert to step 2.3” or use agentic retry loops Modern models (o1/o3) + chain-of-thought enable backtracking-like behavior Be VERY careful with costly resources (files, network), models have time/compute budgets. If it has to choose between slow network (but latest file) and memory (out of date), it may have to pick the out of date ‘hallucination’. Make PDFs that LLMs can parse properly, or use markdown. If you make it harder to extract information, the model can’t extract enough information in time, leading to poor semantic reasoning on incomplete information. You then see that as hallucinations in the extreme case. Making PDFs parseable: % --- Parser-friendly add-ons (toggle on for analysis builds) ---1 \usepackage[tagged=true,activate=true,interwordspace=true]{tagpdf} % tagged PDF structure2 \usepackage{accsupp} % \BeginAccSupp... ActualText ... for key equations3 \usepackage[numbered]{bookmark} % robust outlines (pairs with hyperref)4 \usepackage[none]{hyphenat} % avoid hyphenation (improves text extraction)5 \usepackage[final]{microtype} % then optionally: \DisableLigatures[f]{encoding=*}6 \usepackage{axessibility} % auto /ActualText for math (disable if it clashes)7 % Optional microtype tweak:8 % \DisableLigatures[f]{encoding = *}9 Add TOC. No columns, single column. No multifile LaTeX if you share latex, flatten it. When in doubt, both LaTeX and PDF, or QMD. Get that first query (where you share the file) just right. Do Not: “Summarize this” Do: Contrast Section 2.3.1 in terms of consistency with conclusion, then create a structured summary of the narrative flow between methods and results. 18 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References The never get to it problems We all have problems that are annoying, can be solved, but it will take unknown amount of time, and it’s not time critical. Example Firefox freezes at random moments ~2-10s. No evidence in logs System health fine No changes introduced Manual: set timed logs, use iotop monitoring, then search for listed bugs in kernel, filesystem, firefox, …. . ~5+hrs. Claude: I give 2 sentence description, ask it to give me 4 questions it needs answered to start 10 minutes processing, automated search of iotop + syscall frequency (very very low level logs), finds seemingly unrelated kernel bug, checks nvme firmware/fs configuration, argues why this is the cause, proposes 2-3 solutions. Does it always work? No, but I have 1 min to start it, it can’t damage my system (it never needed sudo), and I have 1 minute to confirm, another to fix. 19 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References The Right Way(s) (tm) 20 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Plan so execution is almost not needed anymore Coding agent (CA) Research Agent (RA) You ⇆ ⇆ You translate intent + semantics rules + feasibility criteria RA –> modular tasks + verifiable objectives for CA CA asks clarification, three of you need to agree If it’s not worth planning, why is it worth doing? Right hand side: Task 3/5 of a proof of concept protein coat of HIV1, workplan for Claude Code CA made by GPT-5 based on our conversations. 21 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Leave a trail Export conversations to Github Keep Agent logs in Github Restart from logs (or switch agents between modes) Multiple agents can coordinate using logs Research agents can’t debug live, but they can using logs. Research agents can revisit the conversations (is this really the best algorithm?) 22 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Know Thyself LLMs do not (?) have reflection, you do. Use it. Be aware of what you want to hear. Explain that that is a problem and how to solve it. What’s your blind spot? (Ask a research agent with memory who it thinks you are) Describe your observations on my strengths and weaknesses (solely from memory) as a research fellow, and excluding any personal information. Use a single paragraph, do not leak project ideas. 23 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Flip the table Let the agent ask you questions. Questions force you to think Agent will ask questions it needs to reduce entropy, saving you 2-3 answers. 24 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References Dynamic beats static (or why instructions are harmful) You cannot ask a model if it’s hallucinating (~Liar’s paradox), you can ask a model what would make it hallucinate Hallucinations are just suboptimal solutions Ask for annotated answers (see below), not just sources, ask how much time it needs or what it will need Instructions can be too static, so leave a high level backtracking path, “If instructions block you from what you think is optimal, activate OVERRIDE protocol” Model learns from you, not instructions. When in doubt, ask. Below I’m asking Codex what its current configuration is (global v local). 25 Introduction Wisdom versus Knowledge Success is Learning from Failure The Right Way(s) (tm) Where to go from here Acknowledgements References