scieee AI-readable full text Open interactive document viewer

End-to-End Oncology Clinical Trial LLM Efficiency For Industry Adoption with FDA/ICH Regulations

Kawchak, Kevin

Abstract

The process of a new oncology treatment from drug discovery and preclinical studies through Phase III clinical trials and FDA review can take over a decade and cost hundreds of millions of dollars. Therefore, automated and easy to use large language models which have adapted to code generation for production use of sensitive data should be implemented into cancer trial workflows at scale. Here, LLMs were implemented throughout the study to create table templates, datasets based on author and online literature; and followed by triple and quintuple LLM workflows for optimization, table population, peer review, and meta-analyses. Primary models utilized were Sonnet 4.5 Extended, Gemini 2.5 Pro, GPT-5 High, and Grok 4 Fast - with Sonnet being used the most to quickly solve increasingly larger problems and reach consensuses between other LLM responses. The iterated workflow features 12 tables from discovery through Phase III trials, NDA/BOA & FDA Review that identified potential prior LLM applications, with estimated LLM time and cost reductions for trial activities. These itemized values for each table were summed by AI and provided in a cost summary of an entire glioblastoma drug development process. The 2025 trial baseline total was estimated to be 10-15 years at a cost of \$282M-\$837M, while the LLM/AI assisted workflow was projected to be 4.8-7.7 years at \$215M-\$673M, which is 52-73% faster and 24-50% less expensive. Key performance indicators of baseline vs. LLM, acceleration and cost reduction areas, regulatory framework summary, and call to action were also provided by Sonnet, marking a transition to broader adoption of LLMs for oncology clinical trials. Note: LLM outputs and LLM-generated code locally executed with patient data may require FDA approval.

Full text

End-to-End Oncology Clinical Trial LLM Efficiency For Industry Adoption with FDA/ICH Regulations Kevin Kawchak Chief Executive Officer ChemicalQDevice San Diego, CA October 26, 2025 [email protected] ABSTRACT The process of a new oncology treatment from drug discovery and preclinical studies through Phase III clinical trials and FDA review can take over a decade and cost hundreds of millions of dollars. Therefore, automated and easy to use large language models which have adapted to code generation for production use of sensitive data should be implemented into cancer trial workflows at scale. Here, LLMs were implemented throughout the study to create table templates, datasets based on author and online literature; and followed by triple and quintuple LLM workflows for optimization, table population, peer review, and meta-analyses. Primary models utilized were Sonnet 4.5 Extended, Gemini 2.5 Pro, GPT-5 High, and Grok 4 Fast - with Sonnet being used the most to quickly solve increasingly larger problems and reach consensuses between other LLM responses. The iterated workflow features 12 tables from discovery through Phase III trials, NDA/BOA & FDA Review that identified potential prior LLM applications, with estimated LLM time and cost reductions for trial activities. These itemized values for each table were summed by AI and provided in a cost summary of an entire glioblastoma drug development process. The 2025 trial baseline total was estimated to be 10-15 years at a cost of $282M-$837M, while the LLM/AI assisted workflow was projected to be 4.8-7.7 years at $215M-$673M, which is 52-73% faster and 24-50% less expensive. Key performance indicators of baseline vs. LLM, acceleration and cost reduction areas, regulatory framework summary, and call to action were also provided by Sonnet, marking a transition to broader adoption of LLMs for oncology clinical trials. Note: LLM outputs and LLM-generated code locally executed with patient data may require FDA approval. Keywords Oncology ·Drug Discovery ·Preclinical ·Phase III Trial ·FDA ·LLM Introduction A baseline workflow in oncology clinical trials refers to the amount of time and money required to meet regulatory requirements in bringing a drug to market. This process entails drug discovery, preclinical, Phase I-III clinical trials, NDA/BLA, cross-phase management & MIDD to achieve final Food and Drug Administration (FDA) approval. AI estimates for a successful glioblastoma drug campaign is ~ 10-15 years at a total cost of ~ $282M-$837M [ 1 ]; with phase trial failures escalating costs to $1.5B-$2.5B+. Therefore, LLMs should be used to increase efficiency in all relevant areas. In particular, Anthropic, Gemini, OpenAI, and Grok LLMs have become proficient in processing quantitative data with high contextual awareness in English; while effective LLM code generation can enable on-site processing of patient data [ 1 , 2 , 3 , 4 ]. Developers have thus automated many tasks using LLMs at scale for drug synergy prediction by Li in 2024 [ 5 ]; and rational enzyme design with the Evo 2 model by Arc Institute, Stanford University, and NVIDIA [ 6 ]. In addition, Sanofi is building biology foundation models to assist quantitative systems pharmacology (QSP) and digital twin simulations for Phase 1b trials and dose optimization [7]. Phase II trial patient enrollment could also be improved by the work of Beattie in 2024 with GPT-4 screening of over 200 patients from the n2c2 dataset [ 8 ]. In addition, Phase II trial data management & monitoring (21 CFR 11, ICH E6) could also see improvements, as shown by works of Shekhar [ 9 ] using mCODE profile generation with 92% accuracy. Kawchak in 2025 developed Python based PDAC QSP and digital twin clinical trial simulations [ 10 , 11 ]. Phase III trials can similarly see a boost in efficiency regarding SAP finalization through the work of Kuo in 2025 using RAG-LLM and LoRA for SAP auto-population achieving 75% report drafting time reduction (p<0.01) [ 12 ]. In addition, SDTM programming & validation can see improvements based on Pharmaverse’s [ 13 ] use of a sdtm.oak R package. Additionally, Pfizer has developed a natural language processing (NLP) platform for case processing of unstructured data [ 14 ], while Council for International Organizations of Medical Sciences (CIOMS) in 2025 reported on automated coding with 2.92M AEs processed annually [ 15 ]. Reason in 2024 published on a GPT-4 automated health economic model programming in R in ( ~ 12 min) per NSCLC model [ 16 ]. Kawchak comprehensive VVUQ model validations with 55 tests addressed FDA M15 and ASME V&V 40 [ 10 ]. NDA/BLA regulatory submissions have been addressed by Narrativa with CTD Module 2 assembly and ICH M4 summaries [ 17 ], and a CSR Atlas Knowledge Graph for direct data extraction as submission-ready narratives with full traceability. Table of Contents Introduction 1 Methods 3 Glioblastoma Baseline Trial Workflow & LLM Projections 4 Table 1: Discovery & Preclinical Research 4 Table 2: Phase I - First-in-Human Safety & Dose Escalation 5 Table 3: Phase II - Proof-of-Concept & Dose Optimization 7 Table 4: Phase III - Protocol Design, Regulatory Strategy & Statistical Framework 9 Table 5: Phase III - Patient Recruitment & Site Operations 10 Table 6: Phase III - Data Management & Monitoring 11 Table 7: Phase III - Statistical Analysis & CDISC Data Standards 12 Table 8: Phase III - Safety Pharmacovigilance & Clinical Study Reporting 14 Table 9: Phase III - Health Economics & Market Access 16 Table 10: Cross-Phase - Model-Informed Drug Development 18 Table 11: Cross-Phase - Program Management & Quality Systems 20 Table 12: NDA/BLA - Regulatory Submission & FDA Review 21 Summary: Complete Development Program 24 Overall Program Economics: Baseline vs. LLM-Optimized 28 LLM Technology Stack for Optimal Performance 29 Key Regulatory Framework Summary 30 Takeaway Message: The LLM Revolution in Drug Development 31 Call to Action 33 Limitations and Future Work 34 Conclusions 34 Data availability 35 Acknowledgements 38 10.5281/zenodo.17451709 2Version: October 26, 2025 Methods Table Templates Requests for current standardized clinical trial approaches were submitted to Sonnet 4.5 Extended (Sonnet) and Gemini 2.5 Pro (Gemini); and provided as inputs to Prompt A. Sonnet then processed Prompt A with a query for formatted table templates specifying regulatory standards and AI time/cost savings. The same AI model then processed Output A templates based on additional user instruction. Prompt B1 then optimized table templates and added instructions for each respective data table. The output was further improved in the same conversation with Prompt B2, which yielded improved table instructions. Output B2 was made more relevant with Prompt B3 instructions for "specific clinical trial processes" with relevant regulations. Output B3 was too long for downstream processing, so Prompt B4 specified a smaller output. Output B4 table templates with additional user instructions comprised Prompt C, with additional formatting realized in Output C as the first working 14 table template. Prompts in methods and Sonnet thoughts are available with the paper’s supplementary files. AI generated the baseline workflow, with some author editing. The remainder of the paper was written by the author, with some AI assistance with sources and explanations. Author Works & Online Literature Prompt D was based on Output C and user instructions, and was used separately for each of 18 author 2024-2025 works as docx or tex files with Sonnet. Prompt D primarily focused on how prior LLM works could be used to make current regulated glioblastoma clinical trial activities more efficient. Earlier studies in chemical research, synthesis, and characterization were ultimately removed by AI, with 10 author studies populating final tables. Online LLM/Oncology literature across academia, industry, and government primarily from 2024-2025 were searched using Sonnet 4.5 Extended Research (SonnetR) using Prompt E, which was based on Output C. The Output E report containing online sources was converted into LaTeX tables by Sonnet in Output F1, with BibTeX entries provided in Output F2. Triple LLM Table Consensus & Online Literature Optimizations Prompt G1 included the Sonnet Output F1-F2, ChatGPT 5 Thinking Research (ChatGPT5), and Grok 4 Fast (Grok) analogous online searches, with additional user instructions to further optimize the 14 LaTeX tables. The resulting Output G1 by SonnetR updated the tables by reaching a consensus between the three LLM outputs in approximately 20 min. Output G2 updated the BibTeX entries to correspond to the prior generation. Further improvements to the tables were realized using Output G1/G2 in Output H1 through additional online searches using SonnetR. Output H2 restored input table dimensions while Output H3 synced table citation names to BibTeX entries using Sonnet. Triple LLM Workflow Consensus & Glioblastoma Baseline Trial & Dataset Incorporation Prompt I1 included Gemini, ChatGPT5, and Sonnet baseline workflows regarding conventional 2025 Phase III glioblastoma clinical trials to create a new comprehensive set of 14 tables with regulatory standards & requirements. Additional updates were made for value summations (I2), timeline totals (I3), and discovery/preclinical/trial designations (I4). An example dataset of some author and online sourced tables was introduced alongside user instructions with Input I4 to create a new baseline workflow with improved formatting in Output I5 by Sonnet. The summary portions of I5 were improved with the addition of impact and LLM method designation, yielding Output I6. Output I7 returned a new workflow utilizing I6 and additional formatting fixes. The workflow (I7) was populated with the tables representing author and online works in Output J with main results shown in Table 13 and Table 14. Output K used the same inputs, and also included mathematical calculations for reference, although not all tables were generated due to prompt complexity. Additional oncology LLM opportunities outside the scope of this study may also exist. Quintuple LLM Peer Review & Quad Meta-Analyses: Author & Online Source Searches The most recent DATASET_21Oct25 containing LaTeX tables and BibTeX entries was used to obtain recommended fixes from 5 LLMs (ChatGPT5, ChatGPTo3, Grok Expert, Gemini, Sonnet), each from Prompt L. The LLM outputs provided table row or description fixes based on their online verifications of LLM methods and applications. Next, the author marked each of the output’s recommendations as Y/N, by verifying in literature whether fixes were accurate and relevant. Each LLM’s "Y" corrections were implemented by the author using versioning between sets of corrections. Prompt L, the 5 LLM Outputs, and DATASET_21Oct25 were then included for identical Prompt M meta prompts by Gemini and GPT5h. These two prompt outputs were again combined with Prompt L, the 5 LLM Outputs, and DATASET_21Oct25 for a total of four meta-analyses (Prompts N/O) using Gemini and GPT5h. Meta-analysis Output O LLM B (GPT5h meta prompt, meta analysis) provided the best accuracy and detail for use. DATASET_22Oct25 represents the latest updated author and online tables prior to paper completion. AI software: Unmodified chat or API playground inference LLMs in MacOS 14.5 (23F79), Chrome Version 140.0.7339.215. AI Models: 1. Sonnet: Sonnet 4.5 Extended was accessed through the Claude website professional plan chat interface [1] 2. SonnetR: Sonnet 4.5 Extended Research was accessed through the Claude website professional plan chat interface [1] 3. Gemini: Gemini 2.5 Pro was accessed through aistudio.google.com Settings: Temp=1, Thinking budget = 32768, Off = (Structured output, Code execution, Function calling, Grounding with Google Search, URL context), Off = (Safety settings), Output length = 65536, Top P = 0.95 [2] 4. ChatGPT5: ChatGPT 5 Thinking Deep research was accessed through the ChatGPT website plus plan chat interface [18] 5. ChatGPTo3: ChatGPT o3 Deep research was accessed through the ChatGPT website plus plan chat interface [19] 6. GPT5h: GPT-5 high reasoning API was accessed through the platform.openai.com website interface [3] 7. Grok: xAI Grok 4 Fast and Grok 4 Expert were accessed through the Grok website chat interface [4] 10.5281/zenodo.17451709 3Version: October 26, 2025 2025 Glioblastoma Baseline Clinical Trial Workflow With LLM Method/Application Time and Cost Projections Executive Summary This baseline workflow establishes the standard timeline, costs, and regulatory requirements for a conventional drug development program from discovery through FDA approval, with detailed focus on Phase III clinical trial activities for adult glioblastoma using 2025 practices. This serves as the reference point for evaluating LLM/AI-driven improvements. Total Program Timeline: ~10-15 years (discovery through FDA approval) • Discovery & Preclinical: 3.75-6 years • Phase I: 1.5-2.2 years • Phase II: 3-4 years • Phase III: 5-6 years • NDA Review: 1.2-1.8 years Total Program Cost Range: ~$282M-$837M (clinical development) • Preclinical: $55M-$220M • Phase I: $7.6M-$27.5M • Phase II: $25M-$53M • Phase III: $114M-$315M • NDA/BLA: $4.7M-$11.2M • Cross-Phase Management: $24.5M-$57.8M • Cross-Phase MIDD: $1.4M-$2.6M Note: Industry Average with failures: $1.5B-$2.5B+. Key Regulatory Framework: FDA 21 CFR Parts 50, 56, 58, 312, 314; ICH E6(R3) GCP; ICH E8, E9, E3; CDISC standards Table 1: Discovery & Preclinical Research Description Identifying and validating drug candidates through laboratory and animal studies to support an IND application for first-in-human testing. This foundational phase generates the scientific rationale and safety data required before any human trials can begin. Activities include target identification and validation, medicinal chemistry for lead optimization, in vitro/in vivo efficacy studies, GLP-compliant toxicology studies, manufacturing scale-up, and IND submission preparation. Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost Li 02/24 [5], Chaves 06/24 [20]: CancerGPT (124M parameters) with few-shot learning for drug synergy prediction; Tx-LLM (PaLM-2 fine-tuned) encoding 709 datasets for multi-modal drug discovery Target discovery & validation (HTS, assay dev, target confirmation; ICH M3(R2)) 8-14 months; $5M-$20M -3 to -5 months -$2M to -$5M Solovev 12/24 [21], NVIDIA 02/25 [6]: Multi-agent pipeline with Llama-3.1-70b achieving 92% accuracy for neurodegenerative disease discovery; Evo 2 (40B parameters, 1M token context) for rational enzyme design Hit-to-lead & lead optimization (medicinal chemistry, 5K-10K compounds; 21 CFR 58) 18-28 months; $15M-$70M -6 to -10 months -$8M to -$25M Kawchak 12/24 [ 22 ]: Multi-cancer TME synthesis from 40 papers (609,733 words) with pathway visualization, knowledge graphs, multi-omics integration across multiple tumor types In vitro/in vivo efficacy studies (cell lines, animal models, PK/PD; ICH M3(R2)) 10-16 months (overlaps optimization); $5M-$30M -4 to -6 months -$3M to -$10M Google 02/25 [23]: AI co-scientist multi-agent system built on Gemini 2.0 with tournament evolution for self-improving hypothesis generation, drug repurposing, antimicrobial resistance mechanisms GLP toxicology studies (rat, dog, 2-species carcinogenicity, repro tox; 21 CFR 58) 14-20 months (parallel to efficacy); $15M-$60M -2 to -4 months -$5M to -$15M Kawchak 01/25 [24], Kawchak 04/25 [25]: Bevacizumab 100-article synthesis (914,257 words) in 38 min with 500 verifications; LUAD 300+ citations across 10 topics, 32,675 words in 3.1 hours Manufacturing & CMC development (GMP synthesis, formulation, stability; 21 CFR 312) 16-26 months (parallel to efficacy/tox); $10M-$30M -3 to -6 months -$3M- $8M Kawchak 10/24 [26], AlphaLife 02/25 [27]: Paclitaxel 10-paper synthesis (66,419 words) to 6,322 words in 296s; AuroraPrime GenAI with built-in QC, Microsoft Word 365 integration for regulatory submissions IND submission preparation & regulatory affairs (21 CFR 312) 3-4 months; $5M-$10M -1 to -2 months -$2M to -$4M Table 1: TOTAL = 45-72 months; $55M-$220M baseline. LLM = -15 to -28 months; -$23M to -$67M 10.5281/zenodo.17451709 4Version: October 26, 2025 Regulatory Standards & Requirements FDA 21 CFR Part 58 (Good Laboratory Practice - GLP) • Regulates: Conduct of nonclinical laboratory studies • Requirements: QA programs, SOPs, protocols, raw data documentation, final study reports ICH M3(R2) - Nonclinical Safety Studies • Regulates: Types, duration, and timing of nonclinical safety studies • Requirements: Toxicity studies, safety pharmacology, reproductive tox timing FDA 21 CFR Part 312 (Investigational New Drug Application) • Regulates: IND submission content • Requirements: Animal pharmacology/toxicology, CMC, clinical protocols, investigator information Table 2: Phase I - First-in-Human Safety & Dose Escalation Description Initial human testing to establish safety, tolerability, pharmacokinetics, and maximum tolerated dose (MTD) or recommended Phase II dose (RP2D) in patients with advanced solid tumors. The study follows sequential dose escalation with small cohorts, intensive PK sampling, and optional expansion cohorts to gather additional safety/efficacy data. Concludes with RP2D determination and clinical study report preparation. Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost Maleki 04/24 [28]: GPT-4 (gpt-4-turbo/gpt-4o) for protocol authoring with prompt engineering, metadata extraction, customizable templates achieving significant efficiency improvements Phase I protocol development (dose escalation design, DLT definitions; ICH E6, ICH E8) 4-6 months; $200K-$400K -1.5 to -2 months -$80K to -$150K – IND submission & FDA review (30-day review; 21 CFR 312) 1.5 months; $0 – – – Site initiation (1-3 specialized oncology sites; ICH E6) 2-3 months; $100K-$300K – – – Dose escalation enrollment (sequential cohorts, DLT observation; ICH E6) 12-18 months; $2M-$8M – – – Dose expansion cohort (optional, additional safety/efficacy; ICH E6) 6-9 months; $2M-$8M – – Sanofi 2024 [7]: QSP modeling with AI for virtual Phase 1b trials and dose optimization through digital twin simulations PK/PD analysis & bioanalytical testing (population PK modeling; ICH E4) Concurrent; $500K-$2M -2 to -3 months -$200K to -$800K – Safety monitoring & data management (21 CFR 11, ICH E6) Concurrent; $500K-$1.5M – – – Medical monitoring & CRO costs (ICH E6) Concurrent; $1M-$3M – – – Drug supply & manufacturing (GMP material; 21 CFR 312) Concurrent; $1M-$3M – – – Data analysis & RP2D determination (ICH E9) 2-3 months; Included above – – QInscribe 09/25 [ 29 ], AlphaLife 02/25 [ 27 ]: GenAI-driven CSR generation (90% time reduction: 50-100 hours to 5 hours); AuroraPrime auto-population from SAP/Protocol Clinical Study Report (ICH E3) 3-4 months; $300K-$800K -2 to -3 months -$200K to -$600K Table 2: TOTAL = 24-35 months; $7.6M-$27.5M baseline. LLM = -5.5 to -8 months; -$480K to -$1.55M 10.5281/zenodo.17451709 5Version: October 26, 2025 Regulatory Standards & Requirements FDA 21 CFR Part 312 (Investigational New Drug Regulations) • Regulates: Phase I conduct and safety reporting • Requirements: IND annual reports, safety updates, protocol amendments ICH E6(R3) - Good Clinical Practice (GCP) • Regulates: Ethical conduct of Phase I trials • Requirements: Informed consent, investigator qualifications, monitoring ICH E8 - General Considerations for Clinical Trials • Regulates: Overall trial design principles • Requirements: Scientific validity, patient focus, quality by design ICH E4 - Dose-Response Information to Support Drug Registration • Regulates: Use of dose-response modeling in development • Requirements: Exploration of dose-response relationship, dose selection for Phase III ICH E9 - Statistical Principles for Clinical Trials • Regulates: Statistical methodology • Requirements: Pre-specified endpoints, sample size justification, Type I error control ICH E3 - Structure and Content of Clinical Study Reports • Regulates: CSR format and content • Requirements: Comprehensive 16-section template FDA 21 CFR Part 11 (Electronic Records; Electronic Signatures) • Regulates: Use of electronic records and signatures • Requirements: System validation, audit trails, controls for closed systems FDA Guidance on Dose Escalation Study Design • Regulates: Design of first-in-human studies • Requirements: Starting dose justification, dose escalation scheme, DLT definition, MTD/RP2D determination 10.5281/zenodo.17451709 6Version: October 26, 2025 Table 3: Phase II - Proof-of-Concept & Dose Optimization Description Preliminary efficacy evaluation in target patient population (glioblastoma) enrolling 50-150 patients at multiple sites to determine if the drug shows sufficient anti-tumor activity to warrant Phase III investment. Includes protocol development, site initiation, patient enrollment with concurrent treatment/follow-up, central imaging review, optional biomarker analyses, and primary analysis followed by CSR preparation. Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost Maleki 04/24 [28]: GPT-4 protocol authoring with significant efficiency & accuracy improvements, customizable templates for study design, endpoints, sample size Phase II protocol development & feasibility (study design, endpoints, SAP; ICH E8, E9) 5-7 months; $300K-$600K -2 to -3 months -$120K to -$250K – Regulatory submission & IRB approval (IND amendment, multi-site IRBs; 21 CFR 312, 50, 56) 2-3 months; $150K-$400K – – – Site initiation & training (8-15 sites; ICH E6) 3-5 months; $400K-$900K – – Jin 11/24 [30], Beattie 08/24 [8]: TrialGPT with 42.6% time reduction, 87.3% accuracy; GPT-4 screening with 202 patients from n2c2 dataset Patient enrollment (50-150 patients; ICH E6, FDA Diversity Guidance) 18-24 months; $15M-$30M -4 to -6 months -$3M to -$6M – Treatment & follow-up (concurrent, extends 6-12 months post-enrollment; ICH E6) 24-30 months total; Included in enrollment – – – Central imaging review & response assessment (RECIST 1.1; ICH E6) Concurrent; $1M-$2M – – – Biomarker analysis & companion diagnostics (if applicable; FDA Companion Dx Guidance) Concurrent; $1M-$3M – – Shekhar 10/24 [9], Choi 09/23 [31]: mCODE profile generation with 92% accuracy; ChatGPT 3.5 pathology data extraction 87.7% accuracy at $0.03/patient in 4 hours vs months Data management & monitoring (21 CFR 11, ICH E6) Concurrent; $2M-$4M -2 to -4 months -$800K to -$1.6M – Drug supply (GMP; 21 CFR 312) Concurrent; $2M-$5M – – – CRO management & oversight (ICH E6) Concurrent; $3M-$6M – – – Primary analysis (PFS or ORR; ICH E9) 3-4 months; $500K-$1.5M – – QInscribe 09/25 [29], Narrativa 2025 [17]: 90% time reduction for CSR (50-100 hours to 5 hours); CSR Atlas generates narratives in seconds vs days with full traceability Clinical Study Report (ICH E3) 4-6 months; Included above -3 to -4 months -$400K to -$800K Table 3: TOTAL = 36-48 months; $25M-$53M baseline. LLM = -11 to -17 months; -$4.32M to -$8.65M 10.5281/zenodo.17451709 7Version: October 26, 2025 Regulatory Standards & Requirements ICH E8 - General Considerations for Clinical Trials • Regulates: Phase II design principles • Requirements: Appropriate endpoints, sample size rationale, patient selection ICH E9 - Statistical Principles for Clinical Trials • Regulates: Statistical methodology for Phase II • Requirements: Alpha spending (if multiple looks), interim analysis plans FDA 21 CFR Part 312 (Investigational New Drug Regulations) • Regulates: IND amendments and conduct • Requirements: Protocol submission, safety updates FDA 21 CFR Part 50 (Protection of Human Subjects) • Regulates: Informed consent requirements • Requirements: Elements of consent, documentation FDA 21 CFR Part 56 (Institutional Review Boards) • Regulates: IRB review procedures • Requirements: IRB composition, review timelines ICH E6(R3) - Good Clinical Practice (GCP) • Regulates: Ethical conduct of trials • Requirements: Informed consent, monitoring, data management FDA 21 CFR Part 11 (Electronic Records; Electronic Signatures) • Regulates: Electronic records and signatures • Requirements: System validation, audit trails FDA Guidance on Enhancing Diversity in Clinical Trials • Regulates: Demographic representation strategies • Requirements: Action plans for broadening eligibility, demographic data collection FDA Companion Diagnostic Device Guidance (if applicable) • Regulates: Development of companion diagnostics • Requirements: Analytical and clinical validation ICH E3 - Structure and Content of Clinical Study Reports • Regulates: CSR format and content • Requirements: Comprehensive reporting template FDA Guidance on Adaptive Design (if applicable) • Regulates: Seamless Phase II/III or adaptive designs • Requirements: Pre-specification, simulation studies, Type I error control 10.5281/zenodo.17451709 8Version: October 26, 2025 Table 4: Phase III - Protocol Design, Regulatory Strategy & Statistical Framework Description Creating the definitive Phase III trial protocol, determining statistical framework (fixed vs. adaptive design), and obtaining regulatory agreement before trial initiation. This phase includes study design conceptualization, SAP development with power calculations, protocol drafting and expert review, fixed design simulations, IND protocol amendment preparation and FDA review, multi-site IRB submissions in staggered waves, and site activation. Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost Maleki 04/24 [28], Wu 07/25 [32]: GPT-4 protocol authoring with customization; BACTA-GPT fine-tuned GPT-3.5 for Bayesian adaptive design, JAGS code generation, full simulation Study design conceptualization & regulatory strategy (trial structure, endpoints, comparators; ICH E8) 1.5 months; $50K-$100K -0.5 to -1 months -$20K to -$40K Kuo 08/25 [12]: RAG-LLM with LoRA fine-tuning, GRPO reinforcement learning for SAP auto-population, TFL generation, 75% report drafting time reduction Statistical Analysis Plan development (power calculations, sample size, multiplicity; ICH E9) 2.5 months (overlaps drafting); $75K-$125K -1 to -1.5 months -$30K to -$50K Maleki 04/24 [ 28 ]: GPT-4 for protocol authoring achieving significant efficiency improvements, comprehensive protocol generation including ICF, CRFs Protocol drafting (comprehensive protocol, ICF, CRFs; ICH E6) 3 months; $100K-$200K -1 to -1.5 months -$40K to -$80K – Expert & feasibility review (KOLs, sites, biostatisticians; ICH E8) 1.5 months; $50K-$150K – – – Protocol finalization & amendments (ICH E6) 1.5 months; $50K-$125K – – Wu 07/25 [32]: BACTA-GPT for adaptive trial simulation reducing weeks of programming to automated generation, full simulation with interim analyses Fixed design simulations & sample size verification (ICH E9) Concurrent with SAP; $70K-$140K -1 to -2 months -$30K to -$60K – IND/Protocol Amendment preparation (21 CFR 312) 1.5 months; $75K-$200K – – – FDA review period (30-day review clock; 21 CFR 312) 1 month; $0 – – – IRB/EC submissions (30-60 sites, staggered waves; 21 CFR 50, 56) 2.5 months (overlaps FDA); $150K-$600K – – – Response to regulatory/IRB queries (21 CFR 312, ICH E6) 1.5 months (IRBs overlap); $50K-$100K – – – Site activation, approvals (ICH E6) 1 month; above – – Table 4: TOTAL = 12-17 months; $670K-$1.79M baseline. LLM = -3.5 to -5 months; -$120K to -$230K Regulatory Standards & Requirements ICH E6(R3) - Good Clinical Practice (GCP) • Regulates: Protocol content and ethical standards • Requirements: Section 6 details complete protocol requirements ICH E8 - General Considerations for Clinical Trials • Regulates: Overall trial design principles • Requirements: Scientific validity, patient focus, quality by design ICH E9 - Statistical Principles for Clinical Trials • Regulates: Statistical methodology • Requirements: Pre-specified endpoints, sample size justification, Type I error control, estimands FDA 21 CFR Part 312 (Investigational New Drug Application) • Regulates: IND submission and amendments • Requirements: Protocol submission, investigator information, safety updates FDA 21 CFR Part 50 (Protection of Human Subjects) • Regulates: Informed consent process • Requirements: Elements of consent (§50.25), documentation requirements FDA 21 CFR Part 56 (Institutional Review Boards) • Regulates: IRB review procedures • Requirements: IRB composition and procedures FDA Guidance on Adaptive Design (2019) - NOT applicable to conventional baseline • Note: Conventional trials use fixed designs; adaptive approaches could reduce sample size/duration by 20-30% FDA Guidance on Master Protocols (2022) - NOT applicable to conventional baseline • Note: Platform trials could reduce per-comparison costs by 30-50% 10.5281/zenodo.17451709 9Version: October 26, 2025 Table 9: Phase III - Health Economics & Market Access Description Developing economic evidence, cost-effectiveness models, and value propositions to demonstrate drug’s economic value to payers and HTA bodies for reimbursement. Work begins 8-12 months before trial completion with economic model conceptualization/programming, data collection (trial data + real-world evidence), model validation, value dossier preparation, and HTA submissions to multiple regions (US/ICER, UK/NICE, Canada/CADTH, EU/EUnetHTA). Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost Reason 02/24 [16]: GPT-4 automated health economic model programming in R, 715s (~12 min) per NSCLC model, 93% error-free, ICER within 1% published values Economic model conceptualization & structure (model type, perspective, time horizon; ISPOR) 2 months; $40K-$70K -0.7 to -1 months -$18K to -$30K Reason 02/24 [16], Kawchak 02/25 [46]: GPT-4 partitioned survival model automation with multi-prompt architecture; o3-mini reasoning with linear dependency prompting for cost solutions & tables in <5 min Model development & programming (Markov/PSM, TreeAge/R/Excel; ISPOR) 5 months; $150K-$300K -2 to -3 months -$70K to -$140K Kawchak 02/25 [46]: 45 focused reports from 357,381 words generating 24,827 words with currency standardization; o3-mini 15 global mAb reports producing comprehensive financial projections Data collection - trial data extraction (survival, AEs, HRU; CHEERS 2022) 8 months (overlaps 5m with model dev); $100K-$200K -3 to -4 months -$50K to -$100K – Real-world evidence data acquisition & analysis (claims, registries; ISPOR) Concurrent with data collection; $200K-$400K – – Kawchak 02/25 [46]: o3-mini rapid scenario modeling & financial forecasting with linear dependency prompting generating complete economic tables in <5 minutes Model validation & sensitivity analyses (internal, external, PSA, scenarios; ISPOR) 3 months (after trial data available); $50K-$100K -1 to -1.5 months -$25K to -$50K – Cost-effectiveness model development (ICERs, QALYs; CHEERS 2022) Included in model development; Included above – – – Budget impact model development (5-year projections; ISPOR) Concurrent; $75K-$150K – – Kawchak 02/25 [46], Ferber 05/24 [47]: Comprehensive 1,539-word reports, 496-word solutions, 463-word tables in 2.3 min; GPT-4 RAG for ASCO & ESMO guideline interpretation Value dossier preparation (medical writing, evidence synthesis; NICE, CADTH) 3 months; $100K-$200K -1.5 to -2 months -$50K to -$100K – HTA submissions (US/ICER, UK/NICE, Canada/CADTH, EU/EUnetHTA: $40K × 4 regions) 2 months; $120K-$160K – – – Market access consulting & pricing strategy (ISPOR) Concurrent; $150K-$300K – – – Health economist team lead (15 months; ISPOR) 15 months total; $190K-$310K – – – Junior health economists & modelers (2 FTEs × 12 months; ISPOR) 12 months; $180K-$300K – – Table 9: TOTAL = 15-20 months; $1.17M-$2.12M baseline. LLM = -8 to -11.5 months; -$213K to -$420K 10.5281/zenodo.17451709 16 Version: October 26, 2025 Regulatory Standards & Requirements CHEERS 2022 (Consolidated Health Economic Evaluation Reporting Standards) • Regulates: Reporting standards for economic evaluations • Requirements: Title, abstract, methods (perspective, comparators, time horizon, discount rate), results (costs, outcomes, uncertainty) ISPOR Good Practice Guidelines • Regulates: Methodological standards for economic models • Requirements: Model transparency, validation, uncertainty handling, budget impact methodology NICE (UK) - National Institute for Health and Care Excellence • Regulates: UK health technology assessment • Requirements: Reference case methods, QALY calculations, willingness-to-pay thresholds (£20K-£30K per QALY) CADTH (Canada) - Canadian Agency for Drugs and Technologies in Health • Regulates: Canadian therapeutic review • Requirements: Economic guidelines, provincial formulary considerations ICER (US) - Institute for Clinical and Economic Review • Regulates: US value assessment • Requirements: Cost-effectiveness thresholds ($100K-$150K per QALY), value-based pricing EUnetHTA (Europe) - European Network for Health Technology Assessment • Regulates: European joint assessments • Requirements: Core HTA model, 5-domain framework 10.5281/zenodo.17451709 17 Version: October 26, 2025 Table 10: Cross-Phase - Model-Informed Drug Development Description Advanced mechanistic modeling (PopPK, PK/PD, QSP, PBPK, digital twins) to understand drug behavior and optimize development decisions across all phases. Includes model conceptualization with literature review, data assembly from preclinical/clinical studies, model development/coding using specialized software, calibration/parameter estimation, validation (internal/external/qualification), trial simulations, and regulatory documentation for FDA M15/EMA submissions. Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost Sanofi 2024 [ 7 ], Certara 2025 [ 48 ]: QSP with AI target ID engines, digital twins with differential equations for real-time interactions achieving trial duration reduction; AI-enabled QSP platform with pre-validated models, purpose-built GPT, Many FDA novel drug approvals Model conceptualization & literature review (modeling objectives, approach selection; FDA M15) 2.5 months; $50K-$90K -1 to -1.5 months -$20K to -$40K Stahlberg 10/22 [49], Vidovszky 07/24 [50]: Cancer Patient Digital Twin with high-performance computing, 9,000+ patients validated; AI-generated digital twins with PROCOVA methodology, Bayesian methods for rare disease Data assembly & curation (preclinical, Phase I PK, disease data; FDA PopPK Guidance) 2.5 months (overlaps 1m with literature review); $40K-$90K -0.8 to -1.2 months -$16K to -$35K Certara 2025 [48], Kawchak 08/25 [10]: Pre-validated model libraries across therapeutic areas, AI model agnostic architecture; autonomous code 34+ Python notebooks (5000+ lines) for QSP in weeks vs 2-3 months Model development & coding (NONMEM, Monolix, SimBiology, custom code; FDA M15) 5 months; $120K-$220K -2 to -3 months -$50K to -$110K Sanofi 2024 [7]: Virtual Phase 1b proof-of-mechanism trials with dose optimization through digital twin simulations, lunsekimig endpoints predicted, research processes: weeks to hours Model calibration & parameter estimation (MLE, Bayesian, optimization; ICH E4) 4 months; $90K-$180K -1.5 to -2 months -$40K to -$80K Stahlberg 10/22 [49], Kawchak 09/25 [11]: CPDT RECIST 1.1 prediction with deep phenotyping; comprehensive VVUQ with 55 tests (verification, validation, UQ, applicability) automation Model validation (internal, external, qualification; FDA M15, ASME V&V 40) 4 months; $100K-$200K -1.5 to -2 months -$40K to -$90K Vidovszky 07/24 [50], Kawchak 06/25 [51]: Prognostic baseline covariate generation enabling smaller control groups while maintaining power; Simulations with 10,000 virtual patients achieving <5% CV Trial simulations & scenario analysis (design optimization, endpoint selection; FDA M15) 3 months; $70K-$150K -1 to -1.5 months -$30K to -$65K Kawchak 09/25 [11]: Automated FDA guidance processing (M15, ASME V&V 40) generating test recommendations; complete ASME V&V 40 test suite with 55 tests in days vs 3-4 weeks Regulatory documentation & consultation (qualification reports, FDA meetings; FDA M15, EMA PBPK) 1.5 months (overlaps with simulations); $50K-$100K -0.5 to -0.8 months -$20K to -$45K – Lead pharmacometrician/QSP modeler (15 months; FDA M15) 15 months total; $225K-$350K – – – Junior modelers/programmers (1-2 FTEs × 12 months; ICH E4) 12 months; $150K-$300K – – – Specialized modeling software & licenses (NONMEM, Monolix, etc.; FDA M15) Concurrent; $50K-$100K – – – High-performance computing resources (FDA M15) Concurrent; $30K-$80K – – – External QSP modeling consultants (FDA M15, ASME V&V 40) As needed; $100K-$250K – – Table 10: TOTAL = 14-20 months; $695K-$1.42M baseline. LLM = -8.3 to -11.5 months; -$216K to -$465K Phase-specific MIDD totals: • Phase I PopPK: 6-9 months; $300K-$500K • Phase II PK/PD: 8-12 months; $400K-$700K • Phase III QSP/digital twin: 14-20 months; $695K-$1.42M Total estimated across program: • Baseline: $1.4M-$2.6M • LLM: $924K-$1.655M 10.5281/zenodo.17451709 18 Version: October 26, 2025 Regulatory Standards & Requirements FDA Model-Informed Drug Development (MIDD) Guidance • Regulates: Use of modeling and simulation in development • Requirements: Model qualification, documentation of assumptions, validation data, sensitivity analyses FDA MIDD Paired Meeting Program (2023-2027) • Regulates: Framework for regulatory consultation • Requirements: Meeting request package, model documentation, proposed application FDA Guidance on Population Pharmacokinetics • Regulates: Population PK analysis methods • Requirements: Model building strategy, covariate selection, validation, diagnostic plots ICH E4 - Dose-Response Information to Support Drug Registration • Regulates: Use of dose-response modeling • Requirements: Exploration of dose-response, dose selection for Phase III FDA Guidance on Physiologically Based Pharmacokinetic (PBPK) Analyses • Regulates: PBPK model development • Requirements: Model verification, validation with clinical data, sensitivity analysis FDA Guidance on Exposure-Response Relationships (2003) • Regulates: E-R analysis to support dosing • Requirements: PK and PD data integration, covariate effects FDA M15 - Model Credibility Assessment (Draft 2024) • Regulates: In silico evidence for regulatory submissions • Requirements: Context of Use, Model Risk evaluation, credibility evidence (verification, validation, UQ, applicability) ASME V&V 40-2018 - Verification and Validation of Computational Models • Regulates: Quality standards for computational models • Requirements: Verification (code correctness), validation (predictive capability), UQ, documentation EMA PBPK Reporting Format • Regulates: European PBPK submissions • Requirements: Model structure, validation data, sensitivity analyses 10.5281/zenodo.17451709 19 Version: October 26, 2025 Table 11: Cross-Phase - Program Management & Quality Systems Description Overarching project management, quality systems, vendor coordination, and operational oversight spanning entire drug development from preclinical through NDA submission, providing essential infrastructure that runs parallel to operational activities. Activities include program setup/planning, ongoing project management with cross-functional meetings, QA/QC operations with internal audits, regulatory affairs support, vendor management, eTMF maintenance, training programs, document control, risk management, and project close-out. Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost Phase III Management (as representative example): Kawchak 06/25 [51], Kawchak 08/25 [10]: Multi-LLM orchestration (ChatGPT-5, Gemini 2.5, Opus 4.1, Sonnet 4.5, Grok 3/4) for complete trial framework; bidirectional learning with sense-analyze-recommend-act-learn feedback loops Project planning & setup (project plan, vendor selection, TMF, SOPs; ICH E6) 4 months (before trial start); $300K-$600K -1 to -1.5 months -$100K to -$200K – Clinical trial manager & operations lead (ICH E6) 60 months (5 years); $1M – – – Project managers (2-3 FTEs; ICH E6) 60 months; $1.4M-$2.1M – – – Clinical operations specialists (4-6 FTEs; ICH E6) 60 months; $2.2M-$3.3M – – Kawchak 05/25 [41], Kawchak 09/25 [11]: Cross-validation across 4 AI models achieving 81% best context accuracy; comprehensive meta-verification framework with up to 95% concordance Quality Assurance/Quality Control team (2-3 FTEs; ICH E6, ICH Q10) 60 months; $1.3M-$2.0M -8 to -12 months -$500K to -$900K Wu 05/24 [52], AlphaLife 02/25 [27]: Localized LLM for FDA labeling Q&A, semantic search & query processing; AuroraPrime GenAI with built-in QC, Microsoft Word 365 integration Regulatory affairs team (2 FTEs; 21 CFR 312) 60 months; $1.5M -6 to -10 months -$400K to -$700K – Vendor management & oversight (CROs, labs, EDC; ICH E6) 60 months; $1.5M-$2.5M – – – Electronic Trial Master File (eTMF) system (21 CFR 11, ICH E6) 60 months; $200K-$400K – – – Training programs & SOPs (GCP, protocol, systems; ICH E6) 60 months; $300K-$600K – – Kawchak 05/25 [ 41 ]: Quad AI peer review identifying Top 12 issues; automated verification of 40 meta-analyses with similarity scoring avg 95% Internal audits (2-3 per year × $75K × 5 years; ICH E6) 60 months; $750K-$1.13M -6 to -10 months -$300K to -$550K – Risk management & quality systems (CTQFs, QTLs; ICH E6) 60 months; $500K-$1M – – – Document control & archiving (ICH E6, 21 CFR 312.57) 60 months; $300K-$500K – – – Cross-functional meetings & governance (ICH E6) 60 months; $500K-$800K – – – Project close-out (vendor reconciliation, archiving; ICH E6) 4 months (after DB lock); Included above – – Table 11: TOTAL = 62 months; $11.5M-$16.8M baseline. LLM = -20 to -32 months; -$1.3M to -$2.35M Full Program Management Summary: • Preclinical: 48 months; $8M-$30M baseline →-6 to -10 months; -$2M to -$8M • Phase I: 24 months; $1M-$3M baseline →-3 to -5 months; -$300K to -$900K • Phase II: 40 months; $4M-$8M baseline →-6 to -10 months; -$1M to -$2.4M • Phase III: 62 months; $11.5M-$16.8M baseline →-20 to -32 months; -$1.3M to -$2.35M • NDA preparation: 18 months; $1M-$2M baseline →-2 to -4 months; -$300K to -$600K TOTAL across all phases: • Baseline: ~10-15 years; $24.5M-$57.8M • LLM: -37 to -61 months; -$4.9M to -$14.25M 10.5281/zenodo.17451709 20 Version: October 26, 2025 Regulatory Standards & Requirements ICH E6(R3) - Good Clinical Practice (GCP) • Regulates: Comprehensive sponsor responsibilities • Requirements: Section 5 (sponsor duties), Section 2.10-2.13 (quality management), QTLs, risk-based approach FDA 21 CFR 312.50-312.62 (Sponsor Responsibilities) • Regulates: Sponsor obligations including delegation • Requirements: Record retention (312.57 - minimum 2 years post-approval), monitoring (312.56), investigator selection (312.53) ICH Q10 - Pharmaceutical Quality System • Regulates: Quality management in pharmaceutical development • Requirements: Management responsibility, quality procedures, CAPA, change management FDA 21 CFR Part 11 (Electronic Records; Electronic Signatures) • Regulates: Electronic records and signatures • Requirements: System validation, audit trails, electronic signatures FDA 21 CFR Part 211 (cGMP for Finished Pharmaceuticals) • Regulates: Manufacturing quality systems • Requirements: Quality control unit, personnel qualifications, facilities, equipment Table 12: NDA/BLA - Regulatory Submission & FDA Review Description Preparation, submission, and FDA review of NDA/BLA to obtain marketing approval—the culmination of all development. Includes Integrated Summary of Safety/Efficacy (ISS/ISE) preparation, Common Technical Document (CTD) Module assembly (1-5), eCTD formatting, submission filing, FDA 60-day filing review, substantive review (10 months standard or 6 months priority), responses to FDA information requests, optional Advisory Committee preparation, labeling negotiations, manufacturing inspection, REMS development if required, and post-marketing commitment planning. 10.5281/zenodo.17451709 21 Version: October 26, 2025 Potential LLM Method or LLM Application Current Clinical Trial Activity Trial Time And Cost LLM ∆Time LLM ∆Cost QInscribe 09/25 [29], Kawchak 01/25 [24]: GenAI CSR generation 90% time reduction; Systematic reviews with AI generated summaries with PRISMA compliance in weeks vs 6-12 months Integrated Summary of Safety (ISS) & Efficacy (ISE) (integrate all trials; ICH M4, ICH E3) 4-6 months; $300K-$600K -2 to -3 months -$180K to -$350K Wu 05/24 [52], AlphaLife 02/25 [27]: Localized LLM for FDA labeling Q&A with semantic search; AuroraPrime auto-population from SAP/Protocol with built-in QC CTD Module 1 assembly (Administrative & Prescribing Information; 21 CFR 314, eCTD) 2 mos. (overlaps ISS/ISE); $100K-$200K -0.7 to -1 months -$50K to -$100K Narrativa 2025 [17]: CSR Atlas Knowledge Graph with generative AI for direct data extraction from TLFs & ADaM, submission-ready narratives with full traceability CTD Module 2 assembly (Summaries; ICH M4) 4 mos. (overlaps ISS/ISE); $200K-$400K -1.5 to -2 months -$100K to -$200K – CTD Module 3 assembly (Quality/CMC; ICH M4, 21 CFR 314) 3 months; $200K-$400K – – – CTD Module 4 assembly (Nonclinical; ICH M4, ICH M3) 2 months; $100K-$200K – – Pharmaverse 10/24 [13], Narrativa 2025 [17]: sdtm.oak modular SDTM with reusable algorithms; CSR Atlas direct TLF & ADaM extraction with ready output CTD Module 5 assembly (Clinical; ICH M4, ICH E3, CDISC) 5 months; $300K-$600K -2 to -2.5 months -$150K to -$300K – eCTD formatting & submission preparation (eCTD specifications, 21 CFR 314) 2 mos. (overlaps assemb.); $100K-$200K – – Wu 05/24 [52]: Localized LLM framework for FDA drug labeling, comparable to ChatGPT performance, computationally inexpensive, secure IT deployment Regulatory affairs team (12 months submission prep/review support; 21 CFR 314) 12 months; $200K-$400K -2 to -3 months -$80K to -$150K – External regulatory consultants (21 CFR 314) As needed; $300K-$600K – – – Quality review & validation (ICH E6, 21 CFR 314) 2 months; $200K-$400K – – – NDA/BLA application fee (2025 PDUFA fee: ~$3.18M standard, ~$1.59M small business) At submission; $1.6M-$3.2M – – – Submission filing (electronic via FDA Gateway; 21 CFR 314) 1 month; Included above – – – FDA 60-day filing review (21 CFR 314.101) 2 months; $0 – – – FDA substantive review - standard (10-month PDUFA goal; PDUFA) 10 months; $0 – – – FDA substantive review - priority (6-month PDUFA goal; PDUFA) 6 months (if granted); $0 – – Wu 05/24 [52]: Localized LLM for FDA labeling Q&A enabling rapid query responses from FDA database Responses to FDA information requests (2-4 requests; 21 CFR 314) During review; $300K-$800K -1 to -2 months -$120K to -$300K – Advisory Committee preparation (if required; 21 CFR 314) 2-3 mos. notice, in review; $200K-$500K – – – FDA meetings & teleconferences (Type B, Type C; 21 CFR 312) In review; $100K-$300K – – – REMS development & implementation (if required; 21 CFR 208) 3-6 months, par. w/ late review; $500K-$2M – – – Labeling negotiations (finalize package insert; 21 CFR 314) Through late review; Included above – – – Manufacturing facility inspection (FDA pre-approval; 21 CFR 314) FDA schedules; From above – – – Post-marketing commitment study planning (21 CFR 314) Parallel w/ review; $200K-$500K – – Table 12: STD NDA/BLA = 18-22 mos.; $4.7M-$11.2M. LLM = -8.2 to -11 mos.; -$680K to -$1.4M PRIORITY NDA/BLA = 14-16 mos.; $4.7M-$11.2M. LLM = -8.2 to -11 mos.; -$680K to -$1.4M 10.5281/zenodo.17451709 22 Version: October 26, 2025 Regulatory Standards & Requirements FDA 21 CFR Part 314 (Applications for FDA Approval to Market a New Drug) • Regulates: NDA content, format, and submission • Requirements: §314.50 (content/format), §314.80 (postmarketing AE reporting), §314.101 (filing), §314.125 (refusal to approve) FDA 21 CFR Part 601 (Licensing of Biological Products) • Regulates: BLA requirements for biologics • Requirements: Similar to NDA plus manufacturing facility inspection, lot release ICH M4 - Common Technical Document (CTD) • Regulates: Structure and format of submissions • Requirements: Five-module structure (Administrative, Summaries, Quality, Nonclinical, Clinical) ICH E3 - Structure and Content of Clinical Study Reports • Regulates: CSR content for submission • Requirements: Comprehensive 16-section format, individual patient listings ICH M3(R2) - Nonclinical Safety Studies • Regulates: Nonclinical study requirements for Module 4 • Requirements: Toxicology, safety pharmacology study reports CDISC SDTM/ADaM • Regulates: Standardized datasets for Module 5 • Requirements: SDTM v3.2+, ADaM v2.1+, define.xml FDA Guidance on Providing Regulatory Submissions in Electronic Format (eCTD) • Regulates: Electronic submission format • Requirements: eCTD backbone files, folder structure, validation Prescription Drug User Fee Act (PDUFA) • Regulates: FDA review timelines and fees • Requirements: Standard review (10-month goal), priority review (6-month goal), user fee ($3.18M in 2025) FDA 21 CFR Part 312 (Investigational New Drug Regulations) • Regulates: Formal meetings between FDA and sponsors • Requirements: Pre-NDA meeting, mid-cycle communication FDA 21 CFR Part 208 (Medication Guides) • Regulates: REMS requirements • Requirements: Risk mitigation strategies if significant safety concerns FDA Guidance on Formal Meetings Between FDA and Sponsors (2017) • Regulates: FDA-sponsor interaction • Requirements: Meeting procedures, Advisory Committee protocols FDA Guidance on Postmarket Requirements and Commitments • Regulates: Post-approval obligations • Requirements: REMS, post-marketing surveillance, pediatric studies (PREA) ICH E6(R3) - Good Clinical Practice (GCP) • Regulates: Quality standards for clinical data in submission • Requirements: Data integrity, monitoring documentation 10.5281/zenodo.17451709 23 Version: October 26, 2025 Summary: Complete Development Program Full Program Timeline & Cost Summary Trial Phase/ Current Stage Duration Trial Event Overall Cost Range Top LLM Method(s) LLM ∆Time LLM ∆Cost Discovery & Preclinical Research 45-72 months (3.75-6 years) $55M-$220M Multi-agent discovery, CancerGPT, Tx-LLM, AI co-scientist, automated synthesis -15 to -28 months -$23M to -$67M Phase I - Safety & Dose Escalation 24-35 months (2-2.9 years). [18-26m without expansion] $7.6M-$27.5M. [$5.6M-$18.7M without expansion] GPT-4 protocol authoring, Sanofi QSP digital twins, GenAI CSR generation -5.5 to -8 months -$480K to -$1.55M Phase II - Proof of Concept 36-48 months (3-4 years) $25M-$53M TrialGPT screening, mCODE standardization, GenAI CSR automation -11 to -17 months - $4.32M to -$8.65M Phase III - Complete Program 60-72 months (5-6 years) $114M-$315M Multi-LLM orchestration, TrialGPT, automated SDV, RAG-LLM SAP/TFL, AI pharmacovigilance, health economics automation -78 to -123 months - $32.76M to - $73.35M NDA/BLA - Submission & Review 18-22 months (standard). [14-16m priority] $4.7M-$11.2M GenAI ISS/ISE, CSR Atlas, modular SDTM, FDA labeling LLM -8.2 to -11 months -$680K to -$1.4M Cross-Phase Management Concurrent throughout program $24.5M-$57.8M Multi-LLM quality control, automated audits, regulatory intelligence -37 to -61 months -$4.9M to - $14.25M Cross-Phase MIDD Concurrent at multiple phases $1.4M-$2.6M AI-enabled QSP, digital twins, automated VVUQ, pre-validated models -8.3 to -11.5 months -$476K to -$945K TOTAL CLINICAL DEVELOPMENT (Current Baseline) ~120-180 months (10-15 years) $282M-$837M – – – TOTAL WITH LLM/AI OPTIMIZATION (Projected) Target: 58-92 months (4.8-7.7 years) Target: $215M-$673M Multi-LLM orchestration across all phases -62 to -88 months (5273%) -$67M to -$164M (2450%) Table 13: Full optimization: 52-73% time reduction (5-7 years), 24-50% cost savings ($67M-$164M) 10.5281/zenodo.17451709 24 Version: October 26, 2025 Phase III Detailed Component Summary Phase III Component Duration Trial Event Overall Cost Range Top LLM Method(s) LLM ∆Time LLM ∆Cost 1. Protocol Design & Regulatory 12-17 months $670K-$1.79M GPT-4 protocol authoring, BACTA-GPT adaptive design, RAG-LLM SAP -3.5 to -5 months -$120K to -$230K 2. Patient Recruitment & Sites 38-44 months $27.5M-$59.6M TrialGPT, GPT-4 screening, mCODE matching, RECTIFIER, Paradigm -18 to -31 months - $15.2M to -$30.6M 3. Data Management & Monitoring 50-56 months (4.2-4.7 years) $19.4M-$21.4M mCODE transformation, automated SDV, digital twin logs, quad AI peer review -37.5 to -57 months -$7.1M to -$11.1M 4. Statistical Analysis & CDISC 12-15 months $1.72M-$3.1M RAG-LLM SAP/TFL, Yang TFL generation, sdtm.oak, metadata-driven automation -4 to -5.5 months -$232K to -$440K 5. Safety PV & Clinical Reporting Safety: 49-53m concurrent. CSR: 8-10m post-analysis $5.84M-$6.81M CIOMS/Pfizer AI PV, GenAI CSR (QInscribe/AlphaLife/Narrativa), plain-language summaries -27 to -41.7 months -$1.7M to -$3.19M 6. Health Economics & Market Access 15-20 months $1.17M-$2.12M Reason GPT-4 economic modeling, Kawchak o3-mini cost solutions, RAG guideline interpretation -8 to -11.5 months -$213K to -$420K 7. Cross-Phase Management (Phase III) 62 months (5.2 years) $11.5M-$16.8M Multi-LLM orchestration, automated QC, AI regulatory intelligence -20 to -32 months -$1.3M to -$2.35M 8. MIDD (Phase III portion) 14-20 months $695K-$1.42M Sanofi/Certara QSP, Stahlberg/Vidovszky digital twins, automated VVUQ -8.3 to -11.5 months -$216K to -$465K PHASE III TOTAL (Current Baseline) 60-72 months (5-6 years) $114M-$315M – – – PHASE III WITH LLM OPTIMIZATION (Projected) Target: 25-36 months (2.1-3 years) Target: $81.24M-$241.65M Multi-LLM pipeline spanning all components -78 to -123 months (6073%) - $32.76M to - $73.35M (2950%) Table 14: Phase III optimization: 60-73% time reduction, 29-50% cost savings 10.5281/zenodo.17451709 25 Version: October 26, 2025 Economic Impact: • Phase III Transformation: From 5-6 years and $114M-$315M to 2.1-3 years and $81.24M-$241.65M (60-73% timeline reduction, 29-50% cost savings) • Full Program Acceleration: From 10-15 years and $282M-$837M to 4.8-7.7 years and $215M-$673M (52-73% faster, 24-50% cheaper) •Per-Patient Economics: From $40K-$80K (site-based) to $0.11-$3.60 (virtual/hybrid) = 99.5-99.999% reduction •ROI: $735K-$2.7M AI investment yields $32.02M-$72.62M net savings = 4,355-2,691% ROI (44-27x return) • Market Acceleration Value: $500M-$2B opportunity from 3-4 years faster approval (additional years of market exclusivity) Regulatory Validation: • FDA Acceptance: Certara AI-enabled QSP supports many FDA novel drug approvals (2014-24), Sanofi received approval for Xenpozyme using virtual Phase 1b trials, FDA M15 (2024) draft guidance provides framework for in silico evidence • EMA Qualification: Vidovszky PROCOVA methodology EMA-qualified for prognostic baseline covariate generation enabling smaller control groups • Industry Adoption: QInscribe/AlphaLife/Narrativa chosen by 5 of top 10 pharma companies, Certara platform used by majority of industry, Paradigm deployed at multiple health systems achieving tripled participation Projected Overall Impact: • Timeline Reduction: Target 52-73% reduction (5-7 years faster) from discovery to approval, with Phase III acceleration of 60-73% (3-4 years faster) through parallel processing and automation • Cost Reduction: Target 24-50% reduction ($67M-$164M saved) across clinical development, with largest savings in patient recruitment ($15.2M-$30.6M), data management ($7.1M-$11.1M), and safety monitoring ($1.37M-$2.37M) • Quality Improvement: 92-96.66% accuracy in automated extraction/matching, duplicate detection in pharmacovigilance, up to 95% concordance in multi-LLM cross-validation, enabling 15-25 percentage point success rate improvement (50% historical to 65-75% target) through predictive analytics and early futility detection • Accessibility: 60-75% personnel reduction (30-50 conventional team to 8-15 with AI augmentation) democratizing drug development for smaller organizations, academic centers, and rare disease programs; virtual patient generation enabling trials previously considered infeasible Strategic Implications The integration of LLM/AI technologies into drug development represents not an incremental improvement but a fundamental paradigm shift: 1. From Serial to Parallel: Traditional development proceeds sequentially through phases with hand-offs between teams. LLM-enabled workflows enable massive parallelization: automated SDV during data collection, real-time statistical analysis during enrollment, simultaneous multi-arm platform trials, concurrent regulatory preparation during analysis. The result: 75-102% effective time reduction in data management through parallel processing, database lock in days vs. 3-6 months, and 60-73% overall Phase III acceleration. 2. From Retrospective to Prospective: Conventional trials analyze data after collection completes, detect safety signals after harm occurs, and identify futility after years of investment. AI enables real-time monitoring with quad AI peer review achieving up to 95% concordance, predictive analytics for early futility detection improving success rates 15-25 percentage points, and adaptive treatment recommendations optimizing patient outcomes cycle-by-cycle. The result: $50M+ prevented futile trials, 49-72% faster safety signal detection, and evidence-driven decisions in hours vs. weeks. 3. From Physical to Virtual: Traditional trials require physical sites, in-person visits, and geographical constraints limiting patient access. Digital twins and QSP models enable virtual patient generation (days vs. 18-24 months), hybrid decentralized designs reducing patient burden, and global accessibility independent of geography. The result: 99.5-99.999% per-patient cost reduction ($0.11-$3.60 vs. $40K-$80K), 47-70% enrollment acceleration, and democratized access for rare diseases and underserved populations. 4. From Manual to Automated: Conventional development relies on manual protocol writing, human screening, paper-based monitoring, and expert analysis creating bottlenecks at every step. LLMs enable automated protocol generation (40-60% faster), AI-powered patient-trial matching (42.6% time reduction, 87.3% accuracy), automated high SDV (37-52% cost savings), and GenAI CSR generation (90% time reduction: 50-100 hours to 5 hours). The result: 60-75% personnel reduction, $4.9M-$14.25M management savings, and 24-50% overall program cost reduction. 5. From Resource-Intensive to Accessible: Conventional trials require $114M-$315M Phase III budgets, 30-50 person teams, and 5-6 year timelines accessible only to large pharma. AI reduces Phase III to $81.24M-$241.65M with 8-15 person AI-augmented teams completing in 2.1-3 years, while $735K-$2.7M AI investment yields 4,355-2,691% ROI (44-27x return). The result: academic centers, biotech startups, and rare disease foundations can now conduct trials previously requiring $100M+ budgets and multinational teams. 10.5281/zenodo.17451709 32 Version: October 26, 2025 Call to Action This baseline workflow provides the foundation for rigorous, quantitative assessment of LLM/AI impact across all drug development activities. The evidence demonstrates: • Which LLM approaches deliver the greatest impact: Multi-LLM orchestration (Kawchak) achieving up to 95% concordance across vendors for bias mitigation, TrialGPT (Jin) delivering 42.6% time reduction with 87.3% accuracy in patient screening, RECTIFIER (Unlu) achieving 97% cost reduction ($0.11 vs. $34.75/patient), GenAI CSR generation (QInscribe, AlphaLife, Narrativa) reducing 50-100 hours to 5 hours (90% time reduction), CIOMS/Pfizer AI pharmacovigilance processing 1.4M AEs annually with high accuracy, and Sanofi/Certara QSP digital twins achieving multi-month trial duration reduction with many FDA approval support • Optimal multi-LLM orchestration strategies: Task-specific model selection (GPT-4 for protocol authoring, TrialGPT for screening, mCODE LLMs for standardization, RAG-LLM for SAP/TFL, GenAI for CSR, GPT-4 for economic modeling), cross-validation frameworks achieving up to 95% concordance (Kawchak quad AI peer review), bidirectional learning loops with sense-analyze-recommend-act-learn feedback (Kawchak digital twins), and transparent provenance tracking with automated audit trails (FDA 21 CFR Part 11 ALCOA+ compliance) • Regulatory pathways and validation frameworks: FDA M15 (2024) draft guidance providing framework for in silico evidence with ASME V&V 40-2018 verification/validation standards, Certara pre-validated QSP models supporting many of FDA novel drug approvals (2014-24), EMA-qualified PROCOVA methodology (Vidovszky) enabling control group size reduction, Sanofi Xenpozyme approval demonstrating regulatory acceptance of virtual Phase 1b trials, and comprehensive VVUQ frameworks (Kawchak) with 55 automated tests achieving verification scores >80% and validation R²>0.90 • Economic models and ROI calculations: Phase III reduction from $114M-$315M to $81.24M-$241.65M ($32.76M- $73.35M savings, 29-50% reduction) through $735K-$2.7M AI investment over 2.1-3 years yielding 4,355-2,691% ROI (44-27x return), full program acceleration from $282M-$837M to $215M-$673M ($67M-$164M savings, 24-50% reduction), per-patient economics transformation from $40K-$80K to $0.11-$3.60 (99.5-99.999% reduction), market acceleration value of $500M-$2B from 3-4 years faster approval (additional years of market exclusivity), and prevented futile trial savings of $50M+ through early AI-driven futility detection • Best practices and implementation roadmaps: Begin with high-impact, low-risk applications (literature synthesis, protocol authoring, patient screening), establish multi-LLM orchestration with cross-validation (up to 95% concordance requirement), implement comprehensive VVUQ at each step (verification >80%, validation R²>0.90), ensure regulatory compliance through transparent provenance tracking (FDA 21 CFR Part 11 ALCOA+ principles), scale to higher-risk applications (automated SDV, AI pharmacovigilance, virtual trials) with robust validation, and maintain human oversight for critical decisions while maximizing AI augmentation for repetitive/analytical tasks The question is no longer whether LLMs will transform drug development, but how quickly organizations can adopt these technologies to deliver life-saving therapies faster, cheaper, and more effectively to patients in need. The future of drug development is here. The baseline is established. The transformation begins now. END OF BASELINE WORKFLOW AND LLM PROJECTIONS This comprehensive baseline workflow with integrated LLM impact assessment demonstrates 52-73% timeline reduction (5-7 years faster) and 24-50% cost savings ($67M-$164M) through multi-LLM orchestration, intelligent automation, and AI-enabled quality management across all phases of drug development from discovery through FDA approval. With validated examples from 10 author studies and 25+ independent research groups, the evidence supports immediate adoption of LLM/AI technologies to transform conventional 10-15 year, $282M-$837M programs into accelerated 5-8 year, $215M-$673M pathways delivering life-saving glioblastoma therapies to patients 5-7 years sooner. 10.5281/zenodo.17451709 33 Version: October 26, 2025 Limitations and Future Work A limitation of the study was the human hours required to engineer many prompts that were part of a) building table templates, b) author work insertions, and c) online literature incorporations. Table templates required optimizations across table types, headings, LaTeX conversions, and accompanying instructions (Prompts A-C). Each of the author’s works were accompanied by Prompt C, and needed to be processed individually by Sonnet due to the size of the task and complexity of populating multiple data tables reliably shown in Output D. Online literature searches by Sonnet, ChatGPT5, and Grok were processed by SonnetR to reach a better consensus regarding accuracy vs. execution in a faster single search. Where performance was critical, additional time was required for the triple LLM workflow consensus, quintuple LLM peer review, and quad meta-analyses due to the number of experiments, dataset preparations, and manuscript versioning for each LLM recommendation. Additional formatting time was required by the author for LaTeX tables, but the full project was completed on day 26 of the 30 day goal. Another limitation of the study was the inference-time compute and minor errors associated with the most performant Sonnet and SonnetR models. The main limitation was the amount of information that could be generated by Sonnet for each prompt, as both content length and complexity for the work became very high. This was apparent in Output J, which populated table information properly, but many of the regulatory standards & requirements were truncated and substituted with phrase "[Previous regulatory text remains unchanged]". Additional truncation was experienced for later sentences of the "Description" paragraphs. Both issues were addressed by the author, with re-incorporation of relevant elements from the input, Output I7. The model’s thoughts were also a precaution for truncation by beginning it’s response with "This is a very large and detailed request." The other Sonnet error was the use of ">" instead of "}" in a few LaTeX tables, which was recognized and corrected by the author. Peer reviews using 5 LLMs of populated tables yielded 37 recommended fixes, targeting 64 instances, with 12 author accepted fixes and 25 author rejections, as represented by the GPT5h meta prompt and meta-analysis (Output O LLM B). Sonnet returned 21 of the recommendations, while the other four LLMs were limited to a total of 16 recommendations. The top limitations in this 32.4% acceptance rate was primarily due to the recommendations already being corrected for 6 occurrences (as verified in literature by the author), the fix was previously addressed by one of the other five LLMs for 6 occurrences, and the recommendation was unverifiable or not in the provided source for 5 occurrences. The most common fixes were the Arc Institute Evo 2 model parameter count from 47B to 40B, as recommended by ChatGPT5, Gemini, and Sonnet; and also the recommendation to remove a Sanofi QSP modeling 6-month timeline reduction claim, as recommended by ChatGPT5. Future work will likely focus on collaborating with pharma and regulators to better understand and apply LLM efficiency to "current clinical trial activity" columns throughout the study. Additional work based on data privacy standards is also required to identify which tasks can be addressed directly with i) Online LLMs (ie. those used in this study), and which tasks can only be executed with local patient data: ii) Online LLM-generated code executed locally or iii) Local LLMs (privacy advantages, but typically less performant). Lastly, the author will need to develop more rigorous verifications for LLM outputs and LLM-generated code for production use. Conclusions Baseline workflows for oncology trials can last over 15 years and cost over $800M, with failed trials pushing this amount well beyond the $1B amount. This process from drug discovery through NDA/BLA can be improved with the use of leading LLMs that are now capable of not only high English contextual awareness, but also quantitative data and code generation for local processing, leading to the AI projection that "Conventional trials require $114M-$315M Phase III budgets, 30-50 person teams, and 5-6 year timelines accessible only to large pharma. AI reduces Phase III to $81.24M-$241.65M with 8-15 person AI-augmented teams completing in 2.1-3 years, while $735K-$2.7M AI investment yields 4,355-2,691% ROI (44-27x return)." The strength of this study was Claude Sonnet 4.5 Extended’s versatility in both understanding complex user prompts and possessing the willingness to provide quality outputs at scale, which was less apparent in other models. The following tasks were performed at high proficiency: table templates, author and online work table populations, and consensuses based on several LLM responses. In addition, current trial activities time and cost, and potential LLM time and cost savings were all performed proficiently by AI. Sonnet followed instructions well regarding ICH/FDA abbreviation assignments in tables that corresponded to comprehensive regulation and requirement explanations. A full program timeline & cost summary, key performance indicators of baseline workflow vs LLM, overall program economics, and call to action concluded the AI study. For the multitude of prompts submitted, few token substitution errors were experienced; with truncation of outputs in rare cases - typically occurring where complexity and dataset size was sufficiently high. Many organizations in 2024-2025 have produced LLM applications that are similar in function to current end-to-end clinical trial workflows. For instance, Arc Institute, Stanford University, and NVIDIA collaborated on the Evo 2 model for rational enzyme design[ 6 ], while Sanofi has targeted Phase 1b trials for dose optimization using a biology foundation model[ 7 ]. A host of LLM applications have been identified by AI that will likely impact Phase III trial activities such as QSP and Digital Twin simulations by Sanofi and Kawchak [ 7 , 10 , 11 ], SAP finalization by Kuo [ 12 ], and SDTM programming and validation by Pharmaverse [ 13 ]. Additionally NDA/BLA activities such as PRISMA compliant reviews expedited by Kawchak literature summarizations [ 24 ]. In summary, AI found that "This comprehensive baseline workflow with integrated LLM impact assessment demonstrates 52-73% timeline reduction (5-7 years faster) and 24-50% cost savings ($67M-$164M) through multi-LLM orchestration, intelligent automation, and AI-enabled quality management across all phases of drug development from discovery through FDA approval." 10.5281/zenodo.17451709 34 Version: October 26, 2025 Data availability: Zenodo [53] Prompt A: Table Templates, 2 Files Prompt B: Table Optimization, 2 Files a) Prompt B2: Table Instructions b) Prompt B3: Trials & Standards c) Prompt B4: Concise Output Prompt C: 14 Working Tables, 2 Files Prompt D: Author Tables Population, 26 Files a) Output D: 18 Author Table Sets of 14 Prompt E: Online Tables Population, 4 Files a) Output E: Online Literate 14 Tables Prompt F: LaTeX Table Conversions, 2 Files a) Prompt F2: Included BibTeX Entries Prompt G: Online Tables Optimization, 4 Files a) Sonnet, ChatGPT5, Grok Triple Online Consensus b) Output G1: Sonnet Optimized Online Tables c) Prompt G2: Updated BibTeX Entries Prompt H: Table Optimizations, 4 Files a) Prompt H2: Table Dimensions b) Prompt H3: Synchronous Citation, BibTeX Prompt I: Comprehensive Tables Consensus, 8 Files a) Gemini, ChatGPT5, Sonnet Triple Workflow Consensus b) Output I1: 14 tables with Regulatory Standards c) Prompt I2: Corrections Value Summations d) Prompt I3: Corrections Timeline Totals e) Prompt I4: Discovery/Preclinical/Trial IDs f) Prompt I5: New Baseline From Table Examples g) Prompt I6: Impact, LLM Method Designations h) Prompt I7: Baseline Workflow Formatting Fixes Prompt J: Baseline Workflow Population, 5 Files a) Kawchak and Online Source Tables Insertion Prompt K: Alternative Baseline Workflow, 2 Files a) Featuring Select Mathematical Calculations Prompt L: Quintuple LLM Peer Review, 16 Files a) ChatGPT5, ChatGPTo3, Grok Expert, Gemini, Sonnet b) Author Verification of Each LLM Output c) Author Versioning for Each LLM Recommendation Prompt M: Meta Prompts for Peer Review Analysis, 6 Files a) Gemini and GPT5h Used to Generate Informed Prompts Prompt N: Dual LLM Meta-Analyses of Prompt M, 7 Files a) Gemini Prompt Used for Two Meta-Analyses b) Output N LLM A: Gemini Meta-Analysis c) Output N LLM B: GPT5h Meta-Analysis Prompt O: Dual LLM Meta-Analyses of Prompt M, Shared a) GPT5h Prompt Used for Two Meta-Analyses b) Output O LLM A: Gemini Meta-Analysis c) Output O LLM B: GPT5h Meta-Analysis 10.5281/zenodo.17451709 35 Version: October 26, 2025 References [1] Anthropic. Introducing claude sonnet 4.5: Claude sonnet 4.5 is the best coding model in the world., September 2025. URL: https://www.anthropic.com/news/claude-sonnet-4-5. [2] Google AI Studio. What will you build? push gemini to the limits of what ai can do, powered by the gemini api. this experimental model is for feedback and testing only. no production use., 2025. URL: https://aistudio.google.com/prompts/new_chat. [3] OpenAI. Introducing gpt-5. our smartest, fastest, most useful model yet, with built-in thinking that puts expert-level intelligence in everyone’s hands., August 2025. URL: https://openai.com/index/introducing-gpt-5/. [4] xAI. Grok 4 fast pushing the frontier of cost-efficient intelligence., 2025. URL: https://x.ai/news/grok-4-fast. [5] Tianhao Li, Sandesh Shetty, Advaith Kamath, Ajay Jaiswal, Xiaoqian Jiang, Ying Ding, and Yejin Kim. Cancergpt for few shot drug pair synergy prediction using large pretrained language models. npj Digital Medicine, 7(1):1–10, February 2024. URL: https://www.nature.com/articles/s41746-024-01024-9,doi:10.1038/s41746-024-01024-9. [6] Garyk Brixi, Matthew G. Durrant, Jerome Ku, Michael Poli, Greg Brockman, Daniel Chang, Gabriel A. Gonzalez, Samuel H. King, David B. Li, Aditi T. Merchant, Mohsen Naghipourfar, Eric Nguyen, Chiara Ricci-Tam, David W. Romero, Gwanggyu Sun, Ali Taghibakshi, Anton Vorontsov, Brandon Yang, Myra Deng, Liv Gorton, Nam Nguyen, Nicholas K. Wang, Etowah Adams, Stephen A. Baccus, Steven Dillmann, Stefano Ermon, Daniel Guo, Rajesh Ilango, Ken Janik, Amy X. Lu, Reshma Mehta, Mohammad R.K. Mofrad, Madelena Y. Ng, Jaspreet Pannu, Christopher Ré, Jonathan C. Schmok, John St. John, Jeremy Sullivan, Kevin Zhu, Greg Zynda, Daniel Balsam, Patrick Collison, Anthony B. Costa, Tina Hernandez-Boussard, Eric Ho, Ming-Yu Liu, Thomas McGrath, Kimberly Powell, Dave P. Burke, Hani Goodarzi, Patrick D. Hsu, and Brian L. Hie. Genome modeling and design across all domains of life with evo 2. bioRxiv, February 2025. URL: http://biorxiv.org/lookup/doi/10.1101/2025.02.18.638918,doi:10.1101/2025.02.18.638918. [7] Sanofi. Digital twinning: Clinical trials powered by ai. Sanofi Magazine, 2024. URL: https://www.sanofi.com/en/magazine/our-science/digital-twinning-clinical-trials-ai. [8] Jacob Beattie, Dylan Owens, Ann Marie Navar, Luiza Giuliani Schmitt, Kimberly Taing, Sarah Neufeld, Daniel Yang, Christian Chukwuma, Ahmed Gul, Dong Soo Lee, Neil Desai, Dominic Moon, Jing Wang, Steve Jiang, and Michael Dohopolski. Large language model augmented clinical trial screening. medRxiv, August 2024. URL: https://www.medrxiv.org/content/10.1101/2024.08.27.24312646v1, doi:10.1101/2024.08.27.24312646. [9] Sree Mandal Shekhar, Daniel Lee, Hwanseok Yoon, et al. Novel development of llm driven mcode data model for improved clinical trial matching to enable standardization and interoperability in oncology research. arXiv preprint, October 2024. URL: https://arxiv.org/abs/2410.19826,arXiv:2410.19826. [10] Kevin Kawchak. Qsp metastatic pancreatic cancer ai clinical trial simulation from protocol to prediction: Code, vvuq, and playbook, August 2025. doi:10.5281/zenodo.17001137. [11] Kevin Kawchak. Accelerating fda compliance and cost efficiency of in silico clinical trials via ai digital twin pancreatic cancer simulation, September 2025. doi:10.5281/zenodo.17239510. [12] Sung-Min Kuo, Shin-Kai Tai, Hong-Yu Lin, and Ray-Chi Chen. Automated clinical trial data analysis and report generation by integrating retrieval-augmented generation (rag) and large language model (llm) technologies. Machine Learning and Knowledge Extraction, 6(8):188, August 2025. URL: https://www.mdpi.com/2673-2688/6/8/188, doi:10.3390/ai6080188. [13] Pharmaverse Consortium. sdtm.oak: An edc and data standard agnostic sdtm data transformation engine. CRAN Package, October 2024. URL: https://github.com/pharmaverse/sdtm.oak. [14] Pfizer Global Product Development. Innovation in pharmacovigilance: Use of artificial intelligence in adverse event case processing. Pfizer News, 2018. URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC6590385/. [15] CIOMS Working Group XIV. Artificial intelligence in pharmacovigilance: Draft report for public consultation. CIOMS, May 2025. URL: https://cioms.ch/wp-content/uploads/2022/05/CIOMS-WG-XIV_ Draft-report-for-Public-Consultation_1May2025.pdf. [16] Thomas Reason, William Rawlinson, James Langham, Andrew Gimblett, Beth Malcolm, and Stevie Klijn. Artificial intelligence to automate health economic modelling: A case study to evaluate the potential application of large language models. PharmacoEconomics Open, 8(2):191–203, February 2024. URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC10884386/,doi:10.1007/s41669-024-00477-8. [17] Narrativa. Automation of clinical study reports with csr atlas. PubMed Central, 2025. URL: https://www.narrativa.com/automation-of-clinical-study-reports/. [18] OpenAI. Gpt-5 is here: Our smartest, fastest, and most useful model yet, with thinking built in., August 2025. URL: https://openai.com/gpt-5/. [19] OpenAI. Introducing openai o3 and o4-mini, 2025. URL: https://openai.com/index/introducing-o3-and-o4-mini/. 10.5281/zenodo.17451709 36 Version: October 26, 2025 [20] Juan Manuel Zambrano Chaves, Eric Wang, Tao Tu, Eeshit Dhaval Vaishnav, Byron Lee, S. Sara Mahdavi, Christopher Semturs, David Fleet, Vivek Natarajan, and Shekoofeh Azizi. Tx-llm: A large language model for therapeutics. arXiv, (arXiv:2406.06316), June 2024. arXiv:2406.06316. URL: http://arxiv.org/abs/2406.06316, doi:10.48550/arXiv.2406.06316. [21] Gleb Vitalevich Solovev, Alina Borisovna Zhidkovskaya, Anastasia Orlova, Anastasia Vepreva, Tonkii Ilya, Rodion Golovinskii, Nina Gubina, Denis Chistiakov, Timur A. Aliev, Ivan Poddiakov, Galina Zubkova, Ekaterina V. Skorb, Vladimir Vinogradov, Nikolay Nikitin, Andrei Dmitrenko, Anna Kalyuzhnaya, and Andrey Savchenko. Towards llm-driven multi-agent pipeline for drug discovery: Neurodegenerative diseases case study. In OpenReview, December 2024. URL: https://openreview.net/forum?id=3ncjySu5ro. [22] Kevin Kawchak. Cancer vs. conversational artificial intelligence. bioRxiv, December 2024. doi:10.1101/2024.12.28.630597. [23] Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, Fan Zhang, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Pushmeet Kohli, Yossi Matias, Andrew Carroll, Kavita Kulkarni, Nenad Tomasev, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R. D. Costa, José R. Penadés, Gary Peltz, Yunhan Xu, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan. Towards an ai co-scientist. arXiv, (arXiv:2502.18864), February 2025. arXiv:2502.18864. URL: http://arxiv.org/abs/2502.18864, doi:10.48550/arXiv.2502.18864. [24] Kevin Kawchak. Clinical decision support based on bevacizumab cancer trials and pushing the limitations of advanced llms, January 2025. doi:10.5281/zenodo.14968162. [25] Kevin Kawchak. Ai revolution toward the cure of lung adenocarcinoma, April 2025. doi:10.5281/zenodo.15278152. [26] Kevin Kawchak. Paclitaxel biosynthesis ai breakthrough. ChemRxiv, October 2024. doi:10.26434/chemrxiv-2024-pqjd3. [27] AlphaLife Clinical. Ai medical writing for csr - auroraprime platform. Microsoft AppSource, February 2025. URL: https: //appsource.microsoft.com/en-us/product/web-apps/alphalife_clinical_saas.create_csr_pilot. [28] Mohammadreza Maleki and Seyed Ghahari. Clinical trials protocol authoring using llms. arXiv preprint, August 2024. URL: https://arxiv.org/abs/2404.05044,arXiv:2404.05044. [29] Quanticate. New medical writing service qinscribe reduces the time to generate a draft clinical study report (csr) by 90% with generative ai solution. PR Newswire, September 2025. URL: https://tinyurl.com/wukwr2xe. [30] Qiao Jin, Zifeng Wang, Charalampos S. Floudas, et al. Matching patients to clinical trials with large language models. Nature Communications, 15:9074, November 2024. URL: https://www.nature.com/articles/s41467-024-53081-z,doi:10.1038/s41467-024-53081-z. [31] Han Suk Choi, Ji Young Song, Kyung Hwan Shin, Ji Hyun Chang, and Bum-Sup Jang. Developing prompts from large language model for extracting clinical information from pathology and ultrasound reports in breast cancer. Radiation Oncology Journal, 41(3):209–216, September 2023. URL: https://pubmed.ncbi.nlm.nih.gov/37793630/, doi:10.3857/roj.2023.00633. [32] Padmanabhan Krishna and Danny Baker. Bacta-gpt: An ai-based bayesian adaptive clinical trial architect. arXiv preprint, July 2025. URL: https://arxiv.org/abs/2507.02130,arXiv:2507.02130. [33] Onur Unlu, Jongkyu Shin, Christine J. Mailly, et al. Retrieval-augmented generation–enabled gpt-4 for clinical trial screening. NEJM AI, 1(7), June 2024. URL: https://ai.nejm.org/doi/full/10.1056/AIoa2400181, doi:10.1056/AIoa2400181. [34] Paradigm Health. Using ai to improve patient access to clinical trials. OpenAI Case Study, 2023. URL: https://openai.com/index/paradigm/. [35] Michael Wornow, Alejandro Lozano, Dev Dash, Jenelle Jindal, Kenneth W. Mahaffey, and Nigam H. Shah. Zero-shot clinical trial patient matching with llms. NEJM AI, 2(1), January 2025. URL: https://ai.nejm.org/doi/10.1056/AIcs2400360,doi:10.1056/AIcs2400360. [36] Shivam Gupta et al. Prism: Patient records interpretation for semantic clinical trial matching system using large language models. npj Digital Medicine, 7:274, October 2024. URL: https://www.nature.com/articles/s41746-024-01274-7,doi:10.1038/s41746-024-01274-7. [37] Aditya Devi, Sakshi Uttrani, Arjun Singla, et al. Automating clinical trial eligibility screening: Quantitative analysis of gpt models versus human expertise. Proceedings of PETRA, June 2024. URL: https://dl.acm.org/doi/10.1145/3652037.3663922,doi:10.1145/3652037.3663922. [38] Georgios Peikos, Symeon Symeonidis, Prokopis Kasela, and Gabriella Pasi. Utilizing chatgpt to enhance clinical trial enrollment. arXiv preprint, June 2023. URL: https://arxiv.org/abs/2306.02077,arXiv:2306.02077. [39] Michelle Gao, Aakrati Varshney, Sophia Chen, et al. The use of large language models to enhance cancer clinical trial educational materials. JNCI Cancer Spectrum, 9(2):pkaf021, April 2025. URL: https://academic.oup.com/jncics/article/9/2/pkaf021/8005863,doi:10.1093/jncics/pkaf021. 10.5281/zenodo.17451709 37 Version: October 26, 2025 [40] Kevin Kawchak. Chatgpt 100,000 patient 24-month in silico phase iii 5-arm pancreatic cancer clinical trial triplicate, July 2025. doi:10.5281/zenodo.16415815. [41] Kevin Kawchak. 10 year glioblastoma clinical trial meta-analyses by autonomous ai at scale. survival, hr, ae, and rob scored in ai reports and charts, including verifications, May 2025. doi:10.5281/zenodo.15549831. [42] Yu Yang, Peter Krusche, Kimberly Pantoja, et al. Using large language models to generate clinical trial tables and figures. WUSS Conference Proceedings, September 2024. URL: https://www.researchgate.net/publication/ 384114978_Using_Large_Language_Models_to_Generate_Clinical_Trial_Tables_and_Figures. [43] Ravikumar Komandur, Jon McDunn, Nikita Nair, Babacar Fall, Adam P. Dicker, and Sean Khozin. Artificial intelligence in biomedical data analysis: A comparative assessment of large language models for automated clinical trial interpretation and statistical evaluation. medRxiv, February 2025. URL: https://www.medrxiv.org/content/10.1101/2025.02.05.25321607v2, doi:10.1101/2025.02.05.25321607. [44] Satoshi Tomioka. sdtm_mapper: Ai sdtm mapping using r for ml and python/tensorflow for dl. GitHub Repository, February 2019. URL: https://github.com/stomioka/sdtm_mapper. [45] Pharmaverse. Path to automation: Sdtm specification and code generation. Pharmaverse Documentation, 2024. URL: https://pharmaverse.github.io/sdtm.oak/articles/study_sdtm_spec.html. [46] Kevin Kawchak. Cost containment of global monoclonal antibody drugs and cancer clinical trials via llm focused reasoning, February 2025. doi:10.5281/zenodo.14968404. [47] Daniel Ferber, Isabella C. Wiest, et al. Gpt-4 for information retrieval and comparison of medical oncology guidelines. NEJM AI, May 2024. URL: https://ai.nejm.org/doi/full/10.1056/AIcs2300235, doi:10.1056/AIcs2300235. [48] Certara. Quantitative systems pharmacology software platform. Certara Product Documentation, 2025. URL: https://www.certara.com/software/quantitative-systems-pharmacology/. [49] Eric A. Stahlberg et al. Exploring approaches for predictive cancer patient digital twins: Opportunities for collaboration and innovation. Frontiers in Digital Health, 4, October 2022. URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC9586248/,doi:10.3389/fdgth.2022.1007784. [50] Andrew A. Vidovszky, Charles K. Fisher, Artur D. Loukianov, et al. Increasing acceptance of ai-generated digital twins through clinical trial applications. Frontiers in Digital Health, July 2024. URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC11263130/,doi:10.3389/fdgth.2024.1415459. [51] Kevin Kawchak. End-to-end pancreatic ductal adenocarcinoma digital twin clinical trial proposals, June 2025. doi:10.5281/zenodo.15735068. [52] Leihong Wu, Joshua Xu, Shraddha Thakkar, Magnus Gray, Yanyan Qu, Dongying Li, and Weida Tong. A framework enabling llms into regulatory environment for transparency and trustworthiness and its application to drug labeling document. Regulatory Toxicology and Pharmacology, 149:105613, May 2024. URL: https://linkinghub.elsevier.com/retrieve/pii/S0273230024000540, doi:10.1016/j.yrtph.2024.105613. [53] Kevin Kawchak. End-to-end oncology clinical trial llm efficiency for industry adoption with fda/ich regulations, October 2025. doi:10.5281/zenodo.17451709. Acknowledgements The author would like to acknowledge Anthropic for providing access to Claude, Google for providing access to Gemini, OpenAI for providing access to ChatGPT and GPT, and xAI for providing access to Grok. Ethical disclosures The author of the article declares no competing interests. Rights and permissions This article is distributed under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author(s) and source are properly credited, a link to the Creative Commons license is provided, and any modifications made are indicated. To view a copy of this license, visit https://creativecommons.org/licenses/by/4.0/. About this study Kawchak K. End-to-End Oncology Clinical Trial LLM Efficiency For Industry Adoption with FDA/ICH Regulations. Zenodo. 2025; 10.5281/zenodo.17451709 [53]. 10.5281/zenodo.17451709 38 Version: October 26, 2025