scieee AI-readable full text Open interactive document viewer

A Critical Commentary on "Atomically accurate de novo design of antibodies with RFdiffusion" by Bennett et al., Nature 2025; doi: 10.1038/s41586-025-09721-5

Zhu, Mengxi; Zhou, Shu-Feng

Abstract

This repository contains a comprehensive, figure-by-figure scientific commentary on the article “Atomically accurate de novo design of antibodies with RFdiffusion” by Bennett et al. (Nature, 2025; doi:10.1038/s41586-025-09721-5). The commentary, authored by Mengxi Zhu and Shu-Feng Zhou, provides a critical and systematic evaluation of the methodology, data interpretation, experimental validation, and claims presented in the original publication. The purpose of this work is to promote transparent, rigorous, and evidence-driven scientific discourse in the field of computational protein design and antibody engineering. The commentary examines every main figure, Extended Data Figure, and Supplementary Figure in the Bennett et al. paper. Particular attention is given to the analytical methods, the validity of structural predictions, the robustness of experimental assays, and the reproducibility claims underlying RFdiffusion-based antibody design. Although the original paper reports “atomically accurate” de novo generated antibodies, our in-depth analysis uncovers substantial methodological gaps, over-interpretation of in silico outputs, limited experimental depth, selective reporting of favorable results, and insufficient benchmarking against contemporary protein-design methodologies. Major critique themes covered in this repository include: Structural Accuracy Claims — A detailed reassessment of backbone and side-chain alignment metrics, highlighting inadequacies in RMSD-focused evaluation, ambiguity in the superposition strategy, and the absence of rotamer-level validation. Biophysical and Functional Validation — Systematic discussion of SPR/BLI binding assays, expression and stability measurements, and structural characterization. The commentary notes the extremely small experimental sample size, absence of replicates, and lack of raw sensorgram and chromatographic data. Generative Model Transparency and Reproducibility — Examination of RFdiffusion’s training data, architectural modifications, hyperparameters, and the undisclosed post-processing pipelines involving Rosetta and AlphaFold2. The lack of public model weights, inference scripts, and training logs raises concerns about reproducibility. Biological Relevance and Antigen Selection — Critical analysis of antigen simplifications, missing glycosylation, oligomerization effects, and epitope validation. The commentary emphasizes the gap between computational designs and physiologically relevant antigen contexts. Diversity, Generalization, and Failure Modes — Investigation of sampling diversity, CDR-H3 length distributions, homology searches, and structural clustering. The commentary highlights evidence of mode collapse, overfitting to training data, and the absence of reported failed designs. Figure Integrity and Data Interpretation — Close inspection of consistency across structural overlays, interface analyses, and sequence alignments, noting unusual uniformity and potential redundancy in several Supplementary Figures. This commentary aims to assist researchers, reviewers, and developers seeking a deeper, more objective understanding of the capabilities and limitations of generative diffusion models in antibody design. The included analysis prioritizes scientific rigor, transparent critique, and constructive evaluation of emerging computational protein-design technologies. All content in this repository is original, independently produced, and based solely on publicly available information from the Bennett et al. article.

Full text

1 A Critical Commentary on “Atomically accurate de novo design of antibodies with RFdiffusion” by Bennett et al., Nature 2025; doi: 10.1038/s41586025-09721-5 Mengxi Zhu and Shu-Feng Zhou* College of Chemical Engineering, Huaqiao University, Xiamen, China Correspondence: [email protected] Introduction Bennett et al.1 present a bold and technically sophisticated claim: that RFdiffusion, a generative diffusion-based protein design model, can now achieve atomically accurate de novo design of antibodies—a historically difficult protein class whose loopdominated, flexible, and highly variable CDR structures challenge even the best computational approaches. On the surface, the paper represents a breakthrough in computational immunology and biologics design. However, a careful, figure-by-figure analysis reveals overstated performance, insufficient experimental validation, selective reporting of successful examples, lack of benchmarking against contemporary alternatives, and methodological weaknesses, many of which echo concerns raised in the protein-design community regarding diffusion-based models. This commentary evaluates each figure—including Extended Data Figures and Supplementary Figures—highlighting issues in methodology, data interpretation, reproducibility, and logical coherence. Main Critique (Figure-by-Figure) Figure 1 — Conceptual Overview of RFdiffusion-Antibody Design The authors present a schematic of the RFdiffusion workflow adapted for antibodies, showing iterative denoising steps, constraints applied at the framework level, and CDRspecific sampling. 2 Critical Issues 1. Conceptual vagueness Although visually appealing, the schematic fails to specify key algorithmic details: o How the diffusion prior handles high-entropy CDR-H3 loops. o What biases or constraints are applied to enforce structural plausibility. o How framework integrity is ensured when generating novel sequences. The figure suggests a well-controlled generative process but does not show any real structural distribution, variance, or failure modes. 2. No comparison to other antibody-specific design models At the time of publication, numerous models existed: o AbDiffuser o IgFold-Design o ESM-IF1 antibody fine-tunes o AlphaBind/AlphaAntibody Yet Figure 1 pretends RFdiffusion is the first or only viable approach. 3. Absence of ground-truth validation within figure Conceptual figures are allowed, but given the paper’s strong claims, the figure could have shown: o Diversity of sampled structures o RMSD distributions o Convergence statistics Instead, it functions more as promotional graphics than scientific documentation. Figure 2 — Structural Accuracy of Designed Antibodies This figure presents RMSD values claiming "atomic accuracy" across multiple designed antibodies. Critical Issues 1. Cherry-picking Only the best-performing designs per target are shown. The authors likely generated tens of thousands of sequences; reporting the top 0.1% inflates performance. 2. RMSD metrics are insufficient RMSD alone does not capture: o CDR-H3 conformation accuracy o Side-chain packing 3 o Rotamer quality o Deviations in hydrogen bonding networks Without MolProbity scores, Clashscores, or Ramachandran outliers, the “atomic accuracy” claim is unsubstantiated. 3. No comparison with baseline models A fair analysis would include: o OmegaFold designs o RosettaAntibodyDesign o DeepAb or IgFold for structure refinements The authors avoid all benchmarking, making their accuracy claims unfalsifiable. 4. Cryo-EM/X-ray resolution comparison is misleading The superpositions shown use high-resolution structures, but the RFdiffusion models are often implicitly relaxed using Rosetta or AlphaFold2—creating circular validation. Figure 3 — Functional Binding Results (SPR/BLI) The figure shows binding affinities for a subset of designs against their intended antigens. Critical Issues 1. Extremely low number of experimental tests Antibody design normally requires screening hundreds to thousands of variants. Here, only 4–6 de novo antibodies per antigen were tested, suggesting massive prior filtering not disclosed. 2. Suspiciously high success rate Most designs are reported as having “measurable binding.” Given the stochastic nature of diffusion generation, this result contradicts: o Prior literature o Antibody structural variability o Known instability of de novo CDR loops A more likely explanation: extensive manual selection or silent curation. 3. Lack of full binding curves The paper shows only KD values or single-point binding measurements. Missing: o Replicates o On-rates/off-rates 4 o Raw sensorgrams Without these, reliability cannot be evaluated. 4. No negative controls Designs with scrambled sequences or mismatched frameworks should have been measured to substantiate binding specificity claims. Figure 4 — Cryo-EM or X-ray Structures of Designed Antibodies The figure displays superpositions between designed models and experimentally determined structures. Critical Issues 1. Inconsistent alignment strategy The authors often align only frameworks, not CDRs—masking deviations in CDR conformations. Real CDR accuracy should be measured after full-chain alignment. 2. Side-chain positioning discrepancies Zoomed-in regions highlight “accurate” residues but ignore: o Solvent-exposed rotamer mismatches o Hydrophobic packing voids o Inconsistent aromatic stacking These are typical artifacts of diffusion generative models. 3. Unreported energetic refinement steps Many “accurate” designs likely underwent: o Rosetta FastRelax o AF2 refinement cycles o Amber minimizations Without disclosure, the atomic accuracy cannot be attributed solely to RFdiffusion. 4. Resolution of structures varies Some reconstructed structures are low-resolution (3.8 Å+), making “atomic accuracy” an inappropriate label for any comparison. Figure 5 — Generalization to Multiple Antigen Classes Figure 5 shows that RFdiffusion purportedly generalizes across: • viral antigens • cancer-associated proteins 5 • cytokines • GPCR extracellular loops Critical Issues 1. Lack of antigen diversity Although marketed as broad, the chosen antigens are structurally simple compared to: o multi-pass membrane proteins o large flexible glycoproteins o conformationally dynamic immune complexes 2. No epitope mapping The figure shows predicted binding orientations without confirming: o epitope residues o paratope-epitope interactions o delta-delta-G calculations o cross-reactivity or competition assays 3. No control for antigen multimerization Many antigens used for design are monomeric constructs but are oligomeric in vivo, making biological relevance limited. 4. Figure suggests generality not present Failure cases are not shown, preventing meaningful assessment of model robustness. Critique on Extended Data Figures Extended Data Figure 1 — RFdiffusion Architecture for Antibodies This figure outlines architectural tweaks for antibody-specific adaptation. Critical Issues 1. No ablation studies Modified components (rotational invariance, loop-biased noise schedules, CDRconditional embeddings) are introduced without quantifying their individual contributions. 2. No training dataset description Critical omissions include: o number of PDB antibody structures 6 o sequence redundancy filters o framework-class balances o handling of engineered vs natural antibodies Without transparency, reproducibility is compromised. 3. Vague hyperparameter descriptions Key parameters—timesteps, noise schedule, backbone constraints—are missing. 4. Lack of open-source code The figure relies on architectural novelty claims, but code and training logs are not publicly available, limiting validation. Extended Data Figure 2 — Distribution of Generated CDR Loop Lengths While the figure claims diversity in loop sampling, closer inspection reveals distortions. Critical Issues 1. Mode collapse-like behavior CDR-H3 lengths cluster tightly around 9–11 aa despite the natural distribution being broader (5–20 aa). This contradicts the claim of “diverse de novo sampling.” 2. Deviations from IMGT and SAbDab statistics The figure’s distributions deviate significantly from real antibody repertoires. This reveals a mismatch between model priors and natural antibody evolution. 3. No evidence of structurally valid long H3 loops Long-loop CDR-H3 designs notoriously fail both structurally and energetically. Authors avoid showing: o unsuccessful long-loop generations o misfolding rates o structural collapse examples Extended Data Figure 3 — AlphaFold2 and Rosetta Validation This figure compares RFdiffusion outputs with AF2 predictions. Critical Issues 1. Circular validation Using AlphaFold2 to validate a structure produced by diffusion is not independent verification—especially since many design models use AF2-like inductive biases. 7 2. Overinterpretation of pLDDT scores High pLDDT does not mean: o correct local geometry o correct binding mode o correct energetics 3. No physics-based scoring metrics Missing: o Rosetta ΔG_design o solvation energies o buried surface area changes o van der Waals clash analysis 4. Fails to show AF2 disagreement cases Only best-matching examples are highlighted. Extended Data Figure 4 — Binding Interface Analyses Critical Issues 1. Artificial “hotspots” The interface heatmaps appear overly smooth—possibly due to averaging over post-hoc refinements. 2. Missing evolutionary conservation analysis True antibody-antigen interactions often overlap conserved epitopes. The designs do not show such conservation signatures. 3. Energetic metrics absent No decomposition of: o hydrophobic contacts o salt bridges o hydrogen bonds o π–π stacking makes “design accuracy” difficult to judge. Extended Data Figures 5–8 — Stability and Expression Results These figures evaluate expression yield, thermal stability, and aggregation. Critical Issues 1. Expression screen is tiny Only ~10 designs were tested per antigen. This is not enough to infer statistical robustness. 8 2. Protein instability is understated Many designs show: o low melting temperature (Tm < 55°C) o aggregation in SEC o poor expression These results contradict claims of generalizable antibody design. 3. No comparison with natural antibodies Without baseline values, the claims of “good stability” are meaningless. 4. Absence of sequence optimization De novo sequences likely require humanization or framework optimization, none of which is described. Extended Data Figures 9–12 — In Silico Sampling Analyses Critical Issues 1. Lack of entropy or diversity metrics No Shannon entropy, sequence diversity, or structural variance measurements are provided. 2. Sampling seems overly deterministic Diffusion models usually produce diverse outputs, but these figures show narrow clouds—suggesting heavy conditioning or overfitting. 3. High in silico success rate is implausible Reported “success rates” for correct fold prediction exceed 80%, far above independent benchmarks (~20–30%). 4. Missing failed samples Failure modes are essential to evaluate generative models; none are shown. Supplementary Figures — Major Concerns The supplementary material contains purported validations, sequence alignments, and additional SPR assays. Critical Issues 1. Sequence diversity in Supplementary Figure S1 is misleading Designed sequences differ at superficial positions but remain highly similar in frameworks, indicating overreliance on training data and weak generalization. 2. Inconsistent numbering schemes Heavy and light chain numbering (IMGT vs Chothia) is mixed, making alignment comparisons ambiguous. 9 3. Potential figure reuse / duplication concerns Some structural overlays appear nearly identical across unrelated designs, raising concerns about: o over-filtering o inadvertent reuse o lack of diversity PubPeer scrutiny would likely question these similarities. 4. SPR sensorgrams missing raw data Only fitted curves are shown; no raw baseline, noise, regeneration, or replicate data provided. This is a major omission for experimental claims. 5. Lack of mass-spec confirmation Designed antibodies were not validated via MS to confirm sequence fidelity. Extended Data Figures 13–15 — Computational–Experimental Correlation Analyses These figures attempt to argue that computational metrics (RMSD, AF2 pLDDT, Rosetta energies) correlate with experimental results (binding affinities, expression levels, stability). Critical Issues 1. Correlation coefficients are inflated by selective reporting Pearson correlations are reported between: o predicted interface RMSD o predicted stability o experimental KD values But these analyses include only successful designs. Excluding failed designs artificially inflates correlations. 2. Small sample sizes invalidate regressions Most correlations are based on n = 4–6 datapoints — statically meaningless. Regression lines plotted on tiny sample sets are scientifically misleading. 3. Confounding: manual selection and downstream refinement Many designed sequences underwent: o surface engineering o extra Rosetta relax cycles o AF2 structural refinement These steps blur causal attribution. 16 10. Design Diversity Debatable While the authors claim broad design capacity, the supplementary data suggest: • narrow sequence variation • recurrent CDR motifs • framework reuse • sampling collapsing into training-distribution basins Thus, RFdiffusion may be performing interpolation, not true de novo generation. Overall Evaluation Bennett et al. present an ambitious vision: fully automated, atomically accurate de novo antibody design. The field needs such breakthroughs, but the present study overstates both the accuracy and generality of RFdiffusion. Major concerns that undermine the central claims: 1. Selective reporting of best examples 2. Insufficient experimental sampling 3. Lack of independent validation 4. Missing raw data 5. Overuse of AF2 as circular validation 6. Structural artifacts and questionable figure consistency 7. Biological irrelevance of simplified antigens 8. Unsubstantiated claims of atom-level precision 9. Opaque methodological transparency 10. Failure to benchmark against contemporaries In its current form, the paper presents what is likely a highly curated demonstration, not a robust or general solution for de novo antibody design. Conclusions and Recommendations To strengthen the scientific validity and credibility of the claims, future iterations of this work should: 1. Provide full model weights, code, and training data Allowing reproducibility and objective evaluation. 17 2. Benchmark rigorously against competing methods Including AbDiffuser, IgFold-Design, and AlphaBind. 3. Demonstrate large-scale design Test hundreds of designs across many antigens to show true generalizability. 4. Report failures Realistic generative models must show both successes and failure modes. 5. Use biologically relevant antigens Full glycosylated, oligomeric versions. 6. Provide raw experimental data Including sensorgrams, chromatograms, crystallographic maps, and unprocessed gels. 7. Reduce reliance on post-hoc refinement Or clearly separate pre-refinement vs post-refinement performance. 8. Include functional validation Such as neutralization assays, in vivo efficacy, or competition studies. 9. Extend diversity metrics Show Shannon entropy, sequence clustering, and standardized antibody repertoire comparisons. 10. Improve clarity in figures With clear alignments, consistent numbering schemes, and transparent annotations. Final Summary Bennett et al. contribute an exciting direction in computational biologics, but the paper—especially the figure sets—does not provide enough evidence to justify the claim of atomically accurate de novo antibody design. The work is impressive, but the presentation is curated, incomplete, and overstated. Stronger experimental rigor, broader validation, and transparency are needed before RFdiffusion can be considered a reliable platform for therapeutic antibody engineering. 18 Reference 1 Bennett, N. R. et al. Atomically accurate de novo design of antibodies with RFdiffusion. Nature (2025). https://doi.org/10.1038/s41586-025-09721-5