Full text
A high-performance and modular protein design pipeline Josh Hardy Lucet Laboratory New Medicines and Diagnostics Division Walter and Eliza Hall Institute of Medical Research Socials/Links ProteinDJ Logo designed by Lyn Deng
Presentation overview 1. Introduction to protein design 2. What ProteinDJ can do for you 3. Getting started in binder design 4. Behind the scenes: Development of ProteinDJ 5. How to install and run ProteinDJ
Introduction to Protein Design
Protein design comes in many flavours •We have been designing and modifying proteins for decades: •Protein fusions – concatenating sequences e.g. expression tags, solubility tags, domains •Mutations/deletions/insertions – to modify function •Grafting of binding loops in antibodies •Many of these sequence changes will also result in a structural change and can affect protein stability and folding •Designing new or ‘de novo’ proteins is the ultimate test of our understanding of the protein structure-sequence relationship •“What I cannot create, I do not understand” – Richard Feynman
The rise of de novo proteins Fox, Taveneau et al. Structure 2025 Woolfson JMB 2021 Deep learning methods have recently accelerated de novo protein design
Fox, Taveneau et al. Structure 2025 The untapped potential of de novo protein binders •A subgenre of protein design is the creation of de novo protein binders that bind target proteins •De novo protein binders have been used to solve solve problems in biology, chemistry and medicine
Software for de novo binder design •A rapidly expanding list with different approaches and availability •Generative methods – trained on PDB database to generate new folds, typically using diffusion from random noise •RFdiffusion •Chai-2 •PXDesign-d •AlphaProteo •Hallucination methods – iterative optimisation of sequences using structure prediction methods •BindCraft •BoltzDesign1 •PXDesign-h Pacesca, Nature, 2025 Watson, Nature, 2023
ProteinMPNN and structure prediction •Some protein design tools (e.g. RFdiffusion) only generate a fold without a sequence or side-chains •ProteinMPNN can design a sequence to match the fold. There are multiple variants e.g. SolubleMPNN, Ligand MPNN, Full-Atom MPNN •We can test the designed sequence using structure prediction •i.e. Does the protein fold the same as the original design? Dauparas et al. Robust deep learning–based protein sequence design using ProteinMPNN. Science. 2022. Figure: Simon Duerr
Challenge 1: Binder design can have a high failure rate •Early de novo design campaigns required expression of hundreds or thousands of designs •Structure prediction methods have improved the chance of success but we still need dozens of designs to test in the lab •Success rates for structure prediction can vary dramatically, often 0-5% •This requires the structure prediction of thousands of designs •Need for high-performance computing Bennet et al. Nature Communications 2023
Fold conditioning mode •Rather than creating ‘de novo’ binders, you can use existing folds as templates •This was a popular approach before RFdiffusion •Advantage in obtaining fold diversity •RFdiffusion has a strong bias to alpha-helical binders •I have collated recommended fold templates published by the Baker lab on our GitHub HHH 3-helical bundle 17,668 templates HHHH 4-helical bundle 3,499 templates HEEHE Synthetic fold 208 templates EHEEHE Ferredoxin fold 1,555 templates + A Target A Binder Design 1 Ꞵ A Binder Design 2 A Ꞵ A Diffusion Diffusion +Noise Secondary structures binder_foldcond
Fold conditioning example Fold conditioning mode with Ferredoxin foldDe novo mode (only helical binders succeeded) Base model success rate = 3.0% ‘Beta’ model success rate = 2.4% Success rate = 8.3% 800 binders against a benchmarking target - Influenza A H1 haemagglutinin (HA)
Modified from Vázquez Torres et al., Nature 2024 Partial diffusion mode •For when you want to diversify a binder e.g. a promising de novo design that needs refinement •Will partially noises/denoises an input fold (will design a new sequence too) •The number of timesteps will determine how much noise will be added: •10-20: Slight variations in fold •30-40: Major variations in fold •50: Full diffusion of fold A Target + Binder A A A Binder Design 1 A Binder Design 2 Diffusion Diffusion +Noise binder_partialdiff t = 2.5 t = 10 t = 20 t = 30 t = 40 t = 50
Motif-Scaffolding mode •The most complex mode •For when you have part of a binder, either: •A binding ‘motif’ e.g. peptide •Or a ‘scaffold’ e.g. framework •Requires careful placement of elements in input PDB file and specification of insertion residues and length A A Target + Scaffold A A A Binder Design 2 Binder Design 1 Diffusion ADiffusion +Noise +Noise binder_motifscaff + Motif Target + Motif + Scaffold Watson et al., Nature 2023
Filtering RFdiffusion outputs •Rfdiffusion sometimes produces low-quality or undesired binder folds •E.g. single alpha-helices •We have implemented several filters to weed these out during the run: •Min/Max Alpha-Helices •Min/Max Beta-Strands •Min/Max Secondary Structures (Beta-Strands+Alpha-Helices) •Min/Max Radius of Gyration 5-helices, RoG = 11.5 Å 3-helices, RoG = 17.8 Å 1-helix, RoG = 25.0 Å 3-helices, RoG = 14.9 Å
Stage 2: Sequence Design •Rfdiffusion diffuses polyglycines without sidechains or sequence information •We need to design new sequences and side-chains, typically with 10-20 attempts per fold •2 options in ProteinDJ: •ProteinMPNN (+FastRelax) •Full-Atom MPNN (GPU accelerated)
Option 1: ProteinMPNN FastRelax •FastRelax is an optional plug-in to ProteinMPNN •It uses Rosetta to relax the backbone and side-chains of the binder •It is designed to be run in cycles e.g. MPNN -> FR -> MPNN -> FR -> MPNN •It does improve success rates, but it is slow (~1-2 minutes per design) compared to a few seconds per design with ProteinMPNN only Before: AGYVPREATPEERAARAAASGAANAARRRARAAEAAAARA After: MGFTPRERTPEERAATGAVRSRRALERRLAELAAEEAERA Cycle 1 Cycle 5 0 0.5 1 1.5 2 2.5 MPNN score with Fast Relax Initial Cycle 1 Cycle 2 Fi n al Cycle 5 Diminishing returns with more than one cycle
Option 2: Full-Atom MPNN •Newer algorithm that predicts sequence and side-chain conformation iteratively •At each iteration, high-scoring side-chains are retained and it will re-attempt other side-chains •GPU-accelerated process (~5 sec per design) •Unlike FastRelax method, the backbone is unchanged – only side-chains atoms are refined
Sequence design method comparison •We compared success rates for five benchmarking targets •Same benchmarking targets as RFdiffusion paper •FastRelax always improved success rates •Full-Atom MPNN was inconsistent, sometimes better than ProteinMPNN, sometimes worse Sequence Design Method HA IL7RαIR PD-L1 TrkA ProteinMPNN(Vanilla) 3.9% 6.0% 21.3% 30.6% 5.4% ProteinMPNN(Vanilla) + FastRelax 7.0% 11.1% 26.4% 47.5% 12.1% ProteinMPNN(Soluble) 5.0% 5.8% 27.6% 33.9% 8.5% ProteinMPNN(Soluble) + FastRelax 6.1% 9.1% 29.6% 52.3% 18.1% FAMPNN 0.1% 10.9% 11.6% 30.9% 7.5%
Sequence design filtering •As with the outputs of RFdiffusion, you can filter the sequence designs by score: •Max Negative Log-Probability (ProteinMPNN) •Max Predicted Side-Chain Error (FAMPNN) •It is unclear if these values are predictive of success •However, it will be useful to capture and retrospectively analyse these scores and see if there are effective cut-offs Dylan Silke – Honours Thesis
Current limitations and future directions •ProteinDJ is (currently) not suited for antibody design or design of enzymes or small molecule binders •These methods require specialised tools and all-atom design software e.g. RFAntibody, LigandMPNN, RFdiffusion2, ProteinDJ2? •Bindcraft is not yet part of our pipeline and gives different binders to RFdiffusion – it’s worth trying both! •Our workflow is linear – no automatic recycling of best designs •With new users and different HPC environments we will inevitably encounter new bugs and things we could not anticipate •Please log an issue on our GitHub •With feedback and testing we hope to make the software easier to use and improve our documentation and tutorials
Getting started in binder design
Understanding your target protein •Software is only part of it – understanding the structure of your target protein is essential •Some target proteins have hydrophobic patches that RFdiffusion will easily generate a binder against – others will be more challenging •Try different hotspots and input models •Or maybe try a new homologue/target •The process is very sensitive to the input PDB model – quality in, quality out •Crystal structure vs. cryo-EM structure vs. Structure prediction
Target structure preparation example Exposed hydrophobic residues Good prep Bad prep Chain break Domain of interest Hotspots RFdiffusion runs very slowly with large input structures so we often crop domains of interest
Uncropped target prediction •Cropping target structures can expose buried interfaces and remove context •We had added a way to provide an ‘uncropped’ target for the structure prediction step and give the prediction more context Chains Residues RFD time AF2 time A149 1 min 0.2 min A + B 298 3 min 0.4 min A + B + C 447 6 min 1 min A + B + C + D + E + F 858 18 min 3 min Binder is flagged as problematic Cropped target Binders that clashes with missing chain Structure Prediction Uncropped target RFdiffusion Merge Domain of interest
What do ‘good’ binders look like? •Good binders come in different shapes and sizes but typically have: •A hydrophobic core i.e. a fold with buried side-chains •Shape-complementarity with the target protein •Sufficient buried surface area (and/or hydrogen bonds) to maintain an interaction •Bigger is not always better – successful mini-binders can be as small as 60 residues. Try a range of sizes •Inspect your best designs in ChimeraX and overlay with relevant structures, check for clashes with full context •Remember it is only a prediction – your binder may not fold that way or even express at all in the lab
Identifying best parameters for binder design •It can be difficult to guess the best parameters for binder design e.g. hotspots, length •And this may vary from target to target too •This requires manual testing of parameters, which takes a lot of time and bookkeeping MPNN Model Ha IL7Ra IR PD-L1 TrkA v_48_010 3.1% 4.9% 11.6% 17.8% 2.1% v_48_020 5.3% 5.4% 26.9% 32.3% 10.6% v_48_030 3.9% 8.4% 26.1% 44.0% 13.5% Manually generated results from 15 runs each Fold Conditioning Program De novo HHH HHHH HEEHE EHEEHE FAMPNN 0.5% 3.8% 3.1% 1.6% 4.1% MPNN 3.0% 9.8% 7.8% 7.1% 8.3% MPNNSol 3.9% 11.3% 7.1% 8.5% 7.6% Comparing MPNN checkpoints Comparing Folds for Binder Design
Bindsweeper: a parameter sweeping tool •We developed a new tool for automated parameter sweeping •Bindsweeper generates an N-dimensional matrix of parameter combinations and performs multiple ProteinDJ runs •Great for testing hotspots or binder folds at small scale e.g. 100 designs, before large scale runs •Command-line only Checkpoint Hotspots Noise scale complex_beta “[B13,B17,B19]” 0.0 complex_beta “[B13,B17,B19]” 0.5 complex_beta “[B13,B17,B19]” 1.0 complex_beta “[B44, B56]” 0.0 complex_beta “[B44, B56]” 0.5 complex_beta “[B44, B56]” 1.0 ProteinDJ run 4 fixed_params: rfd_ckpt_override: complex_beta sweep_params: rfd_hotspots: values: - “[B13,B17,B19]” - “[B44, B56]” rfd_noise_scale: min: 0.0 max: 1.0 step: 0.5 Base parameters Base parameters Base parameters Sweep parametersSweep parametersSweep parameters ProteinDJ run 1 ProteinDJ run 2 ProteinDJ run 3 Compiled designs and success rates nextflow.config Generation of sweep matrix Base parameters Sweep parameters Input YAML configuration Serial Nextflow execution … Parameter sweep bindsweeper.config
Behind the Scenes: Development and Design of ProteinDJ
My first steps into the world of de novo protein design •What I found challenging getting started: •3 programs/python environments to install (took me 2 weeks) •Mixed GPU/CPU utilisation •Different metadata formats •PDB format vs Silent format Rhys Grinter Gavin Knott Hardy et al., ASBMB Newsletter 8/25 Marija Dramicanin Marjan Hadian-Jazi & Dr Richard Birkinshaw As seen previously in the BioCommons Protein Binder Seminar Series! Cyntia Taveneau
Nextflow – Coordinating workflows on HPC •A Java/Groovy-based pipeline language compatible with containers and HPC (including cloud compute) •Widely used for bioinformatics e.g. single-cell sequence data •Data and files can be fed into ‘channels’ and tasks can be performed on each item (or batch of items) https://nf-co.re/rnaseq/3.20.0/
Python scripting - the glue that holds it together •Python is a popular programming language for science and has a huge library of tools and libraries •We used Python and the PyRosetta libraries to tweak input/output files e.g.: •Align two structures together •Extract a sequence from a PDB file •Calculate biophysical metrics •Match metadata files to PDB files •etc. •Note: PyRosetta requires a license for commercial projects (but is used in Bindcraft, dl_binder_design, AlphaFold2 initial guess, etc. anyway) •P.S. Check out ‘FreeBindCraft’ for a PyRosetta workaround for BindCraft https://github.com/cytokineking/FreeBindCraft
Metadata and file formats •This was the most challenging part of development – how to get different programs to cooperate •After several attempts we settled on the following: •Structure files as individual PDB files •Metadata files as individual JSON files •ID tags for folds and sequences in PDB filenames and metadata •Tags for metrics and parameters e.g. rfd_mode, mpnn_temperature •Maximises interoperability and gives us a consistent start/end point for each module •Input: PDB files •Output: PDB files and JSON files
Our Modular Nextflow Pipeline
How to install and use ProteinDJ
Installing ProteinDJ •2 step installation process: 1. Clone the GitHub > git clone https://github.com/PapenfussLab/proteindj 2. Download models/checkpoints for software (~10 mins using our script) > bash scripts/download_models.sh •Our containers will download automatically when you launch the software and will cache locally •But you can build them yourself if you prefer •Note: If not at WEHI your HPC may have different requirements e.g. memory limits, GPU types, queue names etc. that will need to be provided to Nextflow. Ask your system admins for help.
Running ProteinDJ – Configuration •Before running ProteinDJ you need to set the design parameters and provide any input files e.g. hotspots, target PDB files •You can skip stages i.e. run structure prediction and analysis only •Two ways to interact: •Commandline •Seqera (Web server)
Running ProteinDJ - Commandline 1. Edit design parameters in nextflow.config file e.g. number of designs, input PDB 2. Run ProteinDJ > nextflow run main.nf Optionally use profiles: > nextflow run main.nf –profile milton,binder_denovo The nextflow.config file Example profiles:
Running ProteinDJ - Seqera Web Server •Web interface including input form & job tracking •Available for WEHI users through partnership with Australian BioCommons
Conclusion •ProteinDJ is a tool for the present and a framework for the future •The protein design software landscape is rapidly evolving and there is a need for conventions and interoperability •We have made our code open-source to enable collaborative efforts to build similar workflows and incorporate future software ProteinDJ Design Pipeline Hotspots Protein Target RFdiffusion ProteinMPNN or Full-Atom MPNN AlphaFold2 or Boltz-2 Binder Designs and Metadata … Diffusion & Fold Design Sequence Design Structure Prediction Binder Analysis & Reporting