Full text
Co-folding of Carbohydrate-Protein Complexes Francesca Peccati∗,†,‡ †Center for Cooperative Research in Biosciences (CIC bioGUNE), Basque Research and Technology Alliance (BRTA) Bizkaia Technology Park, 48160 Derio, Spain ‡Ikerbasque, Basque Foundation for Science, 48013 Bilbao, Spain E-mail: fp[email protected] Phone: +34 946 572 538 Abstract Deep learning–based methods have transformed in silico protein structure prediction, and co-folding approaches now enable the generation of three-dimensional protein models from sequence in the presence of cofactors and custom ligands, unlocking major opportunities in drug discovery and design. Nonetheless, co-folded models remain limited by factors such as a difficult-to-quantify reliance on pose memorization from training data, and violations of geometric and stereochemical constraints. In this context, protein–carbohydrate interactions are particularly challenging to model accurately owing to the high density of chiral centers of the ligands, the flexibility of glycosidic linkages, and the typically exposed nature of binding sites. In this contribution, we assess the performance of two widely used deep learning models, AlphaFold3 and Boltz-2, in predicting carbohydrate–protein complex structures. Using three carbohydrate-binding protein families, we show that accurate geometries of bound states and insights into protein–carbohydrate recognition specificity can be obtained directly from co-folded models generated using only protein sequence and a simple ligand string notation, paving the way for the high-throughput screening of glycans and glycomimetics. 1
Introduction The latest generation of artificial intelligence models for protein structure prediction enables the generation of full-atom macromolecular models with unprecedented accuracy, simultaneously predicting the geometries of noncovalently interacting species—including nucleic acids, proteins, ligands, and ions—as well as covalently bound post-translational modifications.1–3 Co-folding, which consists of predicting a protein structure in the presence of a ligand from protein sequence information alone, is an attractive concept in the context of drug development, as opposed to classical molecular docking.4Docking approaches require a high-quality three-dimensional structure of the receptor, often prior knowledge of the putative binding site, and typically involve only partial conformational sampling of the receptor. Therefore, they are subject to limitations when the target protein structure is unavailable, of low quality, or only available in the holo form with a different ligand, as induced fit can modify the receptor conformation.5–7 The ability to predict protein–ligand complex structures directly from sequence overcomes these limitations, as the binding pose emerges naturally from the simultaneous generation of the coordinates of the binding partners.4 While it remains under debate to what extent co-folding is generalizable across all ligand classes, ligands and poses most frequently encountered during training are predicted with higher accuracy than rarely seen ones, owing to memorization effects;8,9 therefore, extensive benchmarking is required to explore its limitations, particularly in the context of specific interactions that are challenging to predict using classical methods. Among the several available models for co-folding, AlphaFold3 (AF3) currently sets the accuracy benchmark, especially for difficult tasks such as antibody–antigen structure prediction.1Another recently released model is Boltz-2 (Bo2), which enables both structure and affinity prediction, and offers a high degree of controllability through conditioning and steering.3These include the choice of prediction method (generating ensembles that resemble X-ray crystallography, NMR, or molecular dynamics), the ability to enforce structural templates, and the specification of distance and pocket constraints. Additionally, it allows the application 2
of physics-based potentials during inference to reduce unphysical predictions.3Both models share common drawbacks typical of deep learning–based co-folding approaches, such as steric clashes, incorrect assignment of chiral centers, and distorted bond lengths and angles,3 which can be particularly problematic in the prediction of co-folded carbohydrate–protein complexes. Carbohydrate–protein interactions are highly relevant to a range of biological phenomena, including infection and immunity10,11 and tumor growth and metastasis.12 Lectins are non-enzymatic carbohydrate-binding proteins that decode glycan information on cell surfaces and regulate recognition,13 adhesion,14 and signaling.15 They exert their functions by engaging specific glycan motifs through multivalent, often low-affinity contacts that achieve high avidity at the cell interface.13,16 Because lectin–glycan recognition underlies disease processes in infection, inflammation, and cancer, small-molecule glycomimetics are attractive candidates to inhibit or allosterically modulate lectin function for therapeutic purposes and targeted delivery.17–19 Therefore, studying the interactions of lectins with glycans and glycomimetics is of significant biological and therapeutic relevance. As glycan-lectin affinities are generally weak (in the low millimolar range), biological recognition exploits multivalency, generating large, heterogeneous complexes. This makes the study of their interactions particularly challenging and necessitates the combination of multiple techniques, spanning from X-ray crystallography,20 to NMR,21 to isothermal titration calorimetry.22 Notably, NMR enables the quantification of weak, multivalent glycan–lectin interactions in solution while simultaneously providing information on affinity, kinetics, epitope mapping, conformation, and dynamics through complementary ligandand receptor-observed experiments.21 The low affinity of glycan–lectin interactions is encoded in their structure. Carbohydrates are relatively ordered, hydroxyl-dense ligands that mimic structured water, binding lectins through a combination of hydrogen bonds and interactions with aromatic amino acids.23,24 Carbohydrates exhibit extreme structural diversity, with a high density of stereocenters, high 3
conformational flexibility, variable linkage positions, branching, and extensive post-synthetic modifications.25 Lectins bind specific glycans with high precision via structurally tailored, and typically shallow, carbohydrate-binding sites, where the displacement of structured water molecules is an important entropic drive for binding.24,26,27 All these characteristics make glycan–lectin interactions extremely challenging to predict using traditional computational techniques, thereby limiting the potential impact of computer modeling on their engineering and characterization. In this context, co-folding has the potential to advance the computational study of carbohydrate–protein interactions—provided it can yield accurate binding poses and recapitulate the exquisite selectivity of these recognition processes, particularly with respect to the complex stereochemistry and conformational diversity of glycans. A recent study addresses the use of AF3 to predict the structures of complex glycans, highlighting important limitations in the conformational diversity (flexibility) achievable in glycoprotein prediction, and suggesting a bias originating from the PDB training subset.28 In this work, we build on that knowledge by evaluating the ability of AF3 and Bo2 to predict co-folded carbohydrate–protein complex structures through the collective analysis of multiple co-folding solutions, i.e., ensembles of conformations. We demonstrate that such collective analysis can provide valuable insights into binding propensity and selectivity. By examining three families of carbohydrate-binding proteins, we highlight the strengths and limitations of co-folding. We show not only that these models can generate accurate complex geometries, but also that co-folding represents a powerful approach to explore the specificity of carbohydrate–protein recognition using only protein sequence information and a one-dimensional, text-based notation for ligand representation—thus opening the door to high-throughput screening of protein–carbohydrate interactions. 4
Figure 1: Carbohydrate–protein pairs analyzed in this study. CRD: carbohydrate recognition domain. SLBR: Siglec-like binding region. Crystallographic structures of the modeled proteins are shown, highlighting the binding unit or motif and the key amino acids involved in the interaction. Human galectin-3 CRD: PDB ID 3ZSJ;29 human DC-SIGN CRD: PDB ID 1SL5;30 SLBR Siglec subdomain from S. gordonii strain M99, named GspB: PDB ID 5IUC;31 SLBR Siglec subdomain from S. gordonii strain Challis, named Hsa: PDB ID 6X3Q.32 The systems under study are shown in Figure 1 and include the carbohydrate recognition domain (CRD) of a representative galectin (human galectin-3, henceforth Gal3), that of a representative C-type lectin (human DC-SIGN), and two sialic acid-binding immunoglobulintype lectins (siglec) subdomains of streptococcal siglec-like binding regions (henceforth SLBRs). Gal3, which specifically recognizes galactose, is used in this work to gauge co-folding models’ stereochemical fidelity when presented with disaccharide ligands containing the galactose unit, its C4 epimer glucose, and fluorinated glucose analogs. DC-SIGN, which binds glycans through Ca2+-mediated interactions, is used to investigate whether co-folding can generate higher-order (protein–ion–carbohydrate) binding poses that satisfy the optimal binding epitope33 in both monosaccharides and a tetrasaccharide histo blood group antigen.34 Finally, SLBR domains from two different streptococcal strains are used to evaluate co-folding models’ ability to predict highly challenging interactions, specifically those involving diverse and 5
flexible hypervariable loops within domains homologous to the large and varied immunoglobulin superfamily, which critically influence recognition of a panel of sialylated glycans.32 Methods Ligands preparation Three-dimensional models for all monoand oligosaccharides were prepared using the Glycam Carbohydrate Builder35 (https://glycam.org/cb/) and downloaded in PDB format. The structures were then converted to chiral SMILES format using Open Babel36 version 2.3.1, with the standard -osmiles flag. Although SMILES is not the only acceptable input format for co-folding models, it was selected for its simplicity—requiring no specialized glycobiology expertise—and with the aim of enabling large-scale ligand screening. In contrast, structural formats such as CCD files require manual curation of monosaccharide fragments, making them less amenable to high-throughput workflows.28 All ligand structures in SMILES format are available as supplementary materials. Interestingly, despite the fact that ring conformation information is not present in the SMILES format—which encodes only atomic connectivity—the vast majority of co-folded models correctly predict the most stable ring conformation for each monosaccharide (4C1for Glc, Gal, Man, GlcNAc, and GalNAc; 1C4 for Fuc; and 2C5for Neu5Ac). The only exception is a minority of structures of the histo blood group antigen A type VI, in which the Fuc ring adopts an inverted chair conformation (4C1). Full conformational assignments of all co-folded models are available as supplementary materials. AF3 structure prediction AF3 structure predictions were performed using codebase version 3.0.1 with default settings, generating 100 seeds, with 5 diffusion samples per seed (500 models in total). json data files for all predictions are available as supplementary materials. 6
Bo2 structure prediction Bo2 structure predictions were performed using version 2.2.0, employing the “x-ray diffraction” method with 10 recycling steps and 200 sampling steps. One hundred predictions were generated, with 5 diffusion samples each (500 models in total). No contact or pocket conditioning was applied, so that co-folding was performed without steering the interaction toward any specific region. The --usepotentials flag was activated to apply a guiding potential during the diffusion process, improving the quality of ligand geometries. Input files in yaml format for all predictions are available as supplementary materials. The Multiple Sequence Alignment (MSA) server (https://api.colabfold.com) was used for generating the MSAs. Analysis Geometry analysis of the predicted models—including the assignment of the absolute configuration of stereocenters and ring conformations, ligand RMSD calculations with respect to crystallographic structures, and classification of binding modes—was performed using Python3 scripts provided in the ESI. The configuration of the anomeric carbon was not considered when evaluating stereochemical fidelity. Confidence analysis was carried out by examining the interface predicted template modelling (ipTM) scores of the structure ensembles generated with AF3 and Bo2. Of note, neither model includes an explicit penalty term for violations of the absolute configuration of chiral centers; therefore, ipTM scores should be interpreted as a measure of the confidence in the overall co-folded complex structure, independently of the accuracy of the ligand stereochemistry. 7
Results and discussion Stereochemistry of Gal3 recognition of glycans and glycomimetics The S-face canonical CRD of Gal3 recognizes galactose through a combination of CH-πinteractions with W181 and multiple hydrogen bonds. High selectivity toward galactose arises from a pair of hydrogen bonds involving His158 and Arg162, which engage the axial hydroxyl group at position 4 (O4). For each co-folding model and ligand, we generated an ensemble of 500 structures and analyzed them collectively by assigning the absolute configuration of each chiral center and evaluating structural similarity to available experimental references.29 Both AF3 and Bo2 predict the binding of lactose and N-acetyllactosamine (LacNAc) to Gal3 with high confidence and low geometric dispersion across the co-folded ensembles. All predicted structures display the correct disaccharide stereochemistry and show strong agreement with high-resolution crystallographic reference structures of the complexes (RMSD < 1˚ A; Figures 2A, 2B, and S1).29,37 While Bo2 yields a more structurally diverse ensemble—particularly in the flexible N-terminal region—the predicted binding pose is highly accurate, with a unique conformation of the key interacting residues that closely resembles the crystallographic structure of the complex (Figure 2C). Based on this positive result, we challenged the co-folding models with cellobiose, a disaccharide that differs from lactose exclusively in the configuration of a C4 atom, in which the binding galactose unit is replaced by glucose. Interestingly, neither AF3 nor Bo2 predicts cellobiose with the correct stereochemistry in any of the 500 predicted structures (0% correct chirality, Figure 2A). Instead, both models epimerize the C4 carbon of the non-reducing end Glc unit, yielding structurally indistinguishable predictions from those generated for lactose. This result is significant because it clearly demonstrates that stereochemical violations in these ligands are not random, but are instead induced by the context of the receptor. In other words, the models have learned that galectins bind lactose analogues, and consequently impose a stereochemical distortion that converts glucose into the preferred galactose epimer, 8
favoring binding over stereochemical fidelity during co-folding. As a result, the analysis of stereochemical violations in co-folded structures may carry valuable information about the binding preferences of CRDs. Of note, a symmetric galactose to glucose stereochemical violation was recently reported in a study benchmarking AF3’s ability to predict the structures of large glycans.28 Figure 2: A) Percentage of models predicted with the correct ligand stereochemistry using AF3 and Bo2. B) Overlay of 500 co-folded structures of lactose bound to Gal3, predicted by AF3 and Bo2. C) Crystallographic structure of lactose bound to Gal3.29 D) Overlay of 77 AF3-predicted co-folded models of 4-deoxy-4-fluorocellobiose (left) and of 30 AF3-predicted co-folded models of 2,3,4-deoxy-2,3,4-trifluorocellobiose (right). E) Representative Bo2-predicted co-folded models of 4-deoxy-4-fluorocellobiose (left) and 2,3,4-deoxy2,3,4-trifluorocellobiose (right). Key residues involved in ligand binding are shown as pale cyan sticks. Fluorine atoms are represented as orange spheres. Glc, fluorinated Glc and Gal are shown as blue, pink and yellow sticks, respectively. Motivated by these findings, we sought to further investigate this phenomenon and assess the extent to which stereochemical violations affect the co-folding of ligands that are not strictly natural carbohydrates. As straightforward models of glycomimetics, we examined two analogs: 4-deoxy-4-fluorocellobiose and 2,3,4-deoxy-2,3,4-trifluorocellobiose. With these ligands, we observed divergent behavior between the two co-folding models: while Bo2 9
Figure 4: A) Percentage of models predicted with the correct ligand stereochemistry using AF3 and Bo2. B) Overlay of AF3-predicted co-folded structures of the GspB–sT (left) and GspB–3’sLn (right) complexes. The three hypervariable loops encoding selectivity—CD, EF, and FG—are shown in green, blue, and light brown, respectively. C) Histograms of prediction confidence scores (ipTM) for ensembles of 500 co-folded structures for the GspB–sT, GspB–3’sLn, Hsa–sT, and Hsa–3’sLn complexes. Bars in green indicate models with correct ligand stereochemistry; red bars indicate models with stereochemical violations. Figure 4B shows an overlay of the structures co-folded with AF3. These display negligible geometric dispersion, with the sT ligand predicted at high precision relative to the reference crystallographic structure (Figure S11, RMSD <1˚ A), and accompanied by high confidence scores (ipTM, Figure 4C), indicating that AF3 accurately captures the structural features of this interaction. In contrast, when presented with 3’sLn—a ligand not recognized by GspB—AF3 predicts a broad ensemble of geometries characterized by variable orientations of the GlcNAc unit, which fails to engage the FG and CD loops, while maintaining correct positioning of the Neu5Acα2,3Gal disaccharide. This is associated with a decrease in ipTM scores (Figure 4C) and an increased frequency of stereochemical violations. These results are consistent with the observations made for Gal3, reinforcing the notion that increased 16
structural variability across co-folded ensembles, together with reduced predicted confidence, correlates with diminished binding specificity. This conclusion is further supported by co-folding the same ligands with the more broadly selective Hsa domain, for which prediction confidences increase (Figure 4C). In this case, both ligands are predicted with narrow geometric dispersion (Figures 5 and S11) and minimal stereochemical violations (Figure 4A) by both AF3 and Bo2. These results indicate that cofolding—particularly with AF3—not only produces accurate binding poses but also captures ligand selectivity, even in Ig-like β-sandwich folds where binding interfaces are defined by flexible loop regions. This accurate binding pose description implies that co-folding correctly samples loop geometries in a straightforward manner, which is an open challenge for classical methods.41,42 Finally, motivated by these results, we challenged AF3 with a broader panel of ligands to evaluate its response to structural modifications such as fucosylation and O-sulfation, as shown in Figure 5. The experimentally measured affinities of Hsa for these ligands, reported in Ref. 32, follow the trend: sT >sLeC>3’sLn >sLeX>6S-sLeX, with sT-like ligands being favored over fucosylated and sulfated variants. AF3-predicted co-folded ensembles for the same protein–ligand pairs were ranked based on the percentage of models with correct stereochemistry and associated ipTM confidence scores (Figure 5). These predictions show a clear preference for unfucosylated ligands and yield the following predicted affinity trend: sT >3’sLn >sLeC>sLeX>6S-sLeX, which closely mirrors the experimental ranking, with the exception of a reversal between 3’sLn and sLeC, which differ only in a single glycosidic linkage. This proof-of-concept qualitative in silico affinity screening strongly supports the ability of AF3 co-folding to capture subtle, chemically encoded features of glycan-binding selectivity. It highlights co-folding as a robust approach for both ligand screening and probing the structural basis of glycan recognition. 17
Figure 5: Overlay of AF3 co-folded structures with the correct ligand chirality of the Hsa SLBR siglec domain with a panel of different sialoglycans. The three hypervariable loops encoding selectivity (CD, EF, and FG) are colored in green, blue and light brown, respectively. For each ligand, the percentage of structures with the correct chirality (out of 500 structures) is shown together with the average confidence score (ipTM). Conclusions We have demonstrated that co-folding is a powerful tool for characterizing carbohydrate–protein interactions. A critical challenge in glycan co-folding lies in the large number of chiral centers, which represent a major vulnerability for current structure prediction models. Here, we showed that stereochemistry violations in co-folded structures correlate with binding propensity and thus carry meaningful information about interaction quality, as exemplified by Gal3. Our results also underscore the importance of analyzing not just single predictions, but rather ensembles of co-folded complexes. Because co-folding structure prediction is inherently non-deterministic, ensemble analysis provides insight into both the multiplicity of possible binding modes and the population distribution among them—as shown for DC-SIGN. Furthermore, we demonstrated that AF3 accurately reproduces ligand preferences in siglec-like domains as a function of the composition and geometries of hypervariable loops, indicating remarkable predictive power even in the context of conformationally flex18
ible loop-mediated interfaces. Importantly, all predictions were generated using only the protein sequence and a simple text-based representation of the ligand, highlighting the main limitations and the extraordinary potential of in silico co-folding for large-scale prediction and mechanistic understanding in glycobiology. Future developments will allow direct quantification of binding affinities, whose accuracy is already in par with more time-consuming classical methods,3enabling accurate and efficient characterization of protein-carbohydrate interactions. Data and Software Availability The following data are available through the Zenodo repository https://zenodo.org/records/17618082: AF3 and Bo2 co-folding inputs, AF3 and Bo2 co-folded models and their stereochemistry and ring conformation assignments. AlphaFold open source code can be downloaded from https://github.com/google-deepmind/alphafold3. PyMOL open source code can be downloaded from https:// github.com/schrodinger/pymolopen-source. Sample code used for analysis is provided in the ESI. Acknowledgement This research has been funded by MCIN/AEI/10.13039/501100011033 (grants RYC2022036457-I and EUR2023-143462). Supporting Information Available Distributions of RMSD values for Gal3 bound to lactose and LacNAc relative to reference crystallographic structures; Bo2-predicted ensembles of Gal3–lactose and Gal3–LacNAc complexes; distributions of prediction confidence scores for Gal3–lactose and Gal3–LacNAc 19
ensembles predicted with AF3 and Bo2; AF3and Bo2-predicted ensembles of DC-SIGN in complex with FucOMe and ManOMe; distributions of distances to residues Glu354 and Val351 for FucOMeand ManOMe-bound DC-SIGN in ensembles predicted with AF3 and Bo2; distributions of prediction confidence scores for DC-SIGN–FucOMe and DC-SIGN–ManOMe ensembles predicted with AF3 and Bo2; AF3-predicted ensemble of the DC-SIGN–blood group antigen A type VI complex; comparison of GspB and Hsa SLBRs; distributions of RMSD values for GspB and Hsa bound to sT and 3’sLn relative to reference crystallographic structures; AF3and Bo2-predicted ensembles of Hsa in complex with sT and 3’sLn; all protein sequences and ligand text representations of all co-folded binding partners; sample code for assigning the stereochemistry of co-folding solutions, computing RMSD values, assigning ring conformations, classifying DC-SIGN binding poses, and performing confidence-score analyses. References (1) Abramson, J. et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024,630, 493–500. (2) Fang, Z.; Ran, H.; Zhang, Y.; Chen, C.; Lin, P.; Zhang, X.; Wu, M. AlphaFold 3: an unprecedent opportunity for fundamental research and drug development. Precis. Clin. Med. 2025,8, pbaf015. (3) Passaro, S.; Corso, G.; Wohlwend, J.; Reveiz, M.; Thaler, S.; Somnath, V. R.; Getz, N.; Portnoi, T.; Roy, J.; Stark, H.; Kwabi-Addo, D.; Beaini, D.; Jaakkola, T.; Barzilay, R. Boltz-2: Towards Accurate and Efficient Binding Affinity Prediction. bioRxiv 2025, (4) Bryant, P.; Kelkar, A.; Guljas, A.; Clementi, C.; No´e, F. Structure prediction of proteinligand complexes from sequence information with Umol. Nat. Commun. 2024,15, 4536. 20
(5) Elokely, K. M.; Doerksen, R. J. Docking Challenge: Protein Sampling and Molecular Docking Performance. J. Chem. Inf. Model. 2013,53, 1934–1945. (6) Amaro, R. E.; Baudry, J.; Chodera, J.; Demir, ¨ O.; McCammon, J. A.; Miao, Y.; Smith, J. C. Ensemble Docking in Drug Discovery. Biophys. J. 2018,114, 2271–2278. (7) Nabuurs, S. B.; Wagener, M.; de Vlieg, J. A Flexible Approach to Induced Fit Docking. J. Med. Chem. 2007,50, 6507–6518. (8) ˇ Skrinjar, P.; Eberhardt, J.; Durairaj, J.; Schwede, T. Have protein-ligand co-folding methods moved beyond memorisation? bioRxiv, doi: 10.1101/2025.02.03.636309 2025, (9) Masters, M. R.; Mahmoud, A. H.; Lill, M. A. Do Deep Learning Models for Co-Folding Learn the Physics of Protein-Ligand Interactions? bioRxiv, doi: 10.1101/2024.06.03.597219 2024, (10) Crocker, P. R.; Paulson, J. C.; Varki, A. Siglecs and their roles in the immune system. Nat. Rev. Immunol. 2007,7, 255–266. (11) Schnaar, R. L. Glycobiology simplified: diverse roles of glycan recognition in inflammation. J. Leukoc. Biol. 2016,99, 825–838. (12) Pinho, S. S.; Reis, C. A. Glycosylation in cancer: mechanisms and clinical implications. Nat. Rev. Cancer 2015,15, 540–555. (13) Dam, T. K.; Brewer, C. F. Lectins as pattern recognition molecules: The effects of epitope density in innate immunity. Glycobiol. 2009,20, 270–279. (14) Ludwig, A.-K. et al. Design–functionality relationships for adhesion/growth-regulatory galectins. Proc. Natl. Acad. Sci. USA 2019,116, 2837–2842. (15) Geijtenbeek, T. B. H.; Gringhuis, S. I. Signalling through C-type lectin receptors: shaping immune responses. Nat. Rev. Immunol. 2009,9, 465–479. 21
(16) Mammen, M.; Choi, S.-K.; Whitesides, G. M. Polyvalent Interactions in Biological Systems: Implications for Design and Use of Multivalent Ligands and Inhibitors. Angew. Chem. Int. Ed. 1998,37, 2754–2794. (17) Leusmann, S.; M´enov´a, P.; Shanin, E.; Titz, A.; Rademacher, C. Glycomimetics for the inhibition and modulation of lectins. Chem. Soc. Rev. 2023,52, 3663–3740. (18) Porkolab, V.; Chabrol, E.; Varga, N.; Ordanini, S.; Sutkeviciute, I.; Th´epaut, M.; Garc’ia-Jim´enez, M. J.; Girard, E.; Nieto, P. M.; Bernardi, A.; Fieschi, F. RationalDifferential Design of Highly Specific Glycomimetic Ligands: Targeting DC-SIGN and Excluding Langerin Recognition. ACS Chem. Biol. 2018,13, 600–608. (19) Springer, A. D.; Dowdy, S. F. GalNAc-siRNA Conjugates: Leading the Way for Delivery of RNAi Therapeutics. Nucleic Acid Ther. 2018,28, 109–118. (20) Cordara, G.; Krengel, U. Carbohydrate Chemistry: Chemical and Biological Approaches; The Royal Society of Chemistry, 2013. (21) Quintana, J. I.; Atxabal, U.; Unione, L.; Ard´a, A.; Jim´enez-Barbero, J. Exploring multivalent carbohydrate–protein interactions by NMR. Chem. Soc. Rev. 2023,52, 1591–1613. (22) Dam, T. K.; Talaga, M. L.; Fan, N.; Brewer, C. F. Calorimetry; Methods in Enzymology; Academic Press, 2016; Vol. 567; pp 71–95. (23) Asensio, J. L.; Ard´a, A.; Ca˜nada, F. J.; Jim´enez-Barbero, J. Carbohydrate–Aromatic Interactions. Acc. Chem. Res. 2013,46, 946–954. (24) Peccati, F.; Jim´enez-Os´es, G. Enthalpy–Entropy Compensation in Biomolecular Recognition: A Computational Perspective. ACS Omega 2021,6, 11122–11130. (25) He, X.; Zhao, L.; Tian, Y.; Li, R.; Chu, Q.; Gu, Z.; Zheng, M.; Wang, Y.; Li, S.; 22
Jiang, H.; Jiang, Y.; Wen, L.; Wang, D.; Cheng, X. Highly accurate carbohydratebinding site prediction with DeepGlycanSite. Nat Commun 2024,15, 5163. (26) Bertuzzi, S.; Quintana, J. I.; Ard´a, A.; Gimeno, A.; Jim´enez-Barbero, J. Targeting Galectins With Glycomimetics. Front. Chem. 2020,8. (27) Fadda, E.; Woods, R. J. On the Role of Water Models in Quantifying the Binding Free Energy of Highly Conserved Water Molecules in Proteins: The Case of Concanavalin A. J. Chem. Theory Comput. 2011,7, 3391–3398. (28) Huang, C.; Kannan, N.; Moremen, K. W. Modeling glycans with AlphaFold 3: capabilities, caveats, and limitations. Glycobiol., doi: 10.1093/glycob/cwaf048 2025, (29) Saraboji, K.; H˚akansson, M.; Genheden, S.; Diehl, C.; Qvist, J.; Weininger, U.; Nilsson, U. J.; Leffler, H.; Ryde, U.; Akke, M.; Logan, D. T. The Carbohydrate-Binding Site in Galectin-3 Is Preorganized To Recognize a Sugarlike Framework of Oxygens: Ultra-High-Resolution Structures and Water Dynamics. Biochemistry 2012,51, 296– 306. (30) Guo, Y.; Feinberg, H.; Conroy, E.; Mitchell, D. A.; Alvarez, R.; Blixt, O.; Taylor, M. E.; Weis, W. I.; Drickamer, K. Structural basis for distinct ligand-binding and targeting properties of the receptors DC-SIGN and DC-SIGNR. Nat. Struct. Mol. Biol. 2004, 11, 591–598. (31) Pyburn, T. M.; Bensing, B. A.; Xiong, Y. Q.; Melancon, B. J.; Tomasiak, T. M.; Ward, N. J.; Yankovskaya, V.; Oliver, K. M.; Cecchini, G.; Sulikowski, G. A.; Tyska, M. J.; Sullam, P. M.; Iverson, T. M. A Structural Model for Binding of the Serine-Rich Repeat Adhesin GspB to Host Carbohydrate Receptors. PLoS Pathog. 2011,7, 1–17. (32) Bensing, B. A. et al. Origins of glycan selectivity in streptococcal Siglec-like adhesins suggest mechanisms of receptor adaptation. Nat. Commun. 2022,13, 2753. 23
(33) Mart´ınez, J. D.; N´u˜nez-Franco, R.; Valverde, P.; Delgado, S.; Ard´a, A.; Jim´enezBarbero, J.; Jim´enez-Oses, G.; Ca˜nada, F. J. Glycans and Chirality: Stereoselectivity at the Core of DC-SIGN’s Recognition. A Novel View of the Optimum Minimal Ligand Epitope. Chem. Eur. J. 2025,31, e202501420. (34) Valverde, P.; Delgado, S.; Mart´ınez, J. D.; Vendeville, J.-B.; Malassis, J.; Linclau, B.; Reichardt, N.-C.; Ca˜nada, F. J.; Jim´enez-Barbero, J.; Ard´a, A. Molecular Insights into DC-SIGN Binding to Self-Antigens: The Interaction with the Blood Group A/B Antigens. ACS Chem. Biol. 2019,14, 1660–1671. (35) Grant, O. C.; Wentworth, D.; Holmes, S. G.; Kandel, R.; Sehnal, D.; Wang, X.; Xiao, Y.; Sheppard, P.; Grelsson, T.; Coulter, A.; Miller, G.; Foley, B. L.; Woods, R. J. Generating 3D Models of Carbohydrates with GLYCAM-Web. bioRxiv, doi: 10.1101/2025.05.08.652828 2025, (36) O’Boyle, N. M.; Banck, M.; James, C. A.; Morley, C.; Vandermeersch, T.; Hutchison, G. R. Generating 3D Models of Carbohydrates with GLYCAM-Web. J. Cheminform. 2011,3, 33. (37) S¨orme, P.; Arnoux, P.; Kahl-Knutsson, B.; Leffler, H.; Rini, J. M.; Nilsson, U. J. Structural and Thermodynamic Studies on Cation-Π Interactions in Lectin-Ligand Complexes: High-Affinity Galectin-3 Inhibitors through Fine-Tuning of an Arginine-Arene Interaction. J. Am. Chem. Soc. 2005,127, 1737–1743. (38) Miller, M. C.; Ippel, H.; Suylen, D.; Klyosov, A. A.; Traber, P. G.; Hackeng, T.; Mayo, K. H. Binding of polysaccharides to human galectin-3 at a noncanonical site in its carbohydrate recognition domain. Glycobiol. 2016,26, 88–99. (39) Snyder, G. A.; Colonna, M.; Sun, P. D. The Structure of DC-SIGNR with a Portion of its Repeat Domain Lends Insights to Modeling of the Receptor Tetramer. J. Mol. Biol. 2005,347, 979–989. 24
(40) N´u˜nez-Franco, R.; Muriel-Olaya, M. M.; Jim´enez-Os´es, G.; Peccati, F. AlphaFold2 Predicts Alternative Conformation Populations in Green Fluorescent Protein Variants. J. Chem. Inf. Model. 2024,64, 7135–7140. (41) Fern´andez-Quintero, M. L.; Kokot, J.; Waibl, F.; Fischer, A.-L. M.; Quoika, P. K.; Deane, C. M.; Liedl, K. R. Challenges in antibody structure prediction. mAbs 2023, 15, 2175319. (42) Stachowski, T. R.; Fischer, M. Large-Scale Ligand Perturbations of the Protein Conformational Landscape Reveal State-Specific Interaction Hotspots. J. Med. Chem. 2022, 65, 13692–13704. 25