Proteina-Complexa

Overview. We present Proteina-Complexa, a novel fully atomistic protein binder design framework unifying conditional generative modeling and optimization. This page is partitioned into two parts: Core model development and wet lab validation. We carried out an experimental design campaign with over 1 million binder candidates for 133 targets in collaboration with Manifold Bio, Viva Biotech, Novo Nordisk, Duke University, Cambridge University, LMU Munich and Leipzig University.

No Redesign

Sequences & structures generated end-to-end without inverse folding.

Test-Time Scaling

Latent generative search: More compute at inference → better binders.

Wet Lab Campaign

133 targets. >1M binders screened. Up to 63.5% hit rates. Picomolar affinities.

Carbohydrate Binders

First ever de novo designed carbohydrate binders.

Model Highlights

  • Mini-binder generation for protein and small molecule targets; atomistic motif scaffolding for enzyme design.
  • Generative pretraining with inference-time compute scaling outperforms prior generation and hallucination methods.
  • Teddymer: Synthetic binder-target pairs from domain-domain interactions of predicted monomer structures.
  • No more re-design: Binder sequences generated directly by Proteina-Complexa without additional inverse folding.
  • Interface hydrogen bond optimization for strong biophysical interactions; fold class guidance for control and diversity.
  • State-of-the-art performance on in-silico binder design metrics and in computational enzyme design benchmarks.

Extensive Wet Lab Validation

  • All-to-all binding on 127 target set: Off-target hits for all targets and on-target hits against 115 targets.
  • Method benchmark: Proteina-Complexa generates more experimentally validated hits than prior models.
  • 63.5% hit rates & picomolar affinities for PDGFR; 40%-50% hit rates for kinase mini-protein & peptide binders.
  • First de novo carbohydrate binders — a target class previously thought inaccessible to current methods.
  • Binders to muscle-wasting Activin receptor type IIA with validated blocking of myostatin signaling in cells.
  • Nanomolar binders against Nipah virus and joint structure-sequence re-engineering of existing binders.

1. The Proteina-Complexa Model

Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute

1 NVIDIA    2 University of Oxford    3 Mila - Québec AI Institute    4 Université de Montréal    5 HEC Montréal    6 CIFAR AI Chair    7 AITHYRA    8 School of Biological Sciences, Seoul National University    9 Interdisciplinary Program in Bioinformatics, Seoul National University    10 Institute of Molecular Biology and Genetics, Seoul National University    11 Artificial Intelligence Institute, Seoul National University

* Core contributor.
♢ Equal advising.
† Project lead.
International Conference on Learning Representations (ICLR) 2026
(Oral Presentation)

Abstract

Protein interaction modeling is central to protein design, which has been transformed by machine learning with applications in drug discovery and beyond. In this landscape, structure-based de novo binder design is cast as either conditional generative modeling or sequence optimization via structure predictors ("hallucination"). We argue that this is a false dichotomy and propose Proteina-Complexa, a novel fully atomistic binder generation method unifying both paradigms. We extend recent flow-based latent protein generation architectures and leverage the domain-domain interactions of monomeric computationally predicted protein structures to construct Teddymer, a new large-scale dataset of synthetic binder-target pairs for pretraining. Combined with high-quality experimental multimers, this enables training a strong base model. We then perform inference-time optimization with this generative prior, unifying the strengths of previously distinct generative and hallucination methods. Proteina-Complexa sets a new state of the art in computational binder design benchmarks: it delivers markedly higher in-silico success rates than existing generative approaches, and our novel test-time optimization strategies greatly outperform previous hallucination methods under normalized compute budgets. We also demonstrate interface hydrogen bond optimization, fold class-guided binder generation, and extensions to small molecule targets and enzyme design tasks, again surpassing prior methods. Code, models and new data will be publicly released.




Proteina-Complexa Model Overview

Architecture and Target Conditioning. Proteina-Complexa builds on La-Proteina's partially latent flow matching framework and extends it to conditional binder design. Generation is achieved by iterative denoising (see animation). Key design choices:

  • Continuous latent protein representation. An autoencoder encodes fully atomistic proteins into alpha-carbon backbone coordinates paired with continuous per-residue latent variables that capture amino acid identity and side-chain geometry—and decodes them back. This avoids discretization artifacts and enables accurate atomistic modeling.
  • Target-conditioned denoising. Only the flow model's denoiser conditions on the target; the autoencoder remains shared across all target types. Target features (Atom37 coordinates, amino acid identities, hotspot tokens) are embedded and jointly processed with the binder's noisy representation through a transformer with pair-biased attention.
  • Joint sequence and structure output without redesign. The model generates both sequence and fully atomistic structure simultaneously. Unlike prior approaches, generated sequences are used directly; no separate redesign step is required.

See Figure 1 and our paper for full details.

Figure 1. Proteina-Complexa's architecture. A frozen autoencoder (top) encodes fully atomistic proteins into a partially latent representation (alpha-carbon coordinates + continuous per-residue latents) and decodes them back. The target-conditioned denoiser (bottom) concatenates embedded target features (Atom37 coordinates, amino acid identity, hotspot tokens) with the binder's noisy latent and backbone embeddings, processing them jointly through multi-head pair-biased attention layers.

Test-Time Compute Scaling. Current binder design methods either rely purely on generation (training-time optimization) or purely on hallucination (inference-time optimization without a generative prior). Proteina-Complexa unifies both paradigms—to our knowledge, a first in structure-based binder design. We steer the generative denoising process using rewards from structure prediction confidence (ipAE) and interface hydrogen bond energies (see Figure 2):

  • Best-of-N sampling — scale compute by growing the candidate pool.
  • Beam search — maintain and prune parallel denoising trajectories by reward.
  • Feynman–Kac steering — importance sampling toward the reward-tilted distribution.
  • Monte Carlo tree search — explore the denoising trajectory tree, balancing exploration and exploitation.
  • Generate & Hallucinate — initialize hallucination refinement (e.g., BindCraft) from generative model output rather than from scratch.

This latent generative search framework is general and can incorporate many different objectives; reward differentiability is not necessary. The adjacent animation shows a binder being iteratively refined to increase interface hydrogen bonds. All strategies leverage Proteina-Complexa's fast, fully transformer-based architecture, making repeated rollouts computationally feasible. Scaling inference-time compute via extended search during the generative denoising process allows the model to produce candidates with high scores even for difficult targets.

Figure 2. Proteina-Complexa's generation and inference-time optimization pipeline. (Top) Target-conditioned generation: from noise around the target, the model iteratively denoises a partially latent binder, decodes it, and produces a fully atomistic binder-target complex. (Bottom) Beam search for inference-time optimization: multiple stochastic denoising trajectories (colored paths) are maintained in parallel. At regular intervals, each candidate is rolled out to completion—decoded, co-folded with the target, and scored. Promising trajectories (green checks) are retained while low-scoring ones (red crosses) are pruned, and new branches are launched from the surviving candidates.




The Teddymer Dataset

Binder design requires paired binder-target data, yet experimentally resolved multimers in the PDB are scarce. We exploit the fact that domain-domain interactions within AlphaFold Database (AFDB) monomers closely resemble chain-chain interactions in real multimers. Using TED structural domain annotations, we split AFDB monomers into domains and assemble synthetic dimers: 47M AFDB50 structures → 10M dimers (filtered by proximity and CATH annotations) → 3.5M Foldseek clusters → quality-filtered training set. The resulting Teddymer dataset is an order of magnitude larger than the PDB. Proteina-Complexa trains in stages across four datasets: AFDB monomers, Teddymer dimers, PDB multimers, and PLINDER protein-ligand pairs (Figure 3).

Figure 3. The Teddymer dataset. (Left) Representative Teddymer dimer, constructed by splitting an AFDB monomer into its structural domains (colored chains). The zoom-in highlights interface hydrogen bonds, illustrating that domain-domain interfaces exhibit realistic biophysical interactions. (Right) Overview of the filtered training datasets used by Proteina-Complexa.




Visualizations of Generated Binders

Protein Binders. Proteina-Complexa can generate in-silico mini-binder candidates against single-chain and multi-chain targets. Below, the target is shown with transparent surface and interface hydrogen bonds are highlighted in red. Quantitative evaluations in performance section.

Target: H1

Target: PDL1

Target: Claudin1



Small Molecule Binders. Our model can also design binders to bind small molecules, as shown in the examples below. Generated binders use purple and gold color for alpha helices and beta sheets, respectively. Note that all depicted samples in this section fulfill in-silico success criteria (see our paper for details). Quantitative evaluations below.

Target: FAD

Target: IAI

Target: OQO



Enzyme Design. Proteina-Complexa can also tackle enzyme design, following the Atomic Motif Enzyme (AME) benchmark. An atomistic motif of the enzyme active site is provided together with the substrate molecule, and the model must design a protein that faithfully reconstructs the catalytic residues while accommodating the ligand without steric clashes. See example in Figure 4.

Figure 4. AME task M0157 (glyoxalase II, PDB: 1QH5). Two structurally diverse designs generated by Proteina-Complexa for a challenging 6-residue-island active site. This zinc-dependent metalloenzyme requires precise placement of multiple histidine and aspartate residues coordinating the catalytic zinc ions (blue sphere). Left and right show two independent generations with distinct overall folds, both faithfully reconstructing the full catalytic geometry (given side chain structures are shown as thick red sticks). The zoom-ins reveal the reconstructed active site residues surrounding the zinc center and the bound glutathione-derived substrate, confirming accurate motif reconstruction.



Interface Hydrogen Bond Optimization. Proteina-Complexa's inference-time optimization framework can include interface hydrogen bond energies alongside structure prediction rewards, allowing us to generate binders with enhanced interface hydrogen bonding and extended interaction surfaces. This highlights the generality of our test-time scaling approach: structure-based and physical energy-based rewards can be jointly optimized—a capability not available in prior hallucination methods, which only considered folding model scores. See Figure 5.

Figure 5. Interface hydrogen bond optimization (TrkA target). (a) A binder generated with hydrogen bond energy optimization forms extended interface with 15 hydrogen bonds (red, zoom-in), creating dense network of biophysical interactions across a large contact area. (b) A binder generated without hydrogen bond optimization still passes in-silico success criteria but is smaller, with only 1 interface hydrogen bond.



Fold Class-Conditioned Binder Generation. Previous protein generators often produce primarily alpha-helical outputs. By conditioning Proteina-Complexa on fold class labels, we can explicitly control the secondary structure composition of generated binders. This enables generating structurally diverse binders on demand, providing important control. See Figure 6.

Figure 6. Fold class-conditioned binder generation (IFNAR2 target). Left: mainly alpha-helical binder (purple helices). Center: mainly beta-sheet binder (gold strands). Right: mixed alpha-beta binder combining both secondary structure elements.




In silico Performance and Benchmarking

Generative Base Model. We first evaluate Proteina-Complexa's generative model without test-time optimization against publicly available generative baselines. For each method and target, we generate 200 binders and assess them using established in-silico success criteria based on structure prediction model confidence and alignment scores (see our paper for details), how often a method wins across targets, per-sample generation time, and novelty against PDB. Proteina-Complexa significantly outperforms all baselines on both protein and small molecule targets—even when using its own co-generated sequences directly, without ProteinMPNN-based redesign (Table 1, Table 2).

Table 1. Protein target benchmarking (19 targets, 200 samples each). Self: model-generated sequences. MPNN-FI: ProteinMPNN redesign with fixed interface. MPNN: full backbone redesign. Best in green.

Model # Unique Successes ↑ # Times Best ↑ Time [s] ↓ Novelty ↓
SelfMPNN-FIMPNN SelfMPNN-FIMPNN
RFDiffusion ——4.68 ——3 70.80.87
Protpardelle-1c ——0.73 ——0 8.130.77
APM 0.311.523.15 101 73.10.86
Complexa (ours) 9.1013.614.4 141414 15.60.80

Table 2. Small molecule target benchmarking (4 targets, 200 samples each). Unique successes per molecule. RFDiffusion-AllAtom uses LigandMPNN; Proteina-Complexa uses self-generated sequences.

Model # Unique Successes ↑ Time [s] ↓ Novelty ↓
SAMOQOFADIAI
RFDiffusion-AllAtom 2358 87.40.72
Complexa (ours) 1061719 13.50.71

Inference-Time Compute Scaling. We compare Proteina-Complexa's test-time scaling methods against hallucination baselines (BindCraft, BoltzDesign, AlphaDesign), plotting unique success rate as a function of compute. For easy protein targets, simple best-of-N sampling already outperforms all baselines; for hard targets, structured search (beam search, FKS, MCTS) is required. Across the board, hallucination methods perform poorly under matched compute budgets, while Proteina-Complexa's approaches consistently lead by a large margin. The same pattern holds for small molecule targets, where Proteina-Complexa far outperforms BoltzDesign (Figure 7, Figure 8).

Figure 7. Inference-time compute scaling for protein targets. Unique success rate vs. optimization time (GPU hours) for easy targets (left) and hard targets (right). Proteina-Complexa's search methods (colored curves) consistently outperform hallucination baselines (BindCraft, BoltzDesign, AlphaDesign) under normalized compute budgets.

Figure 8. Inference-time compute scaling for small molecule targets. Unique success rate vs. optimization time, averaged over four molecule targets. Proteina-Complexa's methods again substantially outperform BoltzDesign, the only available hallucination baseline for small molecules.

Enzyme Design (AME Benchmark). We evaluate Proteina-Complexa on the Atomic Motif Enzyme (AME) benchmark, where the model must design a protein that faithfully reconstructs a given active-site motif while accommodating the substrate molecule. The benchmark comprises 41 tasks with 1–7 catalytic residue islands of increasing difficulty. Proteina-Complexa significantly outperforms RFDiffusion2 on nearly all tasks, both with self-generated sequences and LigandMPNN-redesigned sequences (Figure 9).

Figure 9. AME enzyme design benchmark results. Number of unique successes per task (41 tasks, 100 samples each) for Proteina-Complexa vs. RFDiffusion2, comparing self-generated sequences, single LigandMPNN redesign, and best-of-8 LigandMPNN redesigns. Proteina-Complexa outperforms RFDiffusion2 on the vast majority of tasks across all evaluation settings.



2. Experimental Validation of Proteina-Complexa

Latent Generative Search unlocks de novo Design of Untapped Biomolecular Interactions at Scale

1 NVIDIA    2 University of Oxford    3 University of Cambridge    4 Manifold Bio    5 Viva Biotech    6 Duke University    7 LMU Munich    8 Leipzig University    9 Novo Nordisk    10 Seoul National University    11 AITHYRA   

* Equal contribution.
Abstract

De novo protein design has advanced rapidly, yet designing binders to polar, solvent-exposed epitopes and small, flexible ligands remains challenging. Such hydrated surfaces and flexible molecules, including carbohydrates, provide few of the hydrophobic contacts favoured by current methods and have largely resisted de novo binders. To address this challenge, here we introduce latent generative search for binder design, a novel framework that uses reward-guided search at inference time to steer the Proteina-Complexa generative model. The model codesigns sequence and structure—generating them together in a continuous latent space—and thereby removes the inverse-folding step on which current methods rely. In a screen of more than one million designs by multiplexed phage display, latent generative search produced more validated binders than every other method tested, its codesigned sequences surpassing post hoc redesign. It delivered high-affinity binders across therapeutic receptors, a viral attachment protein and intracellular signalling targets. Our approach also accessed previously untapped biology, generating the first de novo proteins that bind a free carbohydrate, including one that discriminates between blood-group antigens—a polar, flexible target class beyond the reach of current design methods.


This large-scale experimental validation effort of Proteina-Complexa is a cross-institutional collaboration led by NVIDIA with Manifold Bio, Viva Biotech, Novo Nordisk, Duke University, Cambridge University, LMU Munich and Leipzig University. Here, we provide highlights; see the preprint for details.




The Latent Generative Search Framework

Our campaigns use the latent generative search framework: Proteina-Complexa codesigns binder sequence and structure in a continuous latent space, and reward-guided search steers generation at inference time. There is no separate inverse folding step (Part 1 above describes the model and the search algorithms in detail). The conditioning input, reward definition, and scoring models change from one target class to the next (Figure 10).

Figure 10. Latent generative search for de novo protein binder design. (a) The conditional codesign generator, Proteina-Complexa, generates binder structure and sequence together using a partially latent representation. A target-conditioned denoiser network iteratively transforms random noise into backbone coordinates and latent variables, after which a decoder maps the codesigned latent binder representation to a fully atomistic output. An auxiliary encoder is used during training of the variational autoencoder. (b) Generative sampling trajectories are guided towards high reward regions of the solution space according to task-relevant scoring models in an inference-time search process. Low-scoring trajectories are terminated, whereas promising trajectories are expanded and diversified, for example through beam search. (c) The complete latent generative search protocol alternates between generative denoising rollouts, reward-based trajectory evaluation, and the subsequent selection and expansion of promising trajectories. The resulting candidates may undergo a final filtering stage using additional scoring functions. The same framework supports diverse design tasks—including monomers, mini-protein binders, peptide binders, and binders to carbohydrate targets—by adapting the conditioning input, reward definition, and scoring models. (d) Confidence metrics from structure prediction models, including ipAE and pLDDT, can serve as rewards during search or as criteria for final filtering. (e) Interface hydrogen bonding scores and molecular force fields provide complementary physics-based scoring functions.




Massive-Scale Benchmark: Testing 1 Million Binders for 127 Targets

We conducted a massive-scale multiplexed phage display campaign with all-to-all binding readout: 1,015,870 designed sequences screened against 127 targets in a single experiment, including diverse, novel, and challenging ones.

Coverage of a 127-target panel. Screening Proteina-Complexa designs across all targets and diverse epitopes reveals broad coverage, with binding strongly enriched toward the intended targets (Figure 11, Table 3). Proteina-Complexa produces on-target hits for 115 of 127 targets (91%); 92 (72%) yielded target-specific binders and 78 yielded poly-specific binders engaging two to four targets. Off-target binding is nonetheless pervasive: every panel target was specifically bound by at least one sequence designed against a different target. These results indicate broad success of Proteina-Complexa across a diverse set of targets. See our preprint for detailed analyses of cross-reactivity, hotspot conditioning effects, and binding specificity patterns.

Figure 11. Proteina-Complexa results across the 127-target panel. On-target specific hit count for each target, ranked; 92 of 127 targets were solved, meaning at least one on-target specific hit.

Table 3. Proteina-Complexa on-target coverage of the 127-target panel. A target counts as solved if at least one binder designed against it bound it. The three categories are not mutually exclusive.

Binders counted Targets solved Share of panel
All 115 / 127 91%
Specific — bound exactly one target 92 / 127 72%
Poly — bound two to four targets 78 / 127 61%

Method comparison: end-to-end codesign outperforms all baselines. Separately, for one top-ranked in silico hotspot per target, we also sampled and evaluated contemporary open baselines—RFDiffusion, RFDiffusion3, BindCraft, and BoltzGen—each given an approximately equal compute budget and testing the same number of selected candidates from each method's generated pool. This compute-matched comparison covers the 83 targets where at least one method produced an on-target hit (Figure 12):

These results establish Proteina-Complexa as a state-of-the-art openly available method for de novo binder design.

Figure 12. On-target specific hit rate (%) across design methods and sequence-design strategies for sequence sets generated by different methods under matched compute (83 targets). Methods are grouped by re-design strategy: no re-design (native model output), partial re-design (fixed interface), and full re-design (dedicated inverse folding model or ProteinMPNN sequences). Each bar is split into monospecific, poly-specific and poly-reactive fractions; the printed percentage is the total, while the solid monospecific portion corresponds to the target-specific hit rates quoted above.




Picomolar PDGFR Binders

Beyond the large-scale 127-target screen, we tested Proteina-Complexa on individual targets with more careful candidate selection and filtering. PDGFR is a polar receptor previously targeted with specialized approaches such as beta-strand pairing due to its challenging surface. We generated 9,000 candidates across multiple hotspot combinations, applied a two-stage filtering pipeline (physicochemical properties, confidence metrics, and monomer structure validation), and selected 192 designs for experimental testing by surface plasmon resonance (SPR).




Binders against ActRIIA blocking Myostatin Signaling in Cells

Activin receptor type IIA (ActRIIA) is a high-affinity receptor for myostatin, activins, and GDF11 that suppresses skeletal muscle growth via Smad2/3 signaling. Blocking ActRIIA is an attractive strategy against muscle wasting—relevant to cancer cachexia, sarcopenia, and lean mass loss during GLP-1 agonist therapies. We designed de novo mini-binders targeting ActRIIA's protein-protein interaction interface (the interactive viewer and Figure 13 show the top 5 binders with functional downstream effects):

Figure 13. De novo binders to the muscle-wasting receptor ActRIIA. (a) Rationale for ActRIIA blockade during semaglutide-induced weight loss, and schematic of binder-mediated inhibition of GDF8/activin A signaling through Smad2/3. (b) Predicted structures of five representative designs (#51, #104, #150, #17 and #101; purple) bound to ActRIIA (grey); the binders share a helical binding mode at low pairwise sequence identity (6–18%). (c) SPR sensorgrams (analyte concentration series, legends in nM) with fitted KD (36–390 nM; tightest design 36 nM). (d) Inhibition of GDF8 (myostatin)-induced Smad2/3 luciferase signaling in cells, shown as GDF8-driven luminescence versus binder concentration, with fitted IC50 (169–1,532 nM).




Designing Peptides and Minibinders against Kinase Targets
PAK1 binders
CK1δ binders

Protein kinases govern nearly all biological signaling through phosphorylation of serine, threonine, and tyrosine residues, yet genetically encoded tools for interrogating their catalytic activity remain limited. We designed de novo binders targeting the catalytic domains of two kinases, PAK1 and CK1δ, spanning two distinct size regimes: conventional mini-protein binders for PAK1 and short peptide binders for CK1δ (see interactive structures and Figure 14):

Figure 14. De novo peptide and mini-protein kinase binders. The left column concerns PAK1 mini-proteins (a, c, e, g, i) and the right column CK1δ peptides (b, d, f, h, j). (a) Signaling context for PAK1, a Rho-family GTPase effector, and the 49–74-aa mini-protein design regime. (b) Signaling context for CK1δ, a constitutive Ser/Thr kinase, and the <31-aa peptide design regime. (c) Split-neomycin-resistance complementation assay for PAK1 mini-proteins, read out by next-generation sequencing (NGS). (d) Streptavidin-bead pull-down assay for biotinylated CK1δ peptides. (e) Mammalian validation of PAK1 binders by representative anti-Myc Western blot and fold enrichment relative to empty vector (mean ± SD, n=3). (f) CK1δ pull-down blots and binding enrichment for CK1–CK18 relative to scrambled control (mean ± SD, n=3–4); 9 of 18 are significant (p<0.05). (g) AlphaFold2-Multimer predicted PAK1 complexes for Pk3–Pk6 (PAK1 surface, binder cartoon). (h) Predicted CK3, CK9 and CK16 peptide–CK1δ complexes. (i) Split-NeoR NGS enrichment profile for the 50 PAK1 designs, ranked by fold change, with the Gaussian-mixture decision boundary and Pk3–Pk6 indicated. (j) Microscale thermophoresis dose–response curves for CK3, CK9 and CK16 (KD = 6.3, 10.8 and 81.1 µM, respectively; mean ± SEM, n=3–4).




De Novo Design and Binder Re-Engineering for the Nipah Virus

De novo binder engaging NiV-G. Hotspot residues in red. Side chains shown for binder and hotspots. Drag to rotate.

The Nipah virus attachment glycoprotein (NiV-G) is a six-bladed beta-propeller that mediates host cell entry by engaging the ephrin-B2 and ephrin-B3 receptors. Its receptor-binding site is a critical neutralization epitope, but the pocket is recessed and partially occluded, making it a challenging design target. As part of the Adaptyv binder competition, we tested Proteina-Complexa in two complementary modes and successfully hit this difficult epitope in both (see interactive structure of the de novo design in the adjacent viewer; quantitative results in Figure 15):

Figure 15. De novo design and binder re-engineering for NiV-G. (a) Generated de novo binder; zoom-ins into binding pocket and interface hydrogen bonding (KD = 56 nM). (b) Re-engineered binder via diffuse-denoise protocol: an existing scaffold is partially noised and re-generated. (c) SPR sensorgrams for all binders, with equilibrium dissociation constants (KD) annotated: the de novo binder at 56 nM and the five re-engineered binders NI1–NI5 at 14, 5.4, 3.5, 7 and 21 nM.




De Novo Binder Design for Carbohydrates

Carbohydrates are small, densely polar, and present hydroxyl-rich surfaces with no hydrophobic character—no computational method had previously designed a protein that binds a free carbohydrate. We targeted the type II A pentasaccharide, the blood group A antigen central to ABO transfusion and transplant compatibility, which differs from the B antigen only in its terminal sugar (N-acetylgalactosamine rather than galactose). See Figure 16 and a confirmed hit in the adjacent animation:

To our knowledge, these are the first de novo designed proteins that bind a free carbohydrate, including one that discriminates between blood group antigens.

Figure 16. De novo binders that recognise a blood group carbohydrate antigen. (a) Principle of ABO blood group compatibility and schematic structures of the A, B and O(H) antigens; the A antigen terminates in N-acetylgalactosamine (GalNAc), the B antigen in galactose (Gal). (b) Fluorescence polarization (FP) screen of designs 1–23 against fluorescently labelled type II A pentasaccharide; 15 designs pass the +10 mP threshold. (c) Excess unlabelled type II A competes design 22 binding back to baseline. (d) Biolayer interferometry (BLI) trace of design 22 against immobilised biotinylated type II A glycan. (e) Predicted glycan-bound structures of designs 3, 4, 9, 22 and 15. (f) Concentration-dependent binding of the same five designs, measured by FP (designs 3, 4, 9 and 22) or BLI equilibrium response (design 15), with fitted KD; the design 3 estimate lies above the measured concentration range. (g) Agglutination assay in which His6-tagged binders are immobilised on Ni-NTA beads to confer multivalency before incubation with red blood cells. (h) Type A2 and type B red blood cell agglutination relative to blank-bead controls (N.D., no agglutination above the negative control). Design 3 agglutinates A2 but not B. (i) BLI of immobilised design 15 against a type II A pentasaccharide–DNA conjugate. (j) Circular dichroism thermal shift of design 15 with and without type B pentasaccharide (106.4 °C → 114.0 °C). (k) FP competition for design 15: 30 mM D-galactose does not compete, whereas 250 µM L-fucose abolishes the signal.




Designed Monomers and a Crystal Structure

2.95 Å crystal structure of the 800-residue design (purple) overlaid on its design model (wheat). Drag to rotate.

Binder design rests on a model that generates designable backbones and realistic sequences. We tested this directly on monomers, asking whether Proteina-Complexa's codesigned sequences express, fold, and stay soluble as well as sequences assigned afterwards by ProteinMPNN (interactive overlay and Figure 17):

Figure 17. Designed monomeric proteins. (a) Outcomes for 48 designed monomers (100–800 aa): soluble expression, thermostable folding and yield, comparing codesigned (Proteina-Complexa) and ProteinMPNN-designed sequences. (b) Expression yield comparison between Proteina-Complexa codesigned and ProteinMPNN-designed sequences for four shared backbones. (c) Amino-acid composition of codesigned sequences (Proteina-Complexa, low and high sampling temperature) versus ProteinMPNN-designed sequences and a UniProt reference. (d) Size-exclusion chromatography (SEC) profile overlaid with the molecular weight determined by multi-angle light scattering (MALS) for a representative 800-residue design. (e) Circular dichroism (CD) spectra at 20 °C and 95 °C for the same design. (f) Crystal structure of the same design (purple) overlaid on its computational model (wheat). The structure was determined at 2.95 Å resolution and resolved 794 of 800 residues; only loop residues 311–316 were unresolved. The backbone RMSD between the resolved experimental structure and the design model is 2.08 Å.

Citations

Proteina-Complexa Model:

@inproceedings{didi2026proteinacomplexa,
  title        = {Scaling Atomistic Protein Binder Design with Generative Pretraining and Test-Time Compute},
  author       = {Kieran Didi and Zuobai Zhang and Guoqing Zhou and Danny Reidenbach and Zhonglin Cao and Sooyoung Cha and Tomas Geffner and Christian Dallago and Jian Tang and Michael M. Bronstein and Martin Steinegger and Emine Kucukbenli and Arash Vahdat and Karsten Kreis},
  booktitle    = {The Fourteenth International Conference on Learning Representations (ICLR)},
  year         = {2026}
}

Experimental Validation:

@article{didi2026latent,
  title        = {Latent generative search unlocks de novo design of untapped biomolecular interactions at scale},
  author       = {Didi, Kieran and Reidenbach, Danny and Penner, Matthew and Ravichandran, Supriya and Case, Marshall and Nichols, Mike and Swanson, Erik and Reis, Alex and Prescott, Maggie and Qian, Yue and Qian, Dongming and Yang, Jingjing and Li, Weiji and Li, Le and Shonai, Daichi and Gay, Sean and Basu Mallik, Bhoomika and Chim, Ho Yeung and Chen, Liurong and Atienza Juanatey, Miguel and Klein, Hubert and Rieger, Dominic and Schlegel, Phillip and Macintyre, Anna U. and Secor, Maxim and Granata, Daniele and Cha, Sooyoung and Cao, Zhonglin and Zhou, Guoqing and Geffner, Tomas and Chen, Xi and Livne, Micha and Zhang, Zuobai and Zhang, Tianjing and Gion, Kyle and Bronstein, Michael M. and Steinegger, Martin and Deibler, Kristine and Soderling, Scott and Schoeder, Clara T. and Khmelinskaia, Alena and Hollfelder, Florian and Dallago, Christian and Kucukbenli, Emine and Vahdat, Arash and Ogden, Pierce and Kreis, Karsten},
  year         = {2026},
  doi          = {10.64898/2026.09.12.751118},
  URL          = {https://www.biorxiv.org/content/early/2026/09/18/2026.09.12.751118},
  journal      = {bioRxiv}
}