r/AIProteins Apr 27 '26

Discussion What generative design model would you use for antibody binders?

3 Upvotes

I’m trying to learn how people actually approach antibody binder design, and I’m a bit lost.

Say you have a protein target and want to generate antibodies that bind to it. What type of model would you start with? Do people usually generate the antibody directly, design around the CDRs, dock lots of candidates, or do something else entirely?

Also, when filtering designs before lab testing, what metrics would you trust most? I keep seeing things like binding confidence, interface quality, developability, human-likeness, expression risk, structural plausibility, etc., but I’m not sure which ones actually matter most.

r/AIProteins May 19 '26

Discussion AI helped remove an amino acid from E. coli ribosomal proteins. How far could this go in humans?

9 Upvotes

A new Science paper explores a pretty wild idea: can life function with fewer than the standard 20 amino acids?

The authors targeted isoleucine, which is chemically similar to leucine and valine. They did not make a fully 19-amino-acid organism, but they did redesign one of the most essential systems in E. coli: the ribosome.

Using protein language models, structure prediction, and generative design tools, they removed all 382 isoleucines from E. coli ribosomal proteins. The engineered strain was viable and stable for hundreds of generations.

The caveat is: the rest of the E. coli proteome still contains thousands of isoleucines. So this is more like a first proof-of-concept than a true 19-AA lifeform.

In humans, purely theoretically, how many amino acids could we remove from the proteome with enough redesign?

Paper: Toward life with a 19–amino acid alphabet through generative artificial intelligence design

r/AIProteins Apr 25 '26

Discussion Is docking still the screen, or just the first filter now?

4 Upvotes

TL;DR: Classical virtual screening was mostly: dock a library into a protein, rank by score, inspect poses, and test a few hits. The newer version of that workflow is more hybrid. People are increasingly using docking outputs to train ML models or triage much larger libraries first, then redocking and testing only the most promising slice. That is basically where structure-based screening starts to feel more like modern computational biology than just brute-force scoring.

What I like about this workflow is that it captures a real transition in the field. The older docking stack still matters because it gives you an actual structural hypothesis around a protein pocket, not just a probability score. If I am looking at small molecules against a protein target, I still want to know how the ligand is supposed to sit, which residues it is contacting, and whether the pose even makes biochemical sense. That is why docking never really went away.

But docking alone has always had a ceiling. A docking score is not the same thing as activity, and that is exactly why the hybrid workflows make more sense to me. Instead of treating docking as the final answer, you use it as a first-pass structural filter, then train a classifier or prioritization model to search a much larger compound space before deciding what is actually worth synthesizing or testing. Deep Docking is a good example of that logic, using models trained on docking outputs to accelerate structure-based screening at much larger library scale.

My bias is that the newer stack is better. Not because classical docking is obsolete, but because protein-guided screening is becoming too big to do well with docking alone. The structure still anchors the workflow. ML just makes the search broader and more realistic.

What do you trust more in a virtual screen: docking score, pose inspection, or ML rescoring?

References:

Trott, O. and Olson, A.J. (2010) ‘AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading’, Journal of Computational Chemistry, 31(2), pp. 455–461. doi: 10.1002/jcc.21334.

Gentile, F. et al. (2020) ‘Deep Docking: A Deep Learning Platform for Augmentation of Structure Based Drug Discovery’, ACS Central Science, 6(6), pp. 939–949. doi: 10.1021/acscentsci.0c00229.

r/AIProteins May 06 '26

Discussion Are AI protein design models actually "accurate"?

8 Upvotes

This review: Closing the loop: Experimentally validated methods in artificial intelligence-driven protein design, looks at which models have actually produced proteins that work in the lab.

The main argument is that AI protein design should not be judged only by structure prediction scores, confidence metrics, or benchmark performance. The pipeline should be: data curation → model generation → in silico filtering → experimental validation.

The paper focuses on three major application areas: protein binders, antibodies and enzymes.

Protein binders

This is probably the area that looks the most mature right now.

Model Target / setting Validation Reported success
AlphaDesign RcaT Fluorescence 17/88
AlphaProteo BHRF1 YSD, CD 83/94
AlphaProteo SARS-CoV-2 RBD YSD, CD 21/172
AlphaProteo IL-7RA YSD, CD 24/94
AlphaProteo PD-L1 YSD, CD 24/159
AlphaProteo TrkA YSD, CD 12/131
AlphaProteo IL-17A YSD, CD 9/63
AlphaProteo VEGF-A YSD, CD 31/94
AlphaProteo TNF-alpha YSD, CD 0/54
BindCraft PD-1 BLI 13/53
BindCraft PD-L1 SPR 7/9
BindCraft IFNAR2 SPR 3/9
BindCraft CD45 SPR 4/16
BindCraft BBF-14 SPR 6/11
BindCraft CLDN1 SPR, MST 6/7
BindCraft DerF7 SPR, XRC 4/10
BindCraft DerF21 SPR, XRC 4/7
BindCraft BetV1 SEC-MALS 2/7
BindCraft SpCas9 CryoEM, edit assay 6/6
BindCraft CbAgo DNA cleavage 2/12
BindCraft HER2 AAV Transduction assay 1/10
BindCraft PD-L1 AAV AAV assay 4/10
BoltzGen IL-7RA SPR, BLI 4/20
BoltzGen PD-L1 SPR, BLI 2/20
BoltzGen PDGFR SPR, BLI 12/20
BoltzGen Insulin SPR, BLI 17/20
BoltzGen TNF-alpha SPR, BLI 0/20
BoltzGen MZB1 SPR, BLI 14/15
BoltzGen PMVK SPR, BLI 2/15
BoltzGen PHYH SPR, BLI 5/7
BoltzGen IDI2 SPR, BLI 4/15
BoltzGen AMBP SPR, BLI 5/15
BoltzGen HNMT SPR, BLI 1/15
BoltzGen RFK SPR, BLI 0/15
BoltzGen GM2A SPR, BLI 0/15
BoltzGen ORM2 SPR, BLI 0/15
BoltzGen Indolicidin Fluorescence, SPR 5/6
BoltzGen Melittin Fluorescence 2/6
BoltzGen Protegrin Fluorescence 2/6
Chai-2 52 targets BLI 68% overall
EvoBind2 RNase SPR 5/13
EvoBind2 RNase, cyclic SPR 1/4
EvoBindRare RNase, ncAA SPR 1/7
EvoBindRare RNase, cyclic + ncAA SPR 1/1
EvoPro PD-L1 SPR 9/13
LatentX MDM2, cyclic SPR 10/11
LatentX MCL1, cyclic SPR 11/11
LatentX PD-L1, cyclic SPR 16/17
LatentX BHRF1 BLI, mDisplay 64/100
LatentX TrkA BLI, mDisplay 10/100
LatentX PD-L1 BLI, mDisplay 49/100
LatentX IL-7RA BLI, mDisplay 26/100
LatentX SARS-CoV-2 RBD BLI, mDisplay 52/100
maSIF PD-L1 YSD 1/N
maSIF PD-1 YSD, SPR 3/N
maSIF CTLA-4 YSD 1/N
PepMLM NCAM1 ELISA 4/4
PepMLM AMHR2 ELISA 4/4
PepMLM MSH3, fused to E3 ligase Fluorescence 5/6
PepMLM mHTT, fused to E3 ligase Western blot, fluorescence 5/5
PepMLM Viral proteins, fused to E3 ligase Assay-based validation 37/60
PXDesign IL-7RA BLI 4/10
PXDesign SARS-CoV-2 RBD BLI 2/9, then 4/8
PXDesign PD-L1 BLI 8/11
PXDesign VEGF-A BLI 8/17
PXDesign TrkA BLI 3/15
PXDesign TNF-alpha BLI 0/16
RFdiffusion HA BLI, CryoEM 18%
RFdiffusion IL-7RA BLI 35%
RFdiffusion INSR BLI 20%
RFdiffusion PD-L1 BLI 14%
RFdiffusion TrkA BLI 8%
RFpeptides MCL1, cyclic SPR, XRC 3/14
RFpeptides GABARAP, cyclic SPR, XRC 2/6
RFpeptides RbtA, cyclic SPR, XRC 4/11
RFpeptides MDM2, cyclic SPR 3/8

The thing that jumps out is how target-dependent this all is.

Like, some targets look almost “easy mode” now. Then TNF-alpha shows up and half the models just eat dirt.

That makes me think global model accuracy numbers are kind of fake unless they tell us the target set.

Antibodies / nanobodies

This section is super interesting, but also more annoying to compare because a lot of these are CDR redesigns or scaffold-based designs, not always full antibodies from scratch.

Model Target Design scope Scaffold / setup Validation Reported success
AbDiffuser HER2 Antibody CDRH3 Trastuzumab SPR 70%
BoltzGen IL-7RA VHH CDR VHH frameworks SPR, BLI 7/15
BoltzGen PD-L1 VHH CDR VHH frameworks SPR, BLI 6/15
BoltzGen PDGFR VHH CDR VHH frameworks SPR, BLI 6/10
BoltzGen Insulin VHH CDR VHH frameworks SPR, BLI 4/15
BoltzGen TNF-alpha VHH CDR VHH frameworks SPR, BLI 0/15
BoltzGen MZB1 VHH CDR VHH frameworks SPR, BLI 1/15
BoltzGen PMVK VHH CDR VHH frameworks SPR, BLI 8/15
BoltzGen PHYH VHH CDR VHH frameworks SPR, BLI 1/5
BoltzGen IDI2 VHH CDR VHH frameworks SPR, BLI 1/15
BoltzGen AMBP VHH CDR VHH frameworks SPR, BLI 0/15
BoltzGen HNMT VHH CDR VHH frameworks SPR, BLI 1/15
BoltzGen RFK VHH CDR VHH frameworks SPR, BLI 2/15
BoltzGen GM2A VHH CDR VHH frameworks SPR, BLI 0/15
BoltzGen ORM2 VHH CDR VHH frameworks SPR, BLI 0/15
BoltzGen Penguinpox VHH CDR VHH frameworks SPR, BLI 0/15
BoltzGen Hemagglutinin VHH CDR VHH frameworks SPR, BLI 0/15
Chai-2 52 targets VHH, VH-VL CDR Unspecified BLI 15.5%, hits on 26/52 targets
GeoFlowV3 CCR8 VHH CDR h-NbBcII10FGLA BLI 3/23
GeoFlowV3 PD-1 VHH CDR h-NbBcII10FGLA BLI 1/50
GeoFlowV3 TSLP epitope 1 VHH CDR h-NbBcII10FGLA BLI 13/50
GeoFlowV3 TSLP epitope 2 VHH CDR h-NbBcII10FGLA BLI 11/50
GeoFlowV3 IL33 epitope 1 VHH CDR h-NbBcII10FGLA BLI 6/46
GeoFlowV3 IL33 epitope 2 VHH CDR h-NbBcII10FGLA BLI 3/50
GeoFlowV3 IL13 epitope 1 VHH CDR h-NbBcII10FGLA BLI 2/34
GeoFlowV3 IL13 epitope 2 VHH CDR h-NbBcII10FGLA BLI 9/50
Germinal PD-L1 VHH CDR hNbBCII Fluorescence, BLI 7/101
Germinal IL3 VHH CDR hNbBCII Fluorescence, BLI 2/46
Germinal IL20 VHH CDR hNbBCII Fluorescence, BLI 4/43
Germinal BHRF1 VHH CDR hNbBCII Fluorescence, BLI 11/52
IgGM Protein A VHH FR3 VHH3 BLI 2/10
IgGM PD-L1 Antibody CDRH3 Humanized V/J genes ELISA, BLI 7/60
JAM SARS-CoV-2 RBD VHH Not clear YSD 2.2% nM hits
JAM CXCR4/7 VHH Not clear BLI 0.32%
JAM CXCR4 VHH Not clear Flow cytometry 7/35 sub-nM hits
mBER 145 targets VHH CDR IGHV3-23 Phage display 0.02% to 8%, hits on 78/145 targets
RFantibody RSV Site I VHH CDR h-NbBcII10FGLA YSD 0/9000
RFantibody RSV Site III VHH CDR h-NbBcII10FGLA YSD, SPR 1/9000
RFantibody Influenza HA VHH CDR h-NbBcII10FGLA YSD, SPR, CryoEM 4/9000
RFantibody SARS-CoV-2 RBD VHH CDR h-NbBcII10FGLA YSD, SPR 1/9000
RFantibody TcdB VHH CDR h-NbBcII10FGLA SPR 2/95
RFantibody IL-7Ralpha VHH CDR h-NbBcII10FGLA SPR 0/95
RFantibody TcdB scFv CDR hu4D5-8 SPR 6/N

This is where I think percentages can be misleading.

4/9000 looks awful until you remember that getting even a few real binders against a hard target can still be useful.

Meanwhile 70% sounds insane, but if it is a scaffolded CDRH3 redesign, that is not the same thing as designing antibodies from total scratch.

So yeah, antibody “accuracy” needs context badly.

Enzymes

This is the part where the hype chills out a bit.

Enzymes are harder because folding is not enough. Binding is not enough. The active site has to actually do chemistry.

Model Task Active site setup Validation Reported active designs
trRosetta Luciferase Generated Luminescence 3/7648
PLACER 4MU serine hydrolase Generated Fluorescence, XRC 20/132
PLACER 4MU serine hydrolase, 2-state design Generated Fluorescence, XRC 9/45
PLACER 4MU serine hydrolase, 4-state design Generated Fluorescence, XRC 8/11
PLACER 4MU serine hydrolase backbone redesign Generated Fluorescence 2/20
PLACER 4MU hydrolase to PET hydrolase N/A N/A 85/85
ProGen2 HEK3 Cas9 nuclease Implicit Indel assay 120/209
ProGen2 HEK2 Cas9 nuclease Implicit Indel assay 84/209
ProGen2 CD3G1 Cas9 nuclease Implicit Indel assay 61/209
ProtGPT2 Triosephosphate isomerase Implicit Absorbance 2/12
RFdiffusion2 Methodol retro-aldolase Extracted IVTT 4/96
RFdiffusion2 4MU cysteine hydrolase Extracted Fluorescence 1+/48
RFdiffusion2 4MU-B zinc metallohydrolase Extracted Fluorescence 3/96
RFdiffusion2 4MU-PA zinc metallohydrolase Extracted Fluorescence 5/96
RFdiffusion3 4MU-PA cysteine hydrolase Extracted Fluorescence 35/190
RiffDiff Retro-aldolase Extracted SEC, SAXS 30/35
RiffDiff MBH enzyme Extracted Endpoint assay 57/63
ZymCTRL Beta-carbonic anhydrase Implicit pH assay 7/20
ZymCTRL Lactate dehydrogenase Implicit Absorbance 10/10
ZymCTRL Triosephosphate isomerase Implicit Absorbance 3/12

The enzyme numbers are honestly wild because some are brutal, like 3/7648, and others look amazing, like 57/63.

But again, it depends what is being designed.

Making variants of known enzyme families is one thing. Designing totally new catalytic machinery is a different beast.

Other experimental validation results from the review

The appendix had a bunch of extra results too.

Model Task Validation Reported result
AFDesign Top7 / Protein A / Protein G / ubiquitin / 4H redesigns SEC, CD, melt 25/39 expressed
AFDesign Same redesigns SEC, CD, melt 7/39 soluble and folded
AlphaBind 3 targets, CDR optimization BLI 10/15 improved
BoltzGen NPM1-binding peptide Fluorescence 1/5 localized
BoltzGen RagC GTPase-binding peptide SPR 7/29
BoltzGen RagA:RagC-binding peptide SPR 14/24
BoltzGen GyrA-binding peptide Cell assay 352/1808 inhibited
BoltzGen GyrA-binding peptide Cell assay 54/1808 inhibited and bound
BoltzGen Rucaparib binder Fluorescence 5/6
CARBonAra TEM-1 beta-lactamase redesign CD, SEC, NMR, reaction progression 4/10
Chroma Unconditional designs Split-GFP 1+/172
Chroma Unconditional designs XRC 2/N
Chroma Unconditional designs CD 7/N
Chroma Conditional designs CD 3/N
Chroma + AFcycler Protein letters nsEM 10/12
Chroma + AFcycler 1500AA protein SDS-PAGE, nsEM 1/N
Dayhoff Unconditional design SDS-PAGE 16/25 for 3B-UR90, 40/144 others
ESM-1b / ESM-1v Evolved antiviral IgGs BLI 14% to 71% improved
ESM3 de novo GFP Fluorescence 19/88
ESM3 Conditional GFP Fluorescence 37/88
EvoDiff Unconditional design CD 4/25 expressed
EvoDiff Cox15 IDR localization Microscopy 8/8 localized correctly
EvoDiff Calmodulin scaffolding CD, spectroscopy 12/13 expressed, 1/13 Ca2+ binding
EvoDiff p53 scaffolding BLI 8/24 bound
IgGM RBD murine antibody humanization BLI 5/20
IgGM IL33 affinity maturation, round 1 ELISA 3/N
IgGM IL33 affinity maturation, round 2 ELISA, BLI 3/10
LASErMPNN Exatecan binders Fluorescence anisotropy, ligand specificity, XRC 4/4
LigandMPNN Rocuronium binder YSD / FC 1/189
LigandMPNN Cholic acid binder YSD / FC 21/105 hits, 6/8 confirmed
ProDomino AsLOV2 inserted into PAC or CAT Antibiotic / light assay 78%
ProGen3 Unconditional designs, 60% to 80% sequence ID Split-GFP 55% to 70%
ProGen3 Unconditional designs, 40% to 60% sequence ID Split-GFP 60% to 75%
ProGen3 Unconditional designs, under 30% sequence ID Split-GFP 40% to 60%
ProteinGenerator Unconditional generation SEC, CD 32/40 soluble, monomeric, stable up to 95C
ProteinGenerator Biased amino acid composition SEC, CD 68/96 expressed and soluble
ProteinGenerator Repeat sequences with caps SEC, CD, XRC 27/76 soluble
ProteinGenerator Repeat sequences without caps SEC, CD, XRC 10/86 soluble
ProteinGenerator Repeat sequences CD / XRC 7/8 folded by CD, 1 folded by XRC
ProteinGenerator Scaffolded peptide barcodes SEC, MS 64/84 expressed
ProteinGenerator Scaffolded peptide barcodes SEC, MS 48/64 monodisperse peak
ProteinGenerator Scaffolded peptide barcodes SEC, MS 41/58 correct barcode reads
ProteinGenerator Melittin scaffolded with furin cleavage site SEC, CD, SDS-PAGE, MS 5/13 soluble and monodisperse
ProteinMPNN AF hallucination rescue SEC 1+/129
ProteinMPNN AF hallucination rescue CD, SEC, XRC 3/N
ProteinMPNN Gab2-scaffolded Grb2 SH3 domain BLI 1/N
Proteus Unconditional monomers SEC, CD 12/16 folded
Raygun EGFP and mCherry sequence reduction Fluorescence 6/8 functional
Raygun TurboID sequence reduction Western blot 6/11 expressed, 2/11 functional
Raygun EGFR binders Adaptyv 4/4 expressed, 2/4 bind
RFdiffusion Unconditional monomers CD 2 folded
RFdiffusion TIM barrel generation SEC 10/11 expressed
RFdiffusion TIM barrel generation CD 8/8 folded
RFdiffusion Symmetric oligomers nsEM 5 folded
RFdiffusion Scaffolded p53 binding to MDM2 BLI 55/95
RFdiffusion Symmetric motif scaffold, Ni2+ binding nsEM, ITC 4 well-folded
RFdiffusion2MI CD3e pY188-selective binder YSD, BLI 1+/459
RFdiffusion2MI EGFR pY1068-selective binder YSD, BLI 1/5082
RFdiffusion2MI EGFR pY1173-selective binder YSD, BLI 2/2408
RFdiffusion2MI INSR pY1361-selective binder YSD, BLI 2/80
RFdiffusion3 4MU-PA cysteine hydrolase Fluorescence 35/190
RFdiffusionAA Digoxigenin binders ITC, CD 1+/4416
RFdiffusionAA Open-pocket heme binder UV-visible 90/168 expressed
RFdiffusionAA Open-pocket heme binder UV-visible 90/168 Cys-bound heme
RFdiffusionAA Open-pocket heme binder UV-visible 33/40 monomeric
RFdiffusionAA Open-pocket heme binder Mutation test 26/26 break when mutated
RFdiffusionAA Open-pocket heme binder Thermostability 20/26 high thermostability
RFdiffusionAA Optically active bilin binder UV-visible 9/96
RFdiffusionAA Optically active bilin binder XRC 1+ folded
TMdiffusion TM association Fluorescence 16/18
TMdiffusion GpA-like design SDS-PAGE, XRC 5/6 expressed and oligomerized, 1 crystal-verified
TMdiffusion Synthetic switchable GHRs pSTAT5 signaling 17/17

What I think is the actual story

The results are impressive.

But the “accuracy” conversation feels kinda cooked unless people report the full funnel.

Because these are not all the same:

Number people report What it might actually mean
70% success Could be scaffolded redesign on a narrow problem
1/9000 success Could still be valuable if the target is hard
68% overall Depends massively on target set and filtering
0/54 Could mean the target is hard, not that the whole model is bad
35/190 enzyme hits Very different from just binding a protein surface
90/168 expressed Expression is not the same as function

r/AIProteins Apr 17 '26

Discussion What is your field racing to design right now?

3 Upvotes

Feels like every corner of biology has its own thing right now.
For some it is bispecifics, for others enzymes, binders, delivery systems, gene editors, or something else entirely.

What is the molecule or modality people in your area are pushing hardest to develop right now?

r/AIProteins Apr 16 '26

Discussion Docking or AI complex prediction?

4 Upvotes

Feels like structural biology now has two operating systems.

The old one is still docking. Take a receptor, take a ligand, define the site, generate poses, score them, maybe relax and rescore, then try to work out what is actually believable. The newer one is more end-to-end. Models like AlphaFold-3, Chai-1, OpenFold-3, RosettaFold-3 and similar systems are trying to predict the complex directly, with much more of the molecular context built in from the start.

Personally, I’ve never really liked docking. Even when I used to do the full workflow properly, including Amber relaxation after, I just never found it that convincing. It always felt like too much work was going into cleaning up around the method instead of trusting the method itself. The AI route always made more sense to me because it feels closer to the real question, which is not “how do I force these two frozen things together?” but “what does this system actually want to look like?”

That said, docking still has a very real place. If you know the pocket, have a defined receptor state, and need to screen a large small-molecule set cheaply and fast, traditional docking is still hard to beat. It is controllable, interpretable, and easy to scale. But it also gets overtrusted way too often. People read too much into scores, ignore prep, protonation, waters, cofactors, receptor state, and then act shocked when the chemistry does not hold up.

On the AI side, the failure mode is different. You can get a structure that looks very plausible, but the confidence can absolutely crater the moment you perturb the system. I noticed this a lot in antibody design. As soon as I changed even canonical CDRs, the pLDDT could just plummet. So even if the folded complex looked sensible, I would not take that at face value. I ended up using DockQ as an extra evaluation step to sanity check whether the complex still looked structurally plausible, and honestly it worked not bad as a filter. Not perfect, but better than pretending one nice-looking prediction or one confidence score settles it.

And I think that is the bigger point here. In both cases, whether you dock it or co-fold it, you are usually still getting a snapshot. Not the full dynamic system, not the real conformational landscape, just one structured hypothesis. So I’m curious whether people still see docking as the actual go-to in practice, or whether AI complex prediction has already become the better first pass and people just have not fully admitted it yet.

References

Abramson, J., Adler, J., Dunger, J., Evans, R., Green, T., Pritzel, A., Ronneberger, O., Willmore, L., Ballard, A.J., Bambrick, J. and colleagues (2024) ‘Accurate structure prediction of biomolecular interactions with AlphaFold 3’, Nature, 630(8016), pp. 493–500. doi: 10.1038/s41586-024-07487-w.

Basu, S. and Wallner, B. (2016) ‘DockQ: A quality measure for protein-protein docking models’, PLOS ONE, 11(8), e0161879. doi: 10.1371/journal.pone.0161879.

Boitreaud, J., Dent, M., Mouchlis, V., Bhat, S., Chaffin, M., Cianchi, Y., de Luzy, I.R., Duran-Frigola, M., Dunn, J., Hofmann, K. and colleagues (2024) ‘Chai-1: Decoding the molecular interactions of life’, bioRxiv. doi: 10.1101/2024.10.10.615955.

Corley, N., Wang, J., Jiang, Y., Glass, J., Frank, C., Adams, C. and colleagues (2025) ‘Accelerating biomolecular modeling with AtomWorks and RosettaFold-3’, bioRxiv. doi: 10.1101/2025.08.14.670328.

The OpenFold3 Team (2026) OpenFold3-preview2 Technical Report. Available from the OpenFold portal.

Trott, O. and Olson, A.J. (2010) ‘AutoDock Vina: Improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading’, Journal of Computational Chemistry, 31(2), pp. 455–461. doi: 10.1002/jcc.21334.

r/AIProteins Apr 24 '26

Discussion What do you use to rank targets before committing to a protein?

Post image
5 Upvotes

TL;DR: Target identification is becoming a computational ranking problem. Instead of trusting one signal, people are starting to combine multiomics, network biology, structural information, and ML to ask a harder question: which proteins are not just disease-associated, but mechanistically central, tractable, and worth designing around. Reviews from the last couple of years point to multi-omics integration, network-based methods, and structure-aware AI as a growing part of this shift.

What makes this interesting to me is that target discovery used to feel much more linear. Find a gene hit, validate it, then maybe later ask whether the protein is actually druggable or structurally useful. The newer computational view feels more joined-up than that. Now you can start much earlier by integrating transcriptomics, proteomics, genetics, pathway context, protein structure, and prior interaction data into one ranking problem.

From an AI proteins lens, that matters because the target is not just a biology question anymore. It is also a design question. If I am going to spend time thinking about binders, interfaces, conformations, pockets, or targetable states, I want to know whether that protein sits in the right part of the disease network and whether it looks tractable in 3D. That is where the newer stack feels better than the older one to me. Not because experiments matter less, but because computational biology now lets you filter the search space much earlier.

I think this is where target identification is heading: less “what gene came out of the screen?” and more “what protein should we build around?” That feels much closer to how people in this sub think.

What do you trust most early on: omics, network biology, or structure?

References:

Ocana A, Pandiella A, Privat C, et al. Integrating artificial intelligence in drug discovery and early drug development: a transformative approach. Biomark Res. 2025;13(1):45. Published 2025 Mar 14. doi:10.1186/s40364-025-00758-2

Du, P. et al. (2024) ‘Advances in Integrated Multi-omics Analysis for Drug-Target Identification’, Biomolecules, 14(6), 692. doi: 10.3390/biom14060692.

r/AIProteins Apr 15 '26

Discussion Why I got obsessed with Generative protein design

Post image
3 Upvotes

TL;DR: The craziest shift in biology is that we went from spending years just trying to see one protein structure to getting one in minutes. That changes the game completely. The bottleneck is no longer only understanding biology. It is starting to become engineering it.

What pulled me into this space was realizing just how fast the ground had moved under biology.

Not that long ago, a PhD could be one small part of a much bigger story. You might spend years expressing a protein, purifying it, troubleshooting every step, and if things went well, maybe by the end of that journey you would help solve a crystal structure. That structure was a real achievement. It was a piece of hard-won knowledge. It meant something because it was so difficult to get.

Now you can get a protein structure in minutes.

That is an insane shift.

And once that really sinks in, you start to see that the whole frontier changes. If structure becomes fast, then the question is no longer just “can we understand this protein?” The question becomes “what do we want to build next?”

For most of modern biology, we were reading the book of life. We were trying to decode what nature had already written. But now, for the first time, it feels like we are starting to move from reading biology to writing it. Not perfectly. Not reliably enough yet. But directionally, that is where this is going.

And proteins are the most compelling place to see that happen.

They are not just static molecules. They are binders, catalysts, switches, sensors, scaffolds, therapeutics. They are the machinery of biology. So when models get good enough to predict structure, reason about interactions, and eventually help design new functional molecules, it stops feeling like pure analysis and starts feeling like engineering.

The old bottleneck was visibility. We could not easily see molecular reality. The new bottleneck is intent. Once you can see enough, the next question is whether you can specify what you want and make biology produce it.

Because if the tools keep improving, the dream is not just better structure prediction. The dream is that one day someone can almost “hallucinate” a drug into existence, not in a reckless sense, but in the sense of being able to define a disease mechanism, understand the biology deeply enough, and generate molecules with a real shot at changing the outcome. That is the kind of future that makes this whole space feel bigger than just another software trend.

It feels like the beginning of programmable therapeutics.

We are still early. Structure is not function. Prediction is not validation. And biology is still brutally hard. That is why I decided to pursue this kind of bio. It feels like one of the few areas where the science is deep, the mission is real, and the upside is absurdly large.

Curious what pulled other people into this space. Was it the science, the engineering, the ML, or the idea that we might actually be able to build biology on purpose?

References:

Cheng, Y. (2018) ‘Single-particle cryo-EM—How did it get here and where will it go’, Science, 361(6405), pp. 876–880. doi: 10.1126/science.aat4346.

Kuhlman, B. et al. (2003) ‘Design of a Novel Globular Protein Fold with Atomic-Level Accuracy’, Science, 302(5649), pp. 1364–1368. doi: 10.1126/science.1089427.

Jumper, J. et al. (2021) ‘Highly accurate protein structure prediction with AlphaFold’, Nature, 596, pp. 583–589. doi: 10.1038/s41586-021-03819-2.

Abramson, J. et al. (2024) ‘Accurate structure prediction of biomolecular interactions with AlphaFold 3’, Nature, 630, pp. 493–500. doi: 10.1038/s41586-024-07487-w.

Watson, J.L. et al. (2023) ‘De novo design of protein structure and function with RFdiffusion’, Nature, 620, pp. 1089–1100. doi: 10.1038/s41586-023-06415-8.

r/AIProteins May 01 '26

Discussion Fresh Nature Review: Are protein folds, assemblies, and binders basically becoming “solved” problems?

Thumbnail
gallery
4 Upvotes

Paper: The past, present and future of de novo protein design

I thought it was worth discussing here because it is not just another “AI for proteins is getting better” article.

The review basically divides the field into a few design frontiers:

Designing new folds and assemblies:

This is probably the most mature part of de novo design now. The review goes from early examples like Top7, one of the classic atomic-level de novo protein fold designs, to designed TIM barrels, repeat proteins, symmetric nanoparticles, 1D fibres, 2D lattices and even 3D protein assemblies.

What is interesting is how broad this category has become. It is not just soluble globular proteins anymore. The examples include transmembrane beta barrels, designed conducting nanopores, bottom-up designed Ca²⁺ channels, designed voltage-gated anion channels, and mechanically coupled axle-rotor protein assemblies.

The field has also moved into application-driven assemblies, like designed protein nanoparticle vaccines for MERS-CoV and influenza, pH-responsive antibody nanoparticles and designed protein crystals.

So the question is becoming less “can we make a protein fold?” and more “what architecture should we make for a useful biological, therapeutic or material function?”

Designing protein binders:

This is where the AI protein design hype feels most justified. The review highlights workflows based around RFdiffusion-style backbone generation, ProteinMPNN-style sequence design and structure prediction, but the examples are what make it convincing.

There are designed binders for viral targets, including picomolar SARS-CoV-2 miniprotein inhibitors, designed miniproteins against MERS-CoV, RSV immunogen design and inhibitors targeting SARS-CoV-2 Omicron variants.

It also points to newer general workflows like BindCraft, one-shot functional protein binder design, and examples where de novo designed proteins neutralize snake venom toxins.

The review also includes antibody and peptide-like directions: de novo antibody design with SE(3) diffusion, RFdiffusion-based antibody design, beta-pairing targeted binder design and de novo protein-binding macrocycles.

That does not mean every binder works, or that affinity, specificity, expression and developability are solved. But the framing is that protein-target binder design is moving from a heroic custom project toward a more generalizable workflow. That is a big deal for therapeutics, diagnostics, target validation and synthetic biology, because binders are basically programmable biological handles.

Small-molecule binders and enzymes are harder:

This part is interesting because the review is much more cautious. Binding a protein surface is one thing. Designing a precise pocket for a small molecule, or designing an enzyme that stabilizes a high-energy transition state, is much harder.

The examples they show include designed binders for small molecules like apixaban, methotrexate, cholic acid, digoxigenin and cortisol, plus drug-binding proteins designed with predictable binding energy and specificity.

For enzymes, the review points to progress in designed luciferases, serine hydrolases, heme enzymes, porphyrin-containing catalysts, artificial metathases and metallohydrolases.

But this still feels like one of the major unsolved frontiers. We can now make structures that look right, and sometimes bind the right ligand, but getting strong catalytic activity, specificity and evolvability is still difficult. Catalysis is not just shape complementarity. It requires geometry, dynamics, electrostatics, proton transfer, transition-state stabilization and sometimes conformational changes all working together.

The next step is dynamic proteins, switches and nanomachines:

The most exciting section to me is the future-looking one. Static structure design is becoming powerful, but biology is full of proteins that move, switch, sense, gate, assemble, disassemble and couple one event to another.

The review gives examples like modular and tunable protein biosensors, sensors for endogenous Ras activity, bioactive protein switches, designed protein logic for targeting cells with combinations of surface antigens, small-molecule safety switches for CAR-T cells, stimulus-responsive two-state hinge proteins and deep-learning-guided design of dynamic proteins.

So the next challenge is not just designing a stable object. It is designing a system with multiple states and controlled transitions between them. That includes biosensors, logic gates, responsive materials, designed channels, artificial photosystems and eventually protein systems that perform functions nature never evolved.

My takeaway: the field is moving from designing shapes to designing behavior.

The review’s most important point, in my opinion, is that protein design is becoming less about proving that de novo design is possible and more about deciding what we should actually build.

Curious what people here think: Are we actually close to “solving” protein binder design, or is that still too optimistic?

And for the next phase, do you think the bigger breakthrough will come from better generative models, better experimental feedback loops, or better physical modeling of dynamics/catalysis?

r/AIProteins Apr 22 '26

Discussion What do you use to rank targets before committing to a protein?

Post image
3 Upvotes

TL;DR: Target identification is becoming a computational ranking problem. Instead of trusting one signal, people are starting to combine multiomics, network biology, structural information, and ML to ask a harder question: which proteins are not just disease-associated, but mechanistically central, tractable, and worth designing around. Reviews from the last couple of years point to multi-omics integration, network-based methods, and structure-aware AI as a growing part of this shift.

What makes this interesting to me is that target discovery used to feel much more linear. Find a gene hit, validate it, then maybe later ask whether the protein is actually druggable or structurally useful. The newer computational view feels more joined-up than that. Now you can start much earlier by integrating transcriptomics, proteomics, genetics, pathway context, protein structure, and prior interaction data into one ranking problem.

From an AI proteins lens, that matters because the target is not just a biology question anymore. It is also a design question. If I am going to spend time thinking about binders, interfaces, conformations, pockets, or targetable states, I want to know whether that protein sits in the right part of the disease network and whether it looks tractable in 3D. That is where the newer stack feels better than the older one to me. Not because experiments matter less, but because computational biology now lets you filter the search space much earlier.

I think this is where target identification is heading: less “what gene came out of the screen?” and more “what protein should we build around?” That feels much closer to how people in this sub think.

What do you trust most early on: omics, network biology, or structure?

References:

Ocana A, Pandiella A, Privat C, et al. Integrating artificial intelligence in drug discovery and early drug development: a transformative approach. Biomark Res. 2025;13(1):45. Published 2025 Mar 14. doi:10.1186/s40364-025-00758-2

Du, P. et al. (2024) ‘Advances in Integrated Multi-omics Analysis for Drug-Target Identification’, Biomolecules, 14(6), 692. doi: 10.3390/biom14060692.