Back to feed
News Story
APriority76
THE DECODER
1 sources

Anthropic says any lab can now let a language model agent run the whole protein design stack

Anthropic announced that its Claude models can autonomously design small proteins that dock onto target structures in the body, achieving a hit rate of up to 35%, far above the industry average of 10-15%. Claude only steered existing specialized tools, and independent review is still pending.

SynthePulse Insight · AI deep reading

Anthropic Lets Claude Take Over the Entire Protein Design Pipeline: 26.8% Hit Rate, but Independent Validation Still Missing

Version 1 · 1 source

Anthropic reports two experiments where Claude models autonomously ran the entire protein design pipeline, succeeding on 14 of 16 targets with a lab hit rate of 26.8%, far exceeding the typical industry rate of 10-15%. However, independent review has not been conducted, some targets failed completely, and confidence scores did not provide early warnings.

  • Claude successfully designed proteins for 14 of 16 targets, with 354 of 1,320 designs binding in lab validation, a 26.8% hit rate.
  • In single-target mode, Mythos Preview's hit rate rose to 35.1%, but compute budget increased 2.8-fold.
  • For industry benchmark: Anthropic cites public database proteinbase.com, with typical hit rates of 10-15%.
  • Claude did not use proprietary protein models but orchestrated open-source tools like PXDesign, RFdiffusion3, etc., excluding AlphaFold-3 due to licensing restrictions.
  • On the RBX1 target, Claude's best design had a binding strength of 3.9 nM, about 10 times stronger than the public competition winner's 45 nM.
  • Two targets failed completely: MBP had no designs that bound, and BBF-14 had only 3 weak binders, with confidence scores failing to warn.
Open section navigationExperimental Design and Core Results

Experimental Design and Core Results

In two experiments, Anthropic had Claude models take on early drug discovery tasks. The first experiment focused on designing minibinders—small proteins that can lock onto target proteins and alter their function. The models Mythos Preview and Opus 4.8 designed against 16 target proteins, 15 of which produced usable measurement data, and Claude succeeded on 14 targets.

In lab tests, 354 of 1,320 designs actually bound to targets, a 26.8% hit rate; when considering only Claude's top-ranked design per target, the hit rate rose to 49%. In multi-target mode (processing all targets within 48 hours), Mythos Preview and Opus 4.8 achieved hit rates of 26.7% and 22.6%, respectively. In single-target mode, Mythos Preview's hit rate improved to 35.1%, but compute budget increased 2.8-fold, and the authors acknowledge that focus and budget cannot be separated.

As a comparison, Anthropic cites typical hit rates of 10-15% from publicly recorded data in the proteinbase.com database.

Technical Stack and Autonomy

Claude did not use proprietary protein models but installed and ran existing open-source specialized software in the field. Protein backbones came from PXDesign (358 designs), RFdiffusion3 (267), Genie 3 (185), etc.; amino acid sequences were primarily computed by SolubleMPNN. Screening and ranking used ESMFold2, ESMFold2-Fast, and Protenix v2. AlphaFold-3 weights, Rosetta, and ESM3 were excluded due to licensing reasons.

The system prompt was about 16,000 words, of which only one-third was scientific guidance; the rest involved scheduling, subagent delegation, validation, and budget discipline. The prompt did not specify any epitopes (attack sites) for any target. Compute budget was $50,000 per multi-target campaign and $10,000 per single-target, run via cloud provider Modal. Humans were only responsible for selecting targets, writing prompts, ordering synthesis, and interpreting data, with only brief non-technical instructions to recover from interruptions.

Anthropic claims that since all models used are open-source, any lab could conduct such activities.

Validation and Comparison

Validation was performed by paid contract labs Adaptyv Bio and Twist Bioscience, who independently measured whether designs bound and binding strength (KD values, in nM). Results varied by target: TREM2 had 72/90 binders, VEGF-A had 54/90. Most notably, for the RBX1 target: in a public design competition, only 9 of 245 new designs bound, whereas 28 of Claude's 90 designs bound. Anthropic reconstructed the competition winner's design and tested it on the same plate; its binding strength was 45 nM, while Claude's best design was 3.9 nM, about 10 times stronger.

TNFα was considered particularly difficult, with multiple previous methods reporting zero hits. Claude produced 12 binders from 150 designs, all from Opus 4.8, but tracing back to only 4 distinct backbones, indicating limited diversity. Additionally, 130 of 233 tested binders also bound to the mouse homolog, which has practical implications for animal studies.

Two targets exposed clear limitations: BBF-14 is a computationally invented barrel protein with no evolutionary history to reference; MBP has a smooth, hydrophilic surface that is difficult to bind. BBF-14 had only 3 weak binders, while MBP had all 90 designs fail. Anthropic stated that folding prediction confidence scores did not warn of these failures; scores for failed targets were nearly identical to those for successful ones.

Chemistry Analysis Experiments

In the second experiment, Opus 5 interpreted raw data from contract labs, including nuclear magnetic resonance spectroscopy (NMR) and liquid chromatography-mass spectrometry (LC-MS). These instruments output proprietary formats that typically require manual analysis in manufacturer software. Claude started from raw files and 1-3 sentence prompts, delivering results in 23 minutes and 19 minutes, respectively.

For LC-MS files, the model did not find a suitable reading program, decoded the format itself, and precisely reproduced the summary values of 2,664 measurement points stored by the instrument. In NMR analysis, Claude proposed the same follow-up experiment that the lab independently ran three days after the initial measurement, and corrected its own error: initially reporting 4 missing signals, then revising to 2 after internal checks.

The report noted that the commonly cited purity values of 96.4% and 96.33% were based on different baselines; by the lab's method, Claude's result was 98.8%.

Significance and Limitations

Anthropic emphasizes that protein structure prediction and generation are not new; AlphaFold, RFdiffusion, and BindCraft are established. The real novelty lies at the upper level: a general language model researches target biology, selects binding sites, installs programs itself, combines 24 different workflows, and gives final rankings without human intervention in design decisions.

However, independent review has not been conducted, and results have not been verified by third parties. Furthermore, failure cases show that confidence scores are unreliable, and some targets failed completely, indicating the method still has limitations.

Anthropic claims any lab could replicate this, but compute budget ($10,000 per target) and contract lab validation costs may pose practical barriers.

Credibility boundary

This report is based on Anthropic's technical report and THE DECODER's summary. All data are self-reported by Anthropic, and independent review has not been conducted. Some comparative data (e.g., industry hit rate of 10-15%) come from the proteinbase.com database cited by Anthropic and have not been independently verified.

Insight takeaway

Anthropic demonstrates the autonomous capability of a general language model in protein design, with hit rates significantly above industry benchmarks, but independent validation is missing, failure cases were not warned, and costs are high, making actual reproducibility questionable.

Primary report

THE DECODER

Primary source