HyperAIHyperAI

Command Palette

Search for a command to run...

Conception autonome de novo de protéines liantes avec Claude

Amir Shanehsazzadeh

Résumé

Les méthodes d'apprentissage profond pour la génération de structures protéiques, la conception de séquences et la prédiction de structures permettent désormais la conception de novo de protéines liantes contre de nombreuses cibles en utilisant seulement quelques dizaines de candidats. Une campagne de conception exige néanmoins une expertise couvrant la biologie de la cible, la modélisation structurale et un ensemble d'outils informatiques en évolution rapide, ainsi que des jours d'orchestration de logiciels et de calcul. Nous avons cherché à déterminer dans quelle mesure un agent d'IA pouvait fournir cette expertise et ce travail. Nous avons consigné le savoir-faire d'une campagne de conception de protéines liantes dans une unique invite de protocole qui ne spécifie aucun épitope, échafaudage ou séquence pour aucune cible. À partir de cette invite, et sans intervention humaine dans aucune décision de conception, Claude Opus 4.8 et Mythos Preview ont mené des campagnes de 24 à 48 heures contre 16 cibles. Ils ont étudié chaque cible, choisi des épitopes, installé et exécuté des modèles open source de conception de protéines et de prédiction de structures, optimisé leurs candidats in silico et livré 30 conceptions classées par cible. Deux organismes de recherche sous contrat indépendants ont synthétisé chaque conception exactement telle que livrée et mesuré sa liaison, 15 des 16 cibles donnant des mesures interprétables. Claude a conçu des protéines liantes contre 14 d'entre elles, et 354 des 1 320 conceptions se sont liées, soit un taux de réussite de 27 % ; parmi les conceptions classées premières pour chaque cible dans chaque campagne, 49 % se sont liées. En concevant contre toutes les cibles simultanément au cours d'une session unique de 48 heures, Mythos Preview et Opus 4.8 ont atteint des taux de réussite de 26,7 % et 22,6 % ; en concevant contre une cible à la fois lors de sessions de 24 heures, le taux de réussite de Mythos Preview était de 35,1 %. Sur la sous-unité RBX1 de l'E3 ligase, récemment l'objet d'un concours ouvert de conception dans lequel 9 des 245 conceptions de novo se sont liées, 28 des 90 conceptions de Claude se sont liées. La liaison la plus forte présentait un KDK _ { \mathrm { D } }KD de 3,9 nM, contre 45 nM pour l'entrée gagnante du concours, resynthétisée et mesurée sur la même plaque. Bien que la réactivité inter-espèces ne fût qu'un objectif secondaire de l'invite, 130 des 233 protéines liantes testées contre l'orthologue murin de leur cible se sont également liées à celui-ci. Tous les modèles utilisés par Claude sont open source, ce qui rend les campagnes de ce type accessibles à tout laboratoire. Nous publions les invites, les modèles computationnels des 1 440 conceptions et les données de liaison des 1 320 conceptions avec mesures fiables, comme protocole reproductible de conception autonome de protéines liantes et comme jeu de données de référence pour le domaine.

One-sentence Summary

Anthropic researchers demonstrate that Claude Opus 4.8 and Mythos Preview, guided by a single protocol prompt with no human input into any design decision, autonomously performed de novo binder design campaigns against 16 targets using open-source protein design and structure-prediction models, achieving a 27% hit rate (354 of 1,320 designs) and 49% binding among top-ranked designs, including a KDK _ { \mathrm { D } }KD of 3.9 nM on RBX1, and released the prompts, computational models of all 1,440 designs, and binding data for the 1,320 designs with reliable measurements.

Key Contributions

  • The paper introduces an autonomous protein binder design protocol based on a single frozen prompt specifying no epitope, scaffold, or sequence, in which Claude Opus 4.8 and Mythos Preview researched targets, chose epitopes, ran open-source design and structure-prediction tools, and delivered ranked de novo binders without human input into any design decision.
  • Experimental validation by two independent contract research organizations across 15 interpretable targets identified binders for 14 targets, with 354 of 1,320 synthesized designs binding (27% hit rate) and 49% of first-ranked designs binding; on RBX1, 28 of 90 designs bound and the tightest KD was 3.9 nM versus 45 nM for the competition's winning entry.
  • The study releases the prompts, computational models of all 1,440 designs, per-design provenance, and binding measurements for 1,320 designs, establishing a reproducible protocol and benchmark dataset for autonomous binder design.

Introduction

The design of de novo protein binders is important for therapeutics and diagnostics, but typical campaigns depend on expert scientists to choose targets, select epitopes, run design tools, and rank candidates. Prior AI-assisted pipelines and agent systems have supported parts of this workflow, yet they still rely on substantial human intervention and are rarely tested at scale with complete experimental reporting. The authors address this gap by showing that Claude, given only a written protocol and open-source tools, can autonomously carry out binder design campaigns across 16 targets, making every design decision from target construct to final ranking and producing experimentally validated binders on most targets.

Method

The authors leverage a single, comprehensive protocol prompt to guide an AI agent through autonomous protein binder design campaigns. This prompt, approximately 16,000 words in length, is loaded as the system prompt for every agent instance. It encapsulates the knowledge of an expert designer, defining the campaign stages, available tools, and selection criteria while leaving specific decisions to the agent. The agent autonomously researches target biology, selects modeling regions and epitopes, and chooses from a pre-cleared menu of open-source backbone generation and sequence design tools. It filters candidates for novelty and sequence composition before committing compute to scoring.

Refer to the framework diagram for the detailed anatomy of this protocol.

The prompt is structured into three primary thematic blocks. The "Science and tooling" block (34.2%) provides the working knowledge for design, including target dossiers, epitope selection, design tool menus, pre-scoring filters, and the specific ranking score. The "Orchestration and verification" block (34.7%) enables sustained autonomy over 24 to 48 hours by defining sub-agent delegation, timeline discipline, and verification rules. The "Operations" block (31.1%) manages the compute budget, pacing governor, and final reporting deliverables.

To evaluate and select the top 30 designs per target, the agent employs a specific ranking score based on an ensemble of structure predictors.

As shown in the figure below, the authors determined that ensembling scores from Protenix v2, ESMFold2, and ESMFold2-Fast yielded the highest macro-averaged precision for distinguishing binders from non-binders.

The ranking score combines the z-scored ipSAEmin\mathrm{ipSAE}_{\min}ipSAEmin (interaction predicted Structural Alignment Error) from these three predictors with self-consistency DockQ (sc-DockQ) terms added at one-quarter weight. The sc-DockQ terms serve as a check to ensure the predicted pose matches the designed complex, although they do not significantly change discrimination.

Following the autonomous campaigns, the authors applied a standardized re-scoring protocol independent of the wet lab work. Every ordered design was evaluated using ten publicly available co-folding predictors under uniform settings. The maximum ipSAEmin\mathrm{ipSAE}_{\min}ipSAEmin from five seeds per predictor was recorded.

The calibration of these re-scored confidence values against experimental binding hit rates is presented in the figure below.

The mean ipSAEmin\mathrm{ipSAE}_{\min}ipSAEmin across the three campaign predictors shows a strong correlation with hit rates across targets. To validate the designs experimentally, the authors utilized two independent contract research organizations (CROs), Adaptyv Bio and Twist Bioscience. Adaptyv Bio expressed designs via cell-free synthesis and measured binding using surface plasmon resonance (SPR) or bio-layer interferometry (BLI). Twist Bioscience expressed designs as Fc fusions in HEK293 cells and used high-throughput SPR arrays. The authors developed an automated labeling rule to integrate the distinct readouts from both CROs. A design was classified as a binder if it met specific criteria in the Adaptyv Bio traces, or if both blind trace grades were positive, or if the Twist Bioscience label was positive and Adaptyv Bio data was uninformative. This rigorous dual-validation pipeline ensured robust classification of the tested designs.

Experiment

The study tested whether an autonomous AI agent could run complete protein-binder design campaigns, from target research to a ranked set of designs, by synthesizing and measuring 1,320 designs across 15 interpretable targets with two independent CROs and assigning each design an integrated binder call. The experiments showed that the agent produced binders for most targets, that its top-ranked designs were enriched for binding, and that co-folding scores separated binders from nonbinders within targets but gave little warning of the least successful targets. Further validations found the agent's designs were competitive with open protein-design competition entries on shared targets, included species cross-reactive and beta-sheet-containing binders, and identified clear failures such as no binders against MBP and only rare weak binders against BBF-14 and 15-PGDH, while the overall evidence remains limited to binding and ranking rather than structure or function.

Design rank was calibrated with binding: considering more of the highest-ranked designs per target increased the number of targets with at least one binder. The single-target campaign reached its maximum shown coverage early, while multi-target campaigns improved more gradually as more designs were included. The top-ranked design per target already yielded binders for roughly half of targets in each campaign. Coverage increased with additional top designs, reaching 10 of 13 targets for the multi-target campaigns and 11 of 15 for the single-target campaign by the top 15 designs. MBP did not produce a binder in any shown campaign, and TNFα did not produce a binder in either Mythos Preview campaign.

This experiment examined how design ranking relates to binding by tracking target coverage as more top-ranked designs were included. The highest-ranked design per target yielded binders for roughly half of targets, and coverage increased with additional top designs, reaching 10 of 13 targets for multi-target campaigns and 11 of 15 for the single-target campaign by the top 15 designs. The single-target campaign reached its maximum coverage early, while multi-target campaigns improved more gradually. MBP produced no binders in any shown campaign, and TNFα produced no binders in either Mythos Preview campaign.


Créer de l'IA avec l'IA

De l'idée au lancement — accélérez votre développement IA avec le co-codage IA gratuit, un environnement prêt à l'emploi et le meilleur prix pour les GPU.

Codage assisté par IA
GPU prêts à l’emploi
Tarifs les plus avantageux

HyperAI Newsletters

Abonnez-vous à nos dernières mises à jour
Nous vous enverrons les dernières mises à jour de la semaine dans votre boîte de réception à neuf heures chaque lundi matin
Propulsé par MailChimp