Field report · evidence cut-off 4 August 2026
The machines can propose.
The clinic decides.
AI is now durable infrastructure for finding structures, narrowing chemical search, engineering proteins, and choosing experiments. It has not yet shown that it can beat pharmaceutical attrition in Phase II and III.
Every consequential claim is tagged by provenance:
01 / Evidence
A plausible molecule is the start of the argument.
Drug discovery headlines often collapse six distinct proofs into one. Reliability rises only as a prediction survives prospective experiments, human testing, and comparison with a credible baseline.
Retrospective benchmark
Useful for debugging; vulnerable to leakage, shortcuts, and friendly splits.
Locked prospective prediction
The model commits before measurement and is compared with simple, strong baselines.
Orthogonal wet-lab validation
Synthesis, target engagement, selectivity, developability, and negative results are disclosed.
Controlled patient evidence
Clinically meaningful endpoints survive randomization, duration, missingness, and replication.
Approval and portfolio advantage
Quality, safety, efficacy—and eventually more useful medicines per dollar and year.
02 / Lifecycle
Where AI is already useful—and where it still breaks.
The closer a task is to structured data and rapid feedback, the stronger the evidence. The closer it gets to causal human biology, the thinner it becomes.
Higher maturity
Structure & mechanism
Protein and complex prediction, pocket discovery, conformational hypotheses.
Limit: a static structure is not affinity, function, or efficacy.Higher maturity
Hit finding
Virtual screening, learned scoring, phenotypic-image search, generative chemistry.
Limit: hit rates vary and domain shift is common.Moderate maturity
Lead optimization
Potency, selectivity, ADME, and physics–ML multi-parameter prioritization.
Limit: rare toxicity and activity cliffs still surprise.Rapidly advancing
Protein & antibody design
Diffusion models, protein language models, binder and product-profile engineering.
Limit: immunogenicity, expression, and benefit need experiments.Moderate maturity
Make & test
Retrosynthesis, reaction optimization, robotics, active-learning loops.
Limit: hardware, safety, scale-up, and negative data.Context-specific
Clinical development
Recruitment, enrichment, endpoints, dose selection, monitoring, operations.
Limit: bias, drift, missingness, privacy, and causality.Early to moderate
CMC & manufacturing
Formulation, process monitoring, fault detection, visual QC, deviation mining.
Limit: GxP validation, change control, cybersecurity, release authority.Increasing adoption
Regulatory & safety
Data extraction, submission drafting, case triage, signal detection.
Limit: provenance, auditability, hallucination, accountability.03 / Timeline
Fourteen years from benchmark wins to pivotal trials.
This is a long arc, not a sudden post-ChatGPT invention. Each milestone expanded what could be predicted or designed; only the newest ones begin to test clinical value.
-
The Merck molecular-activity challenge
Deep neural networks performed strongly on pharmaceutical assay prediction and helped reboot industrial interest in learned molecular representations.
Observed fact -
AtomNet brings 3D convolutions to binding
Structure-based deep learning moved from abstract descriptors toward protein–ligand geometry, foreshadowing learned docking and scoring.
Sponsor-reported -
CASP13 and neural synthesis planning
AlphaFold’s precursor showed a leap in structure prediction; neural networks plus tree search planned synthetic routes rated comparable to literature routes.
Observed fact -
Generation meets wet lab
GENTRL proposed DDR1 inhibitors that were synthesized and tested; AI planning plus modular flow robotics produced 15 drug or drug-like substances.
DDR1 was a known target with published chemistry context. “21 days” described generation—not the complete tested program.
Observed fact -
Halicin and the first prominent AI-designed Phase I entry
A neural screen identified halicin, active in vitro and in mice. DSP-1181 entered Phase I, then was discontinued in 2022 after missing its expected criterion.
Observed fact -
Structure prediction becomes infrastructure
AlphaFold2 reached near-experimental accuracy for many structures; AlphaFold DB expanded to roughly 200 million predictions. AI-assisted repurposing helped nominate baricitinib for COVID-19.
Observed fact -
From predicting proteins to designing them
RFdiffusion generated new structures and binders with experimental validation. ML screening found the preclinical antibiotic candidate abaucin.
Observed fact -
Complexes, clinical translation, and consolidation
AlphaFold3 extended prediction to biomolecular complexes; rentosertib’s target-to-clinic journey was published; the Chemistry Nobel recognized structure prediction and computational protein design.
Observed fact -
The first patient signal—and a pivotal biologic
Rentosertib reported a small Phase IIa FVC signal. FDA qualified AIM-NASH, its first AI drug-development tool. GB-0895 began two Phase III asthma studies.
Observed fact -
The clinic becomes the scoreboard
FDA and EMA issued Good AI Practice principles. Rentosertib entered a 320-patient, 52-week Phase III IPF study. Zasocitinib reported positive pivotal psoriasis data.
Observed fact
04 / Clinical scoreboard
Do not add unlike programs into one “AI drug” count.
Filter by clinical status and by what AI actually did. Stage is not outcome; sponsor-defined provenance is not independent validation.
Showing all 11 programs.
GB-0895
Generate Biomedicines · severe asthma
Affinity, half-life, product profile; known TSLP target.
Phase I PK/biomarker evidence; no Phase II patient-efficacy study. Class precedent helped de-risk the pivotal move.
Rentosertib
Insilico Medicine · idiopathic pulmonary fibrosis
AI prioritized TNIK; generative AI designed and optimized the inhibitor.
Small Phase IIa secondary FVC signal; safety was primary. Short duration, 71 patients, 16 discontinuations.
Zasocitinib / TAK-279
Takeda · plaque psoriasis
FEP+ and structure-guided multiparameter design, with ML in support.
Important computational-design success, but not a strict generative-AI medicine and not approved at cut-off.
ABS-101
Absci · inflammatory bowel disease
Generative-AI-engineered anti-TL1A antibody.
Human safety and PK testing; no patient-efficacy evidence.
ABS-201
Absci · androgenetic alopecia
AI-engineered anti-PRLR antibody.
Favorable blinded aggregate safety and half-life report; no human efficacy yet.
IAM1363
Iambic Therapeutics · HER2-positive cancers
AI-supported design of a brain-penetrant HER2 inhibitor.
No posted results establishing clinical benefit.
DSP-1181
Exscientia / Sumitomo · obsessive-compulsive disorder
Prominent early AI-guided small-molecule program.
Missed the expected Phase I evaluation criterion and was discontinued.
EXS21546
Exscientia · oncology
AI-designed A2A antagonist.
Stopped after information suggested the program was not sufficiently promising.
BEN-2293
BenevolentAI · atopic dermatitis
Platform-supported legacy program; provenance is less clean.
Met safety/tolerability primary endpoint; missed secondary itch and inflammation endpoints.
REC-994
Recursion · cerebral cavernous malformation
ML phenomics linked an existing molecule to disease biology.
An extension did not reproduce early exploratory trends; program was discontinued or offered for partnering.
Isomorphic Labs
Platform company · multiple pharma partnerships
Large partnerships and company-reported preclinical/design-engine progress.
The absence is informative: capital and technical promise should not be counted as clinical validation.
Reading rule: entering Phase I proves a candidate could be made and tested in humans. Starting Phase III proves a sponsor is willing to run the trial. Neither proves patient benefit.
05 / Reliability
Reliable trend. Unreliable extrapolation.
The platform transition is real. The claim that it will raise clinical success is still a forecast—and the available cohort is far too small and selected to settle it.
Calibrated claim confidence
Forecast Percentages are calibrated judgments, not measured probabilities.
The Phase I headline
Jayatunga et al. reported 21 successful Phase I completions among 24 publicly disclosed assets from AI-native companies; Phase II was 4/10.
- Small, overlapping, selected categories
- No matched control for indication, modality, novelty, biomarkers, sponsor, or age
- Phase I tests drug-like properties; Phase II tests human disease biology
Useful signal, not causal proof.
Observed factBenchmark traps to name
- DUD-E decoy bias: ligand shortcuts masquerade as target learning
- PDBbind/CASF similarity: close complexes leak across evaluation
- Homology contamination: protein relatives cross train/test boundaries
- Temporal leakage: future chemistry quietly informs the past
Prefer temporal or cold-target splits, uncertainty, locked predictions, and independent assays.
Observed fact2026 → 2030 scenarios
Three futures worth preparing for
Leading pivotal programs disappoint; few or no strict AI-origin approvals arrive; consolidation and model commoditization continue.
ForecastOne to three approvals have a material AI contribution. Discovery speed improves more than adjusted clinical success; attribution stays contested.
ForecastMultiple clearly AI-origin approvals and matched prospective evidence show better target selection or Phase II transition.
Forecast06 / Economics
Where value is most likely to accrue.
Model architecture is commoditizing. Defensibility shifts toward exclusive data, fast experimental feedback, disciplined clinical execution, and regulated workflows.
Proprietary experimental data
Consistent negative and positive measurements, generated for the decisions the model must make.
Closed-loop systems
Design–make–test–analyze cycles that turn model uncertainty into the next useful experiment.
Translational & clinical execution
Biomarkers, patients, endpoints, operations, and decisions that connect molecules to disease biology.
Regulated operations
Trial operations, CMC, process analytics, model governance, provenance, and audit-ready automation.
Multi-billion-dollar partnership ceilings are mostly contingent “biobucks,” not realized revenue. Faster generation can simply move cost into toxicology and Phase II. Full ROI must include compute, wet labs, failed programs, trials, infrastructure, and capital.
07 / Your entry
Enter through a real decision loop.
The durable career bet is not “be an AI drug discoverer.” Bring one credible discipline to one decision—then show that your work changes what gets tested, made, or advanced.
Software / machine learning · 30 days
Learn to distrust the easy split.
Your wedge
Cheminformatics, uncertainty, scientific software, data pipelines, and model evaluation.
Do next
- Pick one disease area, modality, and decision point.
- Reproduce one RDKit, DeepChem, or TeachOpenCADD workflow.
- Compare a fingerprint baseline with a learned model under scaffold and time splits.
- Write down uncertainty, data version, license, and failure modes.
Ship this
One-page evaluation specification plus a reproducible baseline notebook or report.
Less crowded
High-leverage wedges
- Trial operations and translational biomarkers
- Regulatory model governance and auditability
- CMC and process analytics
- Scientific data infrastructure and assay automation
Crowded
Weak default bets
- Another generic molecular generator
- An unvalidated property predictor
- Foundation-model work without proprietary data
- A generic prediction startup without wet-lab access
08 / Codex × Claude
Two independent reads, one harder conclusion.
Claude Opus 5 Max independently researched the field, then audited the Codex report. Agreement was broad; the useful differences tightened taxonomy, timing, and economic claims.
Where both agreed
The core conclusion survived challenge.
- No strict end-to-end approval identified at the cut-off
- Prospective evidence outranks retrospective benchmarks
- Phase I success is confounded; Phase II is load-bearing
- Protein design may mature faster than end-to-end small molecules
- High-confidence platform transition; low-confidence clinical superiority
Codex contributed
The freshest clinical and regulatory record.
- Two GB-0895 pivotal asthma studies
- Rentosertib’s July 2026 Phase III initiation
- FDA–EMA Good AI Practice principles
- Exact Jayatunga cohort and public-company economics
- Broader CMC and manufacturing coverage
Claude sharpened
The definitions, nulls, and boundary conditions.
- Extend the arc back to 2012
- Make Isomorphic’s absent public clinical asset visible
- Name specific benchmark pathologies
- Treat “AI-designed” as marketing until the funnel is disclosed
- Show where durable economic value may accrue
Strict target + molecule approval by end-2030: about 30%
Claude’s timing challenge lowered the estimate: the pivotal readout and regulatory calendar are tight. A 2030–31 window is more honest.
09 / Prediction ledger
What would change our minds.
Dated, falsifiable predictions are better than vibes. These are the signals to score over the next several years.
Closed-loop labs expand, but human release gates remain.
Watch prospective productivity, safety interventions, failed experiments, and transfer across chemistry—not demo throughput.
Biologics produce the clearer AI engineering wins.
GB-0895 and peer programs should reveal whether designed affinity and half-life translate into patient outcomes and workable products.
Rentosertib provides the cleanest target + molecule pivotal test.
Judge efficacy, safety, discontinuation, geography, and robustness—not merely trial completion or a top-line press release.
One to three approvals carry a material documented AI contribution.
Count separately: AI support, AI-engineered product, physics-first design, and strict AI target + molecule.
Only a matched Phase II cohort can establish clinical advantage.
Preregister definitions and match on indication, modality, target novelty, biomarker strategy, sponsor quality, and program age.
10 / Source ledger
Primary evidence, close to the claim.
Peer-reviewed papers, trial registries, regulators, company filings, and sponsor releases are separated by what each can actually establish.
Foundational science
- DDR1 generative chemistry Nature Biotechnology, 2019
- Automated flow synthesis of 15 compounds Science, 2019
- Halicin antibiotic screen Cell, 2020
- AlphaFold2 Nature, 2021
- AlphaFold Protein Structure Database Nature Structural & Molecular Biology, 2022
- Abaucin Nature Chemical Biology, 2023
- RFdiffusion Nature, 2023
- AlphaFold3 Nature, 2024
Clinical record
- Rentosertib target-to-clinic journey Nature Biotechnology, 2024
- Rentosertib Phase IIa Nature Medicine, 2025
- Rentosertib Phase III ClinicalTrials.gov
- GB-0895 SOLAIRIA-1 ClinicalTrials.gov
- GB-0895 SOLAIRIA-2 ClinicalTrials.gov
- Zasocitinib pivotal results Takeda, 2026
- Zasocitinib discovery Journal of Medicinal Chemistry, 2023
- AI-native clinical success cohort Drug Discovery Today, 2024
Regulation, manufacturing & methods
- FDA–EMA Good AI Practice principles 2026
- FDA’s 500+ AI-component submissions statement 2025
- AIM-NASH qualification FDA, 2025
- AI in pharmaceutical manufacturing FDA discussion paper
- Bayesian reaction optimization Nature, 2021
- LLM-controlled synthesis laboratory Nature Communications, 2024
- Data leakage in machine learning Nature Methods, 2024
- Shortcut learning in drug–target prediction Nature Communications, 2023
Economics & market evidence
- Recursion–Roche collaboration Company release
- Sanofi–Exscientia collaboration Company release
- Isomorphic partnerships Company release
- AI drug-discovery economic model Wellcome / BCG
- Recursion 2025 Form 10-K SEC filing
- Schrödinger 2025 Form 10-K SEC filing
- CACHE blind prospective challenges Benchmark program
- Open Targets Platform Public resource
Method note. “Observed fact” includes peer-reviewed work, regulator records, trial registries, and directly checkable corporate events. Sponsor releases remain tagged when they establish only sponsor-reported results or provenance. Forecasts are explicitly subjective. The negative approval claim is bounded to the reviewed major-regulator and sponsor record as of 4 August 2026.