Abstract
From 2020 through July 2026, artificial intelligence (AI) has transitioned from a preclinical research tool to an operationally embedded component of pharmaceutical R&D. AI-designed or AI-enabled drug candidates have entered human clinical trials across multiple therapeutic areas, and one compound—rentosertib—has advanced to Phase III. However, no AI-discovered drug has yet achieved regulatory approval, Phase II success rates align with historical industry averages (~40%), and the gap between computational promise and demonstrated patient benefit remains substantial. This review synthesizes clinical evidence, adoption patterns, and methodological limitations to support critical evaluation of AI drug discovery claims by medical professionals.
Clinical Successes: From Preclinical Promise to Phase III
The most clinically mature AI-discovered asset is rentosertib (formerly ISM001-055), developed by Insilico Medicine using its Pharma.AI platform. TNIK (Traf2- and Nck-interacting kinase), a novel serine/threonine kinase with no prior selective clinical-stage inhibitors, was identified as a high-priority fibrosis target using PandaOmics, Insilico's AI-powered biology engine, which integrated multi-omics data, causal inference, biological network analysis, and aging-relevant target scoring. The small molecule was subsequently designed and optimized through Chemistry42, a generative chemistry platform. Remarkably, the entire preclinical program—from target hypothesis to preclinical candidate nomination—was completed in approximately 18 months at a reported cost of ~$2.6 million. (By comparison, industry-wide estimates of $430 million represent the fully capitalized average cost per approved drug, encompassing all preclinical, clinical, and regulatory phases as well as the cost of failures—a substantially broader scope than the target-to-candidate window described here.) 1.
Phase IIa results published in Nature Medicine in 2025 (GENESIS-IPF trial; 71 patients, 22 sites in China, 12 weeks) demonstrated safety comparability across arms and a dose-dependent lung-function signal: patients receiving 60 mg once-daily (QD) showed a mean forced vital capacity (FVC) increase of +98.4 mL (95% CI 10.9–185.9) versus −20.3 mL in the placebo group. Exploratory serum proteomic analyses showed dose-dependent downregulation of fibrosis-associated proteins (MMP10, COL1A1, FAP), supporting on-target TNIK pathway modulation 2. On July 7, 2026, Insilico announced initiation of a Phase III trial (320 patients, 47 centers in China, 52-week primary endpoint of annual FVC decline rate), marking the first AI end-to-end discovered drug to enter late-stage development 10.
Several other programs have reached early clinical stages. SGR-1505 (Schrödinger), a MALT1 inhibitor for relapsed/refractory B-cell malignancies, demonstrated no dose-limiting toxicities and an overall response rate of 22% in Phase I (49 patients), with 5/5 Waldenström macroglobulinemia patients responding and ~90% IL-2 inhibition confirming target engagement. BEN-2293 (BenevolentAI), a topical pan-TRK inhibitor for atopic dermatitis, met its primary safety endpoint in Phase IIa but failed primary efficacy endpoints in the intention-to-treat population; a post-hoc subgroup signal in patients with ≥20% body surface area involvement (p=0.0296) requires prospective validation. BEN-8744 (BenevolentAI), a peripherally restricted PDE10 inhibitor for ulcerative colitis, demonstrated safety and tolerability without CNS adverse events—a key differentiator from prior PDE10 inhibitors—in 54 healthy volunteers, with Phase 2-enabling studies underway.
In Japan, DSP-1181 (Sumitomo Dainippon Pharma + Exscientia), a 5-HT1A receptor agonist for obsessive-compulsive disorder (OCD) designed using Exscientia's Centaur Chemist AI platform, entered Phase I in Japan in January 2020, with the exploratory research phase completed in under 12 months; however, Phase I results did not meet expectations and the program was subsequently discontinued 18. For Recursion Pharmaceuticals, REC-4881 (MEK1/2 inhibitor, familial adenomatous polyposis [FAP], Phase II) demonstrated durable total polyp-burden reduction in 9 of 11 evaluable patients (median 53% reduction), with Fast Track and US/EU Orphan Drug designations 1920. REC-1245 (RBM39 degrader, solid tumors and lymphoma, Phase I) received FDA clearance to begin first-in-human studies in 2024 21, with both assets having their targets identified via the Recursion OS phenotypic AI platform.
Across eight leading AI drug discovery companies as of April 2024, 31 drugs were in human clinical trials (17 Phase I, 5 Phase I/II, 9 Phase II/III). Phase I success rates for AI-native programs were estimated at 80–90%, exceeding the historical industry average of 50–60%, suggesting AI effectively identifies molecules with favorable drug-like properties. However, Phase II success rates (~40%) align with historical norms, indicating that target validation and disease biology—not molecular optimization—remain the primary clinical bottlenecks.
Table 1. Representative AI-Discovered or AI-Enabled Clinical-Stage Drug Candidates (2020–2026)
| Company | Asset | Mechanism/Target | Indication | AI Role | Clinical Stage/Status | Key Evidence / Caveat |
|---|---|---|---|---|---|---|
| Insilico Medicine | Rentosertib (ISM001-055) | TNIK inhibitor | Idiopathic pulmonary fibrosis | AI target discovery (PandaOmics) + AI molecular design (Chemistry42) | Phase III initiated July 2026 10 | Phase IIa: +98.4 mL FVC (60 mg QD) vs. −20.3 mL placebo; published Nature Medicine 2025 2; single region (China), 12 weeks, small arms (n~18); Phase III: 320 pts, 52 weeks |
| Schrödinger | SGR-1505 | MALT1 inhibitor | Relapsed/refractory B-cell malignancies | AI computational platform molecular design | Phase I completed; Phase II planned | ORR 22% (10/45 evaluable); 5/5 WM, 3/17 CLL/SLL responses; ~90% IL-2 inhibition; no dose-limiting toxicities; heavily pretreated population |
| BenevolentAI | BEN-2293 | Pan-TRK inhibitor (topical) | Mild-to-moderate atopic dermatitis | AI target identification | Phase IIa completed | Failed primary efficacy endpoints (ITT); post-hoc subgroup signal (≥20% BSA, p=0.0296) lacks prospective validation |
| BenevolentAI | BEN-8744 | PDE10 inhibitor (peripherally restricted) | Ulcerative colitis | AI target discovery (Benevolent Platform) + molecular design | Phase Ia completed; Phase 2-enabling studies underway | Safe in 54 healthy volunteers; no CNS adverse events; no efficacy data yet |
| Sumitomo Dainippon + Exscientia | DSP-1181 | 5-HT1A receptor agonist | OCD | AI molecular design (Centaur Chemist platform) | Phase I (Japan, initiated Jan 2020); discontinued 18 | Exploratory phase <12 months; no Phase II efficacy data in retrieved materials |
| Recursion | REC-4881 | MEK1/2 inhibitor | Familial adenomatous polyposis | AI-enabled phenotypic target-indication discovery (Recursion OS) | Phase II 19 | Durable polyp-burden reduction: 9/11 patients, median 53% reduction 20; Fast Track + US/EU Orphan designation |
| Recursion | REC-1245 | RBM39 degrader | Solid tumors and lymphoma | AI-enabled target discovery (Recursion OS) | Phase I (FDA clearance 2024) 21 | Early-phase; no efficacy data in retrieved materials |
Key Challenges and Limitations
Despite early clinical milestones, several structural barriers limit AI drug discovery productivity. Model generalizability is a primary scientific concern: AI models trained on curated benchmark datasets frequently show degraded performance in prospective, real-world applications. A 2026 risk-tiered validation framework identified four validation tiers—internal ML reproducibility, molecular-science benchmarks, prospective experimental validation, and clinical/translational calibration—and emphasized that retrospective benchmark enrichment is an unreliable predictor of prospective performance 6.
Data quality and fragmentation are compounding factors. Pharmaceutical datasets are proprietary, biased toward well-studied targets and chemical scaffolds, and rarely shared. Models trained on such data may propagate systematic biases into target and molecular predictions. Explainability remains a regulatory and clinical barrier: deep learning and generative models often function as "black boxes," providing limited mechanistic insight into why a particular prediction succeeds or fails—a critical requirement for safety monitoring and regulatory acceptance 7.
Wet-lab validation bottlenecks persist. Computational predictions still require extensive synthesis, cell-based assays, animal models, and toxicology and PK/PD (pharmacokinetics/pharmacodynamics) studies before IND filing; AI does not eliminate this rate-limiting step. Overstatement of AI contribution is widespread: industry timelines citing "discovery in 8–12 months" typically refer to exploratory computational research, not the full preclinical-candidate-to-IND timeline. Regulatory uncertainty is also evolving: in January 2026, the FDA and EMA jointly published 10 guiding principles for responsible AI use across the drug lifecycle, and the FDA's January 2025 draft guidance acknowledged over 500 AI-containing submissions from 2016–2023 but emphasized that approval remains contingent on demonstrated clinical safety and efficacy, not discovery method 45.
Table 3. Key Challenges in AI Drug Discovery and Mitigation Strategies
| Challenge Category | Clinical Relevance | Example Issue | Mitigation Strategy |
|---|---|---|---|
| Model Generalizability | High | Benchmark performance does not predict prospective real-world accuracy; distributional shift between training data and project chemistry 6 | Require prospective experimental validation; report confidence intervals and uncertainty quantification; use independent test sets |
| Data Quality and Bias | High | Proprietary, fragmented datasets biased toward well-studied targets; publication bias in training data | Establish data-sharing consortia; conduct bias audits; use diverse training datasets |
| Explainability | High | Deep learning models cannot explain predictions; limits regulatory and clinical acceptance 7 | Develop interpretable ML methods; require mechanistic functional validation before IND |
| Wet-Lab Bottleneck | High | AI accelerates in silico design but not synthesis, PK/PD, or toxicology | Integrate AI with automated synthesis and high-throughput screening; invest in closed-loop design-make-test cycles |
| Phase II Failure Risk | Critical | AI Phase II success rates (~40%) match historical norms; target biology remains the bottleneck | Prioritize validated targets; use biomarker-driven patient stratification; adaptive trial designs |
| Small Trial Sizes and Short Duration | High | Rentosertib Phase IIa: 71 patients, 12 weeks; insufficient for efficacy generalization 3 | Prospectively power Phase IIb/III for efficacy; longer follow-up; geographic diversity |
| Overstatement of AI Contribution | Moderate | "Discovered in 8 months" conflates computational phase with full preclinical-to-IND timeline | Require transparent disclosure of AI role per step; cite peer-reviewed publications |
| Regulatory Uncertainty | Moderate | No explicit AI drug approval pathway; FDA/EMA guidance is principles-based and non-prescriptive 45 | Engage regulators early (pre-IND meetings); participate in AI-specific guidance development |
| Proprietary Datasets | High | Trade secrets limit external validation and reproducibility | Support open-science initiatives; publish model validation data; encourage data-sharing agreements |
Pharmaceutical Adoption: From Exploratory to Integrated
Large pharmaceutical companies adopted AI through four primary models: (1) equity investments and platform licensing, (2) milestone-based research collaborations, (3) internal AI platform development, and (4) acquisitions of AI-native biotechs. The financial scale of commitments expanded substantially: Sanofi's 2021 collaboration with Exscientia offered up to $5.2 billion in milestone payments for 15 oncology/immunology candidates; Roche/Genentech's 2021 partnership with Recursion provided $150 million upfront and up to $300 million per project across 40 programs; Pfizer extended its PostEra partnership in January 2025 to a total deal value of $610 million, with PostEra reportedly achieving preclinical milestones ~40% faster than Pfizer anticipated; and Iambic's February 2026 multi-year collaboration with Takeda carries potential success-based payments exceeding $1.7 billion, centering on NeuralPLexer, Iambic's protein-ligand structure prediction model, for oncology and gastrointestinal disease 1112. In March 2026, Tempus AI and Daiichi Sankyo announced a collaboration deploying the PRISM2 multimodal foundation model for biomarker discovery and patient stratification in an oncology ADC (antibody-drug conjugate) program 13. Sanofi's 2024 collaboration with Formation Bio and OpenAI reflects a further shift toward sharing proprietary R&D data to build AI-powered infrastructure at scale 8.
Despite the scale of investment, few partnerships from 2012 to 2024 yielded Phase II programs with disclosed efficacy data, and overall partnership success rates remain low. The field is transitioning from exploratory platform licensing toward integrated R&D infrastructure, but this shift has not yet translated into a measurable increase in new drug approvals or a reduction in total development costs.
Table 2. Major Pharmaceutical Adoption Models in AI Drug Discovery (2020–2026)
| Pharma Company | AI Partner/Platform | Deal Type | Therapeutic Focus | Strategic Rationale | Observable Outcome/Status |
|---|---|---|---|---|---|
| Takeda | Iambic (NeuralPLexer) | Multi-year technology + discovery collaboration (Feb 2026) 11 | Oncology, GI/inflammation | Accelerate small-molecule design; protein-ligand structure prediction | Up to $1.7B success payments; design-make-test on weekly cadence; no clinical assets disclosed yet |
| Daiichi Sankyo | Tempus AI (PRISM2) | Strategic collaboration (Mar 2026) 13 | Oncology ADC | Biomarker discovery; AI-driven patient stratification for ADC clinical development | Proof-of-concept models in development; no Phase III readouts yet |
| Sanofi | Exscientia | Research collaboration + license (Dec 2021) | Oncology, immunology | AI target discovery and precision medicine integration | $100M upfront + up to $5.2B milestones; early-stage pipeline; no Phase II efficacy readouts |
| Roche/Genentech | Recursion | Multi-project collaboration (Dec 2021) | Neuroscience, oncology | Merge computational and wet-lab phenotypic screening | $150M upfront + up to $300M per project; REC-4881 Phase II with efficacy signals 1920 |
| Pfizer | PostEra | Multi-year collaboration + equity | Small molecules, ADC payloads | Accelerate medicinal chemistry; achieve milestones 40% faster | Total value $610M (Jan 2025); preclinical milestones achieved faster than internal benchmarks |
| Sanofi | Formation Bio + OpenAI | Strategic collaboration (2024) 8 | Multi-indication | Transform into AI-powered pharma; leverage proprietary data at scale | Ongoing; no marketed products yet |
| Pfizer | CytoReason | Equity investment + platform license (Sept 2022) | Immune-mediated, immuno-oncology | Enhance target understanding across 20+ diseases | $20M equity + up to $110M over 5 years; exploratory stage |
Practical Implications for Medical Professionals
When evaluating AI drug discovery claims in publications, conference presentations, or investor materials, clinicians should apply the following criteria. First, distinguish the AI role precisely: AI target discovery carries different risks than AI molecular optimization; first-in-class AI targets bear the additional uncertainty of unvalidated biology. Second, assess evidence tier: Phase I safety data and small Phase IIa trials (e.g., rentosertib, 71 patients, 12 weeks) are encouraging but insufficient to establish clinical benefit; Phase IIb/III data in geographically diverse, adequately powered populations are required. Third, scrutinize endpoint selection: surrogate endpoints such as FVC change or polyp-burden reduction require correlation with clinical outcomes (mortality, progression-free survival) in longer trials. Fourth, evaluate biological plausibility: AI-selected targets should be supported by independent functional validation in patient-derived tissues, not solely computational prediction. Fifth, demand external validation: proprietary training datasets and lack of peer-reviewed model validation are red flags; assets with published validation in journals such as Nature Biotechnology, Nature Medicine, or Science Advances carry greater evidentiary weight 210.
Red flags include claims of "AI superiority" without Phase II/III comparative data; emphasis on discovery speed without discussion of full development timelines; post-hoc subgroup efficacy claims without prospective validation; and conflation of drug repurposing with de novo AI discovery.
Looking ahead through 2026 and beyond, foundation models (e.g., BioEmu, AlphaFlow, Boltz-2, Chai-1), multimodal biomedical data integration, digital pathology, AI-optimized clinical trial design, and FDA support for New Approach Methodologies including digital twins are likely to accelerate adoption and reduce reliance on animal models 15. However, the central bottleneck—demonstrating that AI-enabled drugs are safer and more efficacious for patients than conventionally discovered drugs—remains unresolved.
Concluding Perspective
AI drug discovery has delivered meaningful early milestones: rentosertib's Phase III initiation, Recursion's REC-4881 efficacy signals in FAP, and Iambic's integration of weekly design-make-test cycles represent genuine innovations. Yet no AI-discovered drug has achieved regulatory approval, and the clinical development challenges that define pharmaceutical R&D—target validation, patient heterogeneity, toxicology, PK/PD, and trial failure—remain unchanged. Medical professionals should welcome AI-enabled discovery as a tool that may accelerate target identification and molecular optimization, while demanding the same rigorous clinical evidence standards applied to any novel therapeutic, and exercising proportionate skepticism toward claims that AI has fundamentally transformed the drug development enterprise.