HyperAIHyperAI

Command Palette

Search for a command to run...

Daphne Koller Warns AI Drug Discovery Needs 1,000x More Data

Daphne Koller, machine learning pioneer and founder of AI drug discovery firm insitro, has argued that the pharmaceutical industry faces a critical data bottleneck that limits the efficacy of artificial intelligence in drug development. In a recent analysis, Koller warns that AI cannot compensate for a fundamental lack of understanding regarding disease mechanisms, estimating that the sector requires approximately 1,000 times more high-quality data to build robust causal models. Koller, who previously held leadership roles at Calico and co-founded Coursera, emphasizes that current AI advancements are largely confined to the mechanism-to-drug phase, where models optimize molecular structures for known targets. However, she notes that over 90 percent of candidates entering clinical trials fail because early therapeutic hypotheses regarding the biological target are incorrect, rendering efficiency gains in molecule design futile. Koller characterizes the industry's trajectory as one of convergence, citing data indicating that while global drug development projects have nearly doubled since the mid-2010s, the number of new biological targets entering pipelines has fallen by roughly 70 percent. Companies are increasingly focusing efforts on a limited set of validated mechanisms, creating a scenario where AI produces increasingly sophisticated interventions for a narrow range of targets. The root cause, according to Koller, is the reliance on observational data that captures correlations rather than causation. Effective drug discovery requires perturbational data derived from interventions that demonstrate how altering specific biological factors affects disease progression. She argues that existing cell atlas data, though massive, is insufficient for mapping causality and that generating standardized, intervention-based experimental data is a prerequisite for meaningful AI application. insitro's operational strategy reflects this priority on data infrastructure. Since its 2018 inception, the company has invested heavily in automated experimental platforms, accumulating over 20 petabytes of integrated data encompassing cell-based experiments, human genetics, and clinical information. Koller advocates for a slow path approach, declining to in-license late-stage external candidates to validate the platform, insisting that internal target discovery is essential to prove efficacy. The company has established partnerships with major pharma firms and is advancing its own AI-discovered candidates from preclinical stages, though clinical proof of concept remains pending. Koller's assessment aligns with broader industry skepticism regarding clinical translation. A recent perspective in Nature Reviews Drug Discovery highlights that while AI benchmarks have advanced, evidence of improved clinical decision-making remains limited. Market developments illustrate the transition challenges: Isomorphic Labs, the AI drug discovery spin-off of Google DeepMind, recently raised $2.1 billion but has delayed its first clinical trials until late 2026, underscoring the difficulties of moving from computational models to human applications. Meanwhile, Insilico Medicine has reported positive signals in Phase IIa trials for rentosertib, an AI-designed treatment for idiopathic pulmonary fibrosis, marking one of the first substantial clinical evaluations of an AI-discovered target. Koller projects that the next phase of AI drug discovery will shift upstream toward target identification and disease classification, requiring a new industry infrastructure characterized by data collaboration, automated perturbational experiments, and tightly integrated teams of computational and life scientists.

Related Links