Venture capital is pouring into artificial intelligence for drug development, but the technology's real impact may not be the splashy new molecule discoveries that grab headlines. Instead, AI is forcing pharmaceutical companies to adopt a "fail-fast" approach, ruthlessly killing unpromising drug candidates earlier in the pipeline to conserve capital and redirect resources.
The data problem beneath the hype
Foundation models in biology promise to draw novel connections from existing scientific literature, identifying new drug targets for diseases with high unmet need. The challenge is that much of the underlying data is unreliable. Multiple analyses, including work from Nature and the MD Anderson Cancer Center, have documented a reproducibility crisis in peer-reviewed research. A large portion of published findings cannot be replicated, even without misconduct by the original teams.
This is distinct from the reproducibility efforts underway within the AI field itself. Machine learning engineers are working to ensure foundation models return the same result for the same query on the same model version. But a reproducible AI output built on irreproducible science does not represent real progress.
When probability models turn paradoxical
Even with clean, reproducible datasets, probability-based models face structural complications. Two classic statistical phenomena illustrate the problem. Bertrand's paradox shows that probabilities shift when using different random selection methods from the same pool. Simpson's paradox demonstrates how trends can reverse depending on how subgroups are combined.
In small molecule drug design, these paradoxes become practical headaches. Minor chemical modifications can produce dramatic changes in how a molecule behaves experimentally. When a chemist applies those same tweaks to a different chemical family, the effects often change in unexpected ways. Avicenna Biosciences has published work showing that chemical changes beneficial in one series can wreck machine learning models trained on another series, even when targeting the same property, because the underlying cause-effect relationships shifted.
"The context needed to model the properties of molecules in a dataset is often not known, especially for complex data like vertebrate toxicology mediators or pharmacokinetic metrics," said Thomas M. Kaiser, co-founder and chief scientific officer at Avicenna Biosciences. Research groups including Avicenna and the Doyle group at MIT are developing methods to address these hidden cause-effect conflicts, though the paradoxical nature of probability makes the work difficult.
Where foundation models can deliver now
Foundation model techniques show genuine promise when applied to tightly focused scientific questions that account for contextual nuance. Identifying novel biological networks as drug targets is one area of potential. Mining clinical data to improve patient selection in translational development represents an even larger opportunity. Companies like Yatiri Bio and a collaboration between Bruker and Noetik are already exploring this space.
The ability to discover translational connections between clinically meaningful endpoints, appropriate patient populations, and preclinical models that reflect human disease could improve drug development success rates. But this hinges on scientists maintaining rigorous experimental standards and careful data evaluation, rather than deferring entirely to model outputs.
Why this matters for IT and development professionals
For developers and engineers working with or adjacent to biotech, the caution around foundation models in drug discovery carries a practical lesson: domain-specific data quality problems are not solved by model scale alone. The reproducibility crisis in scientific literature and the contextual paradoxes in chemical data are not bugs that better architectures will fix. They are intrinsic properties of the data generating process. Teams building AI systems for scientific applications need to invest as heavily in data provenance, experimental validation workflows, and context-aware feature engineering as they do in model selection and training infrastructure.
Your membership also unlocks: