Insilico Medicine has released an open-source AI toolkit for aging research, the centerpiece of a cover study in Cell published September 17, 2026. The release includes a benchmark for evaluating biological reasoning, a family of compact language models, and an autonomous research platform - all designed to help scientists identify and validate therapeutic targets for longevity.
The study follows a September 7 paper in Nature Biotechnology where Insilico reported that its AI-discovered drug candidate rentosertib reduced biological age across six proteomic aging clocks in a Phase IIa trial. Together, the two publications connect clinical evidence with open infrastructure for the next wave of longevity drug discovery.
What the toolkit includes
The Cell paper introduces three interconnected resources. LongevityBench is the first open benchmark built to test whether AI systems can reason across five domains of aging biology: clinical data, genetics, epigenetics, transcriptomics, and proteomics. Longevity-LLMs is a family of compact, open-source language models fine-tuned on aging-specific clinical and multi-omics data. Longevity Claw is an agentic research platform that integrates specialized tools to autonomously identify and prioritize therapeutic targets.
The resources were developed with collaborators from Liquid AI, the Buck Institute for Research on Aging, and Harvard Medical School and Brigham and Women's Hospital. All code, models, training resources, and evaluation tools are publicly available on Hugging Face and GitHub.
Benchmarking frontier models reveals gaps
The team evaluated 18 frontier AI systems on LongevityBench, including models from OpenAI, Google, Anthropic, xAI, DeepSeek, and Moonshot AI. No single model performed best across all five biological domains. Performance also shifted significantly depending on how questions were phrased, pointing to a lack of robustness that could limit reliability in scientific settings.
The hardest task was predicting biological age directly from omics data. Even the largest models struggled, suggesting that scale alone does not produce consistent biological reasoning. LongevityBench was designed to reduce the chance that models could succeed through simple recall of training data, instead testing their ability to interpret new biological measurements and generate evidence-based conclusions.
Smaller specialized models outperform larger systems
Insilico's Longevity-LLMs range from 0.6 billion to 9 billion parameters and were fine-tuned using the company's MMAI Gym for Science framework. They were built on architectures from Liquid AI and Alibaba's Qwen family. Despite their size, these compact models matched or exceeded all 16 evaluated frontier systems on LongevityBench.
The best performer, L-Qwen3.5-9B, scored highest among all 26 AI systems tested, outperforming Google's Gemini 3.1-Pro while using a fraction of the parameters. Even the smallest 0.6-billion-parameter model beat most frontier systems. The results show that curated domain-specific training data and optimization can matter more than raw scale for specialized biological tasks, with practical advantages in computational cost and deployment flexibility.
Autonomous target discovery in action
To test practical utility beyond benchmarks, the researchers embedded L-Qwen3.5-9B into Longevity Claw. The platform combines the language model with tools for gene-set enrichment analysis, aging-clock calculation, population profiling, evidence retrieval, and target evaluation. It can formulate and execute multi-step research workflows rather than simply responding to individual queries.
Deployed across 14 recognized hallmarks of aging, Longevity Claw nominated 328 genes as potential intervention targets. When compared with an independently published reference set of experimentally supported aging-related targets, the candidates showed enrichment of up to 5.6-fold. One nominated gene, KDM1A, was independently validated in a separate study as a dual-purpose aging and cancer target, with modulation extending lifespan in C. elegans.
"Insilico is at the forefront of longevity research," said Alex Zhavoronkov, Ph.D., Founder and Co-CEO of Insilico Medicine. "Our recent work on rentosertib demonstrated that an AI-designed drug can modulate biological aging signatures. In our Cell cover paper, we build on this work by empowering the global scientific community with tools such as LongevityBench, Longevity LLM and Longevity Claw, helping transform AI into an autonomous engine for longevity discovery."
Zhavoronkov added that the longevity community is "moving beyond static aging clocks toward foundation models capable of generating measurable, actionable insights," with the potential to affect drug discovery, personalized interventions, and population-level health outcomes.
The open resources are available on Hugging Face, GitHub, and a public leaderboard at longevitybenchmarks.org.
Why this matters for Science and Research professionals
For researchers working at the intersection of Generative AI and LLM technology and biology, this release provides an open, reproducible foundation for measuring progress in AI-enabled aging research. The benchmark offers a standardized way to distinguish systems that demonstrate genuine biological reasoning from those that mainly reproduce training data. The compact models show that institutions without frontier-scale compute can still deploy competitive AI for Science & Research tools on local infrastructure. The Longevity Claw platform demonstrates a concrete path from benchmark scores to testable biological hypotheses, with candidate targets that have already shown independent experimental validation.
Your membership also unlocks: