AI language models for aging biology

 10
AI language models for aging biology

Insilico Medicine, a clinical-stage biotechnology company powered by generative artificial intelligence (AI), announced the publication of a landmark study in Cell introducing an open AI toolkit, incluiding LongevityBench, Longevity-LLMs and LongevityClaw designed to accelerate research and therapeutic development in aging and longevity..

The publication follows Insilico’s study in Nature Biotechnology, which reported that rentosertib, Insilico’s AI-discovered and AI-designed drug candidate for idiopathic pulmonary fibrosis, reduced biological age across six independent proteomic aging clocks in a Phase IIa clinical trial.

Together, the two studies connect clinical evidence from an AI-discovered therapeutic with an open research infrastructure intended to help scientists identify and develop the next generation of longevity interventions.

Aging is driven by complex and interconnected biological processes that span multiple levels of human biology. Understanding these processes requires researchers to integrate clinical records with genetics, epigenetics, transcriptomics, proteomics and other large-scale biological datasets.

While general-purpose large language models have demonstrated strong performance across many scientific tasks, their ability to reason reliably over real-world aging data has remained largely untested. Existing evaluations also frequently reward the memorization of scientific facts rather than the ability to interpret new biological measurements or generate evidence-based conclusions.

To address this gap, the research team developed LongevityBench, an open benchmark specifically designed to determine whether AI systems can reason across the diverse data modalities that define human aging.

LongevityBench evaluates AI performance across five major biological domains:

·         Clinical data

·         Genetics

·         Epigenetics

·         Transcriptomics

·         Proteomics

The benchmark was designed to reduce the likelihood that models could succeed through simple recall of information encountered during training. Instead, it tests the ability of an AI system to analyze biological data, recognize meaningful patterns and solve problems relevant to aging research.

The researchers evaluated 18 leading frontier AI systems, including models developed by OpenAI, Google, Anthropic, xAI, DeepSeek and Moonshot AI. The analysis revealed substantial performance gaps across biological domains. No single frontier model achieved the strongest results across all five data types. Model performance also changed significantly depending on how questions were phrased, highlighting a lack of robustness that could limit the reliability of general-purpose AI systems in scientific research.

The most difficult challenge was predicting biological age directly from omics measurements. Even the largest frontier models struggled with this task, suggesting that model scale alone is insufficient to produce consistent biological reasoning.

To determine whether these limitations could be overcome without frontier-scale computing resources, the researchers developed a family of five compact, open-source Longevity Large Language Models, or L-LLMs.

The models ranged from 0.6 billion to 9 billion parameters and were fine-tuned on aging-specific clinical and multi-omics data using Insilico’s MMAI Gym for Science, a training and evaluation framework designed to develop specialized AI systems for scientific applications.

The family was built on open model architectures from Liquid AI and Alibaba, including Liquid AI’s LFM2 architecture and Alibaba’s Qwen3 and Qwen3.5 model families.

Despite their comparatively small size, the specialized Longevity-LLMs matched or exceeded all 16 frontier systems evaluated on LongevityBench.

The best-performing model, L-Qwen3.5-9B, achieved the highest overall score among all 26 AI systems included in the study, outperforming every tested frontier model, including Google’s Gemini 3.1-Pro, while using only a fraction of the parameters.

Even the smallest specialized model, with approximately 0.6 billion parameters, outperformed most of the frontier systems tested.

The results demonstrate that carefully curated scientific training data and domain-specific optimization can be more important than model size for specialized biological applications. Compact models may also offer practical advantages for research institutions by lowering computational requirements and enabling more secure, cost-efficient deployment on local infrastructure.

To demonstrate that the models could support practical research beyond benchmark performance, the team embedded L-Qwen3.5-9B into Longevity Claw, a newly released open-source agentic platform for aging research.

Longevity Claw combines the specialized language model with scientific tools for:

·         Gene-set enrichment analysis

·         Biological aging-clock calculation

·         Population-level profiling

·         Evidence retrieval and synthesis

·         Candidate target evaluation and prioritization

The platform was designed to move beyond a conventional chatbot interface. Rather than responding only to individual questions, Longevity Claw can formulate and execute multi-step research workflows, use specialized analytical tools, evaluate intermediate findings and assemble evidence for candidate biological targets.

The researchers deployed the platform across 14 recognized hallmarks of aging, allowing it to investigate multiple interconnected mechanisms implicated in age-related decline. Through this autonomous workflow, Longevity Claw nominated 328 genes as potential targets for aging intervention.

When compared with an independently published reference set of experimentally supported aging-related targets, the candidates nominated by Longevity Claw demonstrated statistically significant enrichment of up to 5.6-fold, supporting the biological relevance of the platform’s results.

One of the nominated genes, KDM1A, was independently validated in a separate published study as a dual-purpose aging and cancer target. In that study, modulation of KDM1A extended lifespan in C. elegans. The independent finding provides an early indication that the platform can surface biologically credible targets that may not previously have been prioritized for longevity drug discovery.

The publication further demonstrates Insilico’s broader commitment to advancing pharmaceutical superintelligence. Insilico is releasing the benchmark, specialized models, training resources, evaluation code and Longevity Claw platform to enable independent testing, validation and further development by researchers worldwide.

By making these resources openly available, Insilico and its collaborators aim to provide a common foundation for measuring progress in AI-enabled aging research. The open framework may also help scientists distinguish systems that demonstrate genuine biological reasoning from those that primarily reproduce information contained in their training data.

https://www.cell.com/cell/fulltext/S0092-8674(26)00999-2

https://sciencemission.com/AI-in-aging-biology