Evo 2 model predicts harmful mutations and generates genome-scale DNA

The open-source AI Evo 2 predicts harmful genetic mutations and generates full DNA sequences. It achieved a 90% success rate designing sequences for human cell lines.

Categorized in: AI News Science and Research
Published on: Jul 28, 2026
Evo 2 model predicts harmful mutations and generates genome-scale DNA

A new artificial intelligence model called Evo 2 can predict whether a genetic mutation is harmful and generate realistic genome-length DNA sequences from scratch, without any task-specific training. The model, developed by a large team at the Arc Institute, Stanford University, and NVIDIA, beat competing tools on several key benchmarks and is fully open source. The research appeared in the journal Nature.

Learning the grammar of DNA across the tree of life

Evo 2 was trained on nearly 9 trillion curated DNA base pairs from bacteria, archaea, eukaryotes, and bacteriophages. This broad dataset lets the model work across virtually every known category of life on Earth. Unlike most biology AI tools that focus on a single domain, Evo 2 learned the underlying logic of DNA from the full library.

The model can process up to 1 million genetic letters at once-the longest context window of any DNA-focused AI model. Many important DNA patterns only become clear across long stretches, so this extended reach helps the model capture signals that shorter-context systems miss. Evo 2 comes in two sizes, with 7 billion and 40 billion parameters, and runs on a new architecture that makes the million-letter window possible.

Zero-shot prediction of harmful mutations

Evo 2 can judge genetic mutations without any prior coaching on what a harmful mutation looks like. This zero-shot prediction relies entirely on patterns learned during training. Tested on ClinVar, a database of human genetic variants with known clinical significance, the model outperformed competing unsupervised methods for mutations outside protein-coding regions-areas that have long been harder to interpret. It also ranked first among zero-shot models on a test focused on mutations that alter how genes are read and processed.

The BRCA1 gene, linked to breast and ovarian cancer, provided a clear example. Evo 2 distinguished harmful mutations from benign ones with high accuracy. When its internal representations were used to train a lightweight classifier, that system reached an area under the precision-recall curve of 0.88, far ahead of simpler baseline approaches.

Generating entire genomes from scratch

Beyond reading DNA, Evo 2 can write it. Given a short starting sequence, the model can produce full genes or genome-length sequences that look biologically realistic. Prompted with a segment of the human mitochondrial genome, it generated sequences about 16,000 letters long that contained the expected protein-coding genes, tRNAs, and rRNAs in the correct order. When prompted with a piece of Mycoplasma genitalium, it produced nearly complete genomes close to 580,000 letters long, with about 70% of predicted genes matching known protein families-an improvement over earlier models.

The team also used Evo 2 to design DNA sequences that produce specific chromatin accessibility patterns, which control whether genes are switched on or off. Paired with guiding prediction tools, the model generated sequences several thousand letters long, synthesized them, and inserted them into mouse and human cells. In mouse embryonic stem cells, the designs hit the intended patterns with AUROC scores between 0.92 and 0.95. In human cell lines, more than 90% of the designs produced the desired result. The generated sequences have not yet been shown to function in living cells, but their structural accuracy across diverse organisms is a notable achievement.

Open source with safety and fairness guardrails

Evo 2 and its training data are fully open source, allowing any researcher to use them. The team deliberately excluded sequences from eukaryote-infecting viruses from the training data, and the model performs poorly on human viral sequences, reducing the risk of misuse. Fairness checks found no systematic disadvantage for people of non-European ancestry, addressing a known problem with many existing genetic tools.

Evo 2 is a direct successor to Evo 1, which was trained only on bacterial genomes. The leap from a single-branch tool to one spanning the entire tree of life represents a major expansion in scope and capability.

Why this matters for science and research

Evo 2 shows how AI for Science & Research is moving beyond narrow, single-purpose tools toward models that can read, interpret, and write genomic blueprints across the full range of living things. For researchers in genetics and molecular biology, a single, open-source model can now tackle tasks from mutation effect prediction to genome design, with no extra training. The ability to generate realistic DNA sequences and predict chromatin accessibility patterns could accelerate experiments in synthetic biology and gene regulation, even if the designs still need to be validated inside living cells. The model's broad scope and zero-shot capabilities change what researchers can attempt, from probing the fundamental grammar of DNA to engineering new biological functions.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)