Researchers at the Stowers Institute for Medical Research have developed a method called PISA that traces what genomic AI models actually learn, not just what they predict. The technique, detailed Aug. 25, 2026, lets scientists separate experimental bias from real biology and could point to the DNA sequences behind gene regulation and genetic disease.
PISA, which stands for pairwise influence by sequence attribution, works by tracing a deep-learning model's prediction at any single DNA base back to every other base that influenced it. The result is a base-pair-resolution map of the model's learned patterns, rather than just its output.
Finding bias in the data
The team tested PISA on nucleosome mapping data, which tracks how DNA wraps around histone proteins. They spotted a technical bias baked into the dataset and mathematically removed it, revealing DNA sequences that position nucleosomes.
Those same sequences unexpectedly marked the boundaries of larger 3D chromatin domains - structures normally identified only through expensive, sequencing-intensive methods. That finding suggests PISA can surface insights that would otherwise require significantly more resources.
From prediction to mechanism
Biologists have long wanted to connect specific DNA sequences to gene regulation and disease. PISA offers a path from AI prediction to mechanistic insight, pointing researchers toward sequence elements that warrant deeper investigation.
Julia Zeitlinger, Ph.D., who led the work, said PISA could help make AI a more powerful discovery tool for biology, giving scientists a clearer path from prediction to mechanistic insight and experimental design.
Why this matters for IT and development professionals
For developers building or deploying AI models in scientific domains, PISA demonstrates a practical approach to model interpretability: instead of accepting a prediction at face value, trace it back to the input features that drove it. The method also shows how interpretability tools can double as data-quality checks - catching technical biases that skew model training.
Teams working on genomic or other high-dimensional data can apply the same principle to their own pipelines, using attribution methods to audit training data and refine model focus.
Your membership also unlocks: