Lab variations in catalyst tests can mislead scientific AI models, study finds

Four labs testing the same rhodium catalyst produced significantly different results, a new study finds-variation that could teach AI models unreliable patterns.

Categorized in: AI News Science and Research
Published on: Aug 15, 2026
Lab variations in catalyst tests can mislead scientific AI models, study finds

Four laboratories testing the same carbon monoxide-producing catalyst produced noticeably different results - a finding that carries direct consequences for researchers training AI models on data that comes from multiple sources.

The round-robin study, published in Nature Catalysis, brought together teams from SLAC National Acceleratory Laboratory, Penn State, Stanford University, and the University of California, Santa Barbara. Each lab followed the same protocols and used the same rhodium-based catalyst to measure how much carbon monoxide and methane the reaction produced. The variation between labs was significant enough that an AI model trained on the combined datasets would have learned unreliable patterns.

"An AI model is only as good as the data used to create it, which will undoubtedly come from multiple sources," said Robert Rioux, Friedrich G. Helfferich Professor of Chemical Engineering at Penn State and co-author on the study. "This study represents the first round-robin study focused on heterogeneous catalysis that quantifies the uncertainty from studies done The between labs, and the impact this uncertainty will have on future data-driven AI modeling."

Researchers are increasingly turning to AI models to accelerate catalyst discovery for converting carbon dioxide into fuels. A well-trained model could predict how a catalyst performs under conditions like temperature and reaction time. That eliminates the need for months of physical testing, much of which occurs under conditions that are hard to recreate in a lab, such as prolonged catalyst use over years in high temperatures with impurities.

But the models require large volumes of consistent training data.

Sources of variation

After the four labs initially shared their results, the teams traced the discrepancies to several controllable factors. The largest variable: how vigorously each lab stirred or shook the reaction mixture.

"It was a bit of an eye-opener," said SLAC staff scientist Adam Hoffman, senior author of the study. "This experience shines light on the practical challenges of including real-world data into machine learning models."

The teams didn't just flag the problem. They standardized reactor design, operating protocols, and experimental conditions, then ran the tests again. The results aligned more closely, demonstrating that reproducibility is achievable when labs coordinate precisely, not just in theory but in practice.

"Our findings are a reminder to exercise caution about what information we feed into a machine-learning model, and how consistency of experimental data can influence the reliability of the outcomes," said Selin Bac, a postdoctoral coefficient at University of California, Santa Barbara and first author on the study.

Practical guidance for experimentalists

The study offers a workflow for researchers building AI models on experimental data: run the same experiments across different labs, measure the variation, identify the known sources of discrepancy, and apply corrective controls before feeding data into a model.

That means an AI model trained on historical snapshot from a single lab is risk of silently overfitting to that lab's specific conditions. Multi-lab data collection has the advantage of broader coverage conditions, but it also introduces the need for documented protocols and round-robin validation.

The findings point toward a practical need for anyone pursuing AI for Science & Research: data reports must be standardized at the source, not corrected post-hoc. And for those looking to refine those skills, that principle is essential across AI Data Analysis Courses as well - models are only as reliable as the consistency of what they are fed.

Why this matters for research scientists

If you work in experimental catalysis, materials science, or any field where an AI model draws on published results from multiple laboratories, the immediate takeaway is direct: quantify cross-lab variation before prioritizing an AI model on your own research decisions.

Adopt round-robin validation in your own multi-site collaborations. Standardize stirring rates, reactor geometry, and reporting protocols. The study's authors suggest exactly that. Their final data sets, after standardized corrections arrived, converged - but only because the labs shared the painstaking work of finding the sources of the initial mismatch.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)