Apodex launches TRACES benchmark to evaluate AI for scientific discovery

Apodex released TRACES, a benchmark testing AI on open-ended scientific problems with 423 cases spanning 561 industries, scoring both process and outcomes across six capabilities.

Categorized in: AI News Science and Research
Published on: Aug 18, 2026
Apodex launches TRACES benchmark to evaluate AI for scientific discovery

Apodex on August 18, 2026, released TRACES, a benchmark for evaluating whether AI systems can handle open-ended scientific problems, use tools, and adapt to feedback while producing verifiable findings. The benchmark is designed around problems where the correct answer may not be known, which distinguishes it from standard AI tests that rely on static datasets and predefined answers.

TRACES places AI systems inside executable environments where they can observe, act, receive feedback, and revise their approach. Depending on the problem, an environment might include scientific literature, structured datasets, code execution, simulators, folding engines, or experimental feedback. The benchmark evaluates both the final outcome and the process the system used to reach it.

"TRACES is a benchmark designed specifically to evaluate progress in discoverative AI. It brings together sophisticated efforts in scouting high-value real-world problems, assembling the tools and data needed to build executable environments, and developing a novel scoring system that evaluates not only outcomes but also the discovery process," said Dr. Sheng Wang, Lead Scientist at Apodex.

Six capabilities under review

Apodex evaluates systems across six categories: Tools, Repair, Alternatives, Coherence, Evidence, and Scope. Tools checks whether a system selects and correctly uses external tools. Repair measures its ability to fix errors after receiving feedback. Alternatives assess whether competing hypotheses are considered and revised as evidence accumulates. Coherence checks logical consistency across longer problem-solving sequences. Evidence verifies that conclusions are grounded in observations or citations, and Scope measures whether a system defines where a conclusion applies and where it does not.

Process outcomes are tied to hundreds of intermediate judgments, which is why TRACES weighs the process itself. Brian Wang, AI Research Scientist at Apodex, said final results alone are insufficient for measuring progress. "In scientific discovery, the answer is one line at the end of hundreds of judgments - what to try next, when the evidence is enough, when to abandon a hypothesis. The capability lives there, and scoring only the last line throws away almost all of it."

The benchmark combines outcome and process verification. The outcome verifier grades a submission against hidden ground truth where available. The process verifier assesses whether the reasoning and evidence meet criteria. Evaluators match observed behaviour against written scoring descriptions, with each finding linked to specific steps in the recorded trajectory. An independent model reviews the evaluation, and disputes trigger re-scoring and adjudication.

What the benchmark is designed to measure

Apodex's broader TRACES framework includes 423 high-value problems drawn from a survey spanning 561 industries across 16 sectors. The company said the benchmark is built to distinguish between systems that can reproduce established scientific knowledge and those that can work through uncertainty, test hypotheses, and adapt to new evidence. Teams developing models, agent systems, or solver frameworks can participate. Researchers can also submit scientific problems to be converted into executable evaluation environments.

For working scientists, the benchmark's approach has immediate relevance. The distinction between systems that retrieve known of thinking about capability evaluation within your own work. Resources you can calibrate your research workflows against these criteria - AI for Science & Research covers advances in this area, and the AI Learning Path for Research Scientists can help you evaluate which a model's final assessment is just the last step in a long chain of decisions.

Why this matters for science researchers

TRACES offers a way to measure AI systems in conditions that resemble real research: open problems, uncertain information, and iterative trial and error. If you are considering AI tools for research, the benchmark provides criteria for evaluating whether a system can handle the unexpected, not just process structured data. Tools that reproduce known work are everywhere. The measurement gap is in how AI systems behave when no one knows the answer, and TRACES is built to test exactly that.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)