AI news ·
CMU researchers release a long-horizon benchmark for AI drug discovery agents
CMU released SMDD-Bench, a 502-task benchmark that tests how well AI agents plan multi-step small-molecule drug design using separate scoring and verification tools.

Researchers at Carnegie Mellon University have released SMDD-Bench, a long-horizon benchmark designed to evaluate AI agents for drug discovery. The dataset contains 502 small-molecule design tasks, each paired with GPU-based oracles and isolated verifiers that score agent performance on multi-step planning problems.
For drug-discovery teams, the benchmark offers a more realistic test of an AI agent's ability to plan across extended sequences of decisions, moving beyond simpler single-step evaluations. The design separates the scoring mechanism from the verification step, which helps isolate different failure modes during agent development. However, success on this benchmark does not replace laboratory validation or clinical evidence.
Source: https://hub.harborframework.com/