Fudan University and Tencent Hunyuan released ExplorationBench on September 25, 2026, a new benchmark designed to measure whether large language models can learn entirely new rules through interaction, rather than simply retrieving facts from pre-training data. The benchmark addresses a critical gap in current AI evaluation: most tests measure a model's existing knowledge, but fail to assess its ability to acquire new capabilities when facing unknown environments.
From static knowledge to active learning
Traditional benchmarks test what a model already knows. The researchers argue this misses a fundamental requirement for reliable AI agents: the ability to improve through experience. They draw a parallel to human learning, where an engineer understands a system better after fixing ten bugs, and a scientist advances through hypothesis and experimentation. For an AI agent to work long-term, it must convert past interactions into future competence.
This work builds on two earlier projects. EvaLearn, published at NeurIPS 2025, challenged the idea that test questions are independent, organizing 648 problems into sequential tasks. It found that models with high static scores didn't always learn effectively from history, sometimes suffering from negative transfer. CL-bench, released in early 2026, tested how well models could learn from complex, self-contained contexts like fictional laws or scientific data. Even with perfect context provided, top models achieved only 17.2% average accuracy, suggesting that simply having information isn't the same as using it.
ExplorationBench takes this further. It removes the safety net of provided knowledge. Models enter a sandbox environment with only a flawed manual and fixed examples. They must design experiments, read feedback, and deduce the underlying rules themselves. To avoid contamination from pre-training data, the team created two "Alien Worlds" with deliberately altered logic and syntax.
Testing in alien worlds
The first world, AlienCode, hides 31 rule changes in a programming language that looks familiar. Operations like SHATTER perform multiplication instead of destruction, and integer literals are XORed with 27. The second world, AlienLogic, modifies 24 natural deduction rules, making standard logical steps invalid while allowing previously non-compliant transformations. Across both worlds, there are 140 test questions verified by deterministic checkers, not by another AI judge.
The evaluation protocol requires four rounds of exploration. After each round, the system's session is copied, tools are disabled, and the model must list the rules it believes are true and answer held-out questions. This separates the ability to discover rules from the ability to apply them. The team ran each model three times to account for variability, reporting the best score from those three trajectories.
Feedback beats thinking alone
The results show that interaction drives learning. In AlienCode, models started with a maximum accuracy of 15.7% after reading examples. After four rounds of exploration, the best score reached 89.0%, with seven of ten models exceeding 60%. In AlienLogic, scores rose from 32.9-51.9% to 58.1-83.8%. However, the researchers tested whether this improvement came from exploration or just extra thinking time. They compared five settings, including one where models had multiple rounds to think but received no environmental feedback.
In the no-feedback setting, AlienCode accuracy stalled between 0.5% and 11.0%. Three models actually declined in AlienLogic. This demonstrates that reasoning alone, without new evidence from the environment, often reinforces incorrect priors. One model abandoned correct rules it had initially deduced after several rounds of unguided thinking. The study highlights that gaining feedback is more effective than extensive internal deliberation.
The source of the experiments also matters. When models had to design their own probes, AlienCode accuracy was significantly higher than when they were given random or pre-selected probes. Autonomous exploration outperformed "post-hoc replay," where the model used the best probes from a previous successful run but didn't decide which to use next. This suggests that the process of hypothesizing and testing is integral to learning, not just the data collected.
Discovering rules vs. using rules
ExplorationBench reveals a disconnect between knowing a rule and applying it. In AlienLogic, providing the complete rule set upfront allowed models to achieve 93-97% accuracy. No model's exploratory best trajectory exceeded this. The bottleneck was clearly discovery. In AlienCode, however, the opposite occurred. For seven of ten models, autonomous exploration yielded higher scores than being given the rules upfront. GPT-5.6 Sol scored 88.1% through exploration versus 69.5% with the rules provided. The exploration process itself trained the models to use the rules correctly.
Even when models could state a rule correctly, they often failed to apply it in complex contexts. If a model's final rule report contained all necessary rules, it still only answered the specific questions correctly 73.4% of the time. One model correctly identified an indexing offset rule but continued to apply it incorrectly inside function bodies, leading to errors. Knowing the rule is distinct from operationalizing it in code or logic.
The study also warns against relying on single-point scores. Exploration is unstable. Kimi K3's final accuracy in AlienCode varied from 5.7% to 79.0% across three identical runs. The ranking of models changes depending on whether you look at the best result, the average, or the median. Furthermore, progress is often non-linear; a single breakthrough in one round can account for nearly half of the total gain. In some cases, continued exploration led to regression, with models convincing themselves of incorrect workarounds after initially finding the right path.
Why this matters for researchers and educators
For those working in science and education, this benchmark signals a shift in how we evaluate and build AI tools. Current models are not just static knowledge bases; their utility depends on their ability to adapt to novel, formalized systems. If you are designing AI for laboratory automation or personalized learning, you cannot assume that providing context is sufficient. The model must be able to test hypotheses and revise its understanding based on strict environmental feedback. Future development should focus on "consolidation"-how agents can turn short-term exploratory success into long-term memory or reusable tools, a direction the authors suggest as the next critical research area.
Your membership also unlocks: