New framework combines language models with a doubt detector to find optimal experimental recipes faster

GOLLuM cuts the number of experiments needed by over 40% while matching Bayesian optimization across 23 chemistry and materials science benchmarks.

Categorized in: AI News Science and Research
Published on: Sep 02, 2026
New framework combines language models with a doubt detector to find optimal experimental recipes faster

Researchers at EPFL's Laboratory of Artificial Chemical Intelligence have built a framework that trains large language models to find optimal experimental conditions faster than traditional methods. The system, called GOLLuM, cuts the number of experiments needed by over 40% while matching the performance of Bayesian optimization across 23 benchmark tasks in chemistry and materials science.

The work addresses a persistent bottleneck in scientific research: exploring vast combinatorial spaces of molecules, materials, and reactions without testing every possibility in the lab. Bayesian optimization helps by learning from previous results and predicting which options are worth testing next. But those models rarely transfer between scientific domains. A model tuned for chemical reactions does not help with materials design, so each new problem starts from scratch.

Large language models offer a different path. They already encode broad scientific knowledge and can work with text-based descriptions of experiments. The problem is reliability. LLMs hallucinate, and their apparent confidence does not track with whether a suggestion is correct. That can send researchers down expensive dead ends.

How GOLLuM combines LLMs with a "doubt detector"

Bojana Ranković and Philippe Schwaller developed a method that pairs an LLM with a Gaussian process, the probabilistic model commonly used in Bayesian optimization. The Gaussian process acts as a doubt detector, scoring both the predicted performance and the uncertainty of each option. Rather than asking the LLM to choose experiments directly, GOLLuM trains it using those uncertainty signals.

"Language models are notoriously bad at knowing when they're wrong. In GOLLuM, that uncertainty becomes the very signal that trains them," said Ranković, who created the method during her doctoral research at EPFL. "The model reorganizes the search space until experiments that behave alike sit close together."

As the system learns, it adjusts an internal map of the search space. Conditions that produce similar results cluster together, while those with different outcomes drift apart. High-yielding conditions separate cleanly from low-yielding ones, giving researchers a human-interpretable view of the experimental landscape.

Performance across 23 benchmarks

The team evaluated GOLLuM on tasks covering organic synthesis, analytical and process chemistry, materials and catalysis, and molecular property optimization. Each run began with just ten low-performing observations. The researchers used the same GOLLuM configuration across all 23 benchmarks without tuning it separately for each problem.

Within a budget of 50 experiments, 36.3% of the conditions GOLLuM tested fell inside the top 5% of all possible outcomes. The traditional method achieved 29.7%. Overall, GOLLuM matched conventional Bayesian optimization while using over 40% fewer experiments, converging faster toward high-performing solutions.

The team also tested what happens when LLMs choose experiments directly, without the Gaussian process coupling. Performance was inconsistent. Failure rates ranged from 10% to roughly 80%, with problems including invented chemical structures, repeated conditions, and suggestions outside the permitted search space.

A different role for language models in the lab

The method, published in Nature Machine Intelligence, shifts how language models participate in experimental research. Instead of relying on their answers alone, researchers can combine their broad representations of scientific information with probabilistic models that explicitly account for uncertainty.

"It's a highly impactful technique that enables us to start experimental optimization campaigns from day one instead of discussing how to best describe experiments and compute descriptors for six months," said Schwaller. "The key is that we optimize directly on a plain English representation of the experimental procedure."

This approach sidesteps the lengthy process of engineering domain-specific descriptors for each new problem. Researchers describe their experiments in plain English, and GOLLuM works directly on that representation. For scientists exploring AI for Science & Research, the framework offers a sample-efficient optimizer that works across scientific domains without custom tuning.

Why this matters for science and research professionals

Experimental optimization campaigns often begin with months of discussion about how to describe experiments and compute meaningful descriptors. GOLLuM removes that upfront engineering step. Researchers can start testing promising conditions immediately, using natural language descriptions rather than hand-crafted numerical representations. For those building deeper expertise, an AI Learning Path for Research Scientists can provide the foundation needed to apply these probabilistic methods in their own labs. The 40% reduction in required experiments translates directly to lower costs, faster project timelines, and more efficient use of limited laboratory resources.


Get Daily AI News

Your membership also unlocks:

700+ AI Courses
700+ Certifications
Personalized AI Learning Plan
6500+ AI Tools (no Ads)
Daily AI News by job industry (no Ads)