Wednesday, 2 September 2026

Aube.

News of progress
LabSingle source

EPFL's GOLLuM needs over 40% fewer experiments in benchmarks

Languages for this article
Original · ENESFRITPT

Originally written in English. 5 languages available; yours is one click away.

At EPFL's Laboratory of Artificial Chemical Intelligence, Bojana Ranković and Philippe Schwaller tested GOLLuM with a narrow budget: 50 experiments to search for strong chemical, materials and molecular conditions. In that evaluation, 36.3% of the conditions it tried landed in the top 5% of all possible outcomes, while the method matched traditional optimization using over 40% fewer experiments.

Scientists already use Bayesian optimization to decide which experiment to run next. It learns from past results, predicts promising options and estimates uncertainty, helping teams avoid spending time and money on weak candidates. The problem is portability: a model that chooses chemical reactions efficiently may not fit materials or molecular design. Large language models carry broad scientific knowledge, but their confidence is a poor guide to whether a suggestion is correct.

GOLLuM adds a Gaussian process — a probabilistic model used to score performance and uncertainty — as a built-in doubt detector. The large language model is not asked to choose experiments on its own; instead, it learns from the Gaussian process's scores. As results accumulate, the system reorganizes its search space so conditions that behave alike sit closer together, while conditions with different results move apart. The method works directly from a plain-English description of the experimental procedure.

The researchers evaluated GOLLuM on 23 benchmark tasks spanning organic synthesis, analytical and process chemistry, materials, catalysis and molecular property optimization. Every run began with 10 low-performing observations, and the team used the same configuration across the tasks rather than tuning it separately. GOLLuM ranked first on average and beat Bayesian optimization with expert-designed descriptors on the reported measures. By contrast, direct prompting produced failure rates from 10% to around 80%, including invented chemical structures, repeated conditions and suggestions outside the permitted search space.

And so, concretely: research teams could use a language model's broad scientific representations without trusting its answers alone, potentially testing promising recipes with fewer laboratory runs.

36.3%Conditions landing in the top 5% within a 50-experiment budget

Sources — read the originals(Paris time)

Phys.org — TechnologyEN
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in