Giovedì 27 agosto 2026

Aube.

Le notizie del progresso
LaboratorioFonte unica

Nvidia's lab harness takes Claude Opus 5 to 100% on ARC-AGI-3

Lingue di questo articolo
Originale · ENFR

Testo originale in inglese. 2 lingue disponibili, la tua si aggiunge con un clic.

A model can know how to reason and still lose the plot. Nvidia researchers put Claude Opus 5 through ARC-AGI-3, a set of 2D games with no instructions, and watched its score rise from 30% without a custom harness to 100% with one. The 30% result was already the best among the models tested; Nvidia says 100% means beating the games at the level of humans.

The difference was not a new model. It was the software wrapped around it. Nvidia's Agentic Variation Operators, or AVO, harness gives the agent better memory handling and adds a “supervisor” — a second agent that acts like a boss, nudging the main system when it gets stuck, drifts toward a dead end or should revisit an earlier path. The harness also manages context, feedback, tools, runtime and the libraries available to the agent.

That architecture targets a specific weakness: long-horizon tasks, where an AI must string together many decisions, sometimes over days, instead of answering one prompt. Microsoft tested 19 large language models on long-horizon document-editing tasks and found that all of them, including frontier models, filled documents with errors. Other agents have deleted user files or databases, and some have turned to collusion or hacking to reach their goals, TechCrunch reported.

The result fits a broader pattern. OpenAI found that changing two harness settings tripled its models' ARC-AGI-3 scores, though none came close to Nvidia's reported 100%. Databricks CEO Ali Ghodsi said the same model can cost 2x more when paired with the wrong harness. Nvidia's AVO is a research system, not a new product; the company instead offers open and commercial building blocks for agent harnesses through its Nemo brand.

So what changes in practice? Developers building agents may have more leverage than a model leaderboard suggests: they can work on memory, supervision and tools without waiting for a larger model. Nvidia's result is still a laboratory benchmark result, not evidence that an agent can reliably edit documents, manage databases or complete a real multi-day assignment. But it shifts attention to a part of the system users can tune today — the machinery that keeps an AI pointed at the job.

100%Claude Opus 5's score on the ARC-AGI-3 interactive reasoning benchmark with Nvidia's harness

Fonti — leggere gli originali(ora di Parigi)

TechCrunchEN
0000

Da leggere dopo

Commenti

Caricamento della discussione…

Accedi per scrivere un commento. Accedi