Tuesday, 1 September 2026

Aube.

News of progress
Single source

Floatboat Harness Beats Claude Opus 4.8 on Five Benchmarks

Languages for this article
Original · ENFR

Originally written in English. 2 languages available; yours is one click away.

A cheap model sat at the center of AOE Tech Labs’ test, but the model was not the variable that moved. The Chinese team ran DeepSeek-V4-Flash-0731 first on DeepSeek’s own minimalist Harness, then placed the same release inside Floatboat’s Harness. Only the execution system changed, and AOE reported that the Floatboat setup beat Claude Opus 4.8 on all five tested agent benchmarks.

The comparison used five third-party public leaderboards, including OpenAI-built BrowseComp. AOE says the Floatboat side ran in isolated physical sandbox environments with identical input parameters, while the control used the provider’s official release rather than a preview checkpoint. The claim is therefore narrower than “a better model”: it is a test of how much the surrounding system can extract from the same model.

The gap widened with the length of the work. Ordered from short to long tasks, AOE reported gains of 1.9%, 9.6%, 12.6%, 19.9% and 23.6%. Short tasks resemble one-shot question answering, where model quality carries more weight. Longer tasks require file access, repeated tool calls, preserved state, self-correction and rollback—the machinery that keeps an agent moving after its first answer is wrong or incomplete.

The price comparison sharpens the result. Claude Opus 4.8 lists at $5 per million input tokens and $25 per million output tokens. At a 3:1 input-output ratio, AOE calculated Floatboat’s blended cost at $0.175 per million tokens, against $10 for Opus 4.8—a reported 57.1 times gap. AOE also introduced the Harness Leverage Ratio, or HLR, which compares the gain from changing the Harness with the gain from moving to a stronger reference model. On DeepSWE, the HLR was 3.57.

So what changes in practice? Teams building daily-use agents may have another lever besides buying a larger model: improve the system that handles tools, files, memory and recovery. AOE Tech Labs has shipped Floatboat Desktop, FloatIM and FloatSchedule, but these benchmark results remain the company’s evaluation of its Harness in controlled environments. They show a strong case for investing in execution design; they do not establish that every real-world workflow will reproduce the same ranking or savings.

57.1 timesReported cost gap between Floatboat’s blended setup and Claude Opus 4.8

Sources — read the originals(Paris time)

PandailyEN
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in