Claude Fable 5.1 leads agent benchmarks
On Terminal-Bench-Science, a test for complex scientific workflows, Claude Fable 5.1 jumps from 24.7% to 52.6%. Anthropic released the model alongside Claude Mythos 5.1; Fable 5.1 is generally available and is designed primarily to handle long-running agent tasks, programming tasks and knowledge work more effectively.
The improvement also shows up in other evaluations. On Terminal-Bench 4.0, which assesses agentic programming tasks in a terminal environment, Fable 5.1 rises from 42.0% to 55.8%. AutomationBench simulates multistep office tasks across different applications and records an increase from 17.1% to 31.4%. Artificial Analysis gives it 66 points in the Intelligence Index at maximum reasoning effort—more than Fable 5 with 62, Opus 5 with 63 and GPT-5.6 Sol with 61 points. On other agentic knowledge-work tasks, however, Fable 5.1 is practically tied with Opus 5.
For developers, compute time remains a major cost factor. The regular prices for the application programming interface (API) are unchanged: $10 per million input tokens and $50 per million output tokens. Cache Reads—meaning the retrieval of previously processed content—are becoming cheaper: They fall from $1 to $0.25 per million tokens. Anthropic estimates that this will result in total costs around 25% lower for typical workloads and up to approximately 45% lower for highly agentic workloads. Artificial Analysis notes, however, that Fable 5.1 generates around 1.7 times as many output tokens as Fable 5; as a result, a task cost an average of $3.76.
Anthropic has also defined the safeguards for cyber tasks more precisely. Fable 5.1 may identify vulnerabilities in source code, while penetration testing, exploit generation and binary vulnerability scanning remain restricted. According to Anthropic, the safeguards in Claude Code are triggered around 60% less often on average than with Fable 5. Mythos 5.1 is available only to vetted users and organizations and has less strict safeguards; both new models carry a text watermark for the first time.
In practical terms, teams with long, multistep programming or knowledge-work tasks get a generally available model that makes significant gains across several agent benchmarks. The price is not automatically lower: Cheaper cache access helps primarily with reused context, while higher output-token consumption can reduce the benefit. Anthropic continues to recommend Opus 5 for most applications and recommends Fable 5.1 mainly for particularly demanding and long-running tasks.
Comments
Loading the thread…
Sign in to leave a comment. Sign in