Both versions are aligned block by block, in reading order: title, the essentials, then paragraph by paragraph. Where the translation merged or split a paragraph, the matching cell stays empty — we never pair two passages by guesswork.
Im Terminal-Bench-Science, einem Test für komplexe wissenschaftliche Arbeitsabläufe, springt Claude Fable 5.1 von 24,7 auf 52,6 Prozent. Anthropic hat das Modell gemeinsam mit Claude Mythos 5.1 veröffentlicht; Fable 5.1 ist allgemein verfügbar und soll vor allem lang laufende Agentenaufgaben, Programmieraufgaben und Wissensarbeit besser bewältigen.
On Terminal-Bench-Science, a test for complex scientific workflows, Claude Fable 5.1 jumps from 24.7% to 52.6%. Anthropic released the model alongside Claude Mythos 5.1; Fable 5.1 is generally available and is designed primarily to handle long-running agent tasks, programming tasks and knowledge work more effectively.
Der Zuwachs zeigt sich auch in anderen Prüfungen. Bei Terminal-Bench 4.0, das agentische Programmieraufgaben in einer Terminal-Umgebung prüft, steigt Fable 5.1 von 42,0 auf 55,8 Prozent. AutomationBench simuliert mehrstufige Büroaufgaben über verschiedene Anwendungen hinweg und verzeichnet einen Anstieg von 17,1 auf 31,4 Prozent. Artificial Analysis misst bei maximaler Denkleistung 66 Punkte im Intelligence Index – mehr als Fable 5 mit 62, Opus 5 mit 63 und GPT-5.6 Sol mit 61 Punkten. Bei anderer agentischer Wissensarbeit liegt Fable 5.1 allerdings praktisch gleichauf mit Opus 5.
The improvement also shows up in other evaluations. On Terminal-Bench 4.0, which assesses agentic programming tasks in a terminal environment, Fable 5.1 rises from 42.0% to 55.8%. AutomationBench simulates multistep office tasks across different applications and records an increase from 17.1% to 31.4%. Artificial Analysis gives it 66 points in the Intelligence Index at maximum reasoning effort—more than Fable 5 with 62, Opus 5 with 63 and GPT-5.6 Sol with 61 points. On other agentic knowledge-work tasks, however, Fable 5.1 is practically tied with Opus 5.
Für Entwickler bleibt die Rechenzeit ein harter Kostenfaktor. Die regulären Preise der Programmierschnittstelle (API) ändern sich nicht: 10 Dollar pro Million Eingabe-Token und 50 Dollar pro Million Ausgabe-Token. Günstiger werden Cache Reads, also das erneute Abrufen bereits verarbeiteter Inhalte: Sie fallen von 1 auf 0,25 Dollar pro Million Token. Anthropic schätzt dadurch rund 25 Prozent niedrigere Gesamtkosten bei typischen und bis zu etwa 45 Prozent niedrigere Kosten bei stark agentischen Arbeitslasten. Artificial Analysis stellt jedoch fest, dass Fable 5.1 rund 1,7-mal so viele Ausgabe-Token wie Fable 5 erzeugt; eine Aufgabe kostete dadurch im Schnitt 3,76 Dollar.
For developers, compute time remains a major cost factor. The regular prices for the application programming interface (API) are unchanged: $10 per million input tokens and $50 per million output tokens. Cache Reads—meaning the retrieval of previously processed content—are becoming cheaper: They fall from $1 to $0.25 per million tokens. Anthropic estimates that this will result in total costs around 25% lower for typical workloads and up to approximately 45% lower for highly agentic workloads. Artificial Analysis notes, however, that Fable 5.1 generates around 1.7 times as many output tokens as Fable 5; as a result, a task cost an average of $3.76.
Anthropic hat außerdem die Schutzmechanismen für Cyberaufgaben präziser gefasst. Fable 5.1 darf Schwachstellen in Quellcode identifizieren, während Penetrationstests, die Erzeugung von Exploits und das Scannen binärer Schwachstellen weiterhin eingeschränkt bleiben. In Claude Code greifen die Schutzmaßnahmen laut Anthropic im Schnitt rund 60 Prozent seltener ein als bei Fable 5. Mythos 5.1 ist nur für geprüfte Nutzer und Organisationen zugänglich und verfügt über weniger strikte Schutzmechanismen; beide neuen Modelle tragen erstmals ein Text-Wasserzeichen.
Anthropic has also defined the safeguards for cyber tasks more precisely. Fable 5.1 may identify vulnerabilities in source code, while penetration testing, exploit generation and binary vulnerability scanning remain restricted. According to Anthropic, the safeguards in Claude Code are triggered around 60% less often on average than with Fable 5. Mythos 5.1 is available only to vetted users and organizations and has less strict safeguards; both new models carry a text watermark for the first time.
Konkret bekommen Teams mit langen, mehrstufigen Programmier- oder Wissensaufgaben ein allgemein verfügbares Modell, das in mehreren Agenten-Benchmarks deutlich zulegt. Der Preis dafür ist nicht automatisch niedriger: Die günstigeren Cache-Zugriffe helfen vor allem bei wiederverwendetem Kontext, während ein höherer Ausgabe-Tokenverbrauch den Vorteil schmälern kann. Für die meisten Anwendungen empfiehlt Anthropic weiterhin Opus 5; Anthropic empfiehlt Fable 5.1 vor allem für besonders anspruchsvolle und langwierige Aufgaben.
In practical terms, teams with long, multistep programming or knowledge-work tasks get a generally available model that makes significant gains across several agent benchmarks. The price is not automatically lower: Cheaper cache access helps primarily with reused context, while higher output-token consumption can reduce the benefit. Anthropic continues to recommend Opus 5 for most applications and recommends Fable 5.1 mainly for particularly demanding and long-running tasks.