Anthropic: Fermat’s Last Theorem formalized in Lean
The proof comprises 13 million lines of Lean code and around 29,500 intermediate theorems.
Models, agents, applications: artificial intelligence put to work on real problems.
Here, artificial intelligence is judged on the evidence: what it does, for whom, at what cost. The section follows models and their measurable progress, agents doing real work, and above all deployments — a hospital screening more patients, a power grid better balanced, a public service answering faster. AI-as-spectacle, productless funding rounds and prophecy get no space.
We favour documented use cases with the numbers attached: how many examinations, what error rate, what savings. Non-Western players — Chinese, Indian, Emirati labs — are covered on the same footing as American ones, because much of the real deployment happens there. Every article cites its sources and separates the prototype from the service in production.
The proof comprises 13 million lines of Lean code and around 29,500 intermediate theorems.
The generally available model improves from 24.7% to 52.6% on Terminal-Bench-Science.
AOE Tech Labs changed only the Harness around DeepSeek-V4-Flash-0731 in its controlled comparison.
Janelia neuroscientist Arco Bast built the Model Hardware Standard with Anthropic and Claude Code.
The expansion launches with 37 preconfigured features for sales analysis and meeting preparation.
Claude will soon be able to read websites, click through them and fill out forms in a separate side panel.
Anthropic’s update lets Claude carry project details from conversations into Cowork without a fresh briefing.
Version v0.1.0-rc.8 adds native image requests and mixed text-image tasks.
The custom harness raised Claude Opus 5 from 30% to 100% on an interactive reasoning benchmark.
Maintaining its 700-billion-parameter scale, GLM-5.3 relies primarily on post-training to improve coding and agent capabilities.