Both versions are aligned block by block, in reading order: title, the essentials, then paragraph by paragraph. Where the translation merged or split a paragraph, the matching cell stays empty — we never pair two passages by guesswork.
Auf OpenRouter konnten Entwickler ab dem 20. August ein anonymes KI-Modell kostenlos ausprobieren. Ox Alpha wurde ausdrücklich um Feedback gebeten und stieg binnen weniger Tage an die Spitze des OpenRouter-Wochenrankings nach verarbeitetem Tokenvolumen: Bis einschließlich 25. August kamen rund 295 Millionen Anfragen zusammen, insgesamt wurden etwa 23,2 Billionen Token verarbeitet.
Starting on August 20, developers could try an anonymous AI model for free on OpenRouter. Ox Alpha explicitly asked for feedback and within a few days rose to the top of OpenRouter's weekly ranking by processed token volume: Through August 25, it received around 295 million requests, while approximately 23.2 trillion tokens were processed in total.
Der Reiz lag nicht nur im Preis. OpenRouter gibt das Kontextfenster mit rund einer Million Token an – genug Raum für große Mengen Quellcode und vorherige Arbeitsschritte. Ox Alpha verarbeitet außerdem Bilder und Videos und kann externe Werkzeuge aufrufen. Für KI-Agenten, die Software schreiben und dabei mit weiteren Programmen interagieren, senkt das die Hürde für lange, komplexe Tests.
The appeal was not just the price. OpenRouter lists a context window of around one million tokens — enough room for large amounts of source code and previous work. Ox Alpha can also process images and videos and call external tools. For AI agents that write software while interacting with other programs, this lowers the barrier for lengthy, complex tests.
Für den Hype sorgte zunächst ein kleiner Vergleich: In zehn Aufgaben des Coding-Benchmarks DeepSWE löste Ox Alpha acht und kam damit auf 80 Prozent. Der vollständig dokumentierte Lauf über alle 113 Aufgaben ergab jedoch nur 58,4 Prozent. Das liegt im Bereich etablierter kommerzieller Modelle, aber unter den Spitzenwerten von Claude Opus 5 mit 74 Prozent und GPT-5.6 Sol mit 73 Prozent. Die Community-Läufe sind wegen unterschiedlicher Testkonfigurationen zudem nur eingeschränkt vergleichbar.
The initial hype was driven by a small comparison: Ox Alpha solved eight of ten tasks in the DeepSWE coding benchmark, giving it a score of 80 percent. The fully documented run across all 113 tasks, however, produced a score of just 58.4 percent. That is within the range of established commercial models, but below the top scores of Claude Opus 5 at 74 percent and GPT-5.6 Sol at 73 percent. The community runs are also only partially comparable because of differences in test configurations.
Konkret bedeutet das: Entwickler konnten ein leistungsfähiges Modell mit langem Kontext und Werkzeugzugriff ohne Zugangshürde an echten Programmieraufgaben testen. Für Z.ai dürfte dieser verdeckte Start zugleich Erkenntnisse über Stärken, Schwächen und typische Einsatzgebiete geliefert haben. Eine unabhängige Einordnung durch Artificial Analysis steht noch aus; die bisherigen Zahlen sind daher ein Signal, kein abschließendes Urteil.
In concrete terms: Developers were able to test a capable model with a long context and tool access on real programming tasks without an access barrier. For Z.ai, the covert launch likely also provided insights into the model's strengths, weaknesses and typical use cases. An independent assessment by Artificial Analysis is still pending; the figures available so far are therefore a signal, not a final verdict.
Die Herkunft ist inzwischen geklärt. Das Open-Source-Werkzeug Modelprint fand technische Übereinstimmungen mit GLM-5.3, unter anderem beim Tokenizer; Z.ai bestätigte Bloomberg gegenüber, dass Ox Alpha eine neue Variante seiner GLM-Reihe ist. Mit dieser Bestätigung wird auch die Datenschutzfrage konkret: Eingaben und Antworten wurden an ein chinesisches Unternehmen übermittelt. Z.ai wollte die Modellgewichte noch am Mittwoch veröffentlichen und den gehosteten Zugang insgesamt eine Woche kostenlos anbieten; ein späterer Preis war noch offen.
The model's origins have now been clarified. The open-source tool Modelprint found technical similarities with GLM-5.3, including in the tokenizer; Z.ai confirmed to Bloomberg that Ox Alpha is a new variant of its GLM series. This confirmation also makes the privacy issue concrete: Inputs and responses were sent to a Chinese company. Z.ai said it would publish the model weights on Wednesday and offer hosted access free of charge for one week in total; a later price had not yet been determined.