Both versions are aligned block by block, in reading order: title, the essentials, then paragraph by paragraph. Where the translation merged or split a paragraph, the matching cell stays empty — we never pair two passages by guesswork.
Ein Modell mit 2,8 Billionen Parametern passt nicht in den üblichen Entwicklerrechner. Moonshots Kimi K3 ist dennoch mit offenen Gewichten verfügbar und führt damit vor, wie weit sich offene Sprachmodelle inzwischen nach oben schieben. In Benchmarks landet es laut heise online regelmäßig auf Spitzenplätzen und antwortete in eigenen Tests ausführlich. Bei heiklen politischen Fragen blieb es allerdings auf Linie der Kommunistischen Partei Chinas.
A model with 2.8 trillion parameters does not fit on a typical developer computer. Moonshot's Kimi K3 is nevertheless available with open weights, showing how far open language models have advanced. According to heise online, it regularly ranks among the top performers in benchmarks and produced detailed answers in its own tests. On sensitive political questions, however, it stayed in line with the Chinese Communist Party.
Die Größe allein erklärt den Sprung nicht. Kimi K3 nutzt Kimi Delta Attention, einen Aufmerksamkeitsmechanismus für Sprachmodelle, bei dem sich lineare und globale Schichten im Verhältnis 3:1 abwechseln. Das soll besonders bei langen Texten Rechenzeit und Speicher sparen. Seine Mixture-of-Experts-Architektur verteilt Aufgaben außerdem auf 896 Experten — so viele waren in offenen Modellen bisher nicht verfügbar.
Size alone does not explain the leap. Kimi K3 uses Kimi Delta Attention, an attention mechanism for language models in which linear and global layers alternate in a 3:1 ratio. This is intended to save computing time and memory, particularly with long texts. Its Mixture-of-Experts architecture also distributes tasks among 896 experts — more than have previously been available in open models.
Auch beim Speicher setzt Moonshot an. Die Gewichte wurden in das kompakte MXFP4-Format überführt; dadurch sinkt der Bedarf gegenüber dem häufig verwendeten bfloat16 auf ein Viertel. Der Engpass verschwindet damit nicht: Für Kimi K3 sind selbst bei kleinen Kontexten mindestens 1,5 TByte RAM erforderlich. Eine lokale Ausführung bleibt damit wenigen Firmen vorbehalten, ist technisch aber möglich. Für die meisten dürfte zunächst der angebotene Clouddienst der praktikablere Zugang sein.
Moonshot is also tackling memory requirements. The weights were converted to the compact MXFP4 format, reducing requirements to one quarter of those for the widely used bfloat16. This does not eliminate the bottleneck: Kimi K3 requires at least 1.5 TByte of RAM, even with small contexts. Local execution therefore remains reserved for relatively few companies, although it is technically possible. For most users, the cloud service on offer is likely to be the more practical option for now.
Parallel kommen weitere große Modelle aus China hinzu. Alibaba stellte Anfang August Qwen3.8-Max mit 2,4 Billionen Parametern vor, zunächst nur per API; die Gewichte sollen folgen. Das Modell soll laut Alibaba fast 16 Tage autonom an einem Coding-Projekt gearbeitet, wissenschaftliche Artikel überprüft sowie den darin verwendeten Programmcode verstanden und optimiert haben. Diese Angaben und die starken Benchmarkwerte müssen unabhängige Tests noch bestätigen. DeepSeek wiederum zeigt mit einem deutlich kleineren Flash-Modell, dass Modellgröße nicht automatisch die beste Leistung bedeutet.
Meanwhile, more large models are emerging from China. At the beginning of August, Alibaba introduced Qwen3.8-Max with 2.4 trillion parameters, initially via API only; the weights are expected to follow. According to Alibaba, the model worked autonomously on a coding project for almost 16 days, reviewed scientific papers, and understood and optimized the code used in them. Independent tests still need to confirm these claims and the strong benchmark results. DeepSeek, in turn, is showing with a significantly smaller Flash model that model size does not automatically mean the best performance.
Und was ändert das konkret? Unternehmen bekommen mehr Auswahl zwischen geschlossenen Diensten und Modellen mit offenen Gewichten. Wer sensible Daten lokal verarbeiten will, kann Kimi K3 grundsätzlich einsetzen — muss dafür aber eine Infrastruktur bezahlen, die laut den genannten Anforderungen für viele Firmen zu teuer ist. Für kleinere Teams verschiebt sich der praktische Vorteil deshalb vorerst in die Cloud. Die technische Richtung ist trotzdem klar: Offene Modelle werden leistungsfähiger, während ihre Hardwarekosten den Zugang weiterhin begrenzen.
So what does this change in practical terms? Companies are gaining more choice between closed services and models with open weights. Those wishing to process sensitive data locally can in principle use Kimi K3 — but they would have to pay for infrastructure that, according to the stated requirements, is too expensive for many companies. For smaller teams, the practical advantage is therefore shifting to the cloud for the time being. The technological direction is nevertheless clear: open models are becoming more capable, while their hardware costs continue to limit access.