Both versions are aligned block by block, in reading order: title, the essentials, then paragraph by paragraph. Where the translation merged or split a paragraph, the matching cell stays empty — we never pair two passages by guesswork.
Auf dem Test-PC stehen eine Radeon RX 9060 XT mit 16 GByte Grafikspeicher und 32 GByte DDR5-RAM. Wird dort Multi Token Prediction, kurz MTP, aktiviert, erzeugen vier getestete Large Language Models (LLMs) laut heise 25 bis 250 Prozent mehr Token pro Sekunde als ohne die Funktion. Das Versprechen lautet: mehr Tempo, ohne zusätzliche Hardware anzuschaffen.
The test PC features a Radeon RX 9060 XT with 16 GB of graphics memory and 32 GB of DDR5 RAM. When Multi Token Prediction, or MTP for short, is enabled, four tested Large Language Models (LLMs) generate 25 to 250% more tokens per second, according to heise, than without the feature. The promise: more speed without buying additional hardware.
Heise prüfte Gemma 4, Qwen3.5 und Qwen3.6 auf der Radeon-Karte. Für weitere Messungen kam eine Nvidia GeForce RTX 3090 zum Einsatz, ebenfalls mit Gemma 4 und Qwen3.6 sowie mit dem neuen Qwen3.8. Die Modelle gibt es laut Test in ausreichend kleinen Varianten, die auf gut ausgestatteten Gaming-PCs mit 16-GByte-Grafikkarten zügig laufen.
Heise tested Gemma 4, Qwen3.5 and Qwen3.6 on the Radeon card. Further measurements were conducted using an Nvidia GeForce RTX 3090, also with Gemma 4 and Qwen3.6, as well as the new Qwen3.8. According to the test, the models are available in variants that are small enough to run quickly on well-equipped gaming PCs with 16 GB graphics cards.
Für den Vergleich nutzte heise die KI-Engine Llama.cpp. Gemessen wurde mit zwei verschiedenen Prompts, jeweils mit und ohne MTP; zusätzlich optimierten die Tester einen Parameter. Die Ergebnisse fielen unterschiedlich aus: Bei einigen Modellen vervielfachte sich die Generierung, mindestens 25 Prozent mehr Tempo waren in den Versuchen aber immer drin.
For the comparison, heise used the AI engine Llama.cpp. Measurements were taken with two different prompts, each with and without MTP; the testers also optimized one parameter. The results varied: generation speed multiplied for some models, but the tests consistently produced at least 25% more speed.
Konkret heißt das: Wer ein unterstütztes lokales Modell auf einem passenden Gaming-PC betreibt, kann möglicherweise deutlich schneller Antworten erzeugen, ohne eine neue Grafikkarte zu kaufen. Das ist vor allem für die lokale Nutzung relevant, bei der die Rechenarbeit auf dem eigenen Rechner stattfindet. Welche Verbesserung ankommt, hängt jedoch vom Modell und weiteren Faktoren ab.
In concrete terms, anyone running a supported local model on a suitable gaming PC may be able to generate responses significantly faster without buying a new graphics card. This is particularly relevant for local use, where the computing takes place on the user's own computer. However, the improvement achieved depends on the model and other factors.
Die Grenze ist klar: MTP ist nicht für jedes lokale KI-Modell verfügbar. Die genannten Werte sind Messungen aus den heise-Versuchen und keine allgemeine Leistungszusage für jede Hardware. Für Nutzer mit Gemma 4, Qwen3.5, Qwen3.6 oder Qwen3.8 liefert der Test dennoch einen konkreten Ansatz, um vorhandene Rechenleistung besser auszunutzen.
The limitation is clear: MTP is not available for every local AI model. The figures cited are measurements from heise's tests and are not a general performance guarantee for every piece of hardware. For users with Gemma 4, Qwen3.5, Qwen3.6 or Qwen3.8, the test nevertheless provides a concrete way to make better use of existing computing power.