Tested: Local LLMs up to 250% faster with MTP
The test PC features a Radeon RX 9060 XT with 16 GB of graphics memory and 32 GB of DDR5 RAM. When Multi Token Prediction, or MTP for short, is enabled, four tested Large Language Models (LLMs) generate 25 to 250% more tokens per second, according to heise, than without the feature. The promise: more speed without buying additional hardware.
Heise tested Gemma 4, Qwen3.5 and Qwen3.6 on the Radeon card. Further measurements were conducted using an Nvidia GeForce RTX 3090, also with Gemma 4 and Qwen3.6, as well as the new Qwen3.8. According to the test, the models are available in variants that are small enough to run quickly on well-equipped gaming PCs with 16 GB graphics cards.
For the comparison, heise used the AI engine Llama.cpp. Measurements were taken with two different prompts, each with and without MTP; the testers also optimized one parameter. The results varied: generation speed multiplied for some models, but the tests consistently produced at least 25% more speed.
In concrete terms, anyone running a supported local model on a suitable gaming PC may be able to generate responses significantly faster without buying a new graphics card. This is particularly relevant for local use, where the computing takes place on the user's own computer. However, the improvement achieved depends on the model and other factors.
The limitation is clear: MTP is not available for every local AI model. The figures cited are measurements from heise's tests and are not a general performance guarantee for every piece of hardware. For users with Gemma 4, Qwen3.5, Qwen3.6 or Qwen3.8, the test nevertheless provides a concrete way to make better use of existing computing power.
Comments
Loading the thread…
Sign in to leave a comment. Sign in