Kimi K3 comes with 2.8 trillion parameters and open weights, but local operation remains expensive
A model with 2.8 trillion parameters does not fit on a typical developer computer. Moonshot's Kimi K3 is nevertheless available with open weights, showing how far open language models have advanced. According to heise online, it regularly ranks among the top performers in benchmarks and produced detailed answers in its own tests. On sensitive political questions, however, it stayed in line with the Chinese Communist Party.
Size alone does not explain the leap. Kimi K3 uses Kimi Delta Attention, an attention mechanism for language models in which linear and global layers alternate in a 3:1 ratio. This is intended to save computing time and memory, particularly with long texts. Its Mixture-of-Experts architecture also distributes tasks among 896 experts — more than have previously been available in open models.
Moonshot is also tackling memory requirements. The weights were converted to the compact MXFP4 format, reducing requirements to one quarter of those for the widely used bfloat16. This does not eliminate the bottleneck: Kimi K3 requires at least 1.5 TByte of RAM, even with small contexts. Local execution therefore remains reserved for relatively few companies, although it is technically possible. For most users, the cloud service on offer is likely to be the more practical option for now.
Meanwhile, more large models are emerging from China. At the beginning of August, Alibaba introduced Qwen3.8-Max with 2.4 trillion parameters, initially via API only; the weights are expected to follow. According to Alibaba, the model worked autonomously on a coding project for almost 16 days, reviewed scientific papers, and understood and optimized the code used in them. Independent tests still need to confirm these claims and the strong benchmark results. DeepSeek, in turn, is showing with a significantly smaller Flash model that model size does not automatically mean the best performance.
So what does this change in practical terms? Companies are gaining more choice between closed services and models with open weights. Those wishing to process sensitive data locally can in principle use Kimi K3 — but they would have to pay for infrastructure that, according to the stated requirements, is too expensive for many companies. For smaller teams, the practical advantage is therefore shifting to the cloud for the time being. The technological direction is nevertheless clear: open models are becoming more capable, while their hardware costs continue to limit access.
Comentários
A carregar a conversa…
Inicie sessão para escrever um comentário. Iniciar sessão