GLM-5.3-Flash runs on more than 100,000 domestic chips
For five days, developers tested an anonymous model called Ox Alpha—“Niu Lai” in Chinese developer circles—on OpenRouter and OpenCode. It processed more than 50 trillion tokens before Zhipu AI confirmed the identity: GLM-5.3-Flash. Usage has since been more than double that of DeepSeek.
The hardware behind that traffic is the real change. Zhipu says its inference service—the system that runs a model and generates responses—used more than 100,000 domestic chips, with hardware efficiency and per-token cost comparable to mainstream Nvidia GPUs. LatePost reported that Huawei, Moore Threads and Hygon may supply the chips, but Zhipu has not confirmed the suppliers or models.
Zhipu also rebuilt the software around the hardware. To scale toward a context of 1 million tokens, its team used SGLang, W8A8 quantization, mixed INT8/FP8/BF16 cache quantization and node-level tensor parallelism. Its production architecture separates multimodal encoding, prompt processing and token-by-token decoding into independently schedulable pools. Zhipu says the result improved end-to-end service performance 3x over its initial baseline on the same hardware.
For developers, the immediate consequence is access to a low-cost model that has already handled unusually heavy public testing. GLM-5.3-Flash has 320 billion total parameters, with 18 billion activated, and scores 57 on the Artificial Analysis Intelligence Index. Zhipu prices it at about one-tenth of GLM-5.3 and one-fortieth of Claude Opus 4.8.
The claim still has boundaries. The chip identities have not been disclosed, and the efficiency and cost comparison comes from Zhipu’s technical documentation; no broader hardware equivalence is established here. But a large domestic cluster carrying real model traffic, rather than a lab demonstration, gives chip designers and model developers a concrete target: more usable AI output for less per-token cost.
Commenti
Caricamento della discussione…
Accedi per scrivere un commento. Accedi