Giovedì 27 agosto 2026

Aube.

Le notizie del progresso
In esercizioFonte unica

GLM-5.3 launches: Zhipu boosts coding and agent capabilities with 700 billion parameters

Lingue di questo articoloConfronta con l’originale

Tradotto dall’IA, lingua originale cinese — vedi il testo originale. 3 lingue disponibili, la tua si aggiunge con un clic.

A 3D webpage depicting the circulation of human blood has been built on a computer screen. Users can switch between visual layers and adjust heart rate and blood pressure, and even simulate hemorrhagic shock. But the human body looks more like a stick figure, with the organs piled together—still far from a detailed medical simulation. This is the first signal after GLM-5.3's launch: it can already turn ideas into interactive webpages quickly, but it has not yet polished every detail to the level of a professional tool.

Zhipu has not continued pushing up the parameter count this time. GLM-5.3 remains in the 700-billion-parameter class, with the upgrade coming mainly from post-training Scaling—that is, continued reinforcement learning and capability training on an existing base model. According to the official technical blog, the model uses a 743B base model and advances reinforcement learning on the same base model as GLM-5.2. The accompanying open-source Slime framework covers the complete workflow required for post-training production-grade models and supports GLM, Qwen, selected DeepSeek models and Llama 3.

The improvements first appear in coding and agent tasks. Across six benchmarks, GLM-5.3 took first place in the GDPVal evaluation, while its results in AutomationBench and Agents’ Last Exam were also near the top. The latter two focus respectively on cross-application workflow orchestration and agent capabilities. Cybersecurity is another area of strength: in ExploitGym, the model faced 898 problems with a 6-hour time limit and completed 130, matching Claude Mythos 5. Zhipu also published a cybersecurity disclosure ledger recording vulnerabilities discovered by the model, with the earliest vulnerability traceable to approximately 40 years ago.

The gaps in hands-on testing are equally clear. The pyramid skateboard game generated by GLM-5.3 offers a good experience, and the 3D scene of a jellyfish lake can also produce relatively complex visuals. But the human-circulation simulation is rough, and complex projects often require longer rounds of repeated testing. In Zcode, the model can call tools, read webpage screenshots, identify problems and continue making revisions. A webpage depicting a planet colliding with Earth ran for more than an hour before completing effects including a cracking crust, erupting lava and debris forming a planetary ring. When Claude Code is used, automatic command approval may also produce a prompt saying it cannot determine whether Bash is safe, a problem believed to be related to third-party models not yet being adapted to the application's mechanism.

Tests show that GLM-5.3 can call tools and continuously make revisions in web applications and automated workflows: it has been integrated into Zhipu's Zcode and AutoClaw and made available to GLM Coding Plan users. Third-party platforms including WorkBuddy, QwenWork and TraeWork are also offering early access. The API is expected to open next Tuesday, and the full weights will be open-sourced within two weeks. It cannot successfully generate every 3D page on the first attempt, but it has demonstrated the corresponding capabilities in coding, tool use and security-task tests. The cost is that the model is being updated so quickly that today's preferred option may soon need to be compared again.

130 questionsNumber of problems GLM-5.3 completed in the ExploitGym cybersecurity test

Fonti — leggere gli originali(ora di Parigi)

爱范儿ZH
0000

Da leggere dopo

Commenti

Caricamento della discussione…

Accedi per scrivere un commento. Accedi