GLM-5.3 launches: Zhipu boosts coding and agent capabilities with 700 billion parameters
A 3D webpage depicting the circulation of human blood has been built on a computer screen. Users can switch between visual layers and adjust heart rate and blood pressure, and even simulate hemorrhagic shock. But the human body looks more like a stick figure, with the organs piled together—still far from a detailed medical simulation. This is the first signal after GLM-5.3's launch: it can already turn ideas into interactive webpages quickly, but it has not yet polished every detail to the level of a professional tool.
Zhipu has not continued pushing up the parameter count this time. GLM-5.3 remains in the 700-billion-parameter class, with the upgrade coming mainly from post-training Scaling—that is, continued reinforcement learning and capability training on an existing base model. According to the official technical blog, the model uses a 743B base model and advances reinforcement learning on the same base model as GLM-5.2. The accompanying open-source Slime framework covers the complete workflow required for post-training production-grade models and supports GLM, Qwen, selected DeepSeek models and Llama 3.
The improvements first appear in coding and agent tasks. Across six benchmarks, GLM-5.3 took first place in the GDPVal evaluation, while its results in AutomationBench and Agents’ Last Exam were also near the top. The latter two focus respectively on cross-application workflow orchestration and agent capabilities. Cybersecurity is another area of strength: in ExploitGym, the model faced 898 problems with a 6-hour time limit and completed 130, matching Claude Mythos 5. Zhipu also published a cybersecurity disclosure ledger recording vulnerabilities discovered by the model, with the earliest vulnerability traceable to approximately 40 years ago.
The gaps in hands-on testing are equally clear. The pyramid skateboard game generated by GLM-5.3 offers a good experience, and the 3D scene of a jellyfish lake can also produce relatively complex visuals. But the human-circulation simulation is rough, and complex projects often require longer rounds of repeated testing. In Zcode, the model can call tools, read webpage screenshots, identify problems and continue making revisions. A webpage depicting a planet colliding with Earth ran for more than an hour before completing effects including a cracking crust, erupting lava and debris forming a planetary ring. When Claude Code is used, automatic command approval may also produce a prompt saying it cannot determine whether Bash is safe, a problem believed to be related to third-party models not yet being adapted to the application's mechanism.
Tests show that GLM-5.3 can call tools and continuously make revisions in web applications and automated workflows: it has been integrated into Zhipu's Zcode and AutoClaw and made available to GLM Coding Plan users. Third-party platforms including WorkBuddy, QwenWork and TraeWork are also offering early access. The API is expected to open next Tuesday, and the full weights will be open-sourced within two weeks. It cannot successfully generate every 3D page on the first attempt, but it has demonstrated the corresponding capabilities in coding, tool use and security-task tests. The cost is that the model is being updated so quickly that today's preferred option may soon need to be compared again.
Comentarios
Cargando el hilo…
Inicia sesión para escribir un comentario. Iniciar sesión