Jueves, 27 de agosto de 2026

Aube.

Las noticias del progreso
LaboratorioCorroborado · 2 fuentes

GLM-5.3 doubles SWE-Marathon score and tops CyberGym

Idiomas de este artículo
Original · ENFR

Texto original en inglés. 2 idiomas disponibles, el tuyo se añade con un clic.

Tang Jie, Zhipu AI’s chief scientist, answered a question about GLM-5.3 on X with one word stretched into a promise: “sooooooon.” A few hours later, GLM-5.3 was officially released. Its target is clear—not every task, but coding, cybersecurity and computer-use agents pushed as far as Zhipu can take them.

The biggest movement is in software engineering. GLM-5.3 lifted SWE-Marathon from 19.4 to 42.5, while FrontierSWE moved from 67.5 to 78.1. On Terminal Bench 3.0, which simulates multi-step work inside a Linux terminal, the score rose from 4.6 to 28.3. Zhipu says the base models have the same parameter count, making post-training the stated source of the improvement.

Cybersecurity is where the release becomes more concrete. CyberGym tests whether a model can find vulnerabilities, build attacks and check their effects—not merely produce secure-looking code. GLM-5.3 scored 84.5, ranking first among the models evaluated. Z.ai’s security ledger lists 2,436 findings across 269 open-source projects, including 1,097 Critical or High severity findings. The oldest dated to 1981, and the vulnerabilities had remained in code for an average of 26.6 years, according to the company’s disclosure work.

The model also lets users choose how much reasoning to spend, through four Effort Level settings ranging from Non-Thinking to Max. Official curves show accuracy rising from about 24.5% at Low to 34.5% at Max, whereas GLM-5.2 moved from about 20% to 23.5%. On agent benchmarks, GLM-5.3 reached 73.0 on Toolathlon Verified and 48.2 on AutomationBench, alongside first-tier results on several other tests.

Concretely, GLM-5.3 is aimed at debugging, terminal operations and vulnerability discovery, with a control for trading speed against reasoning depth. The limits are visible too: the security figures come from Z.ai’s own ledger, and no independent reproduction is cited here. GLM-5.3 may also be a bridge release; Zhipu says the next major model will use a new architecture with doubled parameters and target Fable 5 across the board.

84.5GLM-5.3’s CyberGym score, highest among evaluated models

Fuentes — leer los originales(hora de París)

TechNodeEN
PandailyEN
0000

Para leer a continuación

Comentarios

Cargando el hilo…

Inicia sesión para escribir un comentario. Iniciar sesión