Thursday, 27 August 2026

Aube.

News of progress
Single source

Kakao’s Kanana2 leads Google and Alibaba’s lightweight AI models in safety evaluation

Languages for this articleCompare with the original

Machine-translated from Korean — read the original text. 3 languages available; yours is one click away.

Two lightweight language models from Kakao released on Hugging Face outperformed similarly sized models from Google and Alibaba in a safety evaluation. Both Kanana-2-1.3B-Instruct and Kanana-2-3B-Instruct ranked first with an overall score of 0.70. Kakao released the results on the 18th.

The models used for comparison were Google’s Gemma and Alibaba’s Qwen. In the 1.3B model comparison, Gemma scored 0.68 and Qwen 0.58; in the 3B model comparison, Gemma scored 0.66 and Qwen 0.62. Kanana2 1.3B scored higher than the comparison models in the crime and illegal activity category, while the 3B model scored higher in the sexual content and child protection, and rights violations categories.

The test was AssurAI. Jointly developed last November by the Telecommunications Technology Association, KAIST and Kakao, this Korean-language benchmark reflects Korea’s social and cultural context and contains 35 risk categories and 9,560 evaluation items. It examines risk factors across multimodal conditions such as text, images and audio, as well as malicious prompts and real-world AI service usage scenarios.

Kakao scored the models’ responses against predefined criteria on its own AI safety evaluation platform. It used an LLM-as-a-Judge approach, in which a large language model evaluates another model’s responses. The results are therefore Kakao’s own evaluation findings.

Kakao plans to expand this pre-release validation across its internally developed models and apply it beyond text models to risk assessments for multimodal models and agentic AI.

0.70Overall safety evaluation score for the Kanana2 1.3B and 3B models

Sources — read the originals(Paris time)

전자신문 (ETNews)KO
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in