Among Four Sovereign AI Foundation Models, the Battle Is About Inference Efficiency, Not Size
As a project presentation event for the sovereign AI foundation model project takes place at COEX in Samseong-dong, Seoul, the announcement of the second evaluation results that will determine the four teams’ future advancement is approaching. The focus of this competition involving LG AI Research, Upstage, SK Telecom, and Motif Technologies has shifted from model size to real-world inference performance and efficiency. The results of the second evaluation are expected to be announced soon.
The four teams pursued the same goal in different ways. LG AI Research chose upcycling, using the training results and weights of an existing model to build a larger one, and said it had expanded the size of K-EXAONE 2.0 by more than 3 times compared with the previous model. The approach is intended to reduce the time and cost required to build a new model from scratch.
SK Telecom’s A.X K2 has 688B total parameters, but the number activated during inference is about 33B. A mixture-of-experts architecture maintains the overall knowledge capacity of the model while activating only the parameters needed to process a question. Upstage’s Solar Open2 inherited only about 2.3% of its predecessor’s weights to build a 250B model, targeting agent applications such as office productivity, document work, and coding with long contexts.
Motif Technologies’ Motif 3 is a 314 billion-parameter model, but only 13.2 billion parameters are activated during actual inference. Its proprietary architecture is designed to keep the overall model large while reducing the computational load. The company said it focused on Korean, inference-intensive tasks, and legal and financial data. In the previously released Artificial Analysis Intelligence Index, Motif 3 recorded the highest score among the four models, with 47 points.
What changes in concrete terms is the range of choices available between answer quality and usage costs. SK Telecom said it supports a thinking mode for solving complex problems and a non-thinking mode for short responses with low latency. However, index scores and technical materials show each model’s potential, and whether the same results can be reproduced in actual services requires further confirmation. In the second evaluation, benchmarks account for 40 points out of 100, expert assessments for 35 points, and user assessments for 25 points. Ultimately, three of the four teams will advance to the next stage.
Comments
Loading the thread…
Sign in to leave a comment. Sign in