Thursday, 27 August 2026

Aube.

News of progress
DeployedSingle source

iFLYTEK Uses Engineering to Narrow Domestic AI Compute Gap

Languages for this article
Original · ENFR

Originally written in English. 2 languages available; yours is one click away.

On August 21, iFLYTEK CEO Liu Qingfeng partly explained the company's absence from the latest code- and agent-focused AI wave by pointing to the cost of extremely long context windows on domestic computing infrastructure. His response is not to wait for faster hardware, but to extract more from the chips already available. iFLYTEK says its domestic accelerators can lag Nvidia's H200 by up to roughly fivefold in training efficiency when context extends beyond 256K tokens.

That gap matters less for the workloads iFLYTEK is targeting than it does for frontier experiments. The company says 90% to 95% of use cases across education, healthcare, automotive and state-owned enterprises do not need million-token context windows. It is therefore tuning the whole stack at once: model architecture, operators, communications, memory and training frameworks, while validating techniques for long-sequence inference.

The strategy has produced a concrete deployment. Spark X2-Flash, launched in April, became the first 30-billion-parameter Mixture-of-Experts model that iFLYTEK says was trained and deployed entirely on Huawei Ascend 910B clusters. Spark X2 arrived in February, Spark X2-VL followed in June, and the company said a new all-domestic flagship general model was due in late August, including at its annual 1024 Developers Festival.

The engineering push also reaches the infrastructure underneath the models. In July, iFLYTEK shared a national first prize for the Pengcheng Cloud Brain, a large-scale domestic intelligent compute project led by Pengcheng National Laboratory. On the commercial side, its LLM API and MaaS — model as a service — revenue grew about 70% year on year in the first half. The company is combining trillion-scale cloud models with on-device models and selling outcomes rather than tokens.

For those sectors, the company argues that most deployments can avoid million-token context windows. Liu says the approach protects iFLYTEK's margin and focus as domestic chip ecosystems mature, although the performance figures come from the company and no independent reproduction is cited here. Liu expects the compute disadvantage against domestic rivals to level out within two years; that is a forecast, not a completed result.

90% to 95%Share of target use cases that iFLYTEK says do not need million-token context

Sources — read the originals(Paris time)

PandailyEN
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in