Friday, 4 September 2026Support us

Aube.

News of progress
Original and translation

Gemini 3.8 Flash launches with stronger reasoning, higher costs

Leave comparison

Both versions are aligned block by block, in reading order: title, the essentials, then paragraph by paragraph. Where the translation merged or split a paragraph, the matching cell stays empty — we never pair two passages by guesswork.

Original · Chinese
Gemini 3.8 Flash上线,推理更强也更耗算力
Translation · English
Gemini 3.8 Flash launches with stronger reasoning, higher costs
Original · Chinese
它面向长程软件工程、能连续调用工具的智能代理,以及复杂企业工作流。
Translation · English
It targets long-running software engineering, AI agents that can make continuous tool calls, and complex enterprise workflows.
Original · Chinese
Google给出的成绩逼近旗舰模型,但两家媒体对DeepSWE成绩报道为71.0%和73.7%。
Translation · English
Google’s results approach those of flagship models, but two media outlets reported DeepSWE scores of 71.0% and 73.7%.
Original · Chinese
3.8 Flash Cyber专攻漏洞发现与修复,仅向Fairwind Program中的可信机构开放。
Translation · English
3.8 Flash Cyber focuses on vulnerability discovery and remediation and is available only to trusted organizations in the Fairwind Program.
Original · Chinese

9月2日凌晨,Gemini的版本号又跳了一格:Google上线Gemini 3.8 Flash。它是六周内发布的第三款Flash模型,也是5月宣布的Gemini 3.5 Flash之后、不到四个月内的第四款;与此同时,原定的旗舰Gemini 3.5 Pro至今仍未发布。Flash原本强调速度和成本,如今却被推向长程软件工程、能连续调用工具的智能代理,以及持续数小时的复杂企业工作流。\n\n跑分先把故事推到了旗舰模型附近。Google称,3.8 Flash在DeepSWE v1.1软件开发测试中达到73.7%,Terminal-Bench 2.1从85.8%升至89.4%,略高于Claude Opus 5的89.1%;Artificial Analysis测得它的智能指数为59分,和GPT-5.6 Sol在较高推理档位的成绩相当,但低于Claude Opus 5最高的63分。数字也有需要保留的地方:爱范儿对DeepSWE v1.1的报道记录为71.0%,而非73.7%,两家媒体呈现的同一项成绩并不一致。\n\n更高的能力并非没有代价。Google解释,3.8 Flash面对复杂问题会投入更多推理步骤、反复调用工具并检查工作,这意味着消耗更多Token——模型处理文本时使用的计费单位。Artificial Analysis观察到,它平均生成的输出Token增加约30%,任务平均耗时从2.2分钟变为2.5分钟;即使输出速度约为每秒300个Token,单项任务成本仍升至0.58美元,比3.7 Flash高约40%。目前API价格与前代相同,为每百万输入Token 0.75美元、输出3.75美元,但2027年1月1日起将分别恢复至1.5美元和7.5美元。\n\n真实使用呈现出另一面。爱范儿让模型用Three.js制作空客H145直升机的3D模型,需求拆解不到一分钟完成,109秒后场景和模型生成完毕;速度确实配得上Flash,但机身细节、部件关系和运动方式都很粗糙。第二个地形水流模拟约两分钟完成,功能基本齐全,火山口却出现了漏水。Google官方的魔法城堡演示则由3.8 Flash在Antigravity中循环工作,并调用Nano Banana生成纹理,背后有代理框架、工具环境和多轮流程;一次性提示词得到的,是它单独工作的下限,而不是整套系统托举出的上限。\n\n同日发布的Gemini 3.8 Flash Cyber沿用同一基础模型,但专门优化漏洞发现与修复。Google称它在覆盖20种编程语言的内部测试中成功率超过70%,在CyberGym中一次尝试解决86.2%的任务;不过它没有向普通用户开放,而是通过Fairwind Program提供给经过验证的网络安全机构、政府部门、关键基础设施运营者和主要科技企业,Google称目前已有超过650个合作伙伴。\n\n具体到使用者,开发者现在可以通过Gemini API——让软件调用模型的接口——、Google AI Studio和Gemini Enterprise使用3.8 Flash,部分订阅用户也能在Gemini应用和Google搜索的人工智能模式中调用它。它适合把速度、价格和较长任务能力放在前面的工作:模型拥有一百万Token上下文窗口,可处理文本、图片、音频和视频,最多输出64,000个Token;但若任务更看重计算效率,Google建议降低推理档位或继续使用3.7 Flash。换句话说,3.8 Flash让复杂模型能力更容易被开发者买到,却还没有把人工复核从工作流里拿掉。

Translation · English

On the morning of September 2, Gemini’s version number jumped again: Google launched Gemini 3.8 Flash. It is the third Flash model released in six weeks and the fourth model launched in less than four months since the Gemini 3.5 Flash announcement in May. Meanwhile, the originally planned flagship Gemini 3.5 Pro has still not been released. Flash was originally focused on speed and cost, but it is now being pushed toward long-running software engineering, AI agents that can make continuous tool calls and complex enterprise workflows lasting hours.\n\nThe benchmark results first pushed the story toward flagship-model territory. Google says 3.8 Flash reached 73.7% on the DeepSWE v1.1 software-development test, while its Terminal-Bench 2.1 score rose from 85.8% to 89.4%, slightly above Claude Opus 5’s 89.1%. Artificial Analysis gave it an Intelligence Index score of 59, comparable to GPT-5.6 Sol at higher reasoning settings but below Claude Opus 5’s top score of 63. The figures also require some caution: ifanr reported a DeepSWE v1.1 score of 71.0%, not 73.7%. The two outlets therefore gave different accounts of the same result.\n\nGreater capability comes at a cost. Google explains that, when faced with complex problems, 3.8 Flash uses more reasoning steps, repeatedly calls tools and checks its work. That means consuming more Tokens—the billing unit the model uses to process text. Artificial Analysis observed that its average output increased by about 30%, while average task time rose from 2.2 minutes to 2.5 minutes. Even with an output speed of about 300 Tokens per second, the average cost per task still rose to $0.58, about 40% higher than with 3.7 Flash. API pricing is currently the same as for the previous model, at $0.75 per million input Tokens and $3.75 per million output Tokens, but will rise to $1.50 and $7.50, respectively, on January 1, 2027.\n\nReal-world use presents another side of the picture. ifanr asked the model to create a 3D model of an Airbus H145 helicopter in Three.js. It broke down the requirements in less than a minute, and the scene and model were complete after 109 seconds. The speed certainly lived up to the Flash name, but the body details, component relationships and movement were all crude. A second terrain-and-water-flow simulation took about two minutes and was broadly complete, although water leaked from the volcano crater. Google’s official Magic Castle demonstration, meanwhile, had 3.8 Flash work in a loop in Antigravity and call Nano Banana to generate textures, supported by an agent framework, a tool environment and a multi-round process. A one-shot prompt reveals the lower bound of what the model can do on its own, not the upper bound achieved when an entire system supports it.\n\nGemini 3.8 Flash Cyber, released the same day, uses the same base model but is specially optimized for vulnerability discovery and remediation. Google says it achieved a success rate above 70% in internal tests covering 20 programming languages and solved 86.2% of tasks in a single attempt on CyberGym. It is not available to ordinary users, however. Through the Fairwind Program, it is being provided to verified cybersecurity organizations, government departments, critical-infrastructure operators and major technology companies. Google says it currently has more than 650 partners.\n\nFor users, developers can now access 3.8 Flash through the Gemini API—the interface that lets software call the model—Google AI Studio and Gemini Enterprise. Some subscribers can also use it in the Gemini app and in Google Search’s AI Mode. It is suited to work that prioritizes speed, price and the ability to handle longer tasks: the model has a one-million-Token context window, can process text, images, audio and video, and can output up to 64,000 Tokens. For tasks that place greater emphasis on computational efficiency, however, Google recommends lowering the reasoning setting or continuing to use 3.7 Flash. In other words, 3.8 Flash makes complex-model capabilities more affordable for developers, but it has not eliminated the need for human review from the workflow.

Back to the article