Friday, 4 September 2026Support us

Aube.

News of progress
DeployedCorroborated · 2 sources

Gemini 3.8 Flash launches with stronger reasoning, higher costs

Languages for this article

Machine-translated from Chinese — read the original text. 6 languages available; yours is one click away.

On the morning of September 2, Gemini’s version number jumped again: Google launched Gemini 3.8 Flash. It is the third Flash model released in six weeks and the fourth model launched in less than four months since the Gemini 3.5 Flash announcement in May. Meanwhile, the originally planned flagship Gemini 3.5 Pro has still not been released. Flash was originally focused on speed and cost, but it is now being pushed toward long-running software engineering, AI agents that can make continuous tool calls and complex enterprise workflows lasting hours.\n\nThe benchmark results first pushed the story toward flagship-model territory. Google says 3.8 Flash reached 73.7% on the DeepSWE v1.1 software-development test, while its Terminal-Bench 2.1 score rose from 85.8% to 89.4%, slightly above Claude Opus 5’s 89.1%. Artificial Analysis gave it an Intelligence Index score of 59, comparable to GPT-5.6 Sol at higher reasoning settings but below Claude Opus 5’s top score of 63. The figures also require some caution: ifanr reported a DeepSWE v1.1 score of 71.0%, not 73.7%. The two outlets therefore gave different accounts of the same result.\n\nGreater capability comes at a cost. Google explains that, when faced with complex problems, 3.8 Flash uses more reasoning steps, repeatedly calls tools and checks its work. That means consuming more Tokens—the billing unit the model uses to process text. Artificial Analysis observed that its average output increased by about 30%, while average task time rose from 2.2 minutes to 2.5 minutes. Even with an output speed of about 300 Tokens per second, the average cost per task still rose to $0.58, about 40% higher than with 3.7 Flash. API pricing is currently the same as for the previous model, at $0.75 per million input Tokens and $3.75 per million output Tokens, but will rise to $1.50 and $7.50, respectively, on January 1, 2027.\n\nReal-world use presents another side of the picture. ifanr asked the model to create a 3D model of an Airbus H145 helicopter in Three.js. It broke down the requirements in less than a minute, and the scene and model were complete after 109 seconds. The speed certainly lived up to the Flash name, but the body details, component relationships and movement were all crude. A second terrain-and-water-flow simulation took about two minutes and was broadly complete, although water leaked from the volcano crater. Google’s official Magic Castle demonstration, meanwhile, had 3.8 Flash work in a loop in Antigravity and call Nano Banana to generate textures, supported by an agent framework, a tool environment and a multi-round process. A one-shot prompt reveals the lower bound of what the model can do on its own, not the upper bound achieved when an entire system supports it.\n\nGemini 3.8 Flash Cyber, released the same day, uses the same base model but is specially optimized for vulnerability discovery and remediation. Google says it achieved a success rate above 70% in internal tests covering 20 programming languages and solved 86.2% of tasks in a single attempt on CyberGym. It is not available to ordinary users, however. Through the Fairwind Program, it is being provided to verified cybersecurity organizations, government departments, critical-infrastructure operators and major technology companies. Google says it currently has more than 650 partners.\n\nFor users, developers can now access 3.8 Flash through the Gemini API—the interface that lets software call the model—Google AI Studio and Gemini Enterprise. Some subscribers can also use it in the Gemini app and in Google Search’s AI Mode. It is suited to work that prioritizes speed, price and the ability to handle longer tasks: the model has a one-million-Token context window, can process text, images, audio and video, and can output up to 64,000 Tokens. For tasks that place greater emphasis on computational efficiency, however, Google recommends lowering the reasoning setting or continuing to use 3.7 Flash. In other words, 3.8 Flash makes complex-model capabilities more affordable for developers, but it has not eliminated the need for human review from the workflow.

$0.58Artificial Analysis’s measured average cost per 3.8 Flash task

Sources — read the originals(Paris time)

heise onlineDE
爱范儿ZH
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in