Thursday, 27 August 2026

Aube.

News of progress
DeployedSingle source

Google launches Gemini 3.5 Transcribe for voice input

Languages for this article
Original · ENFR

Originally written in English. 2 languages available; yours is one click away.

On a Pixel 11, a voice note can now arrive as something closer to finished prose than a transcript. Google’s Gemini 3.5 Transcribe already powers the Gboard “Rambler” feature, and the company is preparing to bring the model across its wider ecosystem.

The model does more than recognize words. It can remove “ums” and “uhs,” revise text when a speaker corrects themselves, and consult custom vocabulary for specialized jargon. Google says the path from speech to final text is about 70 percent faster than with Chirp 3, the previous voice-to-text engine.

Accuracy has improved too, though the gap is modest. Google measures Gemini 3.5 Transcribe’s live-speech error rate at 5.5 percent, compared with 7.32 percent for Chirp 3. The system supports 85 languages and up to three speakers in pre-recorded audio.

So what changes in practice? People dictating short messages or notes should spend less time fixing verbal stumbles and transcription errors before sharing the result. That is especially useful when voice input is faster than typing, but it depends on trusting the model’s editorial choices rather than receiving a word-for-word record.

That trade-off remains the limit. Ars Technica found the cleanup effective in testing with Rambler, while noting that Gemini 3.5 Transcribe can alter the wording of what was said. For casual dictation, that may be the point; for situations requiring an exact account, it may be a reason to keep the raw speech transcript.

70 percent fasterGoogle’s claimed improvement from voice to final transcribed text

Sources — read the originals(Paris time)

Ars TechnicaEN
0000

Read next

Comments

Loading the thread…

Sign in to leave a comment. Sign in