Google launches Gemini 3.5 Transcribe for voice input
On a Pixel 11, a voice note can now arrive as something closer to finished prose than a transcript. Google’s Gemini 3.5 Transcribe already powers the Gboard “Rambler” feature, and the company is preparing to bring the model across its wider ecosystem.
The model does more than recognize words. It can remove “ums” and “uhs,” revise text when a speaker corrects themselves, and consult custom vocabulary for specialized jargon. Google says the path from speech to final text is about 70 percent faster than with Chirp 3, the previous voice-to-text engine.
Accuracy has improved too, though the gap is modest. Google measures Gemini 3.5 Transcribe’s live-speech error rate at 5.5 percent, compared with 7.32 percent for Chirp 3. The system supports 85 languages and up to three speakers in pre-recorded audio.
So what changes in practice? People dictating short messages or notes should spend less time fixing verbal stumbles and transcription errors before sharing the result. That is especially useful when voice input is faster than typing, but it depends on trusting the model’s editorial choices rather than receiving a word-for-word record.
That trade-off remains the limit. Ars Technica found the cleanup effective in testing with Rambler, while noting that Gemini 3.5 Transcribe can alter the wording of what was said. For casual dictation, that may be the point; for situations requiring an exact account, it may be a reason to keep the raw speech transcript.
Comentarios
Cargando el hilo…
Inicia sesión para escribir un comentario. Iniciar sesión