Giovedì 27 agosto 2026

Aube.

Le notizie del progresso
In esercizioFonte unica

Google launches Gemini 3.5 Transcribe for voice input

Lingue di questo articolo
Originale · ENFR

Testo originale in inglese. 2 lingue disponibili, la tua si aggiunge con un clic.

On a Pixel 11, a voice note can now arrive as something closer to finished prose than a transcript. Google’s Gemini 3.5 Transcribe already powers the Gboard “Rambler” feature, and the company is preparing to bring the model across its wider ecosystem.

The model does more than recognize words. It can remove “ums” and “uhs,” revise text when a speaker corrects themselves, and consult custom vocabulary for specialized jargon. Google says the path from speech to final text is about 70 percent faster than with Chirp 3, the previous voice-to-text engine.

Accuracy has improved too, though the gap is modest. Google measures Gemini 3.5 Transcribe’s live-speech error rate at 5.5 percent, compared with 7.32 percent for Chirp 3. The system supports 85 languages and up to three speakers in pre-recorded audio.

So what changes in practice? People dictating short messages or notes should spend less time fixing verbal stumbles and transcription errors before sharing the result. That is especially useful when voice input is faster than typing, but it depends on trusting the model’s editorial choices rather than receiving a word-for-word record.

That trade-off remains the limit. Ars Technica found the cleanup effective in testing with Rambler, while noting that Gemini 3.5 Transcribe can alter the wording of what was said. For casual dictation, that may be the point; for situations requiring an exact account, it may be a reason to keep the raw speech transcript.

70 percent fasterGoogle’s claimed improvement from voice to final transcribed text

Fonti — leggere gli originali(ora di Parigi)

Ars TechnicaEN
0000

Da leggere dopo

Commenti

Caricamento della discussione…

Accedi per scrivere un commento. Accedi