Today, we’re launching Gemini 3.5 Transcribe, our new speech-to-text model with sub-second streaming...
TL;DR · AI 摘要
Google推出Gemini 3.5 Transcribe语音识别模型,实现亚秒级实时流处理,非流式WER低至2.6%,支持85种语言并减少70%处理时间。
核心要点
- Gemini 3.5 Transcribe非流式WER达2.6%,流式为4.0%
- 支持85+语言及字母数字标记处理
- 相比Chirp 3模型处理时间缩短70%
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- Gemini 3.5 Transcribe
- 核心特性
- 亚秒级实时流处理
- 智能后处理
- 性能指标
- 2.6%非流式WER
- 4.0%流式WER
- 应用优化
- 85+语言支持
- 70%时间缩减
金句 / Highlights
值得收藏与分享的关键句。
非流式WER 2.6%,流式WER 4.0%,达到行业领先水平
支持清理口语中的"um"、"ah"等不流畅表达
相比Chirp 3模型处理时间缩短70%
Philipp Schmid on X: "Today, we’re launching Gemini 3.5 Transcribe, our new speech-to-text model with sub-second streaming and intelligent post-processing for agent interfaces. 2.6% WER on non-streaming and 4.0% on streaming. Supports 85+ languages, cleans up conversational disfluencies ("um", "ah… / X
Philipp Schmid
@_philschmid
Today, we’re launching Gemini 3.5 Transcribe, our new speech-to-text model with sub-second streaming and intelligent post-processing for agent interfaces. 2.6% WER on non-streaming and 4.0% on streaming. Supports 85+ languages, cleans up conversational disfluencies ("um", "ah", and mid-sentence self-corrections), handles alphanumeric tokens (postal codes or IDs), 70% reduction time for the final transcription compared to Chirp 3. - gemini-3.5-transcribe-live: Sub-second, bidirectional streaming via the Gemini Live API - gemini-3.5-transcribe: Recorded processing via the Interactions API with word-level timestamps and speaker attribution (up to 3 speakers). Having access to Transcribe has made me use voice input way more than I ever did before. The native post-processing is incredible. It knows you mean .json and not a person named "Jason". I sometimes speak 5 minutes into and it perfectly refactors my instruction to my context. Available today in public preview in Gemini API,
@
GoogleAIStudio
and already in the
GeminiApp
and
antigravity
. Try here:
ai.studio/apps/bundled/g…
$
00:00
/$
5:05 PM · Aug 26, 2026
14.5K
Views
22
8
127
27