跳到正文
原文
Google DeepMind·· 2026-08-27

Google DeepMind 发布 Gemini 3.5 Transcribe 语音转文字模型

Intelligent transcription with Gemini 3.5 Transcribe

SI 导读

Google DeepMind 发布 Gemini 3.5 Transcribe,称其为目前最精确的语音转文字模型,可直接把原始音频转成准确、格式化文本。据 Artificial Analysis 测量,流式场景平均 WER 为 4.0%,非流式为 2.6%,最终转写时间较前代 Chirp 3 改善 70%;模型支持 85 种以上语言自动检测、自定义词汇和最多三位说话人区分。

精选SI 评分71
推荐理由

官方给出 WER、延迟与多语言等量化指标,并列出两套 API 与产品落地入口,便于判断语音转写能力边界。

来源:Google DeepMind · deepmind.google

© 2026 SI·Hot · Super Intelligence Hot · 超级智能热点 · 网站数据均来源于网络公开资料,版权归来源方所有