🚀 AuK is officially here. Nano banana🍌 for audio An open-source foundation model for unified spee...
TL;DR · AI 摘要
腾讯发布开源语音生成模型AuK,支持多任务处理并提升推理速度4.5倍。
核心要点
- AuK支持零样本TTS和指令控制生成,无需额外训练数据
- AuK-Flash推理速度较基线模型提升4.5倍
- 模型开源包含代码/权重/演示,支持多说话人和音乐分离
结构提纲
按章节快速跳转。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- 腾讯AuK模型
- 核心功能
- 零样本TTS
- 多模态编辑
- 语音转换
- 加速方案
- 4步推理流程
- 4.5倍加速
- 技术验证
- GitHub开源
- 论文发布
金句 / Highlights
值得收藏与分享的关键句。
Zero-shot TTS能力使模型可直接生成未训练语音的发音
AuK-Flash在保持精度前提下将推理速度提升至原模型的4.5倍
支持音色/风格/情绪多维编辑及口音消除等高级功能
Tencent Hy on X: "🚀 AuK is officially here. Nano banana🍌 for audio An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-… / X
Tencent Hy
@TencentHunyuan
🚀 AuK is officially here. Nano banana🍌 for audio An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation. Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions. Code, weights, and demo are live. Try it and share your feedback. 🤗 Paper & upvote:
huggingface.co/papers/2609.08…
⭐ GitHub & star:
github.com/Tencent-Hunyua…
$
00:00
/$
10:33 AM · Sep 10, 2026
42.4K
Views
42
104
933
669