Google DeepMind Blog

Gemini 3.1 Flash TTS: the next generation of expressive AI speech

3.5内容质量
Gemini 3.1 Flash TTS: the next generation of expressive AI speech

TL;DR · AI 摘要

本文仅为Google DeepMind博客的页面导航与标题占位符,未提供Gemini 3.1 Flash TTS模型的技术架构、性能基准或工程实践细节,属于典型的产品发布公告,缺乏可供工程师参考的实质内容。

核心要点

  • 文章仅提供模型名称与发布声明,缺失核心算法与架构说明。
  • 未包含延迟、音质评测或API调用示例等工程关键指标。
  • 内容实质为官网导航模板,信息密度极低,无技术参考价值。
#TTS#AI语音合成#Google DeepMind#Gemini#产品发布
打开原文

Gemini 3.1 Flash TTS: New text-to-speech AI model

Skip to main content

The Keyword

Gemini 3.1 Flash TTS: the next generation of expressive AI speech

Share

x.comFacebookLinkedIn[Mail](mailto:?subject=Gemini%203.1%20Flash%20TTS%3A%20the%20next%20generation%20of%20expressive%20AI%20speech&body=Check%20out%20this%20article%20on%20the%20Keyword:%0A%0AGemini%203.1%20Flash%20TTS%3A%20the%20next%20generation%20of%20expressive%20AI%20speech%0A%0AGemini%203.1%20Flash%20TTS%20is%20now%20available%20across%20Google%20products.%0A%0Ahttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-tts/)

Copy link

Innovation & AI

Learn more:

See all AI updates

[See all](http://deepmind.google/innovation-and-ai/models-and-research/ "See all Models & Research articles")

[See all](http://deepmind.google/innovation-and-ai/products/ "See all Products articles")

[See all](http://deepmind.google/innovation-and-ai/infrastructure-and-cloud/ "See all Infrastructure & cloud articles")

Learn more:

Google DeepMind blogGoogle Research blogGoogle Developers blogGoogle Cloud blog

See all AI updates

  • Products & platforms

Products & platforms

Learn more:

See all product updates

[See all](http://deepmind.google/products-and-platforms/products/ "See all Products articles")

[See all](http://deepmind.google/products-and-platforms/platforms/ "See all Platforms articles")

[See all](http://deepmind.google/products-and-platforms/devices/ "See all Devices articles")

Learn more:

Google Ads & Commerce blogWaze blog

See all product updates

  • Company news

Company news

[See all](http://deepmind.google/company-news/outreach-and-initiatives/ "See all Outreach & initiatives articles")

[See all](http://deepmind.google/authors/ "See all Leadership articles")

[See all](http://deepmind.google/company-news/inside-google/ "See all Inside Google articles")

Subscribe

["How is Gemini changing Maps?", "What is \"vibe design?\"", "How can I learn new AI skills?"]

Search freely using keywords, or by asking a question

Suggested searches

Subscribe

The Keyword

Innovation & AI

Learn more:

See all AI updates

  • Products & platforms

Products & platforms

Learn more:

See all product updates

  • Company news

Company news

  • [Images](http://deepmind.google/image-library/ "Images")
  • [RSS feed](http://deepmind.google/rss/ "RSS feed")

Subscribe

Breadcrumb

  1. [](https://blog.google/ "The Keyword")
  2. Innovation & AI
  3. Models & research
  4. Gemini Models

Gemini 3.1 Flash TTS: the next generation of expressive AI speech

Apr 15, 2026

· 10 min read

Share

x.comFacebookLinkedIn[Mail](mailto:?subject=Gemini%203.1%20Flash%20TTS%3A%20the%20next%20generation%20of%20expressive%20AI%20speech&body=Check%20out%20this%20article%20on%20the%20Keyword:%0A%0AGemini%203.1%20Flash%20TTS%3A%20the%20next%20generation%20of%20expressive%20AI%20speech%0A%0AGemini%203.1%20Flash%20TTS%20is%20now%20available%20across%20Google%20products.%0A%0Ahttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-tts/)

Copy link

Our newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

V

Vilobh Meshram

Senior Product Manager

M

Max Gubin

Principal Research Engineer on behalf of the Gemini team

Read AI-generated summary

General summary

Gemini 3.1 Flash TTS is here, giving you improved AI speech quality and control. You can now use audio tags to adjust vocal style and pacing in over 70 languages. Test it out in Google AI Studio, Vertex AI, and Google Vids, and know that all audio is watermarked with SynthID to prevent misinformation.

Summaries were generated by Google AI. Generative AI is experimental.

Bullet points

  • "Gemini 3.1 Flash TTS" is a new AI speech model with better control, expressiveness, and quality.
  • This model has improved speech quality, making it sound more natural than previous versions.
  • Audio tags let you control vocal style, pace, and delivery using natural language commands.
  • Developers can use Google AI Studio to fine-tune voices and export settings for consistent use.
  • Gemini 3.1 Flash TTS supports 70+ languages and uses SynthID watermarking to identify AI-generated audio.

Summaries were generated by Google AI. Generative AI is experimental.

Basic explainer

Gemini 3.1 Flash TTS is a new AI that makes computer speech sound more real. It lets people change how the AI talks by using special commands in the text. This AI can speak in over 70 languages and adds a hidden watermark to the audio. This helps people know it's AI-generated and not a real person.

Summaries were generated by Google AI. Generative AI is experimental.

#### Explore other styles:

  • General summary
  • Bullet points
  • Basic explainer

Share

x.comFacebookLinkedIn[Mail](mailto:?subject=Gemini%203.1%20Flash%20TTS%3A%20the%20next%20generation%20of%20expressive%20AI%20speech&body=Check%20out%20this%20article%20on%20the%20Keyword:%0A%0AGemini%203.1%20Flash%20TTS%3A%20the%20next%20generation%20of%20expressive%20AI%20speech%0A%0AGemini%203.1%20Flash%20TTS%20is%20now%20available%20across%20Google%20products.%0A%0Ahttps://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-1-flash-tts/)

Copy link

Image 2: Gemini logo next to the text "3.1 Flash TTS", all over colored dots
Image 2: Gemini logo next to the text "3.1 Flash TTS", all over colored dots

Your browser does not support the audio element.

Listen to article

This content is generated by Google AI. Generative AI is experimental

[[duration]] minutes

Voice Umbriel Speed 1X

Voice Umbriel Gacrux

Speed 0.75X 1X 1.5X 2X

Today, we’re introducing Gemini 3.1 Flash TTS, the latest text-to-speech model that delivers improved controllability, expressivity and quality — empowering developers, enterprises and everyday users to build the next generation of AI-speech applications.

Starting today, 3.1 Flash TTS is rolling out:

Improved speech quality and controllability

We’ve improved the overall speech quality of Gemini 3.1 Flash TTS, making it our most natural and expressive model to date. On the Artificial Analysis TTS leaderboard, a benchmark that captures thousands of blind human preferences, 3.1 Flash TTS achieved an impressive Elo score of 1,211.

Image 3: a gif showing artificial analysis text to speech arena quality elo
Image 3: a gif showing artificial analysis text to speech arena quality elo

Artificial Analysis has also positioned Gemini 3.1 Flash TTS within its “most attractive quadrant” for its ideal blend of high-quality speech generation and low cost. The model stands out further with native multi-speaker dialogue, support for 70+ languages, and granular creative control via natural language.

New audio tags for more expressive speech generation

3.1 Flash TTS also introduces audio tags — an intuitive way to control vocal style, pace and delivery. By embedding natural language commands directly into the text input, you can steer AI-speech output with improved levels of granularity.

Sorry, your browser doesn't support embedded videos, but don't worry, you can download it and watch it with your favorite video player!

3.1 Flash TTS enables enterprises to utilize audio tags within Vertex AI, empowering the next generation of enterprise applications.

Sorry, your browser doesn't support embedded videos, but don't worry, you can download it and watch it with your favorite video player!

3.1 Flash TTS lets you use audio tags to experiment with speech expressivity, pacing and delivery.

Sorry, your browser doesn't support embedded videos, but don't worry, you can download it and watch it with your favorite video player!

3.1 Flash TTS lets you integrate expressive speech capabilities, turning a standard weather app into an engaging experience.

Sorry, your browser doesn't support embedded videos, but don't worry, you can download it and watch it with your favorite video player!

3.1 Flash TTS supports a range of audio tags to add nuanced, engaging delivery to a synonym word hunt app.

Jump to position 1 Jump to position 2 Jump to position 3 Jump to position 4

You can start experimenting with these audio tags along with other updates to the developer experience in Google AI Studio with configurable controls that place the developer in the “director’s chair”:

  • Scene direction: Set the stage by defining the environment and providing specific dialogue instructions. This world-building context helps characters remain “in-character” and react to one another naturally across multiple turns.
  • Speaker-level specificity: Cast characters using unique Audio Profiles, then specify Director’s Notes to toggle pace, tone and accent. Using inline tags, speakers can pivot from these high-level settings to change expression mid-sentence.
  • Seamless export: Once the performance is perfected, these exact parameters can be exported as Gemini API code to ensure consistent, recognizable voices across various projects and platforms.

With these new configurations, developers can enhance precision for specific scenarios, creating memorable characters and immersive audio experiences.

Sorry, your browser doesn't support embedded videos, but don't worry, you can download it and watch it with your favorite video player!

Get started with high-fidelity speech generation in the Google AI Studio Playground.

Built for global scale

Gemini 3.1 Flash TTS delivers high-fidelity speech and more precise control across more than 70 languages. These core optimizations bring advanced style, pacing and accent control to major markets — helping developers create localized, expressive speech experiences for users at global scale.

Early developer and enterprise testers are already seeing the impact of 3.1 Flash TTS, highlighting its impressive controllability and expressivity. They’ve told us how audio tags provide a new level of creative precision, transforming simple text into a high-fidelity vocal performance.

Image 4: Quote from Jay of StyleUAI
Image 4: Quote from Jay of StyleUAI
Image 5: Quote from CTO of AIM Intelligence
Image 5: Quote from CTO of AIM Intelligence
Image 6: Quote from Idan Yonas of Artlist
Image 6: Quote from Idan Yonas of Artlist
Image 7: Quote from Lydia Xu of Sierra
Image 7: Quote from Lydia Xu of Sierra
Image 8: Quote from Shivam Rastogi of Invideo AI
Image 8: Quote from Shivam Rastogi of Invideo AI
Image 9: Quote from Fernanda Bejarano of biia
Image 9: Quote from Fernanda Bejarano of biia
Image 10: Quote from John Wu of HeyGen
Image 10: Quote from John Wu of HeyGen
Image 11: Quote from Soami Kapadia of You learn.AI
Image 11: Quote from Soami Kapadia of You learn.AI
Image 12: Quote from Angel Wen of Sylph.ai
Image 12: Quote from Angel Wen of Sylph.ai
Image 13: Quote from Artugrul Cavusoglu of Mindlid
Image 13: Quote from Artugrul Cavusoglu of Mindlid

Jump to position 1 Jump to position 2 Jump to position 3 Jump to position 4 Jump to position 5 Jump to position 6 Jump to position 7 Jump to position 8 Jump to position 9 Jump to position 10

Watermarked with SynthID

All audio generated by Gemini 3.1 Flash TTS is watermarked with SynthID. This imperceptible watermark is interwoven directly into the audio output, allowing the reliable detection of AI-generated content to help prevent misinformation. For more information on our approach to safety and responsibility, you can review the model card.

Image 14
Image 14
Image 15
Image 15
Image 16
Image 16
Image 17
Image 17

Get more stories from Google in your inbox.Get more stories from Google in your inbox.

Email address

Your information will be used in accordance with Google's privacy policy.

Subscribe

Done. Just one step more.

Check your inbox to confirm your subscription.

You are already subscribed to our newsletter.

You can also subscribe with a different email address .

POSTED IN:

Related stories

![Image 18 Chrome #### Turn your best AI prompts into one-click tools in Chrome By Hafsah Ismail Apr 14, 2026](https://blog.google/products-and-platforms/products/chrome/skills-in-chrome/)

![Image 19 Creating opportunity #### Bringing people together at AI for the Economy Forum By James Manyika Apr 14, 2026](https://blog.google/company-news/outreach-and-initiatives/creating-opportunity/ai-economy-forum/)

![Image 20 Google Workspace #### Create, edit and share videos at no cost in Google Vids By David Nachum Apr 02, 2026](https://blog.google/products-and-platforms/products/workspace/google-vids-updates-lyria-veo/)

![Image 21 Developer tools #### Gemma 4: Byte for byte, the most capable open models By Clement Farabet & Olivier Lacombe Apr 02, 2026](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)

![Image 22 Developer tools #### New ways to balance cost and reliability in the Gemini API By Lucia Loher & Hussein Hassan Harrirou Apr 02, 2026](https://blog.google/innovation-and-ai/technology/developers-tools/introducing-flex-and-priority-inference/)

![Image 23 Google Earth #### We’re creating a new satellite imagery map to help protect Brazil’s forests. Apr 01, 2026](https://blog.google/products-and-platforms/products/earth/satellite-imagery-brazilian-deforestation/)

.

Jump to position 1 Jump to position 2 Jump to position 3 Jump to position 4 Jump to position 5 Jump to position 6

Image 24
Image 24

Let’s stay in touch. Get the latest news from Google in your inbox.

SubscribeNo thanks

Survey

Help us improve The Keyword with a one-question survey

Yes No

This survey is anonymous. All responses will be aggregated and used only for analysis to improve our services.

Did this article provide the level of detail you were looking for?

Yes, I got what I needed No, I wanted more technical depth No, I wanted a simpler overview I was looking for something else entirely

✅ Thank you!

Follow Us

  • [](https://www.instagram.com/google/)
  • [](https://twitter.com/google)
  • [](https://www.youtube.com/google)
  • [](https://www.facebook.com/Google)
  • [](https://www.linkedin.com/company/google)

[](https://www.google.com/)

*