Philipp Schmid(@_philschmid)

Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs!...

8.5内容质量
Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs!...

TL;DR · AI 摘要

Gemini Embedding 2 正式发布,支持文本、图像、视频、音频和 PDF 的统一嵌入模型。

核心要点

  • 单个模型支持 5 种模态的统一嵌入空间
  • 原生支持音频嵌入,无需转录步骤
  • 灵活输出维度,支持多语言和大输入长度
#Gemini#Embedding#多模态#机器学习
打开原文

🖼️ 5 modalities in a single unified embedding space 🌍 Supports up to 8,192 input tokens, 100+ languages 🎧 Embeds audio natively, no transcription step needed 📐 Flexible output https://t.co/WZopA8O4DW" / X

Philipp Schmid on X: "Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs! 🖼️ 5 modalities in a single unified embedding space 🌍 Supports up to 8,192 input tokens, 100+ languages 🎧 Embeds audio natively, no transcription step needed 📐 Flexible output https://t.co/WZopA8O4DW" / X

Don’t miss what’s happening

Image 6
Image 6

Philipp Schmid ![Image 7](http://x.com/_philschmid)

@_philschmid

Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs! Image 8: 🖼️ 5 modalities in a single unified embedding space Image 9: 🌍 Supports up to 8,192 input tokens, 100+ languages Image 10: 🎧 Embeds audio natively, no transcription step needed Image 11: 📐 Flexible output dimensions: 3,072 / 1,536 / 768 via MRL Image 12: 📎Up to 6 images, 120s video, 180s audio and 6-page PDFs per request

Image 13: Image
Image 13: Image

6:11 PM · Apr 22, 2026

·

4,732 Views

4

10

100

56