Philipp Schmid(@_philschmid)
Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs!...
8.5内容质量

TL;DR · AI 摘要
Gemini Embedding 2 正式发布,支持文本、图像、视频、音频和 PDF 的统一嵌入模型。
核心要点
- 单个模型支持 5 种模态的统一嵌入空间
- 原生支持音频嵌入,无需转录步骤
- 灵活输出维度,支持多语言和大输入长度
#Gemini#Embedding#多模态#机器学习
打开原文🖼️ 5 modalities in a single unified embedding space 🌍 Supports up to 8,192 input tokens, 100+ languages 🎧 Embeds audio natively, no transcription step needed 📐 Flexible output https://t.co/WZopA8O4DW" / X
Philipp Schmid on X: "Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs! 🖼️ 5 modalities in a single unified embedding space 🌍 Supports up to 8,192 input tokens, 100+ languages 🎧 Embeds audio natively, no transcription step needed 📐 Flexible output https://t.co/WZopA8O4DW" / X
Don’t miss what’s happening

Philipp Schmid 
Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs! 5 modalities in a single unified embedding space
Supports up to 8,192 input tokens, 100+ languages
Embeds audio natively, no transcription step needed
Flexible output dimensions: 3,072 / 1,536 / 768 via MRL
Up to 6 images, 120s video, 180s audio and 6-page PDFs per request
·
4
10
100
56