Perplexity(@perplexity_ai)
We evaluated 13 retrieval models across three relevance sets, using Recall@1000 as the primary metri...
8.5内容质量

TL;DR · AI 摘要
Perplexity的pplx-embed-v1-4b模型在Web Ranking和Combined数据集上以Recall@1000指标领先,Nemotron-3-Embed-8B在Citation数据集表现最佳。
核心要点
- pplx-embed-v1-4b在Web Ranking数据集取得65.73的Recall@1000成绩
- Nemotron-3-Embed-8B在Citation数据集表现优于其他模型(61.68)
- 13个模型在三个相关性数据集的对比评估揭示了不同场景下的性能差异
结构提纲
按章节快速跳转。
- §研究背景
介绍检索模型评估的行业需求与技术挑战。
- ·评估方法
说明Recall@1000指标选择及三个相关性数据集的定义。
- ›核心结果
展示pplx-embed-v1-4b和Nemotron-3-Embed-8B的领先表现数据。
- ·模型对比
分析不同模型在Web Ranking/Citation/Combined数据集的性能差异。
思维导图
用一张图看清主题之间的关系。
查看大纲文本(无障碍 / 无 JS 友好)
- 检索模型评估结果
- 领先模型
- pplx-embed-v1-4b
- Web Ranking: 65.73
- Combined: 69.11
- Nemotron-3-Embed-8B
- Citation: 61.68
- 评估指标
- Recall@1000
金句 / Highlights
值得收藏与分享的关键句。
pplx-embed-v1-4b在Combined数据集达到69.11的Recall@1000,显著高于其他模型
Nemotron-3-Embed-8B在Citation场景的61.68成绩证明其在学术引用检索的优越性
13个模型的跨数据集评估揭示了检索效果与应用场景的强相关性
#信息检索#模型评估#Recall@1000#Perplexity
打开原文Perplexity 在 X 上的推文:"我们评估了 13 个检索模型在三个相关性数据集上的表现,以 Recall@1000 为主要指标。pplx-embed-v1-4b 在 Web Ranking(65.73)和 Combined(69.11)数据集上表现优于其他模型,而 Nemotron-3-Embed-8B 在 Citation(61.68)数据集上领先。" / X
@perplexity_ai
我们评估了 13 个检索模型在三个相关性数据集上的表现,以 Recall@1000 为主要指标。pplx-embed-v1-4b 在 Web Ranking(65.73)和 Combined(69.11)数据集上表现优于其他模型,而 Nemotron-3-Embed-8B 在 Citation(61.68)数据集上领先。
2026 年 9 月 9 日 下午 8:21
3.6K
浏览量
1
3