xAI actually did it... (Grok 4.6)
xAI发布的Grok 4.6模型在多个基准测试中表现优异,尤其在知识工作和编码任务上超越了OpenAI和Anthropic的模型。
入选理由:Grok 4.6在GDP val基准测试中击败了OpenAI的5.6 Soul和Anthropic的Fable 5 Max
概念
别名:GDPval
OpenAI推出的衡量模型知识工作能力的基准测试。
已跟踪 3 条高相关材料
最近变化
2026-08-13 · Grok 4.6在GDP val基准测试中击败了OpenAI的5.6 Soul和Anthropic的Fable 5 Max
为什么值得关注
GDP val 被反复提及时,通常意味着它正在影响产品路线、开发者工作流或 AI 产业判断。这个页面把分散材料合并成一个可持续更新的观察入口。
xAI actually did it... (Grok 4.6)
Matthew Berman · 8.5 分
xAI发布的Grok 4.6模型在多个基准测试中表现优异,尤其在知识工作和编码任务上超越了OpenAI和Anthropic的模型。
Gemini 3.5 Flash has landed.
Google DeepMind · 8.5 分
Gemini 3.5 Flash在性能和速度上显著提升,相比3.1版本在多个基准测试中表现更优,且速度是其他前沿模型的四倍,现已全面开放使用。
Gemini 3.5 Flash has landed.
Google DeepMind · 7 分
Gemini 3.5 Flash是谷歌推出的新型AI模型,在多项基准测试中表现优于前代3.1 Pro Flash,速度是其他前沿模型的4倍,现已通过产品和API向公众开放。
已收录 3 条与 GDP val 相关的内容,按评分排序。
xAI发布的Grok 4.6模型在多个基准测试中表现优异,尤其在知识工作和编码任务上超越了OpenAI和Anthropic的模型。
入选理由:Grok 4.6在GDP val基准测试中击败了OpenAI的5.6 Soul和Anthropic的Fable 5 Max
Gemini 3.5 Flash delivers significant performance and speed improvements, outperforming 3.1 across nearly all benchmarks with 4x faster output than other frontier models, now fully available through Google products/APIs.
入选理由:Gemini 3.5 Flash在GDP Val等基准测试中表现优异,尤其在编码任务上进步显著
Gemini 3.5 Flash is Google's new AI model that outperforms previous 3.1 Pro Flash in multiple benchmarks, runs 4x faster than other frontier models, and is now available to public through products and APIs.
入选理由:Gemini 3.5 Flash在几乎所有基准测试中都优于3.1 Pro Flash