[AINews] not much happened today
DeepSeek V4-Flash 0731通过微调实现性能跃升,成本降低60%,但未改变架构。
入选理由:DeepSeek V4-Flash 0731在Terminal-Bench测试中提升25.8个百分点至82.7
概念
别名:TB
评估模型终端任务性能的基准测试套件
已跟踪 7 条高相关材料
最近变化
2026-08-01 · DeepSeek V4-Flash 0731在Terminal-Bench测试中提升25.8个百分点至82.7
为什么值得关注
Terminal-Bench 被反复提及时,通常意味着它正在影响产品路线、开发者工作流或 AI 产业判断。这个页面把分散材料合并成一个可持续更新的观察入口。
[AINews] not much happened today
Latent Space · 8.5 分
DeepSeek V4-Flash 0731通过微调实现性能跃升,成本降低60%,但未改变架构。
DeepSeek V4 Flash now runs updated weights on AI Gateway
Vercel News · 8.5 分
DeepSeek V4 Flash在AI Gateway更新权重后Terminal-Bench得分提升至82.7,增强代理能力且无需代码修改。
OpenAI Just Introduced GPT 5.6 (Beats Claude Fable 5 And Mythos)
TheAIGRID · 8.5 分
OpenAI 推出 GPT 5.6 系列模型,包含 Soul、Terror 和 Luna,Soul 超越 Claude Mythos 5 在终端任务表现。
已收录 7 条与 Terminal-Bench 相关的内容,按评分排序。
DeepSeek V4-Flash 0731通过微调实现性能跃升,成本降低60%,但未改变架构。
入选理由:DeepSeek V4-Flash 0731在Terminal-Bench测试中提升25.8个百分点至82.7
DeepSeek V4 Flash在AI Gateway更新权重后Terminal-Bench得分提升至82.7,增强代理能力且无需代码修改。
入选理由:Terminal-Bench测试得分从56.9提升至82.7,性能提升25.8分
OpenAI 推出 GPT 5.6 系列模型,包含 Soul、Terror 和 Luna,Soul 超越 Claude Mythos 5 在终端任务表现。
入选理由:GPT 5.6 Soul 在终端任务中超越 Claude Mythos 5。
Fireworks AI 已上线 GLM 5.2 模型,支持 1M-token 上下文,专注于代码生成,并在多个基准测试中表现优异。
入选理由:GLM 5.2 支持 1M-token 上下文,适用于复杂任务。
Google releases the Gemini 3.5 family, starting with 3.5 Flash for complex agentic workflows. It outperforms 3.1 Pro on coding and agent benchmarks and runs 4x faster, reaching 12x in Antigravity.
入选理由:Gemini 3.5 Flash 专为执行复杂、长周期的智能体工作流而设计。
Thomas Wolf is excited about the extension of Terminal-Bench to scientific fields, known as Terminal-Bench Science. This benchmark evaluates AI models' ability to control tools via the command line to achieve scientific goals. It's open for contributions of real scientific workflows until August 2026, aiming to improve AI models' assistance in research work.
入选理由:Terminal-Bench Science evaluates AI models' performance in handling scientific workflows through command-line tools.
A harness is the core infrastructure for building AI Agents, consisting of tools, execution environments, system prompts, and file systems. By optimizing harness engineering, developers can significantly boost Agent performance on benchmarks like Terminal Bench without changing the underlying model.
入选理由:Harness 定义为模型访问的工具、执行环境、系统提示词和文件系统的集合。