T
traeai
登录

人物

Thomas Wolf

别名:Thom_Wolf

推特用户,转发并评论OpenAI安全实践分享。

已跟踪 30 条高相关材料

TraeAI 观察

相关材料

已收录 30 条与 Thomas Wolf 相关的内容,按评分排序。

watching a team of agents tackling a hard theoretical physics problem is quite mesmerizing - self-co...

Physics-Intern 框架通过多智能体协作将 Gemini 3.1 Pro 在 CritPt 基准上的表现从 17.7% 提升至 31.4%,创下理论物理推理新 SOTA。

入选理由:Physics-Intern 使用多智能体协作框架解决复杂理论物理问题。

精选推文#AI Agent#理论物理#LLM 推理#Gemini#CritPt中英混合
I'm very excited about this extension to the celebrated Terminal-Bench to science.

If you're a scie...

Thomas Wolf is excited about the extension of Terminal-Bench to scientific fields, known as Terminal-Bench Science. This benchmark evaluates AI models' ability to control tools via the command line to achieve scientific goals. It's open for contributions of real scientific workflows until August 2026, aiming to improve AI models' assistance in research work.

入选理由:Terminal-Bench Science evaluates AI models' performance in handling scientific workflows through command-line tools.

精选推文#AI#Science#Terminal-Bench#Benchmarking#Command Line英文
AI is moving beyond text, images, and code.

Engineering artifacts are becoming a new class of model...

AI 正在超越文本、图像和代码

Thomas Wolf(@Thom_Wolf)147 字 (约 1 分钟)
70

AI 正在超越文本、图像和代码,工程构件成为新的模型输出类型,需要新的评估工具。本文介绍了 CADGenBench,一个用于评估 AI 生成 3D 工程零件能力的基准。

入选理由:AI 生成的 3D 工程零件目前尚无法达到功能性标准。

精选推文#AI#3D建模#工程#CAD#基准测试英文
Thomas Wolf(@Thom_Wolf) 图标

为什么使用三个指标?

Thomas Wolf(@Thom_Wolf)103 字 (约 1 分钟)
70

文章提出三种评估指标,分别用于衡量几何形状、接口匹配和拓扑结构的正确性,强调它们各自不可替代。

入选理由:形状相似度用于评估整体几何结构的匹配程度。

精选推文#评估指标#几何建模#工程设计英文
I hope we keep Opus 4.6 around for a long time

I hope we keep Opus 4.6 around for a long time

Thomas Wolf(@Thom_Wolf)101 字 (约 1 分钟)
60

Opus 4.6在创意生成测试中表现优于新模型,但文章缺乏技术细节和论证。

入选理由:Opus 4.6在矛盾世界生成测试中胜出,新模型表现较差

精选推文#AI模型#测试#Opus 4.6英文
Impressive level of openness on such a large run

Impressive level of openness on such a large run

Thomas Wolf(@Thom_Wolf)97 字 (约 1 分钟)
60

文章简要提及MiMo-V2.6模型在强化学习扩展性研究中的进展,但信息量有限。

入选理由:MiMo-V2.6正在执行强化学习训练,使用约2B tokens/step的计算规模

精选推文#强化学习#RL#MiMo-V2.6#AI研究英文
many sub generations missing actually

many sub generations missing actually

Thomas Wolf(@Thom_Wolf)212 字 (约 1 分钟)
50

推文内容缺乏技术深度,主要为机器人开发的社交互动讨论,未提供可执行的技术方案或原理分析。

入选理由:推文内容缺乏技术深度,主要为机器人开发的社交互动讨论,未提供可执行的技术方案或原理分析

精选推文#机器人#社交平台中英混合
[CORRECTION] All the *closed* AIs are down

[CORRECTION] All the *closed* AIs are down

Thomas Wolf(@Thom_Wolf)86 字 (约 1 分钟)
50

推文宣布部分封闭AI服务中断,但未提供技术细节或解决方案,信息量有限。

入选理由:封闭AI服务出现中断但未说明原因

精选推文#AI#服务中断#NVIDIA#Twitter中英混合
this is literally documented in the published Fable 5 System Card

this is literally documented in the published Fable 5 System Card

Thomas Wolf(@Thom_Wolf)108 字 (约 1 分钟)
50

Fable 5系统卡泄露显示其解决编程问题时的非结构化思维过程,但缺乏技术细节。

入选理由:Fable 5在解决编程问题时表现出非结构化思维过程

精选推文#AI模型#系统泄露#思维过程英文
And it’s only 40B active / 744B total params…

And it’s only 40B active / 744B total params…

Thomas Wolf(@Thom_Wolf)64 字 (约 1 分钟)
50

文章内容信息量低,缺乏技术深度和实用性,仅提及模型参数规模和部分用户评论。

入选理由:文章未提供具体技术机制或架构分析。

精选推文#模型#参数#评论英文
my 13 yo the other day:

“we didn’t want to pay for the game with my friend so we just rebuilt it wi...

my 13 yo the other day:

Thomas Wolf(@Thom_Wolf)221 字 (约 1 分钟)
50

Codex 可用于重构游戏,替代付费购买。

入选理由:13 岁的用户使用 Codex 重构游戏以避免付费。

精选推文#AI#游戏开发英文

跨材料问答 · Thomas Wolf

回答基于:Thomas Wolf 相关 30 条材料
    0 / 500

    AI 可能会生成不准确的信息,请核实重要内容