T
traeai
Sign in

人物

Thomas Wolf

别名:Thom_Wolf

推特用户,转发并评论OpenAI安全实践分享。

已跟踪 30 条高相关材料

TraeAI 观察

相关材料

已收录 30 条与 Thomas Wolf 相关的内容,按评分排序。

watching a team of agents tackling a hard theoretical physics problem is quite mesmerizing - self-co...

The Physics-Intern framework boosts Gemini 3.1 Pro's performance on the CritPt benchmark from 17.7% to 31.4% via multi-agent collaboration, setting a new SOTA in theoretical physics reasoning.

入选理由:Physics-Intern 使用多智能体协作框架解决复杂理论物理问题。

FeaturedTweet#AI Agent#Theoretical Physics#LLM Reasoning#Gemini#CritPt中英混合
I'm very excited about this extension to the celebrated Terminal-Bench to science.

If you're a scie...

Thomas Wolf is excited about the extension of Terminal-Bench to scientific fields, known as Terminal-Bench Science. This benchmark evaluates AI models' ability to control tools via the command line to achieve scientific goals. It's open for contributions of real scientific workflows until August 2026, aiming to improve AI models' assistance in research work.

入选理由:Terminal-Bench Science evaluates AI models' performance in handling scientific workflows through command-line tools.

FeaturedTweet#AI#Science#Terminal-Bench#Benchmarking#Command Line英文
AI is moving beyond text, images, and code.

Engineering artifacts are becoming a new class of model...

AI is moving beyond text, images, and code

Thomas Wolf(@Thom_Wolf)147 字 (约 1 分钟)
70

AI is moving beyond text, images, and code. Engineering artifacts are becoming a new class of model outputs and evaluating them requires different tools than we use for text, code, or images. This article introduces CADGenBench, a benchmark for evaluating AI's ability to generate 3D engineering parts.

入选理由:AI 生成的 3D 工程零件目前尚无法达到功能性标准。

FeaturedTweet#AI#3D Modeling#Engineering#CAD#Benchmarking英文
Thomas Wolf(@Thom_Wolf) 图标

Why three metrics?

Thomas Wolf(@Thom_Wolf)103 字 (约 1 分钟)
70

The article introduces three evaluation metrics designed to capture different types of errors, emphasizing their irreplaceability.

入选理由:形状相似度用于评估整体几何结构的匹配程度。

FeaturedTweet#Evaluation Metrics#Geometric Modeling#Engineering Design英文
I hope we keep Opus 4.6 around for a long time

I hope we keep Opus 4.6 around for a long time

Thomas Wolf(@Thom_Wolf)101 字 (约 1 分钟)
60

Opus 4.6在创意生成测试中表现优于新模型,但文章缺乏技术细节和论证。

入选理由:Opus 4.6在矛盾世界生成测试中胜出,新模型表现较差

FeaturedTweet#AI模型#测试#Opus 4.6英文
Impressive level of openness on such a large run

Impressive level of openness on such a large run

Thomas Wolf(@Thom_Wolf)97 字 (约 1 分钟)
60

文章简要提及MiMo-V2.6模型在强化学习扩展性研究中的进展,但信息量有限。

入选理由:MiMo-V2.6正在执行强化学习训练,使用约2B tokens/step的计算规模

FeaturedTweet#强化学习#RL#MiMo-V2.6#AI研究英文
many sub generations missing actually

many sub generations missing actually

Thomas Wolf(@Thom_Wolf)212 字 (约 1 分钟)
50

推文内容缺乏技术深度,主要为机器人开发的社交互动讨论,未提供可执行的技术方案或原理分析。

入选理由:推文内容缺乏技术深度,主要为机器人开发的社交互动讨论,未提供可执行的技术方案或原理分析

FeaturedTweet#机器人#社交平台中英混合
[CORRECTION] All the *closed* AIs are down

[CORRECTION] All the *closed* AIs are down

Thomas Wolf(@Thom_Wolf)86 字 (约 1 分钟)
50

推文宣布部分封闭AI服务中断,但未提供技术细节或解决方案,信息量有限。

入选理由:封闭AI服务出现中断但未说明原因

FeaturedTweet#AI#服务中断#NVIDIA#Twitter中英混合
this is literally documented in the published Fable 5 System Card

this is literally documented in the published Fable 5 System Card

Thomas Wolf(@Thom_Wolf)108 字 (约 1 分钟)
50

Fable 5系统卡泄露显示其解决编程问题时的非结构化思维过程,但缺乏技术细节。

入选理由:Fable 5在解决编程问题时表现出非结构化思维过程

FeaturedTweet#AI模型#系统泄露#思维过程英文
And it’s only 40B active / 744B total params…

And it’s only 40B active / 744B total params…

Thomas Wolf(@Thom_Wolf)64 字 (约 1 分钟)
50

文章内容信息量低,缺乏技术深度和实用性,仅提及模型参数规模和部分用户评论。

入选理由:文章未提供具体技术机制或架构分析。

FeaturedTweet#模型#参数#评论英文
my 13 yo the other day:

“we didn’t want to pay for the game with my friend so we just rebuilt it wi...

my 13 yo the other day:

Thomas Wolf(@Thom_Wolf)221 字 (约 1 分钟)
50

Codex can be used to rebuild games as an alternative to paying for them.

入选理由:13 岁的用户使用 Codex 重构游戏以避免付费。

FeaturedTweet#AI#Game Development英文

跨材料问答 · Thomas Wolf

回答基于:Thomas Wolf 相关 30 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.