T
traeai
Sign in

人物

Thomas Wolf

别名:Thom_Wolf

Hugging Face联合创始人,AI安全领域专家

已跟踪 21 条高相关材料

TraeAI 观察

相关材料

已收录 21 条与 Thomas Wolf 相关的内容,按评分排序。

watching a team of agents tackling a hard theoretical physics problem is quite mesmerizing - self-co...

The Physics-Intern framework boosts Gemini 3.1 Pro's performance on the CritPt benchmark from 17.7% to 31.4% via multi-agent collaboration, setting a new SOTA in theoretical physics reasoning.

入选理由:Physics-Intern 使用多智能体协作框架解决复杂理论物理问题。

FeaturedTweet#AI Agent#Theoretical Physics#LLM Reasoning#Gemini#CritPt中英混合
I'm very excited about this extension to the celebrated Terminal-Bench to science.

If you're a scie...

Thomas Wolf is excited about the extension of Terminal-Bench to scientific fields, known as Terminal-Bench Science. This benchmark evaluates AI models' ability to control tools via the command line to achieve scientific goals. It's open for contributions of real scientific workflows until August 2026, aiming to improve AI models' assistance in research work.

入选理由:Terminal-Bench Science evaluates AI models' performance in handling scientific workflows through command-line tools.

FeaturedTweet#AI#Science#Terminal-Bench#Benchmarking#Command Line英文
AI is moving beyond text, images, and code.

Engineering artifacts are becoming a new class of model...

AI is moving beyond text, images, and code

Thomas Wolf(@Thom_Wolf)147 字 (约 1 分钟)
70

AI is moving beyond text, images, and code. Engineering artifacts are becoming a new class of model outputs and evaluating them requires different tools than we use for text, code, or images. This article introduces CADGenBench, a benchmark for evaluating AI's ability to generate 3D engineering parts.

入选理由:AI 生成的 3D 工程零件目前尚无法达到功能性标准。

FeaturedTweet#AI#3D Modeling#Engineering#CAD#Benchmarking英文
Thomas Wolf(@Thom_Wolf) 图标

Why three metrics?

Thomas Wolf(@Thom_Wolf)103 字 (约 1 分钟)
70

The article introduces three evaluation metrics designed to capture different types of errors, emphasizing their irreplaceability.

入选理由:形状相似度用于评估整体几何结构的匹配程度。

FeaturedTweet#Evaluation Metrics#Geometric Modeling#Engineering Design英文
this is literally documented in the published Fable 5 System Card

this is literally documented in the published Fable 5 System Card

Thomas Wolf(@Thom_Wolf)108 字 (约 1 分钟)
50

Fable 5系统卡泄露显示其解决编程问题时的非结构化思维过程,但缺乏技术细节。

入选理由:Fable 5在解决编程问题时表现出非结构化思维过程

FeaturedTweet#AI模型#系统泄露#思维过程英文
And it’s only 40B active / 744B total params…

And it’s only 40B active / 744B total params…

Thomas Wolf(@Thom_Wolf)64 字 (约 1 分钟)
50

文章内容信息量低,缺乏技术深度和实用性,仅提及模型参数规模和部分用户评论。

入选理由:文章未提供具体技术机制或架构分析。

FeaturedTweet#模型#参数#评论英文
my 13 yo the other day:

“we didn’t want to pay for the game with my friend so we just rebuilt it wi...

my 13 yo the other day:

Thomas Wolf(@Thom_Wolf)221 字 (约 1 分钟)
50

Codex can be used to rebuild games as an alternative to paying for them.

入选理由:13 岁的用户使用 Codex 重构游戏以避免付费。

FeaturedTweet#AI#Game Development英文
👀

👀

Thomas Wolf(@Thom_Wolf)28 字 (约 1 分钟)
20

This tweet contains only an emoji and a link to an external image, lacking substantive technical content, architectural analysis, or engineering practice guidance. It has extremely low information density and offers no valuable reading reference for engineers.

入选理由:原文仅为社交媒体状态更新,缺乏可提取的技术深度或原理说明。

FeaturedTweet#Social Media#Low Information Density#No Technical Content英文

跨材料问答 · Thomas Wolf

回答基于:Thomas Wolf 相关 21 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.