Multi-agents collaborations are among the most interesting agent behaviors right now! We did an exp...
多智能体协作显著提升了 Gemma 4 的推理速度,达到 5 倍提升,并展现出自我监管和协作机制。
入选理由:100+ 智能体协作使 Gemma 4 推理速度提升 5 倍。
人物
也叫:Thom_Wolf
推文发布者
最近变化
2026-07-04 · Fable 5在解决编程问题时表现出非结构化思维过程
Thomas Wolf 被反复提及时,通常意味着它正在影响产品路线、开发者工作流或 AI 产业判断。这个页面把分散材料合并成一个可持续更新的观察入口。
Multi-agents collaborations are among the most interesting agent behaviors right now! We did an exp...
Thomas Wolf(@Thom_Wolf) · 8.5 分
To all the newcomers excited to try Opus 4.8-level models at home: welcome to OpenWeightLand! Thing...
Thomas Wolf(@Thom_Wolf) · 8.5 分
watching a team of agents tackling a hard theoretical physics problem is quite mesmerizing - self-co...
Thomas Wolf(@Thom_Wolf) · 7.8 分
已收录 15 篇与「Thomas Wolf」相关的 AI 资讯和分析。
多智能体协作显著提升了 Gemma 4 的推理速度,达到 5 倍提升,并展现出自我监管和协作机制。
入选理由:100+ 智能体协作使 Gemma 4 推理速度提升 5 倍。
开源模型在OpenWeightLand中提供了更多灵活性和成本优势,用户可自由选择提供商并进行本地部署。
入选理由:开源模型在OpenWeightLand中可自由选择提供商,价格和功能竞争激烈。
The Physics-Intern framework boosts Gemini 3.1 Pro's performance on the CritPt benchmark from 17.7% to 31.4% via multi-agent collaboration, setting a new SOTA in theoretical physics reasoning.
入选理由:Physics-Intern 使用多智能体协作框架解决复杂理论物理问题。
Hugging Face 每周发布 80TB 天体物理数据集,仅需 4GB 内存加载,包含跨光谱的银河图像与时间序列数据。
入选理由:80TB 天体物理数据集包含 30+ 来源的光谱与时间序列数据
Thomas Wolf is excited about the extension of Terminal-Bench to scientific fields, known as Terminal-Bench Science. This benchmark evaluates AI models' ability to control tools via the command line to achieve scientific goals. It's open for contributions of real scientific workflows until August 2026, aiming to improve AI models' assistance in research work.
入选理由:Terminal-Bench Science evaluates AI models' performance in handling scientific workflows through command-line tools.
开源模型将在AGI时代成为文明韧性的重要组成部分,确保人类在任何个体决策下都能保留有意义的智能访问。
入选理由:开源模型将在AGI时代成为文明韧性的重要组成部分。
AI is moving beyond text, images, and code. Engineering artifacts are becoming a new class of model outputs and evaluating them requires different tools than we use for text, code, or images. This article introduces CADGenBench, a benchmark for evaluating AI's ability to generate 3D engineering parts.
入选理由:AI 生成的 3D 工程零件目前尚无法达到功能性标准。
The article introduces three evaluation metrics designed to capture different types of errors, emphasizing their irreplaceability.
入选理由:形状相似度用于评估整体几何结构的匹配程度。
生物技术初创公司正因闭源模型的潜在风险,将90%的技术栈转向开源。
入选理由:生物技术初创公司正在将90%的技术栈迁移到开源。
Fable 5系统卡泄露显示其解决编程问题时的非结构化思维过程,但缺乏技术细节。
入选理由:Fable 5在解决编程问题时表现出非结构化思维过程
文章内容信息量低,缺乏技术深度和实用性,仅提及模型参数规模和部分用户评论。
入选理由:文章未提供具体技术机制或架构分析。
文章内容为社交媒体上的简短列表,缺乏技术深度和实用信息。
入选理由:文章未提供具体技术细节或实用建议。
Codex can be used to rebuild games as an alternative to paying for them.
入选理由:13 岁的用户使用 Codex 重构游戏以避免付费。
The 2026 trend is friends sharing personalized life/work dashboards and AI setups, similar to teenagers bringing Magic: The Gathering binders to school.
入选理由:2026年社交趋势是朋友间分享个性化生活/工作仪表盘和AI配置
This tweet contains only an emoji and a link to an external image, lacking substantive technical content, architectural analysis, or engineering practice guidance. It has extremely low information density and offers no valuable reading reference for engineers.
入选理由:原文仅为社交媒体状态更新,缺乏可提取的技术深度或原理说明。
与「Thomas Wolf」经常一起出现的 AI 术语。
💡 想追踪「Thomas Wolf」的长期趋势?去 实体雷达 · Thomas Wolf 查看详细分析和跨材料问答。