Paper: https://t.co/RbOLdingA0
LLMs在基准测试中表现提升显著,但商业学科应用仍存在明显短板,需更贴近实际场景的评估体系。
入选理由:LLMs在事实记忆类任务准确率已达92%,但商业策略分析仅达67%
人物
也叫:@emollick
推文作者,社交媒体用户。
最近变化
2026-07-22 · LLMs在事实记忆类任务准确率已达92%,但商业策略分析仅达67%
Ethan Mollick 被反复提及时,通常意味着它正在影响产品路线、开发者工作流或 AI 产业判断。这个页面把分散材料合并成一个可持续更新的观察入口。
Paper: https://t.co/RbOLdingA0
Ethan Mollick(@emollick) · 8.5 分
What it feels like to work with Mythos
One Useful Thing · 8.5 分
Cool paper looking at how AIs solve unbounded, complex business problems in many fields by testing how well they can crack the cases we use to teach MBAs in business school: 1) AI already does extremely well across diverse business topics 2) Models are improving rapidly with time
Ethan Mollick(@emollick) · 7.5 分
已收录 10 篇与「Ethan Mollick」相关的 AI 资讯和分析。
LLMs在基准测试中表现提升显著,但商业学科应用仍存在明显短板,需更贴近实际场景的评估体系。
入选理由:LLMs在事实记忆类任务准确率已达92%,但商业策略分析仅达67%
Mythos-class AI模型Claude 5 Fable在多个任务中表现卓越,其能力远超现有模型,并可能改变人与AI的互动方式。
入选理由:Claude 5 Fable在多个任务中表现远超现有模型,包括生成复杂学术论文和创作10页押韵诗。
AI在解决复杂商业问题上表现优异,但缺乏具体技术细节,研究基于MBA案例测试模型能力随时间的提升。
入选理由:AI在跨领域商业案例分析中表现优于人类平均水平
AI is shifting from human-assisted 'co-intelligence' to autonomous agents; Anthropic reports AI now writes 80% of its code with 8x developer productivity gains. The author proposes a 'co-existence' paradigm for thriving alongside AI that sometimes outperforms humans but remains imperfect on the 'jagged frontier'.
入选理由:Anthropic报告AI现编写其80%代码,开发者人均交付量提升8倍,标志自主代理时代来临。
GPT-5.5 Pro excels in fact-checking tasks, accurately tracking key references in complex texts, though its sensitivity to minor details may reduce efficiency.
入选理由:GPT-5.5 Pro 能处理整章内容并准确核查所有关键引用。
Ethan Mollick points out that Mythos proves strong models can not only write code but also discover vulnerabilities. The real threat lies in the fact that general-purpose models are beginning to naturally possess offensive and defensive capabilities, which may mark a turning point in AI security.
入选理由:Mythos 模型展示了强 AI 模型不仅擅长编程,还能发现漏洞。
美国财政部明确反对中国公司通过开源AI进行工业规模的IP盗窃,可能采取制裁措施。
入选理由:美国支持开源AI但反对IP盗窃行为
Ethan Mollick认为Claude的人格化在未来一段时间内将产生重大影响,但具体影响尚不明确。
入选理由:Claude是唯一一个人格化命名的AI。
Marc Andreessen retweeted Ethan Mollick's review of GPT-5.5 Pro, highlighting its strong fact-checking capabilities.
入选理由:GPT-5.5 Pro 能准确识别章节中的关键引用来源。
推文内容空洞,未提供任何技术细节或有价值信息,仅重复模糊表述。
入选理由:推文未包含具体技术内容
与「Ethan Mollick」经常一起出现的 AI 术语。
💡 想追踪「Ethan Mollick」的长期趋势?去 实体雷达 · Ethan Mollick 查看详细分析和跨材料问答。