T
traeai
Sign in

论文

ICML 2026

别名:icml2026

ThunderAgent被接收为Spotlight论文

已跟踪 3 条高相关材料

TraeAI 观察

相关材料

已收录 3 条与 ICML 2026 相关的内容,按评分排序。

7B打败o3、GPT-5!医学AI智能体让模型学会“看哪里、怎么看”

Ophiuchus-7B achieves a mean score of 68.0 on 8 medical VQA benchmarks, surpassing OpenAI-o3 (62.2), Gemini 2.5 Pro (61.8), and GPT-5 (59.9). The core breakthrough is the new ‘Think with Images/Videos’ paradigm: models actively invoke tools like SAM2 and BiomedParse during reasoning to re-examine key regions/moments, making visual evidence an integral part of cognition—not just input.

入选理由:Ophiuchus-7B在8个医学VQA benchmark平均得分68.0,显著高于o3(62.2)、Gemini 2.5 Pro(61.8)与GPT-5(59.9)

FeaturedArticle#Medical AI#Multimodal LLM#Agent#ICML 2026#Visual Reasoning中文
Together AI Blog 图标

ThunderAgent通过优化KV缓存管理,实现单节点吞吐量提升2倍,集群加速2.4倍,解决代理推理中的缓存抖动问题。

入选理由:单节点吞吐量提升2.5倍,P50延迟降低10倍

FeaturedArticle#ThunderAgent#合成数据生成#LLM推理优化#KV缓存管理英文
LLMs are better at predicting what other models will say than what’s actually true. When they’re wro...

斯坦福AI实验室发现大语言模型在预测其他模型输出时表现优异,但事实准确性不足,且自我一致性方法无法提升无验证场景下的真实性。

入选理由:LLMs预测其他模型输出准确率高,但事实判断错误率高达70%(ICML 2026论文数据)

FeaturedTweet#LLM#AI#ICML#Truthfulness中英混合

跨材料问答 · ICML 2026

回答基于:ICML 2026 相关 3 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.