T
traeai
Sign in

模型

Claude Sonnet 4.5

别名:Claude

在 Cua-Bench 上完成 5/25 任务的模型。

已跟踪 3 条高相关材料

TraeAI 观察

相关材料

已收录 3 条与 Claude Sonnet 4.5 相关的内容,按评分排序。

Databricks 图标

Databricks' Instructed-Retriever-1 cuts search latency by 3x and TTFT to ~2s via parallel test-time scaling without quality loss. The unified model handles query generation and reranking in parallel using multi-pivot groupwise reranking, achieving Pareto-optimal recall-precision tradeoffs for enterprise RAG systems.

入选理由:Instructed-Retriever-1使搜索延迟降低3倍以上,TTFT降至约2秒,无需重新配置。

FeaturedArticle#RAG#Test-Time Scaling#Instructed-Retriever-1#Databricks#Retrieval英文
Cua 和 Snorkel AI 联合发布「Cua-Bench」:评测 Agent 在专业软件上的 Computer Use 能力 
@trycua @SnorkelAI 

Cua-Bench 首个...

Cua-Bench 是首个评测 Agent 在专业软件上执行任务能力的基准,测试显示当前模型在复杂 GUI 操作和任务规划上存在明显短板。

入选理由:当前最强模型 GPT-5.5 在 Cua-Bench 上仅完成 6/25 任务,完全通过率仅 24%。

FeaturedTweet#Cua-Bench#Agent#KiCad#评测基准#AI中英混合
The last six months in LLMs in five minutes

The last six months in LLMs in five minutes

Simon Willison's Weblog1128 字 (约 5 分钟)
85

November 2025 was a critical inflection point for LLM development, with model performance changing hands five times among three major vendors in six months, coding agents achieving qualitative leaps to daily usability, and emerging tools like Warelay beginning to appear.

入选理由:2025年11月三大厂商模型性能排名变化5次,Claude Opus 4.5最终胜出

FeaturedArticle#LLM#AI Programming#Model Evaluation#Anthropic#OpenAI英文

跨材料问答 · Claude Sonnet 4.5

回答基于:Claude Sonnet 4.5 相关 3 条材料
    0 / 500

    AI may generate inaccurate information. Please verify important content.