返回 AI 情报
技巧精选 682026-09-24 09:24Artificial Analysis (@ArtificialAnlys)

Artificial Analysis:Claude Opus 5.5 登顶 Coding Agent Index,但单任务成本升至 $13.04

Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all t…

精选理由

原文给出 Opus 5.5 在三项编码评测的具体分数、token 用量和任务成本,读者可以据此权衡性能与价格。

AI 摘要

Artificial Analysis 测评显示,Claude Opus 5.5 在 Claude Code max effort 下以 66 分登顶 Coding Agent Index,较 Opus 5(60)高 6 分,三项评测 Terminal-Bench 4.0(63.1%)、DeepSWE v1.1(68.4%)、SWE-Atlas-QnA(66.4%)全部提升。

正文 · AI 翻译

Claude Opus 5.5 成为 Artificial Analysis 编程智能体指数的新晋第一,在三项评测中均有提升,但每任务成本更高

在 Claude Code 中以最大努力运行时,Opus 5.5 在编程智能体指数上得分 66,这是我们测得的最高分。相比 Opus 5(60)提升 6 分,相比 Claude Fable 5.1(62)提升 4 分。

Anthropic 已将 Opus 定价下调至每百万输入/输出 token 4 美元/20 美元,此前 Opus 5 为 5 美元/25 美元,缓存读取从 0.50 美元降至 0.20 美元。即便有这些降价,Opus 5.5 的每任务成本仍为 13.04 美元,高于 Opus 5 的 10.79 美元,因为它使用的 token 数量大幅增加。

关键要点:

➤ 在三项编程智能体指数评测中均有提升:Terminal-Bench 4.0 从 Opus 5 的 54.5% 升至 63.1%,DeepSWE v1.1 从 62.5% 升至 68.4%,SWE-Atlas-QnA 从 62.1% 升至 66.4%。提升最大的是 Terminal-Bench,达 +8.6 个百分点。

➤ 最高分伴随着最高的每任务成本:Opus 5.5 的每任务成本为 13.04 美元,相比 Opus 5 的 10.79 美元上涨 21%。它每个任务使用约 1560 万 token,而 Opus 5 为 1140 万,其中输出 token 约为后者的 2.4 倍。

➤ 扩展了编程智能体指数与每任务成本之间的帕累托前沿:在我们的对比中,没有成本更低的模型能匹配 Opus 5.5 的得分。它在其高成本端将前沿向上推进。

其他模型细节:

➤ 定价:每百万输入/输出 token 分别为 $4/$20,较 Opus 5 下降 20%。缓存读取每百万 $0.20,较 $0.50 下降 60%。

➤ 评估设置:Claude Code 以最大 effort 运行,在 DeepSWE v1.1、Terminal-Bench 4.0 和 SWE-Atlas-QnA 上测量。Coding Agent Index 对每项评估赋予相同权重。

原文

Original Title

Claude Opus 5.5 is the new #1 in the Artificial Analysis Coding Agent Index, with gains across all t…

Source

Artificial Analysis (@ArtificialAnlys)

Site

x.com

Published

2026-09-24 09:24

阅读原文· x.com

继续阅读