返回 AI 情报
模型2026-09-28T17:00:36.000ZX:Arena (@arena)

Claude Opus 5.5 (High) 进入 Agent Arena 第 2 名,成本比 Opus 5 (Max) 低 56%

Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement score of +12.15%. Only Fable 5.1 (Max) ra...

AI 摘要

Arena 公布 Claude Opus 5.5 (High) 进入 Agent Arena 排名第 2,净改进分 +12.15%,仅次于 Fable 5.1 (Max)。

正文 · AI 翻译

Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement score of +12.15%. Only Fable 5.1 (Max) ranks higher. However, at a $1.31 median price per task, Opus 5.5 (High) comes in at 64% less cost, pushing out the Pareto frontier.

Opus 5.5 (High) posts a higher net improvement score than both prior Opus 5 variants, while costing 40% less than Opus 5 (High) and 56% less than Opus 5 (Max).

By signal, Opus 5.5 (High) ranks: - #1 Steerability (+14.50%) - #2 Confirmed Success (+15.50%) - #3 Praise vs Complaint (+19.80%) - #4 Bash Recovery (+10.64%)

Congrats to the @AnthropicAI on another frontier model release!

原文

Original Title

Claude Opus 5.5 (High) from @AnthropicAI just entered Agent Arena at #2, with a net improvement score of +12.15%. Only Fable 5.1 (Max) ra...

Source

X:Arena (@arena)

Site

x.com

Published

2026-09-28T17:00:36.000Z

阅读原文· x.com

继续阅读