Claude Sonnet 5.5 发布,Terminal-Bench 4.0 得分 70.6% 且价格不变
Claude Sonnet 5.5 is out and it scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at unchanged prices.
AI 摘要
Anthropic 发布 Claude Sonnet 5.5,Terminal-Bench 4.0 得分 70.6%,远高于 Sonnet 5 的 10.3%,价格维持 $2/$10 每百万输入/输出 token。
正文 · AI 翻译
Claude Sonnet 5.5 is out and it scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at unchanged prices.
Overall, 30% cost reduction per-task due to faster speeds and fewer tool calls.
Keeps Sonnet 5's $2/$10 per million input/output tokens, half Opus 5.5's rates.
Its savings instead come from doing less work per job, since Anthropic says fewer tokens and tool calls cut the total cost of a task by up to 30%.
Output also arrives more than 30% faster than Sonnet 5's, making Sonnet 5.5 Anthropic's quickest Sonnet yet.
Users can dial an effort setting, trading longer reasoning and more self-checking for a higher cost per task.
The economics might count for more than the leaderboard. Sonnet 5.5 operating at Low or Medium effort is able to exceed Sonnet 5's top score at about one-tenth the cost per task. On FrontierCode, it says Sonnet 5.5 at High effort scores roughly 10 points above Sonnet 5 at the same setting while costing approximately one-fifteenth as much per task.
Anthropic characterizes Sonnet 5.5 as ideal for comparatively well-defined routine work, such as software debugging, coding, document production, building presentations and spreadsheets, and designing or refining interfaces.
原文
Original Title
Claude Sonnet 5.5 is out and it scores 70.6% on Terminal-Bench 4.0, up from Sonnet 5's 10.3%, at unchanged prices.
Source
X:Rohan Paul (@rohanpaul_ai)
Site
x.com
Published
2026-09-28T19:08:15.000Z