Artificial Analysis 评测:GPT-Live-1 以 81.5 分登顶 Speech to Speech Index
OpenAI's GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81….
精选理由
原文给出语音到语音模型的完整分数、速度和成本对比,读者可据此在不同后端配置间做选型判断。
AI 摘要
Artificial Analysis 发布 Speech to Speech Index,OpenAI 的 GPT-Live-1 以 81.5 分(Astra 后端,medium 推理强度)排名第一,超过 Grok Voice Think Fast 2.0 High 的 81.3;Sol 后端配置得 80.1 排第三。
正文 · AI 翻译
OpenAI 的 GPT-Live-1 首次亮相便以 81.5 分登顶 Artificial Analysis 语音到语音指数,其委托后端模型为 Astra,领先于 Grok Voice Think Fast 2.0
GPT-Live-1 是 @OpenAI 推出的全新全双工语音到语音模型,它可以将推理和工具调用委托给后端文本模型,同时继续对话。开发者通过 API 流式输入音频并接收语音返回,后端文本模型可单独配置。我们评估了两种后端配置:Astra 在中等推理强度下,以及 Sol 在低推理强度下。
关键要点: ➤ 语音到语音指数:GPT-Live-1(Astra,medium)达到 81.5,排名第 1,而 GPT-Live-1(Sol,low)得分为 80.1,排名第 3。Grok Voice Think Fast 2.0 High 以 81.3 位居两者之间 ➤ 语音智能体竞技场:GPT-Live-1(Sol,low)以 1,053 Elo 在偏好度上排名第 3,任务成功率为 90.9%,而 GPT-Live-1(Astra,medium)以 1,048 Elo 排名第 4,任务成功率为 87.4%。Gemini 3.1 Flash Live Minimal 以 1,096 Elo 在偏好度上领先,而 Grok Voice Think Fast 2.0 High 以 94.6% 在任务成功率上领先 ➤ Tau Voice:GPT-Live-1(Astra,medium)和 GPT-Live-1(Sol,low)在我们的智能体性能基准上占据前两名,分别为 67.9% 和 59.3%,领先于 Grok Voice Think Fast 2.0 High 的 56.5%。 ➤ Big Bench Audio:GPT-Live-1(Astra,medium)得分为 90.1%,GPT-Live-1(Sol,low)得分为 89.0%,在音频推理上落后于 Grok Voice Think Fast 2.0 High 的 97.2% 和 Qwen Audio 3.0 Realtime Plus 的 99.2% ➤ 速度:在 Big Bench Audio 上,GPT-Live-1(Astra,medium)的平均首次音频时间为 1.34 秒,GPT-Live-1(Sol,low)为 1.24 秒,而 Grok Voice Think Fast 2.0 High 为 0.70 秒 ➤ 成本:GPT-Live-1(Astra,medium)每小时输入音频成本为 $5.83,而 GPT-Live-1(Sol,low)为 $4.47,包含委托后端模型使用量,相比之下,在我们固定的 Big Bench Audio 定价子集上,Grok Voice Think Fast 2.0 High 为 $4.80
详见下方了解更多 ⬇️
原文
Original Title
OpenAI's GPT-Live-1 debuts at #1 on the Artificial Analysis Speech to Speech Index at a score of 81….
Source
Artificial Analysis (@ArtificialAnlys)
Site
x.com
Published
2026-09-15 11:14
继续阅读
Trail of Bits 批评 1Password 的 AI 补丁基准存在误导,并发布两个补丁验证 Agent 技能
技巧 · 09/15 11:00
Anthropic 与 OpenAI 提议协调放缓前沿 AI 开发,Cohere CEO 等批评者质疑其真实动机
技巧 · 09/15 09:04
阶跃星辰发布 StepAudio 3 系列语音大模型,多款在 Artificial Analysis 榜单全球第一
模型 · 09/15 07:29
Fireworks 评测 DeepSeek-V4.1-Flash:DeepSWE 达 GPT-6 Astra 水准、成本仅 1/15
模型 · 09/15 00:36