返回 AI 情报
产品2026-09-30T03:10:36.000ZX:Artificial Analysis (@ArtificialAnlys)

Artificial Analysis 开源 AA-AgentPerf-Local 本地模型智能体推理测试工具并公布首批结果

Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own ...

AI 摘要

Artificial Analysis 发布开源工具 AA-AgentPerf-Local,通过在笔记本和工作站硬件上回放 8 个真实智能体任务共 168 轮模型交互(上下文增长至约 56K tokens)来测试本地推理性能。

正文 · AI 翻译

Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own laptop or workstation, and browse our list of serving configurations to plan your next agent setup

Key points:

➤ We’re open sourcing AA-AgentPerf-Local, which replays real agent trajectories on laptop & workstation hardware to test inference performance

➤ We’re releasing initial results for NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro

➤ The tool and leaderboard will soon expand to cover more hardware, frameworks and models, and will stay updated over time as new releases launch

We already benchmark inference performance on mobile phones and datacenter-scale hardware, and are expanding our coverage to include laptops and workstations. Alongside hosting a leaderboard displaying results from popular model and hardware combinations, we are open-sourcing all code and data required to run AA-AgentPerf-Local to ensure that individuals and companies are able to run their own trials, informing their local AI serving decisions.

AA-AgentPerf-Local replays real agent sessions. Our default workload is 8 recorded agentic tasks, spanning 168 model turns. Each request carries the full conversation so far, as a real agent's would, so context grows to ~56K tokens. Every turn generates exactly its recorded number of tokens, so every system does identical work. Tool execution is skipped by default to isolate inference speed, but when benchmarking your own system, you can also replay the real recorded tool delays or run tool calls live on your CPU.

At launch, we are focusing on the performance a single agent can achieve when able to utilize the entire system; we plan to expand this coverage over time to cover multi-agent systems and scenarios where an agent must run alongside other regular processes.

The initial set of hardware covered on our official page is: NVIDIA DGX Spark (128 GB), AMD Ryzen AI Halo (128 GB), MacBook Pro M5 Pro (64 GB), and NVIDIA GeForce RTX 5090 (32 GB). These have been selected to cover a range of platforms (CUDA, ROCm, Vulkan, Metal), memory capacities, and bandwidths. We will be expanding the featured hardware to include x86 (and other) laptops, AI-focused graphics cards such as the RTX PRO 6000 Blackwell, and more, to give consumers a well-rounded view of performance across different hardware types.

The models featured at launch are: Qwen3.5-9B, Qwen3.8-27B, Qwen3.6-35B-A3B, and Ling 3.0 Flash (124B / 5B active). These initial models span a range of memory requirements and dense/MoE architectures, and are each benchmarked at 4-bit quantizations to reflect realistic serving conditions. The core set of models we feature will shift over time as new open weights models are released. Beyond the featured models, AA-AgentPerf-Local is able to benchmark performance of any OpenAI-compatible inference server the user runs, meaning that any model and config can be tested locally.

Choice of serving configuration is important, with the runtime, quantization and speculative decoding changing results substantially. Where an official off-the-shelf config was available for a system and model pair, we used the published config. We developed our own configs for all other cases. Every config uses speculative decoding (MTP, DFlash or DSpark), and all 14 are published in the repo and on the configs page on our website.

Initial results:

➤ Completion time mapped most closely to each model’s active parameter count: Qwen3.6-35B-A3B (3B active) was the fastest model on every system, e.g. 2.5-3.3x faster than the dense Qwen3.8-27B. However, active parameters are not the whole story, with Ling 3.0 Flash (124B total, 5B active) still finishing behind Qwen3.5-9B (nearly double the active parameter count) on all hardware that can support it.

➤ The GeForce RTX 5090 was the fastest system for every model that fits in its 32 GB, achieving completion times >3.5x faster than the other systems. Single-user decoding is heavily influenced by memory bandwidth, and the GeForce RTX 5090 has 1,792 GB/s against 256-307 GB/s for the unified-memory systems.

➤ The DGX Spark and Ryzen AI Halo are overall similar systems, with the same amount of unified memory, comparable memory bandwidth, and the same launch MSRP of $4,000. On our default trajectory set, the DGX Spark was 1.4-1.7x faster on three of four models, with the Ryzen AI Halo tying it on Qwen3.5-9B. The gap is far larger than their 7% bandwidth difference - the Spark has greater low-precision compute than the Ryzen AI Halo, enabling it to prefill faster, and its more mature CUDA software likely plays a role in enabling better MoE and speculative decoding performance.

➤ The MacBook Pro (M5 Pro, 64 GB, 20-core GPU) is the only laptop tested so far, and exhibited competitive results, finishing within 2–9% of the Ryzen AI Halo on Qwen3.6-35B-A3B and Qwen3.8-27B (though 21% slower on Qwen3.5-9B). It has the most memory bandwidth of the three unified-memory systems (307 GB/s) and its current price of $3,700 is the lowest of the systems tested so far. Its results are likely held back by software maturity and compute available for prefill.

➤ Despite the agentic trajectories serving 73-93% of prompt tokens from the KV cache, simply reading each turn’s new input (during prefill) used up substantial proportions of the end-to-end completion time, e.g. 22-41% for Qwen3.8-27B. This was especially impactful on systems with low compute FLOP/s relative to their memory bandwidth, such as the Ryzen AI Halo and potentially the MacBook Pro (MacBook FLOP/s are unpublished).

➤ Speculative decoding was implemented on all of the most successful configs so far, e.g. raising Qwen3.8-27B decode speeds ~30-120% above the bandwidth-constraint roofline.

原文

Original Title

Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own ...

Source

X:Artificial Analysis (@ArtificialAnlys)

Site

x.com

Published

2026-09-30T03:10:36.000Z

阅读原文· x.com

继续阅读

Artificial Analysis 开源 AA-AgentPerf-Local 本地模型智能体推理测试工具并公布首批结果 - AI 情报频道 - 小黑丸