Skip to main content
Back to Blog
AI/MLInnovationProduct Development
23 September 20267 min readUpdated 30 September 2026

AA-AgentPerf-Local benchmarks local AI agents on laptops and workstations

AA AgentPerf Local benchmarks local AI agents on laptops and workstations AA AgentPerf Local is an open source inference testing tool designed to measure how local AI agents per...

By Software Development Team

AA-AgentPerf-Local benchmarks local AI agents on laptops and workstations

AA-AgentPerf-Local is an open-source inference testing tool designed to measure how local AI agents perform on laptop and workstation hardware. It replays recorded agent trajectories, publishes results for selected hardware and model combinations, and provides serving configurations that can be used for local testing.

How the benchmark works

The default workload contains eight recorded agentic tasks spanning 168 model turns. Each request includes the complete conversation history accumulated up to that point, allowing the context to grow to approximately 56K tokens. Every turn generates exactly the number of tokens recorded in the original session, so each tested system performs the same work.

Tool execution is disabled by default to isolate inference performance. Users can alternatively replay the recorded tool delays or execute tool calls live on the CPU when testing their own systems.

The initial benchmark measures the performance available to a single agent using the entire system. Future versions are planned to cover multi-agent workloads and cases where an agent runs alongside other regular processes.

Hardware and models

The initial official hardware list includes:

  • NVIDIA DGX Spark, 128 GB
  • AMD Ryzen AI Halo, 128 GB
  • MacBook Pro M5 Pro, 64 GB
  • NVIDIA GeForce RTX 5090, 32 GB

These systems represent different platforms, including CUDA, ROCm, Vulkan, and Metal, as well as varying memory capacities and bandwidths. Additional systems are planned, including x86 laptops and AI-focused graphics cards such as the RTX PRO 6000 Blackwell.

The launch model set includes:

  • Qwen3.5-9B
  • Qwen3.8-27B
  • Qwen3.6-35B-A3B
  • Ling 3.0 Flash, 124B total parameters and 5B active parameters

The models cover different memory requirements and dense or mixture-of-experts architectures. Each is tested using 4-bit quantization to reflect practical serving conditions. The featured model set will change as new open-weight models become available.

AA-AgentPerf-Local can also test any model served through an OpenAI-compatible inference server, allowing users to benchmark their own model and configuration locally.

Serving configurations

The inference runtime, quantization method, and speculative decoding implementation can substantially affect results. When an official, off-the-shelf configuration was available for a hardware and model combination, it was used. Other configurations were developed for the benchmark.

All 14 launch configurations use speculative decoding through MTP, DFlash, or DSpark. The configurations are published in the project repository and on the associated configurations page.

Initial findings

The first results show several performance patterns:

  • Completion time generally tracked active parameter count. Qwen3.6-35B-A3B, with 3B active parameters, was the fastest model on every system. It completed tasks 2.5 to 3.3 times faster than the dense Qwen3.8-27B. Active parameters were not the only factor: Ling 3.0 Flash, with 124B total parameters and 5B active parameters, finished behind Qwen3.5-9B on every system that could run it, despite Qwen3.5-9B having nearly twice as many active parameters.

  • The GeForce RTX 5090 was the fastest system for every model that fit within its 32 GB of memory. Its completion times were more than 3.5 times faster than those of the other systems. Single-user decoding is strongly influenced by memory bandwidth, and the RTX 5090 provides 1,792 GB/s, compared with 256 to 307 GB/s for the unified-memory systems.

  • The DGX Spark and Ryzen AI Halo produced broadly similar results overall. Both have 128 GB of unified memory, comparable memory bandwidth, and a launch MSRP of $4,000. Across the default trajectories, the DGX Spark was 1.4 to 1.7 times faster on three of the four models and tied the Ryzen AI Halo on Qwen3.5-9B. The difference was larger than the 7% bandwidth gap. The DGX Spark provided greater low-precision compute, which enabled faster prefilling, while its more mature CUDA software may have contributed to stronger mixture-of-experts and speculative-decoding performance.

  • The MacBook Pro M5 Pro was the only laptop tested. With 64 GB of memory and a 20-core GPU, it finished within 2% to 9% of the Ryzen AI Halo on Qwen3.6-35B-A3B and Qwen3.8-27B, although it was 21% slower on Qwen3.5-9B. Its 307 GB/s memory bandwidth was the highest among the three unified-memory systems. At $3,700, it also had the lowest current price among the systems tested. Software maturity and available prefill compute may have limited its results.

  • Prefill remained a major part of completion time. The agentic trajectories served 73% to 93% of prompt tokens from the KV cache, but reading each turn's new input still consumed substantial end-to-end time. For Qwen3.8-27B, prefill accounted for 22% to 41% of completion time. This was especially significant on systems with relatively low compute throughput compared with their memory bandwidth, including the Ryzen AI Halo and potentially the MacBook Pro, whose FLOPS are unpublished.

  • The strongest configurations used speculative decoding. For example, speculative decoding increased Qwen3.8-27B decoding speeds by approximately 30% to 120% above the bandwidth-constrained roofline.

Price and performance observations

Most hardware prices are substantially higher than their launch MSRPs. At current market prices, the four initial hardware types become more expensive as their end-to-end completion times decrease. The MacBook Pro M5 Pro, 64 GB, is currently the least expensive system, while a GeForce RTX 5090 system is by far the most expensive. The benchmark page uses launch MSRP by default and allows users to enter custom prices.

Decode speeds for Qwen3.8-27B were broadly similar across the three unified-memory systems, which offer between 256 and 307 GB/s of memory bandwidth. The GeForce RTX 5090 delivered more than five times the decode speed of those systems, largely because its memory bandwidth is more than five times higher at 1,792 GB/s.

Prefill speeds followed a similar pattern, but the DGX Spark's higher compute-to-bandwidth ratio resulted in faster context reading. This matters for the agentic trajectories because, even when most of each prompt is served from cache, each session contains approximately 196K new input tokens and 31K output tokens. When a tool returns a large file, the agent can take up to 24 seconds to respond on the Ryzen AI Halo and up to 39 seconds on the MacBook Pro, compared with approximately 2 seconds on the GeForce RTX 5090.

Planned expansion

AA-AgentPerf-Local and the associated laptops and workstations results page are expected to be updated as new hardware and software become available. Planned additions include:

  • A broader selection of hardware and models
  • A more extensive repository of local AI serving configurations
  • Additional inference frameworks, including Inco Splash, which is currently being tested
  • Leaderboards for user-submitted results
  • Leaderboards using live CPU tool calls
  • Multi-agent scenarios

The project is open source, and its code and data support local trials across OpenAI-compatible inference servers.