NVIDIA launches AIPerf to reliably test LLM speed at scale

NVIDIA has launched AIPerf, a new tool designed to reliably measure the speed of large language model (LLM) inference at scale. Unlike its predecessor, GenAI-Perf, AIPerf is a ground-up rewrite and doesn’t run on top of Perf Analyzer; this architectural overhaul addresses limitations encountered when benchmarking under realistic loads. The tool utilizes a multiprocessed system coordinated over ZMQ to bypass the Python Global Interpreter Lock, a common concurrency bottleneck, ensuring the client itself doesn’t become the limiting factor.

AIPerf Architecture: Scaling LLM Benchmarking Beyond GenAI-Perf

GenAI-Perf, the tool it replaces, relied on a single-process architecture that became constrained by the GIL under realistic workloads and request rates; this created a bottleneck that masked true server performance. The architectural overhaul represented by AIPerf directly addresses this issue, allowing for more accurate benchmarking by preventing the benchmarking tool itself from becoming a limiting factor in the measurement process.

This design choice ensures that observed latency accurately reflects the server’s capabilities, rather than the client’s processing constraints. According to the developers, this separation allows AIPerf to scale more effectively and supports a broader range of workload types.

The tool currently supports over fifteen endpoint types, including chat, responses, and NIM rankings, and accommodates public datasets like ShareGPT, as well as trace replay formats from Mooncake, Baseten, WEKA (AgentX), and others. AIPerf’s benchmarking capabilities extend beyond simple throughput measurements, offering detailed breakdowns of latency distributions to identify potential performance bottlenecks.

The system can pinpoint long-tail distributions, revealing scenarios where a server might exhibit healthy average latency but suffer from significant delays for a small percentage of requests. This detailed analysis helps identify and address performance issues that might not be apparent from aggregate metrics alone.

Beyond latency, AIPerf integrates telemetry data from DCGM or pynvml, providing insights into GPU power draw, utilization, and memory consumption during benchmark runs. This allows for correlation between latency spikes and resource constraints, eliminating the need for separate profiling sessions. The ability to simultaneously monitor both performance metrics and resource utilization streamlines the debugging process and facilitates more comprehensive performance analysis. For example, correlating a latency spike with a memory pressure event can quickly identify memory-related bottlenecks without requiring additional profiling tools.

Qwen3-0.6B Setup: Initial Synthetic ISL/OSL Benchmark with vLLM

Establishing a reproducible measurement loop is central to benchmarking large language model inference, and the team began initial tests using the Qwen3-0.6B model served through vLLM. This model was deliberately chosen for its size, allowing it to run on a single GPU, and its speed, facilitating rapid iteration during the setup of the measurement process; the focus was not on benchmarking Qwen3-0.6B itself, but on creating a reliable and repeatable testing framework.

Once established, the framework allows for easy swapping of different models or endpoint types with a single flag modification. To begin, the team instructed users to pull and start vLLM with the reasoning parser enabled using a specific Docker command, configuring it to host Qwen3-0.6B on port 8000.

AIPerf can be installed using the uv package manager, or within a virtual environment using pip, though users on aarch64 systems may encounter installation issues due to the crick dependency requiring a C toolchain, NVIDIA says. With the server operational and AIPerf installed, a first benchmark profile can be executed, specifying the Qwen3-0.6B model, a “chat” endpoint type, and the local server address. These flags ensure a consistent and controlled baseline for initial testing, limiting variability in request and response lengths.

Additional flags, such as –extra-inputs min_tokens:128 and –extra-inputs ignore_eos:true, further refine the benchmark parameters. Moving beyond static benchmarks, AIPerf allows for the configuration of more dynamic traffic patterns to simulate real-world inference loads. The team demonstrated this by switching to a Poisson arrival pattern with a request rate, using the –arrival-pattern poisson and –request-rate 10 flags.

This configuration introduces variability in the inter-arrival times of requests, drawing from an exponential distribution, more closely mirroring the unpredictable nature of production traffic, according to NVIDIA. This work falls within the areas of agentic AI, generative AI, developer tools, and techniques, and is relevant to users of AGX, Jetson, TensorRT, and TensorRT-LLM platforms.

AIPerf Metrics: Measuring TTFT, ITL, and Throughput at Scale

AIPerf delivers a detailed breakdown of latency metrics, reporting Time to First Token (TTFT), Inter-Token Latency (ITL), and output token throughput alongside percentile breakdowns, minimums, maximums, averages, and standard deviations for comprehensive analysis. These granular measurements move beyond simple averages, allowing engineers to pinpoint performance bottlenecks within the language model inference pipeline and assess the impact of varying request loads. The tool’s output includes metrics to both CSV and JSON formats, facilitating integration into existing data analysis pipelines and enabling automated performance monitoring.

Beyond the core latency and throughput metrics, AIPerf also reports request latency, the end-to-end time for a full response, combining prefill and decode costs into a single value for a complete view of performance. The developers note the depth of data provided for detailed analysis and reproducible results.

This contrasts with the single-concurrency case, an idealized scenario that minimizes TTFT but sacrifices throughput, highlighting the trade-offs between latency and capacity. The developers provide a migration guide for those porting existing workflows, acknowledging the changes from GenAI-Perf and facilitating a smooth transition to the new benchmarking methodology.

Dynamic Load Testing: Configuring Poisson Patterns & Variable Token Lengths

AIPerf allows users to move beyond static benchmarks by configuring traffic patterns that more closely resemble real-world inference loads, offering options including constant, Poisson, and gamma arrival patterns with tunable burstiness. The tool’s capabilities extend to gradual ramping of concurrency and request rate, alongside synthetic distributions for both input and output token lengths, giving operators control over load shape rather than simply volume. A demonstration utilized the Qwen3-0.6B model served through vLLM to illustrate these dynamic testing features. Without these settings, throughput numbers can be lower and lack reproducibility across multiple runs.

Measuring time-to-first-token (TTFT) and inter-token latency (ITL) necessitates the inclusion of the –streaming flag, as it is not optional for these metrics. Moving beyond this fixed pattern, AIPerf enables the creation of more realistic scenarios through synthetic workload knobs that introduce variability into requests. A command line configuration employing –model Qwen/Qwen3-0.6B, –endpoint-type chat, –url localhost:8000, –arrival-pattern poisson, –synthetic-input-tokens-mean 512, –synthetic-input-tokens-stddev 128, –output-tokens-mean 128, and –output-tokens-stddev 32 demonstrates this functionality.

This configuration simulates a server experiencing bursts and gaps in traffic, mirroring the queuing behavior observed under real-world conditions. Introducing –synthetic-input-tokens-stddev 128 adds variance around the 512-token mean, creating a mix of short and long prompts, forcing the server to handle variable prompt lengths during prefill.

Similarly, –output-tokens-stddev 32 introduces variance on the output side, removing the need for the min_tokens and ignore_eos flags used in the static benchmark. Analysis of LLM metrics from the Poisson run reveals noticeably wider distributions compared to the static baseline, an expected outcome given the increased competition for GPU access and varying prefill lengths. The team observed that the Poisson command line introduced a request rate centered around 10 requests per second.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of The Neuron

The Neuron

With a keen intuition for emerging technologies, The Neuron brings over 5 years of deep expertise to the AI conversation. Coming from roots in software engineering, they've witnessed firsthand the transformation from traditional computing paradigms to today's ML-powered landscape. Their hands-on experience implementing neural networks and deep learning systems for Fortune 500 companies has provided unique insights that few tech writers possess. From developing recommendation engines that drive billions in revenue to optimizing computer vision systems for manufacturing giants, The Neuron doesn't just write about machine learning—they've shaped its real-world applications across industries. Having built real systems that are used across the globe by millions of users, that deep technological bases helps me write about the technologies of the future and current. Whether that is AI or Quantum Computing.

Latest Posts by The Neuron: