A server can lead a tokens per second comparison and still keep your users waiting. AI inference benchmarks predict production latency only when they reproduce your request mix, arrival patterns, and service-level objectives.
For platform teams, the useful result is sustainable capacity at acceptable latency, with enough headroom for bursts and failures. Start by defining what users experience, then test representative LLM inference requests, including DeepSeek R1, across the complete serving path under load.
Key Takeaways
- Benchmark representative production workloads, including prompt and output lengths, arrival patterns, cache behavior, and the complete serving path.
- Evaluate time to first token, intertoken latency, end-to-end latency, errors, and latency-qualified throughput against workload-specific SLOs; averages and peak throughput alone can hide poor user experience.
- Use open-loop load tests to expose queue growth, and identify sustainable capacity where queues remain stable and latency objectives are met.
- Choose tools according to the measurement boundary, and make hardware comparisons reproducible by pinning configurations and disclosing optimizations.
- Compare cost at the same workload, quality threshold, and latency target, then carry benchmark metrics into production monitoring to detect when traffic moves beyond the tested envelope.
Read AI inference benchmarks through your SLOs
Aggregate throughput describes how much work a deployment completes. However, it doesn’t tell you whether individual requests arrive quickly enough or stream smoothly.
For interactive applications, separate the first response from the rest of generation. Time to first token includes the client-observed wait before output starts. Intertoken latency measures the gaps between subsequent output tokens.
Use these metrics together to assess inference performance, rather than choosing a winner from one column of throughput metrics.
| Metric | What it measures | Production interpretation |
|---|---|---|
| Time to first token (TTFT) | Request start to first output token | Initial responsiveness |
| Intertoken latency (ITL) | Time between successive output tokens | Streaming smoothness |
| End-to-end latency | Request start to completed response | Total task duration |
| Token throughput | Generated tokens per second across requests | Serving capacity |
| Error and timeout rates | Failed or unfinished requests | Whether capacity is usable |
Track p50, p95, and p99 latency where the sample size supports them. Also preserve the measurement boundary: client-side TTFT includes network and gateway delays that engine-only measurements omit.
Don’t substitute average time per output token for individual intertoken latency gaps and their tail distribution. An average can conceal pauses that users notice.
Define SLOs by workload. A chat interface needs responsive streaming; a nightly classification job may prioritize completion deadlines. Report latency-qualified throughput, counting successful work that meets defined latency limits. Keep failures and timeouts visible instead of dropping them from latency summaries.
Build a workload that matches production
A benchmark with one prompt length and one output length describes a narrow operating point. Production traffic usually spans several sequence lengths and serving conditions, including batch size.
Preserve prompt and output distributions
Replay sanitized request traces when possible. Retain input-token lengths, output lengths, arrival timestamps, and workload categories. Because tokenizers differ, measure lengths with the tested model’s tokenizer, especially for models such as DeepSeek R1.
Long prompts increase work in the prefill stage and consume more KV-cache space. Meanwhile, longer outputs extend decoding and keep requests resident in memory.
Include retrieval-augmented prompts and conversation histories if the application uses them. Synthetic prompts help isolate scaling behavior, but they don’t reproduce semantic prefix reuse or realistic speculative-decoding acceptance rates.
Measure actual generated lengths, too. A maximum output-token setting doesn’t guarantee that every response reaches that length.
Separate cold, warm, and cached traffic
Model loading, kernel initialization, and compilation can distort a short test. Warm the deployment before steady-state measurement, then test startup and recovery separately.
Prefix caching needs its own controls. Record hit rates, reused token counts, and cache state at the beginning of each run. Otherwise, repeated prompts can make prefill appear cheaper than fresh production traffic.
The KV cache stores attention keys and values for active sequences. Its memory demand grows with retained tokens and model architecture, so concurrency and context length jointly determine capacity.
Test a realistic mixture of cache hits and misses. Keep uncached results available so teams can understand performance when traffic changes or replicas restart.
Test arrival rates, not just concurrency
Performance benchmarking characterizes a configuration and its inference performance. Load testing checks whether the deployed service maintains its objectives as demand rises. Production planning needs both.
Use open-loop traffic to expose queues
A closed-loop test keeps a fixed number of requests active and replaces each completed request. It’s useful for concurrency sweeps, but response slowdowns also reduce the arrival rate.
An open-loop generator schedules arrivals independently of completions. As a result, it exposes queue growth when demand exceeds serving capacity.
GuideLLM’s workload testing framework supports configurable traffic profiles and endpoint-level measurements. Use fixed-rate traffic for controlled comparisons and bursty or Poisson arrivals when they match observed demand.
Check the generator itself for CPU, connection, and network limits. An overloaded client can make the server appear underutilized.
Find the sustainable operating point
Increase offered load in steps. At each step, record achieved throughput, token throughput, latency distributions, queue depth, errors, and outstanding requests.
Hold each level long enough to reveal accumulating queues. Then repeat around the point where tail latency starts rising sharply. Sustainable capacity requires stable queues and acceptable completion behavior, not a brief throughput peak.
A test that stops generating traffic while requests remain queued needs a drain period. Otherwise, unfinished requests can disappear from the reported latency distribution.
Finally, run overload and recovery tests. Verify admission control, timeout handling, and whether canceled requests release resources promptly. A deployment that recovers poorly can remain slow after the burst ends.
Choose benchmark tools by measurement boundary
No single suite answers every procurement and deployment question.
MLPerf Inference Datacenter provides standardized workloads and evaluation rules for hardware comparisons across datacenters. Its Offline scenario measures bulk-processing performance, while Server tests exercise arrivals under benchmark-defined latency constraints. MLPerf Inference results help screen hardware, but don’t guarantee performance for your application.
GuideLLM supports open source benchmarking of serving endpoints with configurable requests and traffic. As an endpoint tool, it measures the API path and can compare deployments behind compatible interfaces, regardless of their inference framework.
For engine-level experiments, vLLM and TensorRT-LLM offer benchmarking tools. The benchmark CLI documentation recommends GuideLLM for production-server benchmarking. Use engine tests to investigate runtime changes, then validate them through the deployed endpoint.
Apply the same principle to TensorRT-LLM and SGLang: internal measurements help explain performance, while client measurements establish application-facing behavior.
Record tool versions and metric definitions alongside results. Different harnesses may calculate token timing, throughput, or request duration differently.
Keep the load generator outside the serving host when measuring a remote service. Place it where actual callers run if regional network latency matters. Explicitly label whether tokenization, authentication, routing, and response serialization fall inside the timed boundary.
Make hardware comparisons reproducible
Comparing NVIDIA Blackwell systems, AMD Instinct MI355X deployments, and CPU servers requires a common workload contract. Hardware labels and hardware optimization settings alone don’t establish an apples-to-apples comparison.
Pin the model and execution configuration
Record the model revision, such as DeepSeek R1, tokenizer, precision, quantization method, context limit, runtime and TensorRT-LLM versions, drivers, container digest, and startup arguments.
Also capture accelerator count, CPU allocation, memory, interconnect, and network topology. For multi-node serving, report tensor and pipeline parallelism, communication overhead, and the number of replicas.
CPU trials need equivalent discipline. Record NUMA placement, thread settings, instruction-set support, and memory configuration. A small model such as Llama 3.2 3B Instruct can be a candidate, but its exact implementation and task quality must pass evaluation.
The MLPerf Inference rules illustrate the controls standardized comparisons require. Application tests need similarly explicit configuration records.
Disclose optimizations and validate quality
Quantization, speculative decoding, and multi-token prediction can change throughput and latency. Test them as named configurations rather than invisible tuning differences.
Speculative decoding proposes tokens that a target model verifies. Its gains depend on acceptance rates, verification cost, and workload. Multi-token prediction likewise needs disclosure of the implementation and enabled settings.
Report draft-model resource use and accepted output tokens. Proposed tokens aren’t delivered output, so counting them inflates useful throughput.
Run task-level quality checks after precision or decoding changes. Then repeat performance tests with fixed datasets and seeds where supported.
For purchased infrastructure, make workload-testing access part of your GPU cloud contract requirements. A provider’s peak benchmark doesn’t establish the performance of your allocated cluster.
Compare cost at the same latency target
A lower hourly price doesn’t guarantee a lower cost per token. A slower deployment may need more replicas to meet the same tail-latency target.
First, size each candidate to satisfy the same workload, quality threshold, and SLOs. Then calculate cost using that fleet size, including spare capacity and expected utilization.
For generated output, calculate cost per token using:
Cost per million acceptable output tokens = total serving cost ÷ acceptable output tokens × 1,000,000.
Define “acceptable” before testing. Count output tokens from successful requests that meet the chosen quality and latency requirements. Report input-token volume separately because long prompts can drive substantial compute cost without increasing generated output.
Include host instances, accelerators, networking, storage, and platform overhead within the stated cost boundary. On-premises comparisons also need depreciation, power, cooling, operating costs, and total cost of ownership.
Don’t compare a fully occupied benchmark server with a production fleet carrying idle headroom. FinOps for AI workloads depends on demand patterns as much as raw efficiency.
Publish both peak efficiency and expected operating cost per token. Buyers need to see the price of meeting demand, not merely the cheapest moment during a test.
Carry benchmark metrics into production monitoring
Use the same metric definitions to compare benchmark results with production telemetry. Otherwise, measures of inference performance can describe different experiences.
Monitor client-observed time to first token (TTFT), token gaps, completion latency, and errors by model and workload class. Alongside them, collect queue time, active requests, cache occupancy, CPU pressure, accelerator memory, and network activity.
When TTFT rises, queue growth and prefill pressure help narrow the cause. If token gaps worsen while queues remain stable, investigate decoding contention and resource limits. Neither pattern proves a root cause without supporting telemetry.
Compare observed traffic with the benchmark envelope. Longer prompts, lower cache-hit rates, or additional generation can invalidate the original capacity estimate.
Use those changes to update inference capacity planning. Re-run representative tests after model, runtime, quantization, or routing changes.
Finally, measure the application path separately when retrieval, tools, or safety checks sit outside inference. A healthy model endpoint can still support a slow user experience.
Frequently Asked Questions
Which metrics matter most in an AI inference benchmark?
Track time to first token, intertoken latency, end-to-end latency, throughput, and errors or timeouts. Use latency distributions such as p50, p95, and p99 where sample size supports them, and interpret the results against workload-specific SLOs.
Why use open-loop traffic instead of only testing concurrency?
Open-loop tests schedule requests independently of completions, so they reveal queue growth when demand exceeds serving capacity. Closed-loop tests are useful for concurrency sweeps, but slow responses also reduce their request arrival rate.
How should I choose a benchmark tool?
Use standardized suites such as MLPerf to screen hardware, endpoint tools such as GuideLLM to measure the deployed API path, and engine tools to investigate runtime changes. Validate engine-level results through the application-facing endpoint, and record each tool’s metric definitions and versions.
How can I compare the cost of different deployments fairly?
First size each deployment to meet the same workload, quality threshold, and latency objectives, then calculate cost using the required fleet size and operating headroom. Count only successful output that meets the defined quality and latency requirements, and state which infrastructure and operating costs are included.
Choose capacity backed by evidence
AI inference benchmarks support deployment decisions when workload, measurement boundary, and configuration are explicit. Their strongest result is sustainable, SLO-compliant capacity, supported by repeatable tests and realistic costs.
Treat every benchmark as evidence within a tested operating envelope, not a production guarantee. Carry its metrics into monitoring so you can recognize when real traffic no longer fits that envelope.

