AI Inference Capacity Planning for Enterprise Platform Teams

Reading Time: 8 minutes

A GPU can have spare compute and still reject another request. Its KV cache may be full, or new requests may queue behind long generations. That mismatch makes AI inference capacity planning more complex than dividing forecast tokens by a GPU’s advertised throughput.

For an enterprise platform team, capacity planning should start with measured request demand and end with tested service level objectives (SLOs). Along the way, account for concurrency, memory, scaling delays, shared tenants, and infrastructure capacity held in reserve.

Key Takeaways

  • Forecast request-level demand during busy intervals, including workload mix, input and output lengths, tenant traffic, and bursts; daily token totals can hide important differences.
  • Define usable capacity as throughput a measured serving configuration can sustain while meeting latency SLOs, then benchmark representative traffic and concurrency.
  • Check KV-cache and accelerator memory separately from throughput: a replica count that passes a load test can still fail when active requests exhaust cache capacity.
  • Plan for scaling delays, reserved capacity, and shared-fleet costs, and validate the plan with peak, growth, burst, and replica-loss tests.

Why inference needs a different capacity model

Training has a finish line; serving has a queue

Training jobs usually have a known dataset, a target run time, and some flexibility in when they run. An inference endpoint must serve requests as they arrive. A busy hour can include short prompts, long documents, tool calls, and responses that generate thousands of tokens.

Daily token totals hide both traffic shape and workload mix. Two services can produce the same number of tokens but need different GPU fleets: one gets steady traffic, while the other handles sharp bursts.

Latency constrains usable throughput

A server’s maximum tokens per second isn’t necessarily its production capacity. As concurrency rises, batching can improve total throughput while increasing queue time and time-to-first-token (TTFT). Users also experience inter-token latency during generation.

Set separate latency targets for TTFT, time per output token, and request completion. Measure them at a stated traffic level and percentile. This makes capacity claims testable under defined conditions.

Build a forecast from request-level demand

Preserve the mix, not just the average

Start with gateway and serving logs to define the inference workload and guide demand forecasting. For each request, retain arrival time, tenant, model, input tokens, output tokens, completion status, and latency. Keep sensitive prompt content out of planning datasets unless there’s an approved reason to store it.

Group multi-tenant demand by model and workload class. Interactive chat, document analysis, and agent workflows have different token distributions and acceptable wait times. Forecast requests per second during busy intervals, then model input and output lengths separately. Reasoning models can generate more tokens, increasing decode work and keeping cache blocks occupied longer.

An agent may turn one user action into several model calls and retries. For workflow economics, cost per API call for AI products helps connect that fan-out to the service a customer actually uses.

Track the resources around generation

Inference demand also reaches retrieval indexes, object storage, logs, and network links. Use storage capacity planning to forecast document ingestion and index growth separately from serving traffic. A retrieval bottleneck may inflate request latency even when GPUs have room.

Keep at least three demand cases: normal peak, expected growth, and a credible burst. Use trace history for predictive modeling, but treat forecast scenarios as assumptions to validate, not guarantees. Include model upgrades and changes to context-window or output limits. Those product decisions can alter the plan before request volume changes.

Convert demand into a first replica estimate

Use throughput measured at the SLO

Consider an illustrative service peak of 20 requests per second. Its typical request has 900 input tokens and 250 output tokens. That implies 18,000 input tokens per second and 5,000 output tokens per second at that traffic level.

Suppose a representative load test finds that one replica sustains three requests per second while meeting the agreed latency targets. The initial calculation is ceil(20 / 3) = 7 active replicas. Keeping one more ready would make the deployed fleet eight replicas. That eighth replica is a proposed availability and burst allowance, not proof that every failure scenario is covered.

Request concurrency is a separate check. If mean end-to-end service time at peak is 30 seconds, Little’s law gives an average of 20 × 30 = 600 in-flight requests across the fleet. Seven active replicas would carry about 86 each with even resource allocation. Confirm that queues, cache capacity, and tenant placement remain acceptable at that load.

Check memory before approving the count

The KV cache stores attention keys and values for active sequences. Its footprint grows with retained tokens and requests in flight. Model weights, runtime overhead, and cache blocks must all fit in accelerator memory. Longer generations, especially from reasoning models, can keep cache blocks occupied longer.

For a conventional full-context cache, a rough estimate per token is 2 × layers × KV heads × head dimension × bytes per element. Treat that as a screening calculation. The actual footprint depends on the model’s attention design, cache dtype, block allocation, prefix reuse, and serving engine.

A replica count that passes the throughput test can still fail when long prompts and concurrent generations exhaust available cache blocks.

If that happens, revise context limits, routing, batching, memory strategy, or replica count. Don’t assume that adding GPU compute alone fixes a memory limit.

Benchmark a representative serving configuration

Server racks stand beside subtle light trails in a modern data center.

Test the traffic shape you expect

Replay a mix of short and long prompts, output lengths for reasoning models, tenants, and arrival bursts. Match the prefill and decode phase mix expected in production. Hold the model, precision, serving engine, batching and request scheduling, and accelerator setup constant while comparing load levels. Run long enough to observe queue buildup, cache churn, and steady-state behavior.

Record throughput at the SLO, not only the server’s peak throughput. Repeat tests after changing batching limits, quantization, context length, or machine learning model versions. A single tokens-per-second figure cannot stand in for these tests.

Read serving metrics alongside GPU metrics

Track P50 and P95 time-to-first-token (TTFT), inter-token latency, and request completion time against latency targets. Also measure queue time, running and waiting requests, output throughput, and failures. Pair those with GPU memory, utilization, and KV-cache occupancy. vLLM’s serving metrics documentation describes request and cache signals that help separate queue pressure from cache pressure.

GPU utilization alone can mislead: a device may be busy yet deliver poor latency, or show low utilization while memory limits admission. Elastic’s vLLM metric-tuning example illustrates how serving and device measurements can be examined together. Account for hardware heterogeneity when comparing candidate accelerator configurations, and benchmark each before treating its capacity as a planning input.

Choose GPU placement and deployment to match the bottleneck

Separate prefill and decode only when tests support it

Prefill processes the input prompt and can demand substantial compute. Decode generates tokens sequentially and is often sensitive to memory bandwidth and KV-cache capacity. A shared pool is simpler, but a long-prompt surge may compete with ongoing generations.

Separating prefill and decode can let teams tune hardware and resource allocation for each phase. Disaggregation also adds cache-transfer traffic, routing work, and another queue. Benchmark both designs under the same request mix and SLOs; disaggregation isn’t an automatic win.

Two GPU server clusters exchange glowing tokens, with one receiving input and the other producing outputs.

Treat fractional GPUs as a measured fit

Small models or lightly loaded tenants may not need an entire accelerator. Fractional GPU allocation can help, but the memory slice must fit the weights, cache, and concurrency target.

NVIDIA Multi-Instance GPU (MIG) can create isolated GPU instances on supported hardware, while other sharing approaches have different isolation and scheduling properties. Test isolation and noisy-neighbor effects before pooling tenants. Also test whether one large model would perform better on an unpartitioned device. Partitioning can reduce stranded capacity, but it can also leave a workload short of memory.

Compare the full deployment cost

Managed model APIs shift hardware operations to a provider, but teams still need to forecast tokens, rate limits, and data-handling requirements. Cloud GPU instances offer more serving control; bare metal may suit predictable, sustained demand when the team can operate the fleet. A sovereign cloud may fit data residency or control requirements, while on-premises deployments must account for rack density, power, and cooling.

Compare useful requests per dollar at the SLO, including infrastructure costs, idle reserve, networking, storage, support, and deployment time. GPU utilization alone doesn’t show useful requests delivered at the SLO, so assess token economics against service output. For reserved cloud supply, check substitution rights and telemetry in the GPU cloud contract buyer checklist before assuming a quoted GPU-hour covers the required service.

Scale on serving pressure, not a single device gauge

Connect application demand to Kubernetes scaling

For Kubernetes deployments, export serving metrics to a monitoring system and make a suitable demand signal available to the Horizontal Pod Autoscaler or another autoscaling controller. Queue depth and waiting requests can reveal pressure before latency breaches the SLO. Pair pod scaling with node scaling, because a new replica can’t run without suitable accelerator capacity.

Google Cloud’s LLM autoscaling guidance cautions against relying on CPU, memory, or GPU utilization as standalone scaling signals for GPU inference. Preallocated GPU memory is particularly poor as a scale-down signal: traffic can fall while reported memory remains occupied.

Budget for the time it takes capacity to arrive

Reactive scaling can’t prevent every burst-related delay. Measure image pull, model load, cache warmup, cluster provisioning, and traffic-routing time. Compare that lead time with how quickly the forecast peak can arrive and whether it fits latency targets.

Keep warm replicas for latency-sensitive services where justified. For predictable peaks, schedule capacity ahead of time; for surprises, set queue limits and admission policies. Test scale-down as carefully as scale-up so terminating replicas don’t drop active generations. Multi-tenant fleets also need quotas and priorities to guide resource allocation and keep one tenant’s surge from consuming another’s reserved headroom.

Make idle capacity visible in unit economics

A low cost per GPU-hour can still make each completed request expensive when devices sit idle between bursts. Track cost per completed request and output token at the SLO. Include idle reserve, failures, and retries, which consume capacity without delivering results.

Shared fleets need an agreed method for assigning reserved headroom and shared infrastructure costs. FinOps chargeback and showback for AI offers a way to distinguish service consumption from shared overhead. For cluster-level reporting, Kubernetes cost allocation with OpenCost covers workload, idle, and shared-cost attribution.

Revisit unit costs when the request mix changes. More reasoning tokens, longer context, or a new tenant can shift token economics and increase costs, even if headline request volume stays flat.

Validate the plan before reserving capacity

Use a short evidence check at each model release and major traffic change:

  • Compare forecast request rates and token distributions with recent gateway traces, including bursts and retries.
  • Confirm the tested model, precision, serving engine, context limit, and accelerator match the proposed deployment.
  • Load-test the normal peak, growth case, and burst against latency targets for time-to-first-token, generation, and completion SLOs.
  • Verify KV-cache occupancy, waiting requests, and errors while the fleet carries the calculated concurrency.
  • Measure cluster provisioning time for pods and nodes, then test the loss of a replica during peak traffic.
  • Reconcile billed GPU, storage, and network usage with tenant-level cost reports to validate resource allocation, storage capacity planning assumptions, and reserved idle capacity.

Record the range of results, not only the best run. If demand or benchmark results vary widely, reserve headroom against a stated risk tolerance and set a date to retest.

Frequently Asked Questions

What is AI inference capacity planning?

It is the process of estimating the serving resources needed to handle expected model requests while meeting latency and availability targets. It accounts for workload mix, concurrency, memory, scaling delays, and cost—not just peak GPU throughput.

How do I estimate the number of replicas?

Divide peak requests per second by the requests per second one replica sustains at the agreed SLO, then round up. Check that the resulting fleet can also handle expected concurrency and KV-cache use, and add a tested allowance for availability or bursts.

Why isn’t GPU utilization enough to measure capacity?

Utilization does not show whether requests are waiting, latency targets are being missed, or memory limits are preventing admission. Pair GPU measurements with queue, latency, throughput, and KV-cache metrics.

What should trigger autoscaling for an inference service?

Queue depth and waiting requests can reveal serving pressure before latency breaches an SLO, so they can complement other scaling signals. Measure how long pods and accelerator nodes take to become ready, and keep warm capacity when reactive scaling would be too slow.

Conclusion

AI inference capacity planning starts with the work requests create, not a GPU’s peak specification. Usable capacity is the throughput a measured configuration can sustain while meeting latency and memory limits.

The initial replica calculation gives you a starting point. Production traces, representative benchmarks, and scaling tests tell you whether it will hold when traffic arrives.

Scroll to Top