GPU pilot & procurement

Inference GPU sizing: test latency and throughput before comparing cost

Build a representative inference test with model, precision, context, concurrency and quality targets. Connect the measured result to a transparent GPU budget.

· Arvica Cloud · Analysis & buying guide

Memory fit is the first gate

Record the exact model revision and serving framework, then estimate weights at the intended precision. Leave room for KV cache and runtime overhead; parameter count alone is not the required GPU memory. Validate the intended prompt and output lengths on the actual configuration before using any cost estimate as a purchasing decision.

Use representative request traffic

Create a small, permitted test set covering short and long inputs, expected output lengths and the intended request mix. Record both request arrival rate and concurrent in-flight requests; they are different quantities. Increase load in explicit stages and record completed requests, failures and queueing. Keep model quality checks identical across configurations.

Define the latency measurements

Measure time to first token, streaming gaps and full response time at the client location relevant to users. vLLM distinguishes time to first token, inter-token latency and time per output token; tools may aggregate these differently. Write down the measurement points and use the same method for each comparison. [1]

Control cache and warm-up conditions

Record whether the test starts cold or after warm-up and whether prefix caching is intended. vLLM warns that repeated prompts can reuse prefix cache and inflate throughput. Compare a defined cache condition, save the random seed or test-set revision and report any cache reuse. Do not present a warm repeated-prompt test as unseen production traffic. [1]

Compare the cost of accepted service

Choose a configuration that meets the agreed quality, error and tail-latency criteria at the required load. Calculate cost using actual billable runtime and complete charges, then divide by accepted requests or an explicitly defined token volume. Do not compare an output-only denominator with combined input-and-output tokens. Our calculator is a planning tool; it does not prove that a selected configuration can serve the assumed traffic.

A planning example, not a benchmark

Suppose a quoted configuration costs A$6 per billable hour and runs 100 hours. GPU charges are A$600. If that run completes 200,000 requests meeting the agreed criteria, the GPU-only cost is A$3 per 1,000 accepted requests, before other charges. Every number in this example is hypothetical. If those requests miss the target, the low unit cost is not an accepted result.

Sources

Reviewed: 2026-09-13

  1. vLLM benchmark CLI and latency definitions ↗

Apply this to your workload

See the workload solution →

No charge before quote acceptance; the infrastructure provider is disclosed before activation.

Buying guides and pilot worksheet