Introduction
Selecting the right GPU for AI inference is a multi-dimensional optimization problem. The wrong choice can result in 3-5x higher cost per token or latency that fails to meet application SLAs. The right choice depends on model size, latency requirements, throughput targets, and budget constraints.
The GPU market for inference has diversified significantly. H100 dominates for large model serving, but L40S offers better cost efficiency for mid-size models, and A10G remains the most cost-effective option for small model inference. Specialized inference chips from AWS (Inferentia), Google (TPU), and others offer compelling economics for specific workloads.
This guide provides a systematic framework for hardware selection, with empirical performance data across GPU types and model sizes. The goal is to match hardware capabilities to workload requirements rather than defaulting to the most powerful (and expensive) option.
Best for 70B+ model inference
Best cost/token for 7B-70B
Best for sub-7B inference
Cost advantage of specialized chips
Selection framework
Hardware selection should follow a structured decision process. Start with model size, which determines minimum VRAM requirements. Then evaluate latency requirements, which constrain batch size and parallelism strategy. Finally, apply cost constraints to select the most efficient option within the performance envelope.
The VRAM requirement is the hard constraint. A model requires approximately 2 bytes per parameter in FP16 (e.g., 70B model = 140GB minimum). Add 20-30% overhead for KV cache and activations. This determines the minimum GPU configuration.
VRAM sizing rule
Latency requirements determine whether you can use large batches (high throughput, higher latency) or must use small batches (lower throughput, lower latency). Interactive applications typically require p99 latency under 2 seconds for first token; batch processing can tolerate much higher latency.
GPU comparison
Inference GPU comparison
| GPU | VRAM | Memory BW | FP16 TFLOPS | INT8 TOPS | TDP | Cost/hr (cloud) | Best Model Size |
|---|---|---|---|---|---|---|---|
| H100 SXM5 | 80GB HBM3 | 3.35 TB/s | 989 | 1979 | 700W | ~$3.50 | 70B+ |
| H100 PCIe | 80GB HBM3 | 2.0 TB/s | 756 | 1513 | 350W | ~$2.80 | 70B+ |
| L40S | 48GB GDDR6 | 864 GB/s | 362 | 733 | 350W | ~$1.80 | 7B-70B |
| A100 80GB | 80GB HBM2e | 2.0 TB/s | 312 | 624 | 400W | ~$2.50 | 13B-70B |
| A10G | 24GB GDDR6 | 600 GB/s | 125 | 250 | 150W | ~$0.75 | <7B |
| T4 | 16GB GDDR6 | 320 GB/s | 65 | 130 | 70W | ~$0.35 | <3B |
Memory bandwidth matters more than TFLOPS
GPU deep dive
H100 SXM5
The H100 SXM5 is the current performance leader for LLM inference. Its 80GB HBM3 memory, 3.35 TB/s memory bandwidth, and NVLink 4.0 interconnect make it optimal for large model serving (70B+). The H100's FP8 support enables near-lossless quantization with 2x throughput improvement over FP16.
L40S
The L40S is the most cost-effective option for mid-size model inference (7B-70B). At roughly 60% of H100 cost with 80% of the throughput for 7B-13B models, it offers the best cost-per-token for the most common production workloads. Its GDDR6 memory (rather than HBM) reduces cost but also reduces memory bandwidth.
A10G
The A10G is purpose-built for inference workloads. With 24GB GDDR6 and 250W TDP, it fits in standard server configurations without liquid cooling. For models under 7B parameters, A10G delivers excellent cost efficiency. It is the recommended choice for high-volume small model serving.
Specialized inference chips
Inference-specific chips (AWS Inferentia2, Google TPU v4, Qualcomm Cloud AI 100) offer 2-5x better cost efficiency than GPUs for specific model architectures. The tradeoff is reduced flexibility — these chips require model compilation and may not support all model architectures or operators.
Specialized inference chips
Specialized inference accelerators
| Chip | Provider | TOPS | Memory | Power | Cost/hr | Best Use Case |
|---|---|---|---|---|---|---|
| Inferentia2 | AWS | 380 | 32GB | 110W | ~$0.40 | Transformer inference on AWS |
| TPU v4 | 275 | 32GB HBM | 170W | ~$1.20 | JAX/TF models on GCP | |
| Gaudi2 | Intel | 432 | 96GB HBM2e | 600W | ~$1.50 | PyTorch models on Intel |
| Cloud AI 100 | Qualcomm | 400 | 32GB | 75W | On-prem | Edge and on-prem inference |
| MI300X | AMD | 1307 | 192GB HBM3 | 750W | ~$2.80 | Large model inference |
Specialized chip tradeoffs
Hardware selection architecture
Hardware Selection Decision Framework
Deployment Pattern
Single GPU, multi-GPU, or cluster
GPU Selection
Optimal hardware for requirements
Cost Constraint
Filters hardware options by budget
Throughput Requirement
Determines GPU count and configuration
Latency Requirement
Constrains batch size and parallelism
Model Size
Determines minimum VRAM requirement
Cost analysis
Cost per token is the primary economic metric for inference hardware selection. It combines hardware cost (amortized), utilization rate, and throughput. A GPU running at 80% utilization costs half as much per token as the same GPU at 40% utilization.
On-premises hardware amortizes over 3-5 years. At 80% utilization, an H100 server (8x H100, ~$350,000) costs approximately $0.30-0.80 per million tokens for LLaMA-70B inference. Cloud instances for the same workload cost $2-5 per million tokens — a 3-10x premium for flexibility.
The break-even point for on-premises vs cloud inference depends on utilization. At 60%+ utilization, on-premises is almost always cheaper. Below 30% utilization, cloud flexibility typically wins on economics.
Hardware ROI calculator
Inference Hardware Cost Calculator
Estimate monthly inference cost based on hardware selection and workload.
Estimated results
Monthly hardware cost
Monthly tokens generated
Cost per 1M tokens
Effective throughput
Annual hardware cost
vs OpenAI API ($15/M)
Deployment patterns
Hardware selection determines deployment topology. Single-GPU deployments are simplest but limited to models that fit in one GPU's VRAM. Multi-GPU configurations use tensor parallelism (splitting model layers across GPUs) or pipeline parallelism (splitting model stages across GPUs).
Tensor vs pipeline parallelism
For 70B models, the recommended configuration is 4x H100 SXM5 with NVLink (4-way tensor parallelism). This provides 320GB VRAM (ample for 70B + KV cache), 13.4 TB/s aggregate memory bandwidth, and NVLink communication that keeps tensor parallelism overhead under 5%.
Frequently asked questions
What GPU is best for LLM inference?
It depends on model size. For models 70B+, H100 SXM5 provides the best throughput. For 7B-70B models, L40S offers the best cost-per-token. For models under 7B, A10G is the most cost-effective. Always benchmark your specific model and workload before committing to hardware.
When should I use L40S instead of H100?
L40S is preferred when serving 7B-70B models where cost efficiency matters more than peak throughput. L40S costs roughly 60% of H100 while delivering 80-90% of the throughput for these model sizes. For 70B+ models or workloads requiring maximum throughput, H100 is the better choice.
What is the cost per token for different GPU options?
For LLaMA-70B at 80% utilization: H100 on-prem ~$0.40/M tokens, L40S on-prem ~$0.55/M tokens, A100 on-prem ~$0.60/M tokens. Cloud equivalents run 3-10x higher. OpenAI API for comparable models runs $10-30/M tokens — 25-75x on-premises cost at scale.
Are inference-specific chips worth it?
For high-volume, stable workloads with supported model architectures, yes. AWS Inferentia2 and Google TPU v4 offer 2-5x better cost efficiency than H100 for specific models. The tradeoff is reduced flexibility, longer deployment cycles, and potential compatibility issues with newer model architectures.