Introduction

Selecting the right GPU for AI inference is a multi-dimensional optimization problem. The wrong choice can result in 3-5x higher cost per token or latency that fails to meet application SLAs. The right choice depends on model size, latency requirements, throughput targets, and budget constraints.

The GPU market for inference has diversified significantly. H100 dominates for large model serving, but L40S offers better cost efficiency for mid-size models, and A10G remains the most cost-effective option for small model inference. Specialized inference chips from AWS (Inferentia), Google (TPU), and others offer compelling economics for specific workloads.

This guide provides a systematic framework for hardware selection, with empirical performance data across GPU types and model sizes. The goal is to match hardware capabilities to workload requirements rather than defaulting to the most powerful (and expensive) option.

H100

Best for 70B+ model inference

L40S

Best cost/token for 7B-70B

A10G

Best for sub-7B inference

2-5x

Cost advantage of specialized chips

Selection framework

Hardware selection should follow a structured decision process. Start with model size, which determines minimum VRAM requirements. Then evaluate latency requirements, which constrain batch size and parallelism strategy. Finally, apply cost constraints to select the most efficient option within the performance envelope.

The VRAM requirement is the hard constraint. A model requires approximately 2 bytes per parameter in FP16 (e.g., 70B model = 140GB minimum). Add 20-30% overhead for KV cache and activations. This determines the minimum GPU configuration.

VRAM sizing rule

Required VRAM = (model parameters × 2 bytes for FP16) × 1.25 overhead factor. A 70B model needs 70B × 2 × 1.25 = 175GB minimum. This requires 3x H100 80GB (240GB) or 4x A100 40GB (160GB — tight). Always add buffer for KV cache growth.

Latency requirements determine whether you can use large batches (high throughput, higher latency) or must use small batches (lower throughput, lower latency). Interactive applications typically require p99 latency under 2 seconds for first token; batch processing can tolerate much higher latency.

GPU comparison

Inference GPU comparison

GPUVRAMMemory BWFP16 TFLOPSINT8 TOPSTDPCost/hr (cloud)Best Model Size
H100 SXM580GB HBM33.35 TB/s9891979700W~$3.5070B+
H100 PCIe80GB HBM32.0 TB/s7561513350W~$2.8070B+
L40S48GB GDDR6864 GB/s362733350W~$1.807B-70B
A100 80GB80GB HBM2e2.0 TB/s312624400W~$2.5013B-70B
A10G24GB GDDR6600 GB/s125250150W~$0.75<7B
T416GB GDDR6320 GB/s6513070W~$0.35<3B

Memory bandwidth matters more than TFLOPS

LLM inference is memory-bandwidth bound, not compute bound. A GPU with higher memory bandwidth will outperform a GPU with higher TFLOPS for most inference workloads. This is why H100 SXM5 (3.35 TB/s) significantly outperforms H100 PCIe (2.0 TB/s) despite similar compute specs.

GPU deep dive

H100 SXM5

The H100 SXM5 is the current performance leader for LLM inference. Its 80GB HBM3 memory, 3.35 TB/s memory bandwidth, and NVLink 4.0 interconnect make it optimal for large model serving (70B+). The H100's FP8 support enables near-lossless quantization with 2x throughput improvement over FP16.

L40S

The L40S is the most cost-effective option for mid-size model inference (7B-70B). At roughly 60% of H100 cost with 80% of the throughput for 7B-13B models, it offers the best cost-per-token for the most common production workloads. Its GDDR6 memory (rather than HBM) reduces cost but also reduces memory bandwidth.

A10G

The A10G is purpose-built for inference workloads. With 24GB GDDR6 and 250W TDP, it fits in standard server configurations without liquid cooling. For models under 7B parameters, A10G delivers excellent cost efficiency. It is the recommended choice for high-volume small model serving.

Specialized inference chips

Inference-specific chips (AWS Inferentia2, Google TPU v4, Qualcomm Cloud AI 100) offer 2-5x better cost efficiency than GPUs for specific model architectures. The tradeoff is reduced flexibility — these chips require model compilation and may not support all model architectures or operators.

Specialized inference chips

Specialized inference accelerators

ChipProviderTOPSMemoryPowerCost/hrBest Use Case
Inferentia2AWS38032GB110W~$0.40Transformer inference on AWS
TPU v4Google27532GB HBM170W~$1.20JAX/TF models on GCP
Gaudi2Intel43296GB HBM2e600W~$1.50PyTorch models on Intel
Cloud AI 100Qualcomm40032GB75WOn-premEdge and on-prem inference
MI300XAMD1307192GB HBM3750W~$2.80Large model inference

Specialized chip tradeoffs

Specialized chips require model compilation and may not support all PyTorch operators. Evaluate compatibility with your model architecture before committing. Compilation time can be 30-60 minutes per model version, which impacts deployment velocity.

Hardware selection architecture

Hardware Selection Decision Framework

Deployment Pattern

Single GPU, multi-GPU, or cluster

Single GPUTensor ParallelPipeline ParallelCluster

GPU Selection

Optimal hardware for requirements

A10GL40SH100 PCIeH100 SXM5

Cost Constraint

Filters hardware options by budget

Dev/TestProductionEnterpriseHyperscale

Throughput Requirement

Determines GPU count and configuration

<1K tok/s1K-10K10K-100K100K+

Latency Requirement

Constrains batch size and parallelism

<100ms100-500ms500ms-2sBatch

Model Size

Determines minimum VRAM requirement

<7B7B-70B70B-405B405B+
Stack layers — top to bottom: highest to lowest abstraction

Cost analysis

Cost per token is the primary economic metric for inference hardware selection. It combines hardware cost (amortized), utilization rate, and throughput. A GPU running at 80% utilization costs half as much per token as the same GPU at 40% utilization.

On-premises hardware amortizes over 3-5 years. At 80% utilization, an H100 server (8x H100, ~$350,000) costs approximately $0.30-0.80 per million tokens for LLaMA-70B inference. Cloud instances for the same workload cost $2-5 per million tokens — a 3-10x premium for flexibility.

The break-even point for on-premises vs cloud inference depends on utilization. At 60%+ utilization, on-premises is almost always cheaper. Below 30% utilization, cloud flexibility typically wins on economics.

Hardware ROI calculator

Inference Hardware Cost Calculator

Estimate monthly inference cost based on hardware selection and workload.

8 GPUs
1256
2 $/hr
0.354
70 %
10100
500 tok/s
505,000

Estimated results

$11,520

Monthly hardware cost

7B

Monthly tokens generated

$1.59

Cost per 1M tokens

2,800 tok/s

Effective throughput

$138,240

Annual hardware cost

9.5x cheaper

vs OpenAI API ($15/M)

Deployment patterns

Hardware selection determines deployment topology. Single-GPU deployments are simplest but limited to models that fit in one GPU's VRAM. Multi-GPU configurations use tensor parallelism (splitting model layers across GPUs) or pipeline parallelism (splitting model stages across GPUs).

Tensor vs pipeline parallelism

Tensor parallelism splits individual layers across GPUs — all GPUs work on every token, reducing latency. Pipeline parallelism assigns different layers to different GPUs — latency increases but memory scales linearly. For inference, tensor parallelism is preferred when GPUs are connected via NVLink; pipeline parallelism is used for cross-node scaling.

For 70B models, the recommended configuration is 4x H100 SXM5 with NVLink (4-way tensor parallelism). This provides 320GB VRAM (ample for 70B + KV cache), 13.4 TB/s aggregate memory bandwidth, and NVLink communication that keeps tensor parallelism overhead under 5%.

Frequently asked questions

What GPU is best for LLM inference?

It depends on model size. For models 70B+, H100 SXM5 provides the best throughput. For 7B-70B models, L40S offers the best cost-per-token. For models under 7B, A10G is the most cost-effective. Always benchmark your specific model and workload before committing to hardware.

When should I use L40S instead of H100?

L40S is preferred when serving 7B-70B models where cost efficiency matters more than peak throughput. L40S costs roughly 60% of H100 while delivering 80-90% of the throughput for these model sizes. For 70B+ models or workloads requiring maximum throughput, H100 is the better choice.

What is the cost per token for different GPU options?

For LLaMA-70B at 80% utilization: H100 on-prem ~$0.40/M tokens, L40S on-prem ~$0.55/M tokens, A100 on-prem ~$0.60/M tokens. Cloud equivalents run 3-10x higher. OpenAI API for comparable models runs $10-30/M tokens — 25-75x on-premises cost at scale.

Are inference-specific chips worth it?

For high-volume, stable workloads with supported model architectures, yes. AWS Inferentia2 and Google TPU v4 offer 2-5x better cost efficiency than H100 for specific models. The tradeoff is reduced flexibility, longer deployment cycles, and potential compatibility issues with newer model architectures.