Introduction

Cost per token is the fundamental economic metric for LLM inference. It determines whether a product is economically viable, how pricing should be structured, and whether on-premises infrastructure makes sense compared to cloud APIs. Yet most organizations lack rigorous cost benchmarks for their specific workloads.

This guide presents empirical cost benchmarks across GPU platforms and cloud providers, with methodology for applying these benchmarks to your specific workload. The benchmarks cover LLaMA-7B through LLaMA-405B on H100, A100, L40S, and major cloud providers.

The key finding: on-premises H100 inference at scale costs $0.20-0.80 per million tokens, compared to $1-15 for cloud GPU instances and $10-60 for managed API services. The economics strongly favor on-premises at scale, but require sufficient volume to justify capital investment.

$0.40

On-prem H100 cost per 1M tokens (70B)

$3-8

Cloud GPU cost per 1M tokens (70B)

$30+

Managed API cost per 1M tokens (GPT-4)

75x

API vs on-prem cost ratio at scale

Benchmark methodology

Benchmarks were conducted using vLLM with continuous batching enabled, INT8 quantization (W8A8), and batch sizes optimized for maximum throughput. All measurements represent steady-state throughput, not burst performance.

Cost calculations include hardware amortization (3-year straight-line), power costs ($0.10/kWh), cooling overhead (PUE 1.3), and 10% for maintenance and operations. Cloud costs use on-demand pricing; reserved instance pricing reduces cloud costs by 30-40%.

Token counts use the standard definition: one token approximately equals 0.75 words in English. Benchmark prompts use a mix of short (200 token) and long (2000 token) inputs with 500-token average outputs.

Benchmark conditions

All benchmarks use vLLM with continuous batching, INT8 quantization, and batch sizes optimized for maximum throughput. On-premises costs assume 80% utilization, 3-year amortization, $0.10/kWh power, and PUE 1.3. Cloud costs use on-demand pricing as of January 2025.

On-premises benchmarks

On-premises H100 SXM5 achieves 2,000-8,000 tokens per second for LLaMA-7B and 200-800 tokens per second for LLaMA-70B, depending on batch size and quantization. At 80% utilization, this translates to $0.20-0.40 per million tokens for 7B models and $0.40-0.80 for 70B models.

L40S provides better cost efficiency for 7B-13B models. At $1.80/hour vs $3.50/hour for H100, and 80% of the throughput for small models, L40S achieves $0.15-0.30 per million tokens for 7B models — the lowest cost option for this model size.

A100 80GB sits between H100 and L40S in both performance and cost. For organizations with existing A100 infrastructure, it remains cost-competitive for 13B-70B models at $0.35-0.70 per million tokens.

Cloud benchmarks

Cloud GPU instances (AWS p4d, Azure NDv4, GCP A3) cost 3-8x more per token than on-premises for equivalent hardware, reflecting cloud provider margins and flexibility premium. AWS p4d (A100) costs approximately $2-4 per million tokens for LLaMA-70B.

Managed inference APIs (OpenAI, Anthropic, Google) cost $10-60 per million tokens — 25-150x on-premises cost at scale. The premium reflects model quality, reliability, and the absence of infrastructure management. For low-volume use cases, the convenience premium is justified.

The break-even point between cloud APIs and on-premises inference is typically 500M-2B tokens per month. Below this volume, cloud APIs are more economical. Above this volume, on-premises infrastructure pays back within 12-18 months.

Cost comparison table

Cost per million tokens by platform and model size

Costs in USD per million tokens. On-premises costs include amortization, power, and operations. Cloud costs are on-demand; reserved pricing reduces by 30-40%.
PlatformLLaMA-7BLLaMA-70BLLaMA-405BGPT-4 classNotes
On-Prem H100 SXM5$0.15-0.25$0.40-0.80$1.50-3.00N/A80% util, 3yr amort
On-Prem A100 80GB$0.20-0.35$0.50-0.90$2.00-3.50N/A80% util, 3yr amort
On-Prem L40S$0.12-0.22$0.55-1.00N/AN/ABest for 7B-13B
AWS p4d (A100)$1.00-1.80$2.50-4.50$8-15N/AOn-demand pricing
Azure NDv4 (A100)$1.20-2.00$3.00-5.00$9-16N/AOn-demand pricing
GCP A3 (H100)$1.50-2.50$3.50-6.00$10-18N/AOn-demand pricing
OpenAI API$0.10-0.30$0.90-1.80N/A$10-30GPT-3.5 to GPT-4
Anthropic APIN/A$0.80-2.40N/A$15-75Claude Haiku to Opus

Cost benchmark framework

Cost Benchmark Framework

TCO Analysis

3-year total cost including ops, power, facilities

3yr TCOBreak-evenNPV

vs Cloud API

Comparison against managed API pricing

OpenAIAnthropicGoogle

Cost per Token

Hardware cost / (throughput × utilization × time)

$/M tokensvs Cloud APIvs Competitors

Throughput

Tokens per second per GPU at target utilization

tok/s/GPUBatch EfficiencyQuantization

Utilization Rate

Effective GPU utilization determines cost efficiency

Batch SizeQueue DepthUtilization %

Hardware Cost

GPU purchase price amortized over 3 years

CapEx3yr AmortDepreciation
Stack layers — top to bottom: highest to lowest abstraction

On-premises ROI calculator

On-Premises vs API Cost Calculator

Calculate the break-even point and 3-year savings for on-premises inference vs cloud API.

1,000 M tokens
1100,000
15 $
160
8 GPUs
1100
25,000 $
5,000500,000

Estimated results

$15,000

Monthly API cost

$25,000

Monthly on-prem cost

$-10,000

Monthly savings

$-120,000

Annual savings

320000.0 months

Break-even

$-680,000

3-year NPV

Break-even analysis

The break-even point for on-premises inference depends on three variables: API cost per token, on-premises infrastructure cost, and monthly token volume. Higher API costs and higher token volumes accelerate break-even.

When on-premises wins

On-premises inference becomes economically superior when monthly token volume exceeds 500M tokens at $15/M API pricing, or 2B tokens at $5/M API pricing. At these volumes, on-premises typically breaks even within 12-18 months and generates 3-5x ROI over 3 years.

Organizations should also factor in non-economic benefits of on-premises inference: data privacy (no tokens sent to third-party APIs), latency (no internet round-trip), customization (fine-tuned models), and independence from API pricing changes.

Optimization impact on cost

Cost reduction from optimization techniques

OptimizationCost ReductionImplementation EffortQuality ImpactRecommended
Continuous batching60-75%Low (use vLLM)NoneAlways
INT8 quantization40-50%Low<1%Always
INT4 quantization60-70%Low1-3%When memory-constrained
Speculative decoding30-50% latencyMediumNoneLatency-sensitive apps
Prefix caching20-40%LowNoneShared system prompts
Flash Attention 215-25%Low (built-in)NoneAlways

Frequently asked questions

How much does LLM inference cost on H100?

On-premises H100 SXM5 inference costs approximately $0.20-0.40 per million tokens for 7B models and $0.40-0.80 per million tokens for 70B models at 80% GPU utilization. These costs include hardware amortization, power, cooling, and operations. Cloud H100 instances cost 3-5x more due to cloud provider margins.

How does on-premises compare to cloud API costs?

On-premises inference costs $0.20-0.80 per million tokens at scale. OpenAI API costs $10-60 per million tokens for comparable models. This represents a 25-150x cost difference at scale. The on-premises advantage grows with volume — the capital cost is fixed while the per-token cost decreases with utilization.

At what scale does on-premises inference become cheaper?

The break-even point depends on the specific model and cloud API pricing. For GPT-4-class models at $30/M tokens, on-premises becomes cheaper at approximately 200-500M tokens per month. For cheaper APIs ($5/M tokens), the break-even is 1-2B tokens per month. Always model your specific workload and growth trajectory.

How do you calculate cost per token?

Cost per token = (monthly infrastructure cost) / (monthly tokens generated). Monthly infrastructure cost includes hardware amortization (purchase price / 36 months), power (GPU TDP × hours × $/kWh × PUE), and operations (10-15% of hardware cost annually). Monthly tokens = average throughput (tokens/second) × utilization rate × seconds per month.