Introduction
Cost per token is the fundamental economic metric for LLM inference. It determines whether a product is economically viable, how pricing should be structured, and whether on-premises infrastructure makes sense compared to cloud APIs. Yet most organizations lack rigorous cost benchmarks for their specific workloads.
This guide presents empirical cost benchmarks across GPU platforms and cloud providers, with methodology for applying these benchmarks to your specific workload. The benchmarks cover LLaMA-7B through LLaMA-405B on H100, A100, L40S, and major cloud providers.
The key finding: on-premises H100 inference at scale costs $0.20-0.80 per million tokens, compared to $1-15 for cloud GPU instances and $10-60 for managed API services. The economics strongly favor on-premises at scale, but require sufficient volume to justify capital investment.
On-prem H100 cost per 1M tokens (70B)
Cloud GPU cost per 1M tokens (70B)
Managed API cost per 1M tokens (GPT-4)
API vs on-prem cost ratio at scale
Benchmark methodology
Benchmarks were conducted using vLLM with continuous batching enabled, INT8 quantization (W8A8), and batch sizes optimized for maximum throughput. All measurements represent steady-state throughput, not burst performance.
Cost calculations include hardware amortization (3-year straight-line), power costs ($0.10/kWh), cooling overhead (PUE 1.3), and 10% for maintenance and operations. Cloud costs use on-demand pricing; reserved instance pricing reduces cloud costs by 30-40%.
Token counts use the standard definition: one token approximately equals 0.75 words in English. Benchmark prompts use a mix of short (200 token) and long (2000 token) inputs with 500-token average outputs.
Benchmark conditions
On-premises benchmarks
On-premises H100 SXM5 achieves 2,000-8,000 tokens per second for LLaMA-7B and 200-800 tokens per second for LLaMA-70B, depending on batch size and quantization. At 80% utilization, this translates to $0.20-0.40 per million tokens for 7B models and $0.40-0.80 for 70B models.
L40S provides better cost efficiency for 7B-13B models. At $1.80/hour vs $3.50/hour for H100, and 80% of the throughput for small models, L40S achieves $0.15-0.30 per million tokens for 7B models — the lowest cost option for this model size.
A100 80GB sits between H100 and L40S in both performance and cost. For organizations with existing A100 infrastructure, it remains cost-competitive for 13B-70B models at $0.35-0.70 per million tokens.
Cloud benchmarks
Cloud GPU instances (AWS p4d, Azure NDv4, GCP A3) cost 3-8x more per token than on-premises for equivalent hardware, reflecting cloud provider margins and flexibility premium. AWS p4d (A100) costs approximately $2-4 per million tokens for LLaMA-70B.
Managed inference APIs (OpenAI, Anthropic, Google) cost $10-60 per million tokens — 25-150x on-premises cost at scale. The premium reflects model quality, reliability, and the absence of infrastructure management. For low-volume use cases, the convenience premium is justified.
The break-even point between cloud APIs and on-premises inference is typically 500M-2B tokens per month. Below this volume, cloud APIs are more economical. Above this volume, on-premises infrastructure pays back within 12-18 months.
Cost comparison table
Cost per million tokens by platform and model size
| Platform | LLaMA-7B | LLaMA-70B | LLaMA-405B | GPT-4 class | Notes |
|---|---|---|---|---|---|
| On-Prem H100 SXM5 | $0.15-0.25 | $0.40-0.80 | $1.50-3.00 | N/A | 80% util, 3yr amort |
| On-Prem A100 80GB | $0.20-0.35 | $0.50-0.90 | $2.00-3.50 | N/A | 80% util, 3yr amort |
| On-Prem L40S | $0.12-0.22 | $0.55-1.00 | N/A | N/A | Best for 7B-13B |
| AWS p4d (A100) | $1.00-1.80 | $2.50-4.50 | $8-15 | N/A | On-demand pricing |
| Azure NDv4 (A100) | $1.20-2.00 | $3.00-5.00 | $9-16 | N/A | On-demand pricing |
| GCP A3 (H100) | $1.50-2.50 | $3.50-6.00 | $10-18 | N/A | On-demand pricing |
| OpenAI API | $0.10-0.30 | $0.90-1.80 | N/A | $10-30 | GPT-3.5 to GPT-4 |
| Anthropic API | N/A | $0.80-2.40 | N/A | $15-75 | Claude Haiku to Opus |
Cost benchmark framework
Cost Benchmark Framework
TCO Analysis
3-year total cost including ops, power, facilities
vs Cloud API
Comparison against managed API pricing
Cost per Token
Hardware cost / (throughput × utilization × time)
Throughput
Tokens per second per GPU at target utilization
Utilization Rate
Effective GPU utilization determines cost efficiency
Hardware Cost
GPU purchase price amortized over 3 years
On-premises ROI calculator
On-Premises vs API Cost Calculator
Calculate the break-even point and 3-year savings for on-premises inference vs cloud API.
Estimated results
Monthly API cost
Monthly on-prem cost
Monthly savings
Annual savings
Break-even
3-year NPV
Break-even analysis
The break-even point for on-premises inference depends on three variables: API cost per token, on-premises infrastructure cost, and monthly token volume. Higher API costs and higher token volumes accelerate break-even.
When on-premises wins
Organizations should also factor in non-economic benefits of on-premises inference: data privacy (no tokens sent to third-party APIs), latency (no internet round-trip), customization (fine-tuned models), and independence from API pricing changes.
Optimization impact on cost
Cost reduction from optimization techniques
| Optimization | Cost Reduction | Implementation Effort | Quality Impact | Recommended |
|---|---|---|---|---|
| Continuous batching | 60-75% | Low (use vLLM) | None | Always |
| INT8 quantization | 40-50% | Low | <1% | Always |
| INT4 quantization | 60-70% | Low | 1-3% | When memory-constrained |
| Speculative decoding | 30-50% latency | Medium | None | Latency-sensitive apps |
| Prefix caching | 20-40% | Low | None | Shared system prompts |
| Flash Attention 2 | 15-25% | Low (built-in) | None | Always |
Frequently asked questions
How much does LLM inference cost on H100?
On-premises H100 SXM5 inference costs approximately $0.20-0.40 per million tokens for 7B models and $0.40-0.80 per million tokens for 70B models at 80% GPU utilization. These costs include hardware amortization, power, cooling, and operations. Cloud H100 instances cost 3-5x more due to cloud provider margins.
How does on-premises compare to cloud API costs?
On-premises inference costs $0.20-0.80 per million tokens at scale. OpenAI API costs $10-60 per million tokens for comparable models. This represents a 25-150x cost difference at scale. The on-premises advantage grows with volume — the capital cost is fixed while the per-token cost decreases with utilization.
At what scale does on-premises inference become cheaper?
The break-even point depends on the specific model and cloud API pricing. For GPT-4-class models at $30/M tokens, on-premises becomes cheaper at approximately 200-500M tokens per month. For cheaper APIs ($5/M tokens), the break-even is 1-2B tokens per month. Always model your specific workload and growth trajectory.
How do you calculate cost per token?
Cost per token = (monthly infrastructure cost) / (monthly tokens generated). Monthly infrastructure cost includes hardware amortization (purchase price / 36 months), power (GPU TDP × hours × $/kWh × PUE), and operations (10-15% of hardware cost annually). Monthly tokens = average throughput (tokens/second) × utilization rate × seconds per month.