Skip to main content
DCS Global

AI Training Infrastructure Guides

Quick Reference

AI Training Infrastructure — Quick Reference

Definitions, specifications, and decision frameworks for AI model training infrastructure.

Definition: AI Training Infrastructure

AI training infrastructure is the compute, storage, and networking systems optimised for running machine learning training workloads — the process of adjusting model weights by repeatedly processing training data through a neural network. Training is compute-intensive (days to weeks of continuous GPU operation), memory-intensive (model weights + gradients + optimizer states must fit in GPU memory), and I/O-intensive (training data must be streamed to GPUs faster than they can consume it).

  • ▸Compute: GPU clusters with NVLink (within nodes) and InfiniBand (between nodes) for distributed training.
  • ▸Storage: parallel file systems (WEKA, GPFS, Lustre) capable of 200+ GB/s aggregate throughput to feed GPUs.
  • ▸Networking: InfiniBand NDR (400 Gbps) or 400G Ethernet with RoCE v2 for all-reduce gradient synchronisation.
  • ▸Memory: GPU HBM memory for model weights + gradients + optimizer states; CPU RAM for data preprocessing.

Training Infrastructure: On-Premises vs. Cloud

FactorOn-premisesCloud (AWS/Azure/GCP)
Cost at 40%+ utilisationLower 3-year TCOHigher at sustained use
Cost at <40% utilisationHigher (fixed capital)Lower (pay-per-use)
Data sovereigntyComplete controlData leaves perimeter
CustomisationFull hardware/software controlLimited to provider SKUs
Deployment time3–18 monthsHours to days
Spot/preemptible pricingN/AAvailable (with interruption risk)
Interconnect performanceInfiniBand NDR (400 Gbps)EFA/InfiniBand (varies by instance)

AI Training Infrastructure Sizing Rules of Thumb

  • GPU memory rule: model parameters × 2 bytes (FP16) × 4 (weights + gradients + 2× optimizer states) = minimum GPU memory for full fine-tuning.
  • Storage throughput rule: GPU cluster should never be I/O-bound; target 1 GB/s of storage throughput per GPU.
  • Network bandwidth rule: all-reduce bandwidth should be at least 10% of GPU compute bandwidth to avoid communication bottlenecks.
  • Power rule: 10.2 kW per H100 SXM5 server; add 30–40% for networking, storage, and cooling overhead.
  • Checkpoint storage rule: save checkpoints every 500–1,000 steps; each checkpoint = model size × 4 (FP32) or × 2 (BF16).
  • Utilisation target: aim for 80%+ MFU (Model FLOP Utilisation) on well-optimised training runs; below 50% indicates a bottleneck.
6 Articles

AI Training Infrastructure

Training infrastructure, frameworks, data pipelines, and MLOps. Complete technical guides for building the infrastructure that trains large language models and other AI systems.

Reference

Key Concepts

Distributed Training

Data, tensor, and pipeline parallelism — the strategies for distributing training across hundreds or thousands of GPUs efficiently.

Checkpointing

Saving model state during training — frequency trade-offs, storage requirements, and fast checkpoint libraries like FSDP and Megatron.

MLOps

Experiment tracking, model registry, pipeline orchestration, and the operational practices that make training reproducible and manageable.

Cost Optimization

Mixed precision, gradient checkpointing, spot instances, and the infrastructure strategies that reduce training cost per model.

Ready to Build Your AI Infrastructure?

Our certified engineers design and deploy enterprise AI infrastructure — from single GPU servers to 1,000+ GPU clusters.