Skip to main content
DCS Global

Generative AI Infrastructure Guides

Quick Reference

Generative AI Infrastructure — Quick Reference

Definitions, comparisons, and infrastructure requirements for generative AI at enterprise scale.

Definition: Generative AI Infrastructure

Generative AI infrastructure is the compute, storage, networking, and software stack required to train, fine-tune, and serve generative AI models — including large language models (LLMs), diffusion models, and multimodal models. It differs from standard AI infrastructure in its extreme memory requirements (frontier LLMs require terabytes of GPU memory), high-throughput storage needs (training datasets of 1–100+ TB), and serving complexity (managing thousands of concurrent inference requests).

  • ▸LLM training: requires GPU clusters with NVLink/InfiniBand, parallel storage at 200+ GB/s, and weeks of continuous compute.
  • ▸LLM inference: requires GPU memory sufficient to hold model weights (70B model ≈ 140 GB in FP16), plus KV cache for context.
  • ▸Fine-tuning: smaller GPU footprint than pre-training; LoRA/QLoRA techniques reduce memory requirements significantly.
  • ▸RAG (Retrieval-Augmented Generation): requires vector database infrastructure alongside the LLM serving stack.

Generative AI Model Infrastructure Requirements

Representative values for FP16 precision. BF16 and quantisation reduce memory requirements. Actual needs vary by architecture.
Model sizeGPU memory (FP16)Min. GPUs (H100)Training GPUsInference GPUs
7B parameters14 GB1× H100 (80 GB)8–16 H1001 H100
13B parameters26 GB1× H100 (80 GB)16–32 H1001 H100
70B parameters140 GB2× H100 (80 GB)64–128 H1002–4 H100
405B parameters810 GB11× H100 (80 GB)512+ H1008–16 H100
1T+ parameters (MoE)2+ TB25+ H100 (80 GB)1,024+ H10032+ H100

Generative AI Infrastructure Stack — Layer by Layer

  1. 1.Compute layer: GPU clusters (NVIDIA H100/H200 SXM5 for training; L40S or A10G for inference) with NVLink within nodes and InfiniBand between nodes.
  2. 2.Storage layer: parallel file system (WEKA, GPFS, Lustre) for training data; NVMe-oF or local NVMe for checkpoints; object storage for model artifacts and datasets.
  3. 3.Network layer: InfiniBand NDR (400 Gbps) or 400G Ethernet with RoCE v2 for training; 25–100G Ethernet for inference serving.
  4. 4.Serving layer: vLLM, TensorRT-LLM, or Triton Inference Server for LLM serving; continuous batching for throughput optimisation.
  5. 5.Orchestration layer: Kubernetes with GPU operator for container scheduling; Slurm for HPC-style training jobs.
  6. 6.MLOps layer: MLflow or Weights & Biases for experiment tracking; model registry; CI/CD for model deployment; monitoring and drift detection.
6 Articles

Generative AI Infrastructure

LLMs, diffusion models, RAG, and the infrastructure for generative AI at scale. From fine-tuning to production serving — complete technical guides for every layer of the generative AI stack.

Reference

Key Concepts

LLM Serving

The infrastructure and software stack for serving large language models at production scale — vLLM, TensorRT-LLM, and serving frameworks.

Fine-Tuning Methods

LoRA, QLoRA, full fine-tuning, and the infrastructure requirements for each approach to adapting foundation models.

RAG Systems

Retrieval-augmented generation architecture — vector stores, embedding pipelines, and the infrastructure for grounding LLM outputs in enterprise data.

Cost Optimization

Quantization, batching, caching, and the infrastructure strategies that reduce generative AI cost per token or per generation.

Ready to Build Your AI Infrastructure?

Our certified engineers design and deploy enterprise AI infrastructure — from single GPU servers to 1,000+ GPU clusters.