Why GPU Networking Is Different

GPU cluster networking is fundamentally different from conventional data center networking. In a conventional server cluster, network traffic is primarily client-server (north-south) — servers communicate with clients and storage. In a GPU training cluster, the dominant traffic pattern is all-to-all (east-west) — every GPU communicates with every other GPU during gradient synchronization. This all-reduce communication pattern requires high bandwidth, low latency, and RDMA (Remote Direct Memory Access) to minimize CPU overhead.

400 Gb/s

InfiniBand NDR

<1 µs

IB Latency

20–40%

All-Reduce Impact

1–5 µs

RoCEv2 Latency

InfiniBand for AI Clusters

InfiniBand is the dominant interconnect for large-scale AI training clusters. It provides native RDMA (bypassing the CPU for data transfers), sub-microsecond latency, and the highest available bandwidth. NVIDIA's acquisition of Mellanox (the primary InfiniBand vendor) has tightly integrated InfiniBand with NVIDIA GPU infrastructure.

HDR InfiniBand (200 Gb/s)
Current widely-deployed generation. 200 Gb/s per port. Sub-microsecond latency. Supported by NVIDIA ConnectX-6 and Quantum-2 switches. Standard for H100 and A100 deployments. DGX H100 includes 8× HDR200 ports (one per GPU).
NDR InfiniBand (400 Gb/s)
Latest generation. 400 Gb/s per port — 2× HDR bandwidth. NVIDIA Quantum-3 switches. Required for frontier model training (100B+ parameters) where HDR bandwidth becomes a bottleneck. DGX H200 and next-generation systems use NDR.
NVIDIA Quantum-X800 (800 Gb/s)
Next-generation InfiniBand announced for 2025–2026 deployments. 800 Gb/s per port. Targets exascale AI training clusters. Not yet widely deployed.

Ethernet for AI Workloads

Ethernet is the dominant networking technology for inference clusters and is increasingly used for training clusters where InfiniBand's cost premium is not justified. Modern 400GbE with RoCEv2 provides near-InfiniBand performance at lower cost, with the advantage of a broader ecosystem and simpler integration with existing network infrastructure.

Inference Clusters
Standard 100GbE or 400GbE Ethernet is sufficient for most inference workloads. Inference does not require all-reduce communication — GPUs operate independently. Network requirements are primarily for loading model weights from storage and serving inference requests. Standard Ethernet with ECMP load balancing is appropriate.
Training Clusters (Ethernet)
For training clusters where InfiniBand cost is prohibitive, 400GbE with RoCEv2 is a viable alternative. Requires careful network configuration (ECN, PFC, DCQCN) to achieve low latency and prevent congestion. Performance is 10–20% lower than equivalent InfiniBand for large all-reduce operations.

RoCE (RDMA over Converged Ethernet)

RoCEv2 enables RDMA over standard Ethernet infrastructure, providing near-InfiniBand performance at lower cost. However, RoCE requires careful network configuration to achieve low latency — standard Ethernet congestion control mechanisms are insufficient for RDMA traffic.

Priority Flow Control (PFC)
PFC prevents packet drops on RDMA traffic by pausing transmission when buffers fill. Required for RoCEv2 — packet drops cause RDMA retransmissions that severely degrade performance. Configure PFC on all switches in the RoCE fabric.
Explicit Congestion Notification (ECN)
ECN signals congestion before packet drops occur, allowing senders to reduce transmission rate proactively. Combined with DCQCN (Data Center Quantized Congestion Notification), ECN provides effective congestion control for RoCE traffic without PFC pauses.
Lossless Network Configuration
RoCE requires a lossless network — no packet drops on RDMA traffic. Requires sufficient buffer capacity at all switches, proper QoS configuration to prioritize RDMA traffic, and careful traffic engineering to prevent congestion. Significantly more complex to configure than standard Ethernet.

Network Topology for GPU Clusters

Fat-Tree (Clos Network)
The standard topology for GPU training clusters. Provides full bisection bandwidth — any two nodes can communicate at full link speed simultaneously. Scales to thousands of nodes. Requires 2× as many switches as a simpler topology but eliminates network bottlenecks. Used in all major AI supercomputers.
Rail-Optimized (DGX SuperPOD)
NVIDIA's recommended topology for DGX clusters. Each GPU in a DGX server connects to a dedicated leaf switch (one switch per GPU rail). Optimizes for the all-reduce communication pattern in tensor-parallel training. More efficient than fat-tree for NVIDIA GPU clusters.
Dragonfly+
Used in some large-scale HPC and AI clusters. Lower cost than fat-tree for very large clusters (1,000+ nodes) but with higher latency for some traffic patterns. Less common than fat-tree for AI workloads.

InfiniBand vs. Ethernet Comparison

GPU Cluster Networking Technology Comparison

TechnologyBandwidthLatencyRDMACostEcosystemBest For
InfiniBand HDR (200Gb)200 Gb/s per port<1 µsNativeHighNVIDIA/MellanoxLarge-scale training clusters
InfiniBand NDR (400Gb)400 Gb/s per port<1 µsNativeVery highNVIDIA/MellanoxHyperscale training, frontier models
RoCEv2 (100GbE)100 Gb/s per port1–5 µsYes (with ECN)MediumBroadMid-scale training, inference clusters
RoCEv2 (400GbE)400 Gb/s per port1–5 µsYes (with ECN)HighBroadLarge-scale training alternative to IB
Standard Ethernet (100GbE)100 Gb/s per port10–100 µsNoLow-MediumUniversalInference, management, storage
NVLink (intra-server)900 GB/s per GPU<1 µsN/A (GPU-GPU)Included in SXMNVIDIA onlyIntra-server GPU communication

Frequently Asked Questions

Do I need InfiniBand for my GPU cluster?

It depends on your workload. For large-scale training (16+ GPUs, large models): InfiniBand HDR or NDR is strongly recommended. The all-reduce communication overhead with standard Ethernet can reduce training throughput by 20–40% compared to InfiniBand. For inference clusters: standard Ethernet (100GbE or 400GbE) is sufficient — inference workloads do not require all-reduce communication. For small training clusters (8 GPUs within a single DGX): NVLink handles intra-server communication; InfiniBand is only needed for inter-server communication. For mid-scale training (16–64 GPUs): RoCEv2 with 400GbE is a cost-effective alternative to InfiniBand with 10–20% performance penalty.

What is the cost difference between InfiniBand and Ethernet for a GPU cluster?

InfiniBand infrastructure costs 2–3× more than equivalent Ethernet. For a 16-node DGX H100 cluster: InfiniBand HDR fabric (switches + cables + NICs): ~$200K–$400K. Equivalent 400GbE Ethernet fabric: ~$80K–$150K. The performance premium of InfiniBand (10–20% higher training throughput) must be weighed against the cost premium. For a $5M GPU cluster, the $200K networking premium is ~4% of total cost — often justified. For smaller deployments, RoCEv2 Ethernet may be more cost-effective.

How many network ports does each GPU server need?

DGX H100: 8× HDR200 InfiniBand ports (one per GPU) + 2× 10GbE management ports + 2× 100GbE storage/data ports. HGX H100 8-GPU: similar configuration. For inference servers with PCIe GPUs: 2× 100GbE or 400GbE for inference traffic + 1× 10GbE management. The key principle: each GPU should have its own dedicated network port for training workloads — sharing ports between GPUs creates network bottlenecks during all-reduce operations.