11 Resources in This Cluster
Frequently Asked Questions
What network bandwidth does AI training require?
AI training bandwidth requirements depend on model size and cluster scale. An 8-GPU H100 cluster requires 3.2 Tbps of bisectional bandwidth (400 Gbps per GPU). A 64-GPU cluster requires 25.6 Tbps. InfiniBand NDR (400 Gbps) or 400G Ethernet with RDMA (RoCE v2) are the standard fabrics. Insufficient bandwidth causes GPU idle time during all-reduce operations, reducing training throughput by 30–70%.
What is spine-leaf architecture and why is it used in data centers?
Spine-leaf is a two-tier network topology where every leaf switch connects to every spine switch, creating a non-blocking fabric with predictable, low-latency paths between any two endpoints. It replaced three-tier (core-distribution-access) architectures because it scales horizontally without oversubscription, supports east-west traffic patterns (server-to-server), and provides consistent latency regardless of traffic path.
What structured cabling standard should data centers use?
Data centers should follow TIA-942-B (Telecommunications Infrastructure Standard for Data Centers) for overall infrastructure design, and ANSI/TIA-568.3-D for optical fiber cabling. For AI and high-performance computing, OM4 or OM5 multimode fiber supports 100G–400G over distances up to 150m. Single-mode OS2 fiber is required for distances above 300m or for 800G+ links.
Key Terms
A two-tier data center network topology where every leaf switch connects to every spine switch, providing non-blocking, low-latency connectivity between all endpoints.
A networking technology that allows data to be transferred directly between the memory of two computers without involving the CPU, dramatically reducing latency for AI training workloads.
The routing protocol that manages how data is routed between autonomous systems on the internet and increasingly within large data center fabrics.
The ratio of potential bandwidth demand to available bandwidth in a network. A 4:1 oversubscription means four times more traffic could be generated than the network can carry simultaneously.
Buyer\'s Guide
Questions to ask when evaluating networking solutions.
Why it matters: For AI training, oversubscription above 1:1 causes GPU stalls during all-reduce operations. For general enterprise workloads, 4:1 is acceptable. Vendors who cannot provide oversubscription ratios are hiding performance limitations.
Red flag: Network designs that do not specify oversubscription ratios.
Related Topic Clusters
Ready to implement networking?
Design high-performance networks for AI, cloud, and enterprise workloads Our certified engineers are ready to help you design and deploy the right solution.