2N Redundancy
Power SystemsA redundancy configuration where two complete, independent systems are available, each capable of supporting the full load. The highest level of redundancy for mission-critical infrastructure.
Authoritative definitions for 80+ terms across data centers, AI infrastructure, critical power, networking, cybersecurity, and cloud. Written and maintained by DCS Global's engineering team.
Showing 71 of 71 terms
A redundancy configuration where two complete, independent systems are available, each capable of supporting the full load. The highest level of redundancy for mission-critical infrastructure.
Mechanical unit that conditions and circulates air within a data center. Larger than CRAC units, typically connected to chilled water systems. Used in larger facilities where centralized cooling is more efficient.
A storage system using only NAND flash (SSD) storage, providing dramatically higher IOPS and lower latency than hybrid or spinning disk arrays.
ASHRAE's Technical Committee 9.9 publishes thermal guidelines for data center equipment. Defines A1–A4 equipment classes with operating temperature and humidity ranges.
A device that automatically transfers electrical load from a primary power source to a backup source (generator or alternate utility feed) upon detecting a failure. Transfer time: 4ms (static) to 30ms (mechanical).
An isolated location within a cloud region or data center campus with independent power, cooling, and network infrastructure. Designed to be failure-independent from other zones.
The routing protocol of the internet. Used in data centers for external connectivity and increasingly for internal routing in large-scale spine-leaf fabrics.
A 1U or 2U filler panel installed in empty rack spaces to prevent hot air recirculation between the front and rear of the rack. Critical for maintaining hot/cold aisle separation.
A physically secured, fenced area within a colocation facility allocated to a single tenant. Provides dedicated space with controlled access.
Upfront investment in physical assets (servers, network equipment, facility). On-premises infrastructure is primarily CapEx.
A device that conditions and distributes cooling liquid to server cold plates in a direct liquid cooling system. Manages temperature, pressure, and flow rate.
The gold standard cybersecurity certification issued by (ISC)². Requires 5 years of experience and covers 8 security domains.
A data center facility where multiple customers house their own servers and networking equipment in a shared facility. The facility provides power, cooling, physical security, and network connectivity.
A Tier III data center characteristic where any component can be maintained or replaced without shutting down the IT load. Requires redundant paths for power and cooling.
A self-contained precision cooling unit with its own refrigeration circuit. Cools air by passing it over a direct expansion (DX) coil. Common in smaller data centers and edge deployments.
A precision cooling unit that uses chilled water (from a central chiller plant) rather than a self-contained refrigeration circuit. More efficient at scale than CRAC units.
NVIDIA's parallel computing platform and programming model that enables GPU acceleration for general-purpose computing. The foundation of the NVIDIA AI software ecosystem.
Software platform that monitors, manages, and optimizes data center infrastructure including power, cooling, space, and assets in real time.
NVIDIA's purpose-built AI infrastructure systems. DGX H100 contains 8x H100 SXM5 GPUs with NVLink interconnect, 640GB HBM3 memory, and dual 400GbE/InfiniBand connectivity.
A cooling method where cold plates are attached directly to heat-generating components (CPUs, GPUs) and liquid is circulated through them. Supports 100+ kW/rack. Required for NVIDIA H100 SXM5 and GB200.
A UPS topology where incoming AC power is converted to DC, then back to AC. Provides complete electrical isolation and the highest power quality. The standard for mission-critical data centers.
A fiber optic technology that multiplexes multiple optical signals onto a single fiber using different wavelengths. Enables 100+ Gbps over long distances.
A routing strategy that distributes traffic across multiple equal-cost paths simultaneously, increasing bandwidth and providing redundancy. Essential in spine-leaf architectures.
A cooling mode where outdoor air or water is used directly for cooling without mechanical refrigeration, reducing energy consumption. Air-side economizers use outdoor air directly; water-side use cooling towers.
A Tier IV data center characteristic where any single failure — including a complete path failure — does not interrupt IT operations. Requires 2N+1 minimum redundancy.
U.S. federal law requiring federal agencies and contractors to implement information security programs. Compliance is required for all federal IT systems.
The maximum weight per square foot (or kg/m²) a raised floor or concrete slab can support. Critical for high-density AI racks. Standard raised floor: 1,000–2,000 lbs/sq ft. AI racks may require 3,000+ lbs/sq ft.
A diesel or natural gas engine-driven alternator that provides backup power during utility outages. Typically starts within 10 seconds and reaches full load within 30 seconds.
A massively parallel processor originally designed for graphics rendering, now the primary compute engine for AI/ML workloads. Contains thousands of cores optimized for matrix operations.
A high-speed, high-capacity memory technology stacked directly on the GPU die. HBM3e in H200 provides 4.8 TB/s memory bandwidth, critical for large model inference.
A data center layout where server racks are arranged so cold air intakes face a cold aisle and hot air exhausts face a hot aisle. Prevents hot and cold air mixing, improving cooling efficiency.
A physical enclosure around the hot aisle that captures server exhaust air and directs it back to cooling units, preventing mixing with cold supply air.
An IT architecture combining on-premises infrastructure with public cloud services, connected by a private network or VPN. Enables workload portability and burst capacity.
Data centers exceeding 100MW of IT load, typically operated by cloud providers (AWS, Azure, GCP, Meta, Google). Characterized by extreme standardization, automation, and economies of scale.
A cooling method where IT equipment is submerged in a thermally conductive but electrically non-conductive liquid. Single-phase uses mineral oil; two-phase uses fluorocarbon fluids that boil and condense. Supports 200+ kW/rack.
A high-performance networking technology providing ultra-low latency (600ns) and high bandwidth (NDR: 400Gb/s per port). The dominant interconnect for large GPU training clusters.
The international standard for information security management systems (ISMS). Provides a framework for establishing, implementing, and maintaining information security.
Unit of apparent power in AC circuits. Actual power (kW) = kVA × Power Factor. UPS systems are rated in kVA. A 100 kVA UPS at 0.9 PF delivers 90 kW.
A deep learning model trained on massive text datasets with billions to trillions of parameters. Examples: GPT-4, LLaMA, Claude, Gemini. Infrastructure requirements scale with parameter count.
Statistical measure of the average time between system failures. Used in reliability engineering to predict component and system failure rates.
Average time required to restore a failed component or system to operational status. Directly impacts overall system availability.
A strategy using services from multiple cloud providers (AWS + Azure + GCP) to avoid vendor lock-in, optimize cost, or meet data sovereignty requirements.
A redundancy configuration where N components are required for operation and one additional component is available as a spare. If any single component fails, the spare takes over.
NVIDIA's library for multi-GPU and multi-node collective communication operations (AllReduce, AllGather, etc.) used in distributed training.
InfiniBand's current generation standard providing 400 Gb/s per port (800 Gb/s bidirectional). Successor to HDR (200 Gb/s). The standard for large AI training clusters.
The organization that establishes standards for electrical testing and maintenance. NETA MTS (Maintenance Testing Specifications) is the standard for electrical acceptance testing in data centers.
NIST's catalog of security and privacy controls for federal information systems. The foundation of U.S. government cybersecurity compliance.
NVIDIA's high-speed GPU-to-GPU interconnect within a single node. NVLink 4.0 provides 900 GB/s bidirectional bandwidth between GPUs, far exceeding PCIe 5.0's 128 GB/s.
A storage protocol designed specifically for flash storage, providing dramatically lower latency (microseconds vs milliseconds for SATA/SAS) and higher IOPS.
NVIDIA's switching chip that enables all-to-all NVLink connectivity between all GPUs in a DGX system. NVSwitch 3.0 in DGX H100 provides 900 GB/s per GPU.
A storage architecture that manages data as objects (vs files or blocks). Highly scalable, ideal for unstructured data (AI datasets, backups, archives). Examples: AWS S3, MinIO, Ceph.
Ongoing costs for running a business (cloud subscriptions, maintenance contracts, staffing). Cloud infrastructure is primarily OpEx.
A distributed file system that stores data across multiple storage nodes simultaneously, providing aggregate bandwidth that scales with the number of nodes. Examples: GPFS (IBM Spectrum Scale), Lustre, WEKA. Required for large AI training clusters.
A device that distributes electrical power to multiple IT equipment outlets within a rack or row. Types range from basic (passive) to intelligent (monitored, switched, with outlet-level metering).
The ratio of total data center power consumption to IT equipment power consumption. PUE = Total Facility Power ÷ IT Equipment Power. A PUE of 1.0 is theoretical perfection; world-class facilities achieve 1.05–1.12.
A modular floor system elevated above the structural slab, creating a plenum for cable management and underfloor air distribution. Standard height: 12–24 inches.
A technology allowing direct memory access from one computer's memory to another's without involving the CPU. Critical for low-latency GPU cluster communication.
/ROH-see-vee-two/
RDMA protocol running over standard Ethernet. Requires lossless network (PFC + ECN). Lower cost than InfiniBand but higher latency.
A network architecture approach that separates the control plane from the data plane, enabling centralized, programmable network management.
An auditing standard developed by the AICPA that evaluates a service organization's controls related to security, availability, processing integrity, confidentiality, and privacy over a period of time (typically 6–12 months).
A two-tier network topology where every leaf switch connects to every spine switch, providing predictable latency and easy horizontal scaling. The standard architecture for modern data centers.
An electronic switching device that transfers load between two power sources in less than 4 milliseconds — fast enough to prevent IT equipment from detecting the transfer.
The complete cost of an asset over its useful life, including acquisition, operation, maintenance, and disposal. Essential for build vs. buy vs. cloud decisions.
Specialized processing units within NVIDIA GPUs designed specifically for matrix multiply-accumulate operations (the core computation in neural networks). H100 has 528 Tensor Cores.
A telecommunications infrastructure standard for data centers published by the Telecommunications Industry Association. Defines Rated-1 through Rated-4 classifications (similar to but distinct from Uptime Institute Tiers).
The Uptime Institute's data center classification system based on redundancy, availability, and fault tolerance. Tier I = 99.671% availability; Tier IV = 99.995% availability.
A device that provides emergency power to IT equipment when the main power source fails. Bridges the gap between utility failure and generator startup (typically 10–30 seconds).
The global authority on data center performance and reliability. Publishes the Tier Standard and provides independent certification of Tier I–IV data centers.
The dedicated memory on a GPU. Determines maximum model size for inference (model must fit in VRAM). H100: 80GB HBM3; H200: 141GB HBM3e; B200: 192GB HBM3e.
A network virtualization technology that encapsulates Layer 2 frames within UDP packets, enabling Layer 2 networks to span Layer 3 boundaries. Used for multi-tenant data center networks.
A security model based on the principle "never trust, always verify." Eliminates implicit trust based on network location; every access request is authenticated, authorized, and continuously validated.