RPO and RTO: The Two Numbers That Define Your Requirements
RPO
Recovery Point Objective
How much data loss is acceptable?
RPO of 4 hours means the organization can tolerate losing up to 4 hours of data. Backups must run at least every 4 hours.
Drives: Backup frequency and replication lag
RTO
Recovery Time Objective
How long can recovery take?
RTO of 2 hours means the organization must be able to restore operations within 2 hours of a failure. DR architecture must support this timeline.
Drives: DR architecture and recovery automation
RPO and RTO must be defined per workload
Backup vs. Disaster Recovery
Backup
Protects against
Data loss: corruption, accidental deletion, ransomware encryption
Does not protect against
Facility loss: fire, flood, power failure, natural disaster
Recovery method
Restore data to existing or new infrastructure
Typical RTO
Hours to days (depending on data volume and infrastructure)
Disaster Recovery
Protects against
Facility loss, primary data center unavailable
Does not protect against
Data corruption that has replicated to the DR site
Recovery method
Fail over to secondary site with replicated data
Typical RTO
Minutes to hours (depending on DR architecture)
Backup Architecture: The 3-2-1 Rule
The 3-2-1 rule is the minimum standard for enterprise backup architecture: 3 copies of data, on 2 different media types, with 1 copy offsite. Modern ransomware attacks target backup systems specifically: the 3-2-1 rule must be extended to 3-2-1-1-0: 3 copies, 2 media types, 1 offsite, 1 immutable (air-gapped or WORM), 0 errors verified by recovery testing.
Ransomware targets backup systems
DR Architecture Options
Cold Standby
RTO: 24–72 hours
Cost: Low
Secondary site with hardware available but not running. Data restored from backup. Lowest cost, highest RTO.
Warm Standby
RTO: 4–24 hours
Cost: Medium
Secondary site with infrastructure running but not serving production traffic. Data replicated periodically. Moderate cost and RTO.
Hot Standby
RTO: 15 minutes–4 hours
Cost: High
Secondary site fully operational with real-time data replication. Failover is rapid. High cost, requires duplicate infrastructure.
Active-Active
RTO: Near-zero
Cost: Very High
Both sites serve production traffic simultaneously. Failover is transparent. Highest cost, requires full infrastructure at both sites.
Testing and Validation
Backups and DR plans that have not been tested are not reliable. The only way to verify that a backup can be restored is to restore it. The only way to verify that a DR plan works is to execute it. Organizations that discover their backup or DR plan does not work during an actual incident face recovery times that are far longer than their stated RTO.
Backup recovery testing should be performed quarterly at minimum: monthly for mission-critical workloads. DR failover testing should be performed annually at minimum: with a full failover test that verifies the complete recovery process, not just individual components.