AWS Cross-Region Disaster Recovery Strategies
Design and implement robust cross-region disaster recovery architectures for business continuity on AWS.
Designing effective disaster recovery (DR) strategies is essential for business continuity. This guide covers cross-region DR patterns on AWS, from backup-restore to multi-region active-active architectures.
Understanding DR Concepts
Key Metrics
Two critical metrics define DR requirements:
- Recovery Time Objective (RTO): Maximum acceptable downtime
- Recovery Point Objective (RPO): Maximum acceptable data loss
DR Strategies Overview
| Strategy | RTO | RPO | Cost |
|---|---|---|---|
| Backup & Restore | Hours | Hours | $ |
| Pilot Light | 10s of minutes | Minutes | $ |
| Warm Standby | Minutes | Seconds | $$ |
| Active-Active | Real-time | Zero | $$ |
Backup and Restore
Architecture
Lowest cost, longest recovery:
BackupConfiguration:
Primary:
Region: us-east-1
Resources:
- Database: RDS PostgreSQL
- Storage: S3 buckets
- Compute: EC2 AMIs
DR:
Region: us-west-2
ReplicationTargets:
- S3 Cross-Region Replication
- RDS Automated Snapshots (copied)
- AMI copies
Implementation
S3 Cross-Region Replication
S3ReplicationRule:
Type: AWS::S3::Bucket
Properties:
BucketName: production-data
ReplicationConfiguration:
Role: !GetAtt ReplicationRole.Arn
Rules:
- Id: ReplicateAllObjects
Status: Enabled
Destination:
Bucket: !Sub arn:aws:s3:::dr-production-data
StorageClass: STANDARD_IA
DeleteMarkerReplication:
Status: Disabled
RDS Snapshot Replication
import boto3
def copy_rds_snapshot_cross_region():
source_rds = boto3.client('rds', region_name='us-east-1')
target_rds = boto3.client('rds', region_name='us-west-2')
# Get latest snapshot
snapshots = source_rds.describe_db_snapshots(
DBInstanceIdentifier='production-db',
SnapshotType='automated'
)['DBSnapshots']
latest = sorted(snapshots, key=lambda x: x['SnapshotCreateTime'])[-1]
# Copy to DR region
target_rds.copy_db_snapshot(
SourceDBSnapshotIdentifier=latest['DBSnapshotArn'],
TargetDBSnapshotIdentifier=f"dr-{latest['DBSnapshotIdentifier']}",
SourceRegion='us-east-1',
CopyTags=True,
KmsKeyId='alias/dr-encryption-key'
)
Pilot Light
Architecture
Minimal core infrastructure ready to scale:
PilotLightArchitecture:
Primary:
Region: us-east-1
Components:
- ALB with Auto Scaling Groups
- RDS Multi-AZ
- ElastiCache cluster
DR:
Region: us-west-2
PilotLight:
- RDS Read Replica (can be promoted)
- Minimal EC2 capacity (1-2 instances)
- Pre-configured Auto Scaling Groups (scaled to 0)
- Route 53 health checks configured
Terraform Implementation
# Primary region resources
module "primary" {
source = "./modules/application"
providers = {
aws = aws.primary
}
environment = "production"
instance_count = var.primary_instance_count
rds_multi_az = true
}
# DR region pilot light
module "dr_pilot_light" {
source = "./modules/application"
providers = {
aws = aws.dr
}
environment = "dr"
instance_count = 1 # Minimal capacity
rds_multi_az = false
# RDS read replica from primary
rds_source_instance = module.primary.rds_instance_arn
}
# Route 53 failover
resource "aws_route53_health_check" "primary" {
fqdn = module.primary.alb_dns_name
port = 443
type = "HTTPS"
resource_path = "/health"
failure_threshold = 3
request_interval = 30
}
resource "aws_route53_record" "app" {
zone_id = var.route53_zone_id
name = "app.example.com"
type = "A"
failover_routing_policy {
type = "PRIMARY"
}
set_identifier = "primary"
health_check_id = aws_route53_health_check.primary.id
alias {
name = module.primary.alb_dns_name
zone_id = module.primary.alb_zone_id
evaluate_target_health = true
}
}
resource "aws_route53_record" "app_dr" {
zone_id = var.route53_zone_id
name = "app.example.com"
type = "A"
failover_routing_policy {
type = "SECONDARY"
}
set_identifier = "dr"
alias {
name = module.dr_pilot_light.alb_dns_name
zone_id = module.dr_pilot_light.alb_zone_id
evaluate_target_health = true
}
}
Failover Procedure
def initiate_pilot_light_failover():
# 1. Promote RDS read replica
rds = boto3.client('rds', region_name='us-west-2')
rds.promote_read_replica(
DBInstanceIdentifier='dr-production-db'
)
# 2. Scale up Auto Scaling Groups
autoscaling = boto3.client('autoscaling', region_name='us-west-2')
autoscaling.update_auto_scaling_group(
AutoScalingGroupName='dr-web-asg',
MinSize=4,
DesiredCapacity=8,
MaxSize=16
)
# 3. Warm up ElastiCache if needed
elasticache = boto3.client('elasticache', region_name='us-west-2')
# Cache warming logic
# 4. Update Route 53 if not automatic
# (Automatic with health checks, manual otherwise)
return {'status': 'failover_initiated'}
Warm Standby
Architecture
Scaled-down but functional DR environment:
WarmStandbyArchitecture:
Primary:
Region: us-east-1
Capacity: 100%
Components:
- ALB (6 targets)
- RDS Multi-AZ (db.r5.2xlarge)
- ElastiCache (3 nodes)
DR:
Region: us-west-2
Capacity: 25% # Ready to scale
Components:
- ALB (2 targets, ASG can scale)
- RDS Multi-AZ (db.r5.large, can be scaled)
- ElastiCache (1 node, can add)
- Active traffic handling
Global Accelerator Setup
GlobalAccelerator:
Type: AWS::GlobalAccelerator::Accelerator
Properties:
Name: production-app-accelerator
Enabled: true
Listener:
Type: AWS::GlobalAccelerator::Listener
Properties:
AcceleratorArn: !Ref GlobalAccelerator
PortRanges:
- FromPort: 443
ToPort: 443
Protocol: TCP
EndpointGroup:
Type: AWS::GlobalAccelerator::EndpointGroup
Properties:
ListenerArn: !Ref Listener
EndpointGroupRegion: us-east-1
TrafficDialPercentage: 100
EndpointConfigurations:
- EndpointId: !Ref PrimaryALB
Weight: 100
DREndpointGroup:
Type: AWS::GlobalAccelerator::EndpointGroup
Properties:
ListenerArn: !Ref Listener
EndpointGroupRegion: us-west-2
TrafficDialPercentage: 0 # Increase during failover
EndpointConfigurations:
- EndpointId: !Ref DRALB
Weight: 100
Active-Active
Architecture
Full capacity in multiple regions:
ActiveActiveArchitecture:
Regions:
- Region: us-east-1
Capacity: 100%
Traffic: 50%
- Region: us-west-2
Capacity: 100%
Traffic: 50%
GlobalServices:
- DynamoDB Global Tables
- Aurora Global Database
- S3 Cross-Region Replication (bidirectional)
- Route 53 with latency-based routing
Aurora Global Database
GlobalCluster:
Type: AWS::RDS::GlobalCluster
Properties:
GlobalClusterIdentifier: production-global
Engine: aurora-postgresql
EngineVersion: "14.6"
StorageEncrypted: true
PrimaryCluster:
Type: AWS::RDS::DBCluster
Properties:
GlobalClusterIdentifier: !Ref GlobalCluster
Engine: aurora-postgresql
DBClusterIdentifier: production-primary
MasterUsername: !Ref DBUsername
MasterUserPassword: !Ref DBPassword
EnableIAMDatabaseAuthentication: true
SecondaryCluster:
Type: AWS::RDS::DBCluster
Properties:
GlobalClusterIdentifier: !Ref GlobalCluster
Engine: aurora-postgresql
DBClusterIdentifier: production-secondary
# No master credentials - read replica
DynamoDB Global Tables
resource "aws_dynamodb_table" "sessions" {
name = "user-sessions"
billing_mode = "PAY_PER_REQUEST"
hash_key = "session_id"
attribute {
name = "session_id"
type = "S"
}
stream_enabled = true
stream_view_type = "NEW_AND_OLD_IMAGES"
replica {
region_name = "us-west-2"
}
replica {
region_name = "eu-west-1"
}
}
Testing DR
Regular DR Drills
DRTestingSchedule:
FullFailover:
Frequency: Annually
Duration: 4-8 hours
Scope: Complete region failover
PartialFailover:
Frequency: Quarterly
Duration: 2-4 hours
Scope: Database failover only
BackupRestore:
Frequency: Monthly
Duration: 1-2 hours
Scope: Restore and validate
HealthChecks:
Frequency: Weekly
Duration: 30 minutes
Scope: Verify DR readiness
Automated Testing
def run_dr_test():
results = {
'timestamp': datetime.now().isoformat(),
'tests': []
}
# Test 1: RDS failover
rds_result = test_rds_failover()
results['tests'].append({
'name': 'RDS Failover',
'success': rds_result['success'],
'duration': rds_result['duration']
})
# Test 2: Application health in DR
app_result = test_dr_application_health()
results['tests'].append({
'name': 'DR Application Health',
'success': app_result['success'],
'response_time': app_result['response_time']
})
# Test 3: Data consistency
data_result = verify_data_consistency()
results['tests'].append({
'name': 'Data Consistency',
'success': data_result['success'],
'lag': data_result['replication_lag']
})
# Store results and alert if issues
store_dr_test_results(results)
if not all(t['success'] for t in results['tests']):
alert_dr_team(results)
return results
Working with Warqline
We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.
Best Practices Checklist
Planning
- Define RTO/RPO for each workload
- Document recovery procedures
- Identify critical dependencies
- Calculate DR costs
Implementation
- Automate failover procedures
- Implement health checks
- Configure cross-region replication
- Set up monitoring
Testing
- Schedule regular DR drills
- Document test results
- Update procedures based on findings
- Train operations team
Conclusion
Cross-region disaster recovery is essential for business continuity. Choose the appropriate strategy based on your RTO/RPO requirements and budget, then implement comprehensive testing to ensure readiness. our reviews help validate your DR architecture against Well-Architected best practices.