AWS Cross-Region Disaster Recovery Strategies

Design and implement robust cross-region disaster recovery architectures for business continuity on AWS.

Designing effective disaster recovery (DR) strategies is essential for business continuity. This guide covers cross-region DR patterns on AWS, from backup-restore to multi-region active-active architectures.

Understanding DR Concepts

Key Metrics

Two critical metrics define DR requirements:

  • Recovery Time Objective (RTO): Maximum acceptable downtime
  • Recovery Point Objective (RPO): Maximum acceptable data loss

DR Strategies Overview

Strategy RTO RPO Cost
Backup & Restore Hours Hours $
Pilot Light 10s of minutes Minutes $
Warm Standby Minutes Seconds $$
Active-Active Real-time Zero $$

Backup and Restore

Architecture

Lowest cost, longest recovery:

BackupConfiguration:
  Primary:
    Region: us-east-1
    Resources:
      - Database: RDS PostgreSQL
      - Storage: S3 buckets
      - Compute: EC2 AMIs
  
  DR:
    Region: us-west-2
    ReplicationTargets:
      - S3 Cross-Region Replication
      - RDS Automated Snapshots (copied)
      - AMI copies

Implementation

S3 Cross-Region Replication

S3ReplicationRule:
  Type: AWS::S3::Bucket
  Properties:
    BucketName: production-data
    ReplicationConfiguration:
      Role: !GetAtt ReplicationRole.Arn
      Rules:
        - Id: ReplicateAllObjects
          Status: Enabled
          Destination:
            Bucket: !Sub arn:aws:s3:::dr-production-data
            StorageClass: STANDARD_IA
          DeleteMarkerReplication:
            Status: Disabled

RDS Snapshot Replication

import boto3

def copy_rds_snapshot_cross_region():
    source_rds = boto3.client('rds', region_name='us-east-1')
    target_rds = boto3.client('rds', region_name='us-west-2')
    
    # Get latest snapshot
    snapshots = source_rds.describe_db_snapshots(
        DBInstanceIdentifier='production-db',
        SnapshotType='automated'
    )['DBSnapshots']
    
    latest = sorted(snapshots, key=lambda x: x['SnapshotCreateTime'])[-1]
    
    # Copy to DR region
    target_rds.copy_db_snapshot(
        SourceDBSnapshotIdentifier=latest['DBSnapshotArn'],
        TargetDBSnapshotIdentifier=f"dr-{latest['DBSnapshotIdentifier']}",
        SourceRegion='us-east-1',
        CopyTags=True,
        KmsKeyId='alias/dr-encryption-key'
    )

Pilot Light

Architecture

Minimal core infrastructure ready to scale:

PilotLightArchitecture:
  Primary:
    Region: us-east-1
    Components:
      - ALB with Auto Scaling Groups
      - RDS Multi-AZ
      - ElastiCache cluster
      
  DR:
    Region: us-west-2
    PilotLight:
      - RDS Read Replica (can be promoted)
      - Minimal EC2 capacity (1-2 instances)
      - Pre-configured Auto Scaling Groups (scaled to 0)
      - Route 53 health checks configured

Terraform Implementation

# Primary region resources
module "primary" {
  source = "./modules/application"
  
  providers = {
    aws = aws.primary
  }
  
  environment    = "production"
  instance_count = var.primary_instance_count
  rds_multi_az   = true
}

# DR region pilot light
module "dr_pilot_light" {
  source = "./modules/application"
  
  providers = {
    aws = aws.dr
  }
  
  environment    = "dr"
  instance_count = 1  # Minimal capacity
  rds_multi_az   = false
  
  # RDS read replica from primary
  rds_source_instance = module.primary.rds_instance_arn
}

# Route 53 failover
resource "aws_route53_health_check" "primary" {
  fqdn              = module.primary.alb_dns_name
  port              = 443
  type              = "HTTPS"
  resource_path     = "/health"
  failure_threshold = 3
  request_interval  = 30
}

resource "aws_route53_record" "app" {
  zone_id = var.route53_zone_id
  name    = "app.example.com"
  type    = "A"
  
  failover_routing_policy {
    type = "PRIMARY"
  }
  
  set_identifier = "primary"
  health_check_id = aws_route53_health_check.primary.id
  
  alias {
    name                   = module.primary.alb_dns_name
    zone_id               = module.primary.alb_zone_id
    evaluate_target_health = true
  }
}

resource "aws_route53_record" "app_dr" {
  zone_id = var.route53_zone_id
  name    = "app.example.com"
  type    = "A"
  
  failover_routing_policy {
    type = "SECONDARY"
  }
  
  set_identifier = "dr"
  
  alias {
    name                   = module.dr_pilot_light.alb_dns_name
    zone_id               = module.dr_pilot_light.alb_zone_id
    evaluate_target_health = true
  }
}

Failover Procedure

def initiate_pilot_light_failover():
    # 1. Promote RDS read replica
    rds = boto3.client('rds', region_name='us-west-2')
    rds.promote_read_replica(
        DBInstanceIdentifier='dr-production-db'
    )
    
    # 2. Scale up Auto Scaling Groups
    autoscaling = boto3.client('autoscaling', region_name='us-west-2')
    autoscaling.update_auto_scaling_group(
        AutoScalingGroupName='dr-web-asg',
        MinSize=4,
        DesiredCapacity=8,
        MaxSize=16
    )
    
    # 3. Warm up ElastiCache if needed
    elasticache = boto3.client('elasticache', region_name='us-west-2')
    # Cache warming logic
    
    # 4. Update Route 53 if not automatic
    # (Automatic with health checks, manual otherwise)
    
    return {'status': 'failover_initiated'}

Warm Standby

Architecture

Scaled-down but functional DR environment:

WarmStandbyArchitecture:
  Primary:
    Region: us-east-1
    Capacity: 100%
    Components:
      - ALB (6 targets)
      - RDS Multi-AZ (db.r5.2xlarge)
      - ElastiCache (3 nodes)
      
  DR:
    Region: us-west-2
    Capacity: 25%  # Ready to scale
    Components:
      - ALB (2 targets, ASG can scale)
      - RDS Multi-AZ (db.r5.large, can be scaled)
      - ElastiCache (1 node, can add)
      - Active traffic handling

Global Accelerator Setup

GlobalAccelerator:
  Type: AWS::GlobalAccelerator::Accelerator
  Properties:
    Name: production-app-accelerator
    Enabled: true

Listener:
  Type: AWS::GlobalAccelerator::Listener
  Properties:
    AcceleratorArn: !Ref GlobalAccelerator
    PortRanges:
      - FromPort: 443
        ToPort: 443
    Protocol: TCP

EndpointGroup:
  Type: AWS::GlobalAccelerator::EndpointGroup
  Properties:
    ListenerArn: !Ref Listener
    EndpointGroupRegion: us-east-1
    TrafficDialPercentage: 100
    EndpointConfigurations:
      - EndpointId: !Ref PrimaryALB
        Weight: 100

DREndpointGroup:
  Type: AWS::GlobalAccelerator::EndpointGroup
  Properties:
    ListenerArn: !Ref Listener
    EndpointGroupRegion: us-west-2
    TrafficDialPercentage: 0  # Increase during failover
    EndpointConfigurations:
      - EndpointId: !Ref DRALB
        Weight: 100

Active-Active

Architecture

Full capacity in multiple regions:

ActiveActiveArchitecture:
  Regions:
    - Region: us-east-1
      Capacity: 100%
      Traffic: 50%
      
    - Region: us-west-2
      Capacity: 100%
      Traffic: 50%
      
  GlobalServices:
    - DynamoDB Global Tables
    - Aurora Global Database
    - S3 Cross-Region Replication (bidirectional)
    - Route 53 with latency-based routing

Aurora Global Database

GlobalCluster:
  Type: AWS::RDS::GlobalCluster
  Properties:
    GlobalClusterIdentifier: production-global
    Engine: aurora-postgresql
    EngineVersion: "14.6"
    StorageEncrypted: true

PrimaryCluster:
  Type: AWS::RDS::DBCluster
  Properties:
    GlobalClusterIdentifier: !Ref GlobalCluster
    Engine: aurora-postgresql
    DBClusterIdentifier: production-primary
    MasterUsername: !Ref DBUsername
    MasterUserPassword: !Ref DBPassword
    EnableIAMDatabaseAuthentication: true

SecondaryCluster:
  Type: AWS::RDS::DBCluster
  Properties:
    GlobalClusterIdentifier: !Ref GlobalCluster
    Engine: aurora-postgresql
    DBClusterIdentifier: production-secondary
    # No master credentials - read replica

DynamoDB Global Tables

resource "aws_dynamodb_table" "sessions" {
  name         = "user-sessions"
  billing_mode = "PAY_PER_REQUEST"
  hash_key     = "session_id"
  
  attribute {
    name = "session_id"
    type = "S"
  }
  
  stream_enabled   = true
  stream_view_type = "NEW_AND_OLD_IMAGES"
  
  replica {
    region_name = "us-west-2"
  }
  
  replica {
    region_name = "eu-west-1"
  }
}

Testing DR

Regular DR Drills

DRTestingSchedule:
  FullFailover:
    Frequency: Annually
    Duration: 4-8 hours
    Scope: Complete region failover
    
  PartialFailover:
    Frequency: Quarterly
    Duration: 2-4 hours
    Scope: Database failover only
    
  BackupRestore:
    Frequency: Monthly
    Duration: 1-2 hours
    Scope: Restore and validate
    
  HealthChecks:
    Frequency: Weekly
    Duration: 30 minutes
    Scope: Verify DR readiness

Automated Testing

def run_dr_test():
    results = {
        'timestamp': datetime.now().isoformat(),
        'tests': []
    }
    
    # Test 1: RDS failover
    rds_result = test_rds_failover()
    results['tests'].append({
        'name': 'RDS Failover',
        'success': rds_result['success'],
        'duration': rds_result['duration']
    })
    
    # Test 2: Application health in DR
    app_result = test_dr_application_health()
    results['tests'].append({
        'name': 'DR Application Health',
        'success': app_result['success'],
        'response_time': app_result['response_time']
    })
    
    # Test 3: Data consistency
    data_result = verify_data_consistency()
    results['tests'].append({
        'name': 'Data Consistency',
        'success': data_result['success'],
        'lag': data_result['replication_lag']
    })
    
    # Store results and alert if issues
    store_dr_test_results(results)
    if not all(t['success'] for t in results['tests']):
        alert_dr_team(results)
    
    return results

Working with Warqline

We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.

Talk to an engineer

Best Practices Checklist

Planning

  • Define RTO/RPO for each workload
  • Document recovery procedures
  • Identify critical dependencies
  • Calculate DR costs

Implementation

  • Automate failover procedures
  • Implement health checks
  • Configure cross-region replication
  • Set up monitoring

Testing

  • Schedule regular DR drills
  • Document test results
  • Update procedures based on findings
  • Train operations team

Conclusion

Cross-region disaster recovery is essential for business continuity. Choose the appropriate strategy based on your RTO/RPO requirements and budget, then implement comprehensive testing to ensure readiness. our reviews help validate your DR architecture against Well-Architected best practices.