Spot Fleet Strategies for Cost Savings

Implement Spot Fleet strategies to reduce EC2 costs by up to 90% while maintaining application availability and performance.

EC2 Spot Instances offer up to 90% savings compared to On-Demand pricing. This guide covers Spot Fleet strategies for maximizing savings while maintaining availability.

Spot Fleet Configuration

Diversified Allocation

SpotFleet:
  Type: AWS::EC2::SpotFleet
  Properties:
    SpotFleetRequestConfigData:
      IamFleetRole: !GetAtt SpotFleetRole.Arn
      TargetCapacity: 10
      AllocationStrategy: capacityOptimized
      InstanceInterruptionBehavior: terminate
      ReplaceUnhealthyInstances: true
      TerminateInstancesWithExpiration: true
      
      LaunchTemplateConfigs:
        - LaunchTemplateSpecification:
            LaunchTemplateId: !Ref SpotLaunchTemplate
            Version: !GetAtt SpotLaunchTemplate.LatestVersionNumber
          Overrides:
            - InstanceType: m5.large
              SubnetId: !Ref PrivateSubnet1
              WeightedCapacity: 1
            - InstanceType: m5a.large
              SubnetId: !Ref PrivateSubnet1
              WeightedCapacity: 1
            - InstanceType: m5n.large
              SubnetId: !Ref PrivateSubnet2
              WeightedCapacity: 1
            - InstanceType: m5.xlarge
              SubnetId: !Ref PrivateSubnet2
              WeightedCapacity: 2

Mixed Instances Policy

AutoScalingGroup:
  Type: AWS::AutoScaling::AutoScalingGroup
  Properties:
    MixedInstancesPolicy:
      InstancesDistribution:
        OnDemandBaseCapacity: 2
        OnDemandPercentageAboveBaseCapacity: 20
        SpotAllocationStrategy: capacity-optimized
        SpotInstancePools: 0  # Use all pools with capacity-optimized
        
      LaunchTemplate:
        LaunchTemplateSpecification:
          LaunchTemplateId: !Ref LaunchTemplate
          Version: !GetAtt LaunchTemplate.LatestVersionNumber
        Overrides:
          - InstanceType: m5.large
          - InstanceType: m5a.large
          - InstanceType: m5n.large
          - InstanceType: m5d.large
          - InstanceType: m4.large
          - InstanceType: r5.large
          - InstanceType: c5.large

Interruption Handling

Termination Notice Handler

import boto3
import requests
import time

def check_spot_interruption():
    # Check instance metadata for termination notice
    try:
        response = requests.get(
            'http://169.254.169.254/latest/meta-data/spot/termination-time',
            timeout=2
        )
        if response.status_code == 200:
            termination_time = response.text
            return {'interrupted': True, 'termination_time': termination_time}
    except:
        pass
    
    return {'interrupted': False}

def graceful_shutdown():
    # Drain connections
    drain_load_balancer()
    
    # Checkpoint application state
    checkpoint_state()
    
    # Deregister from service discovery
    deregister_service()
    
    # Signal readiness for termination
    print("Ready for termination")

EventBridge Integration

SpotInterruptionRule:
  Type: AWS::Events::Rule
  Properties:
    EventPattern:
      source:
        - aws.ec2
      detail-type:
        - EC2 Spot Instance Interruption Warning
    Targets:
      - Id: SpotInterruptionHandler
        Arn: !GetAtt SpotInterruptionLambda.Arn

SpotInterruptionLambda:
  Type: AWS::Lambda::Function
  Properties:
    Runtime: python3.9
    Handler: index.handler
    Code:
      ZipFile: |
        import boto3
        import json
        
        def handler(event, context):
            instance_id = event['detail']['instance-id']
            
            # Notify ASG to launch replacement
            autoscaling = boto3.client('autoscaling')
            
            # Detach instance from target group
            elbv2 = boto3.client('elbv2')
            
            # Log for analysis
            print(json.dumps({
                'event': 'spot_interruption',
                'instance_id': instance_id,
                'time': event['time']
            }))
            
            return {'statusCode': 200}

Capacity Rebalancing

def setup_capacity_rebalancing():
    ec2 = boto3.client('ec2')
    autoscaling = boto3.client('autoscaling')
    
    # Enable capacity rebalancing on ASG
    autoscaling.update_auto_scaling_group(
        AutoScalingGroupName='my-spot-asg',
        CapacityRebalance=True
    )

# Handle rebalancing events
def handle_rebalance_recommendation(event):
    instance_id = event['detail']['instance-id']
    
    # Proactively drain and replace
    drain_instance(instance_id)
    
    # ASG will launch replacement before termination
    print(f"Rebalancing: draining {instance_id}")

Savings Analysis

def calculate_spot_savings(on_demand_hours, spot_hours, instance_type):
    pricing = boto3.client('pricing', region_name='us-east-1')
    
    # Get On-Demand price
    od_response = pricing.get_products(
        ServiceCode='AmazonEC2',
        Filters=[
            {'Type': 'TERM_MATCH', 'Field': 'instanceType', 'Value': instance_type},
            {'Type': 'TERM_MATCH', 'Field': 'operatingSystem', 'Value': 'Linux'},
            {'Type': 'TERM_MATCH', 'Field': 'preInstalledSw', 'Value': 'NA'},
            {'Type': 'TERM_MATCH', 'Field': 'tenancy', 'Value': 'Shared'}
        ]
    )
    
    # Get Spot price history
    ec2 = boto3.client('ec2')
    spot_response = ec2.describe_spot_price_history(
        InstanceTypes=[instance_type],
        ProductDescriptions=['Linux/UNIX'],
        MaxResults=1
    )
    
    on_demand_price = 0.0968  # Example for m5.large
    spot_price = float(spot_response['SpotPriceHistory'][0]['SpotPrice'])
    
    on_demand_cost = on_demand_hours * on_demand_price
    spot_cost = spot_hours * spot_price
    
    return {
        'on_demand_cost': on_demand_cost,
        'spot_cost': spot_cost,
        'savings': on_demand_cost - spot_cost,
        'savings_percentage': ((on_demand_cost - spot_cost) / on_demand_cost) * 100
    }

Working with Warqline

We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.

Talk to an engineer

Conclusion

Spot Fleet strategies can dramatically reduce EC2 costs. Implement diversification across instance types and AZs, handle interruptions gracefully, and use capacity rebalancing for maximum availability.