Spot Fleet Strategies for Cost Savings
Implement Spot Fleet strategies to reduce EC2 costs by up to 90% while maintaining application availability and performance.
EC2 Spot Instances offer up to 90% savings compared to On-Demand pricing. This guide covers Spot Fleet strategies for maximizing savings while maintaining availability.
Spot Fleet Configuration
Diversified Allocation
SpotFleet:
Type: AWS::EC2::SpotFleet
Properties:
SpotFleetRequestConfigData:
IamFleetRole: !GetAtt SpotFleetRole.Arn
TargetCapacity: 10
AllocationStrategy: capacityOptimized
InstanceInterruptionBehavior: terminate
ReplaceUnhealthyInstances: true
TerminateInstancesWithExpiration: true
LaunchTemplateConfigs:
- LaunchTemplateSpecification:
LaunchTemplateId: !Ref SpotLaunchTemplate
Version: !GetAtt SpotLaunchTemplate.LatestVersionNumber
Overrides:
- InstanceType: m5.large
SubnetId: !Ref PrivateSubnet1
WeightedCapacity: 1
- InstanceType: m5a.large
SubnetId: !Ref PrivateSubnet1
WeightedCapacity: 1
- InstanceType: m5n.large
SubnetId: !Ref PrivateSubnet2
WeightedCapacity: 1
- InstanceType: m5.xlarge
SubnetId: !Ref PrivateSubnet2
WeightedCapacity: 2
Mixed Instances Policy
AutoScalingGroup:
Type: AWS::AutoScaling::AutoScalingGroup
Properties:
MixedInstancesPolicy:
InstancesDistribution:
OnDemandBaseCapacity: 2
OnDemandPercentageAboveBaseCapacity: 20
SpotAllocationStrategy: capacity-optimized
SpotInstancePools: 0 # Use all pools with capacity-optimized
LaunchTemplate:
LaunchTemplateSpecification:
LaunchTemplateId: !Ref LaunchTemplate
Version: !GetAtt LaunchTemplate.LatestVersionNumber
Overrides:
- InstanceType: m5.large
- InstanceType: m5a.large
- InstanceType: m5n.large
- InstanceType: m5d.large
- InstanceType: m4.large
- InstanceType: r5.large
- InstanceType: c5.large
Interruption Handling
Termination Notice Handler
import boto3
import requests
import time
def check_spot_interruption():
# Check instance metadata for termination notice
try:
response = requests.get(
'http://169.254.169.254/latest/meta-data/spot/termination-time',
timeout=2
)
if response.status_code == 200:
termination_time = response.text
return {'interrupted': True, 'termination_time': termination_time}
except:
pass
return {'interrupted': False}
def graceful_shutdown():
# Drain connections
drain_load_balancer()
# Checkpoint application state
checkpoint_state()
# Deregister from service discovery
deregister_service()
# Signal readiness for termination
print("Ready for termination")
EventBridge Integration
SpotInterruptionRule:
Type: AWS::Events::Rule
Properties:
EventPattern:
source:
- aws.ec2
detail-type:
- EC2 Spot Instance Interruption Warning
Targets:
- Id: SpotInterruptionHandler
Arn: !GetAtt SpotInterruptionLambda.Arn
SpotInterruptionLambda:
Type: AWS::Lambda::Function
Properties:
Runtime: python3.9
Handler: index.handler
Code:
ZipFile: |
import boto3
import json
def handler(event, context):
instance_id = event['detail']['instance-id']
# Notify ASG to launch replacement
autoscaling = boto3.client('autoscaling')
# Detach instance from target group
elbv2 = boto3.client('elbv2')
# Log for analysis
print(json.dumps({
'event': 'spot_interruption',
'instance_id': instance_id,
'time': event['time']
}))
return {'statusCode': 200}
Capacity Rebalancing
def setup_capacity_rebalancing():
ec2 = boto3.client('ec2')
autoscaling = boto3.client('autoscaling')
# Enable capacity rebalancing on ASG
autoscaling.update_auto_scaling_group(
AutoScalingGroupName='my-spot-asg',
CapacityRebalance=True
)
# Handle rebalancing events
def handle_rebalance_recommendation(event):
instance_id = event['detail']['instance-id']
# Proactively drain and replace
drain_instance(instance_id)
# ASG will launch replacement before termination
print(f"Rebalancing: draining {instance_id}")
Savings Analysis
def calculate_spot_savings(on_demand_hours, spot_hours, instance_type):
pricing = boto3.client('pricing', region_name='us-east-1')
# Get On-Demand price
od_response = pricing.get_products(
ServiceCode='AmazonEC2',
Filters=[
{'Type': 'TERM_MATCH', 'Field': 'instanceType', 'Value': instance_type},
{'Type': 'TERM_MATCH', 'Field': 'operatingSystem', 'Value': 'Linux'},
{'Type': 'TERM_MATCH', 'Field': 'preInstalledSw', 'Value': 'NA'},
{'Type': 'TERM_MATCH', 'Field': 'tenancy', 'Value': 'Shared'}
]
)
# Get Spot price history
ec2 = boto3.client('ec2')
spot_response = ec2.describe_spot_price_history(
InstanceTypes=[instance_type],
ProductDescriptions=['Linux/UNIX'],
MaxResults=1
)
on_demand_price = 0.0968 # Example for m5.large
spot_price = float(spot_response['SpotPriceHistory'][0]['SpotPrice'])
on_demand_cost = on_demand_hours * on_demand_price
spot_cost = spot_hours * spot_price
return {
'on_demand_cost': on_demand_cost,
'spot_cost': spot_cost,
'savings': on_demand_cost - spot_cost,
'savings_percentage': ((on_demand_cost - spot_cost) / on_demand_cost) * 100
}
Working with Warqline
We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.
Conclusion
Spot Fleet strategies can dramatically reduce EC2 costs. Implement diversification across instance types and AZs, handle interruptions gracefully, and use capacity rebalancing for maximum availability.