AWS Cost Anomaly Detection: Automated Spending Alerts
Learn how to implement AWS Cost Anomaly Detection to automatically identify and alert on unusual spending patterns.
AWS Cost Anomaly Detection uses machine learning to identify unusual spending patterns and alert you before costs spiral out of control. This guide covers implementing comprehensive cost monitoring.
Understanding Cost Anomaly Detection
How It Works
Cost Anomaly Detection:
- Learns Patterns: ML models learn your normal spending behavior
- Identifies Anomalies: Detects deviations from expected patterns
- Alerts Teams: Sends notifications through multiple channels
- Provides Context: Shows root cause analysis
Key Benefits
- Early Warning: Catch issues before bills arrive
- ML-Powered: Adapts to your spending patterns
- Zero Maintenance: No manual threshold management
- Integration: Works with existing alerting systems
Setting Up Monitors
Cost Monitor Types
Create monitors for different scopes:
AWS Services Monitor
{
"MonitorType": "DIMENSIONAL",
"MonitorName": "All-Services-Monitor",
"MonitorDimension": "SERVICE",
"MonitorSpecification": {
"DimensionType": "DIMENSION",
"DimensionValues": []
}
}
Linked Account Monitor
For organizations:
CostMonitor:
Type: AWS::CE::AnomalyMonitor
Properties:
MonitorName: Production-Accounts-Monitor
MonitorType: DIMENSIONAL
MonitorDimension: LINKED_ACCOUNT
ResourceTags:
- Key: Environment
Value: production
Cost Category Monitor
For custom groupings:
{
"MonitorType": "CUSTOM",
"MonitorName": "Platform-Costs",
"MonitorSpecification": {
"And": null,
"CostCategories": {
"Key": "Team",
"Values": ["platform", "infrastructure"]
},
"Dimensions": null,
"Not": null,
"Or": null,
"Tags": null
}
}
Terraform Configuration
resource "aws_ce_anomaly_monitor" "all_services" {
name = "all-services-monitor"
monitor_type = "DIMENSIONAL"
monitor_dimension = "SERVICE"
}
resource "aws_ce_anomaly_monitor" "production" {
name = "production-costs"
monitor_type = "CUSTOM"
monitor_specification = jsonencode({
Tags = {
Key = "Environment"
Values = ["production"]
}
})
}
resource "aws_ce_anomaly_subscription" "alerts" {
name = "cost-anomaly-alerts"
frequency = "IMMEDIATE"
monitor_arn_list = [
aws_ce_anomaly_monitor.all_services.arn,
aws_ce_anomaly_monitor.production.arn
]
subscriber {
type = "EMAIL"
address = "finops@company.com"
}
subscriber {
type = "SNS"
address = aws_sns_topic.cost_alerts.arn
}
threshold_expression {
dimension {
key = "ANOMALY_TOTAL_IMPACT_PERCENTAGE"
values = ["20"]
match_options = ["GREATER_THAN_OR_EQUAL"]
}
}
}
Configuring Alert Thresholds
Threshold Types
Configure multiple threshold strategies:
Percentage-Based
Alert when cost exceeds normal by percentage:
ThresholdExpression:
Dimensions:
Key: ANOMALY_TOTAL_IMPACT_PERCENTAGE
MatchOptions:
- GREATER_THAN_OR_EQUAL
Values:
- "25" # Alert if 25%+ above expected
Absolute Amount
Alert when dollar impact exceeds threshold:
ThresholdExpression:
Dimensions:
Key: ANOMALY_TOTAL_IMPACT_ABSOLUTE
MatchOptions:
- GREATER_THAN_OR_EQUAL
Values:
- "100" # Alert if $100+ above expected
Combined Thresholds
Use both for nuanced alerting:
ThresholdExpression:
And:
- Dimensions:
Key: ANOMALY_TOTAL_IMPACT_PERCENTAGE
MatchOptions: ["GREATER_THAN_OR_EQUAL"]
Values: ["10"]
- Dimensions:
Key: ANOMALY_TOTAL_IMPACT_ABSOLUTE
MatchOptions: ["GREATER_THAN_OR_EQUAL"]
Values: ["50"]
Alert Subscriptions
Email Notifications
Simple email alerts:
{
"SubscriptionName": "email-alerts",
"Frequency": "IMMEDIATE",
"MonitorArnList": ["arn:aws:ce::123456789012:anomalymonitor/abc123"],
"Subscribers": [
{
"Type": "EMAIL",
"Address": "finops@company.com"
},
{
"Type": "EMAIL",
"Address": "engineering-leads@company.com"
}
]
}
SNS Integration
For advanced workflows:
CostAlertsTopic:
Type: AWS::SNS::Topic
Properties:
TopicName: cost-anomaly-alerts
KmsMasterKeyId: !Ref AlertsKMSKey
AnomalySubscription:
Type: AWS::CE::AnomalySubscription
Properties:
SubscriptionName: production-alerts
Frequency: IMMEDIATE
MonitorArnList:
- !Ref ProductionMonitor
Subscribers:
- Type: SNS
Address: !Ref CostAlertsTopic
ThresholdExpression:
Dimensions:
Key: ANOMALY_TOTAL_IMPACT_ABSOLUTE
MatchOptions:
- GREATER_THAN_OR_EQUAL
Values:
- "200"
Slack Integration
Route alerts to Slack:
import json
import urllib3
def lambda_handler(event, context):
message = event['Records'][0]['Sns']['Message']
anomaly = json.loads(message)
slack_message = {
"blocks": [
{
"type": "header",
"text": {
"type": "plain_text",
"text": "Cost Anomaly Detected"
}
},
{
"type": "section",
"fields": [
{
"type": "mrkdwn",
"text": f"*Service:*\n{anomaly['dimensionValue']}"
},
{
"type": "mrkdwn",
"text": f"*Impact:*\n${anomaly['impact']['totalImpact']:.2f}"
},
{
"type": "mrkdwn",
"text": f"*Expected:*\n${anomaly['impact']['totalExpectedSpend']:.2f}"
},
{
"type": "mrkdwn",
"text": f"*Actual:*\n${anomaly['impact']['totalActualSpend']:.2f}"
}
]
},
{
"type": "actions",
"elements": [
{
"type": "button",
"text": {
"type": "plain_text",
"text": "View in Cost Explorer"
},
"url": f"https://console.aws.amazon.com/cost-management/home#/anomaly-detection/anomalies/{anomaly['anomalyId']}"
}
]
}
]
}
http = urllib3.PoolManager()
http.request(
'POST',
SLACK_WEBHOOK_URL,
body=json.dumps(slack_message),
headers={'Content-Type': 'application/json'}
)
return {'statusCode': 200}
Root Cause Analysis
Understanding Anomalies
When an anomaly is detected, analyze:
- Time Range: When did spending increase?
- Service: Which AWS service is affected?
- Account: Which account(s) are involved?
- Region: Is it localized to a region?
- Usage Type: What specific usage changed?
API Analysis
import boto3
ce = boto3.client('ce')
def analyze_anomaly(anomaly_id):
# Get anomaly details
anomaly = ce.get_anomalies(
AnomalyId=anomaly_id
)['Anomalies'][0]
# Get detailed cost breakdown
start_date = anomaly['AnomalyStartDate']
end_date = anomaly['AnomalyEndDate']
cost_details = ce.get_cost_and_usage(
TimePeriod={'Start': start_date, 'End': end_date},
Granularity='DAILY',
Metrics=['BlendedCost'],
GroupBy=[
{'Type': 'DIMENSION', 'Key': 'SERVICE'},
{'Type': 'DIMENSION', 'Key': 'USAGE_TYPE'}
],
Filter={
'Dimensions': {
'Key': 'SERVICE',
'Values': [anomaly['DimensionValue']]
}
}
)
return {
'anomaly': anomaly,
'breakdown': cost_details
}
Automated Remediation
Common Remediation Actions
Implement automatic responses:
def remediate_anomaly(event, context):
anomaly = json.loads(event['Records'][0]['Sns']['Message'])
service = anomaly['dimensionValue']
impact = anomaly['impact']['totalImpact']
# Auto-remediation for known patterns
if service == 'Amazon EC2' and impact > 500:
# Find and stop idle instances
ec2 = boto3.client('ec2')
instances = ec2.describe_instances(
Filters=[
{'Name': 'tag:Environment', 'Values': ['development']},
{'Name': 'instance-state-name', 'Values': ['running']}
]
)
for reservation in instances['Reservations']:
for instance in reservation['Instances']:
# Check CPU utilization
cpu = get_instance_cpu(instance['InstanceId'])
if cpu < 5:
ec2.stop_instances(InstanceIds=[instance['InstanceId']])
log_action(f"Stopped idle instance: {instance['InstanceId']}")
elif service == 'Amazon S3':
# Check for unusual data transfer
analyze_s3_transfer(anomaly)
elif service == 'AWS Lambda':
# Check for runaway functions
analyze_lambda_invocations(anomaly)
return {'statusCode': 200, 'remediated': True}
Dashboard and Reporting
CloudWatch Dashboard
CostAnomalyDashboard:
Type: AWS::CloudWatch::Dashboard
Properties:
DashboardName: cost-anomaly-overview
DashboardBody: !Sub |
{
"widgets": [
{
"type": "metric",
"properties": {
"metrics": [
["AWS/Billing", "EstimatedCharges", "ServiceName", "AmazonEC2"],
["...", "AmazonS3"],
["...", "AWSLambda"]
],
"period": 86400,
"stat": "Maximum",
"title": "Daily Spending by Service"
}
},
{
"type": "text",
"properties": {
"markdown": "## Recent Anomalies\nView details in [Cost Anomaly Detection](https://console.aws.amazon.com/cost-management/home#/anomaly-detection)"
}
}
]
}
Working with Warqline
We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.
Best Practices
Implementation Checklist
- Create monitors for each account tier
- Set appropriate thresholds per environment
- Configure multiple notification channels
- Implement automated remediation for common issues
- Regular threshold review and adjustment
- Train teams on anomaly response procedures
Threshold Guidelines
| Environment | Percentage | Absolute |
|---|---|---|
| Production | 10% | $200+ |
| Staging | 25% | $50+ |
| Development | 50% | $25+ |
Conclusion
AWS Cost Anomaly Detection provides essential visibility into spending patterns. Combined with Warqline's comprehensive cost analytics, organizations can proactively manage cloud costs and prevent budget overruns.