AWS Cost Anomaly Detection: Automated Spending Alerts

Learn how to implement AWS Cost Anomaly Detection to automatically identify and alert on unusual spending patterns.

AWS Cost Anomaly Detection uses machine learning to identify unusual spending patterns and alert you before costs spiral out of control. This guide covers implementing comprehensive cost monitoring.

Understanding Cost Anomaly Detection

How It Works

Cost Anomaly Detection:

  1. Learns Patterns: ML models learn your normal spending behavior
  2. Identifies Anomalies: Detects deviations from expected patterns
  3. Alerts Teams: Sends notifications through multiple channels
  4. Provides Context: Shows root cause analysis

Key Benefits

  • Early Warning: Catch issues before bills arrive
  • ML-Powered: Adapts to your spending patterns
  • Zero Maintenance: No manual threshold management
  • Integration: Works with existing alerting systems

Setting Up Monitors

Cost Monitor Types

Create monitors for different scopes:

AWS Services Monitor

{
  "MonitorType": "DIMENSIONAL",
  "MonitorName": "All-Services-Monitor",
  "MonitorDimension": "SERVICE",
  "MonitorSpecification": {
    "DimensionType": "DIMENSION",
    "DimensionValues": []
  }
}

Linked Account Monitor

For organizations:

CostMonitor:
  Type: AWS::CE::AnomalyMonitor
  Properties:
    MonitorName: Production-Accounts-Monitor
    MonitorType: DIMENSIONAL
    MonitorDimension: LINKED_ACCOUNT
    ResourceTags:
      - Key: Environment
        Value: production

Cost Category Monitor

For custom groupings:

{
  "MonitorType": "CUSTOM",
  "MonitorName": "Platform-Costs",
  "MonitorSpecification": {
    "And": null,
    "CostCategories": {
      "Key": "Team",
      "Values": ["platform", "infrastructure"]
    },
    "Dimensions": null,
    "Not": null,
    "Or": null,
    "Tags": null
  }
}

Terraform Configuration

resource "aws_ce_anomaly_monitor" "all_services" {
  name              = "all-services-monitor"
  monitor_type      = "DIMENSIONAL"
  monitor_dimension = "SERVICE"
}

resource "aws_ce_anomaly_monitor" "production" {
  name         = "production-costs"
  monitor_type = "CUSTOM"
  monitor_specification = jsonencode({
    Tags = {
      Key    = "Environment"
      Values = ["production"]
    }
  })
}

resource "aws_ce_anomaly_subscription" "alerts" {
  name      = "cost-anomaly-alerts"
  frequency = "IMMEDIATE"
  
  monitor_arn_list = [
    aws_ce_anomaly_monitor.all_services.arn,
    aws_ce_anomaly_monitor.production.arn
  ]
  
  subscriber {
    type    = "EMAIL"
    address = "finops@company.com"
  }
  
  subscriber {
    type    = "SNS"
    address = aws_sns_topic.cost_alerts.arn
  }
  
  threshold_expression {
    dimension {
      key           = "ANOMALY_TOTAL_IMPACT_PERCENTAGE"
      values        = ["20"]
      match_options = ["GREATER_THAN_OR_EQUAL"]
    }
  }
}

Configuring Alert Thresholds

Threshold Types

Configure multiple threshold strategies:

Percentage-Based

Alert when cost exceeds normal by percentage:

ThresholdExpression:
  Dimensions:
    Key: ANOMALY_TOTAL_IMPACT_PERCENTAGE
    MatchOptions:
      - GREATER_THAN_OR_EQUAL
    Values:
      - "25"  # Alert if 25%+ above expected

Absolute Amount

Alert when dollar impact exceeds threshold:

ThresholdExpression:
  Dimensions:
    Key: ANOMALY_TOTAL_IMPACT_ABSOLUTE
    MatchOptions:
      - GREATER_THAN_OR_EQUAL
    Values:
      - "100"  # Alert if $100+ above expected

Combined Thresholds

Use both for nuanced alerting:

ThresholdExpression:
  And:
    - Dimensions:
        Key: ANOMALY_TOTAL_IMPACT_PERCENTAGE
        MatchOptions: ["GREATER_THAN_OR_EQUAL"]
        Values: ["10"]
    - Dimensions:
        Key: ANOMALY_TOTAL_IMPACT_ABSOLUTE
        MatchOptions: ["GREATER_THAN_OR_EQUAL"]
        Values: ["50"]

Alert Subscriptions

Email Notifications

Simple email alerts:

{
  "SubscriptionName": "email-alerts",
  "Frequency": "IMMEDIATE",
  "MonitorArnList": ["arn:aws:ce::123456789012:anomalymonitor/abc123"],
  "Subscribers": [
    {
      "Type": "EMAIL",
      "Address": "finops@company.com"
    },
    {
      "Type": "EMAIL", 
      "Address": "engineering-leads@company.com"
    }
  ]
}

SNS Integration

For advanced workflows:

CostAlertsTopic:
  Type: AWS::SNS::Topic
  Properties:
    TopicName: cost-anomaly-alerts
    KmsMasterKeyId: !Ref AlertsKMSKey

AnomalySubscription:
  Type: AWS::CE::AnomalySubscription
  Properties:
    SubscriptionName: production-alerts
    Frequency: IMMEDIATE
    MonitorArnList:
      - !Ref ProductionMonitor
    Subscribers:
      - Type: SNS
        Address: !Ref CostAlertsTopic
    ThresholdExpression:
      Dimensions:
        Key: ANOMALY_TOTAL_IMPACT_ABSOLUTE
        MatchOptions:
          - GREATER_THAN_OR_EQUAL
        Values:
          - "200"

Slack Integration

Route alerts to Slack:

import json
import urllib3

def lambda_handler(event, context):
    message = event['Records'][0]['Sns']['Message']
    anomaly = json.loads(message)
    
    slack_message = {
        "blocks": [
            {
                "type": "header",
                "text": {
                    "type": "plain_text",
                    "text": "Cost Anomaly Detected"
                }
            },
            {
                "type": "section",
                "fields": [
                    {
                        "type": "mrkdwn",
                        "text": f"*Service:*\n{anomaly['dimensionValue']}"
                    },
                    {
                        "type": "mrkdwn",
                        "text": f"*Impact:*\n${anomaly['impact']['totalImpact']:.2f}"
                    },
                    {
                        "type": "mrkdwn",
                        "text": f"*Expected:*\n${anomaly['impact']['totalExpectedSpend']:.2f}"
                    },
                    {
                        "type": "mrkdwn",
                        "text": f"*Actual:*\n${anomaly['impact']['totalActualSpend']:.2f}"
                    }
                ]
            },
            {
                "type": "actions",
                "elements": [
                    {
                        "type": "button",
                        "text": {
                            "type": "plain_text",
                            "text": "View in Cost Explorer"
                        },
                        "url": f"https://console.aws.amazon.com/cost-management/home#/anomaly-detection/anomalies/{anomaly['anomalyId']}"
                    }
                ]
            }
        ]
    }
    
    http = urllib3.PoolManager()
    http.request(
        'POST',
        SLACK_WEBHOOK_URL,
        body=json.dumps(slack_message),
        headers={'Content-Type': 'application/json'}
    )
    
    return {'statusCode': 200}

Root Cause Analysis

Understanding Anomalies

When an anomaly is detected, analyze:

  1. Time Range: When did spending increase?
  2. Service: Which AWS service is affected?
  3. Account: Which account(s) are involved?
  4. Region: Is it localized to a region?
  5. Usage Type: What specific usage changed?

API Analysis

import boto3

ce = boto3.client('ce')

def analyze_anomaly(anomaly_id):
    # Get anomaly details
    anomaly = ce.get_anomalies(
        AnomalyId=anomaly_id
    )['Anomalies'][0]
    
    # Get detailed cost breakdown
    start_date = anomaly['AnomalyStartDate']
    end_date = anomaly['AnomalyEndDate']
    
    cost_details = ce.get_cost_and_usage(
        TimePeriod={'Start': start_date, 'End': end_date},
        Granularity='DAILY',
        Metrics=['BlendedCost'],
        GroupBy=[
            {'Type': 'DIMENSION', 'Key': 'SERVICE'},
            {'Type': 'DIMENSION', 'Key': 'USAGE_TYPE'}
        ],
        Filter={
            'Dimensions': {
                'Key': 'SERVICE',
                'Values': [anomaly['DimensionValue']]
            }
        }
    )
    
    return {
        'anomaly': anomaly,
        'breakdown': cost_details
    }

Automated Remediation

Common Remediation Actions

Implement automatic responses:

def remediate_anomaly(event, context):
    anomaly = json.loads(event['Records'][0]['Sns']['Message'])
    service = anomaly['dimensionValue']
    impact = anomaly['impact']['totalImpact']
    
    # Auto-remediation for known patterns
    if service == 'Amazon EC2' and impact > 500:
        # Find and stop idle instances
        ec2 = boto3.client('ec2')
        instances = ec2.describe_instances(
            Filters=[
                {'Name': 'tag:Environment', 'Values': ['development']},
                {'Name': 'instance-state-name', 'Values': ['running']}
            ]
        )
        
        for reservation in instances['Reservations']:
            for instance in reservation['Instances']:
                # Check CPU utilization
                cpu = get_instance_cpu(instance['InstanceId'])
                if cpu < 5:
                    ec2.stop_instances(InstanceIds=[instance['InstanceId']])
                    log_action(f"Stopped idle instance: {instance['InstanceId']}")
    
    elif service == 'Amazon S3':
        # Check for unusual data transfer
        analyze_s3_transfer(anomaly)
    
    elif service == 'AWS Lambda':
        # Check for runaway functions
        analyze_lambda_invocations(anomaly)
    
    return {'statusCode': 200, 'remediated': True}

Dashboard and Reporting

CloudWatch Dashboard

CostAnomalyDashboard:
  Type: AWS::CloudWatch::Dashboard
  Properties:
    DashboardName: cost-anomaly-overview
    DashboardBody: !Sub |
      {
        "widgets": [
          {
            "type": "metric",
            "properties": {
              "metrics": [
                ["AWS/Billing", "EstimatedCharges", "ServiceName", "AmazonEC2"],
                ["...", "AmazonS3"],
                ["...", "AWSLambda"]
              ],
              "period": 86400,
              "stat": "Maximum",
              "title": "Daily Spending by Service"
            }
          },
          {
            "type": "text",
            "properties": {
              "markdown": "## Recent Anomalies\nView details in [Cost Anomaly Detection](https://console.aws.amazon.com/cost-management/home#/anomaly-detection)"
            }
          }
        ]
      }

Working with Warqline

We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.

Talk to an engineer

Best Practices

Implementation Checklist

  1. Create monitors for each account tier
  2. Set appropriate thresholds per environment
  3. Configure multiple notification channels
  4. Implement automated remediation for common issues
  5. Regular threshold review and adjustment
  6. Train teams on anomaly response procedures

Threshold Guidelines

Environment Percentage Absolute
Production 10% $200+
Staging 25% $50+
Development 50% $25+

Conclusion

AWS Cost Anomaly Detection provides essential visibility into spending patterns. Combined with Warqline's comprehensive cost analytics, organizations can proactively manage cloud costs and prevent budget overruns.