AWS DevOps Guru: Proactive Operational Insights

Leverage AWS DevOps Guru's ML-powered insights to detect operational issues before they impact your applications.

AWS DevOps Guru uses machine learning to detect abnormal application behavior, identify operational issues, and recommend remediation actions. This guide covers implementing DevOps Guru for proactive operations.

Understanding DevOps Guru

Core Capabilities

DevOps Guru provides:

  • Anomaly Detection: ML identifies unusual behavior patterns
  • Root Cause Analysis: Correlates metrics to find issue sources
  • Recommendations: Actionable guidance for remediation
  • Integration: Works with CloudWatch, X-Ray, and CloudFormation

How It Works

  1. Data Collection: Gathers CloudWatch metrics and logs
  2. ML Analysis: Applies machine learning models
  3. Correlation: Links related anomalies
  4. Notification: Alerts through SNS or EventBridge
  5. Guidance: Provides specific recommendations

Enabling DevOps Guru

Basic Setup

Enable DevOps Guru via console or CLI:

aws devops-guru update-service-integration \
  --service-integration '{"opsCenter":{"optInStatus":"ENABLED"}}'

CloudFormation Configuration

DevOpsGuruNotificationChannel:
  Type: AWS::DevOpsGuru::NotificationChannel
  Properties:
    Config:
      Sns:
        TopicArn: !Ref OpsAlertsTopic

DevOpsGuruResourceCollection:
  Type: AWS::DevOpsGuru::ResourceCollection
  Properties:
    ResourceCollectionFilter:
      CloudFormation:
        StackNames:
          - "production-*"
          - "shared-infrastructure"

Terraform Setup

resource "aws_devopsguru_notification_channel" "main" {
  sns {
    topic_arn = aws_sns_topic.devops_alerts.arn
  }
}

resource "aws_devopsguru_resource_collection" "production" {
  type = "CLOUD_FORMATION"
  
  cloudformation {
    stack_names = ["production-*", "api-*"]
  }
}

resource "aws_devopsguru_service_integration" "main" {
  logs_anomaly_detection {
    opt_in_status = "ENABLED"
  }
  
  kms_server_side_encryption {
    kms_key_id     = aws_kms_key.devops_guru.arn
    opt_in_status  = "ENABLED"
    type           = "CUSTOMER_MANAGED_KEY"
  }
}

Resource Coverage

Coverage Types

CloudFormation Stack Coverage

Monitor all resources in stacks:

ResourceCollection:
  CloudFormation:
    StackNames:
      - production-api
      - production-web
      - shared-database

Tag-Based Coverage

Monitor resources by tags:

ResourceCollection:
  Tags:
    - AppBoundaryKey: DevOpsGuruCoverage
      TagValues:
        - "enabled"
        - "critical"

Account-Wide Coverage

Monitor everything:

{
  "ResourceCollectionFilter": {
    "All": {}
  }
}

Insight Types

Reactive Insights

Issues detected after they occur:

  • Application latency spikes
  • Error rate increases
  • Resource exhaustion
  • Service failures

Proactive Insights

Predictions before issues occur:

  • Capacity limits approaching
  • Configuration drift detection
  • Performance degradation trends
  • Cost optimization opportunities

Understanding Insight Details

import boto3

devops_guru = boto3.client('devops-guru')

def analyze_insight(insight_id):
    insight = devops_guru.describe_insight(Id=insight_id)
    
    analysis = {
        'id': insight['ProactiveInsight']['Id'],
        'severity': insight['ProactiveInsight']['Severity'],
        'status': insight['ProactiveInsight']['Status'],
        'prediction': insight['ProactiveInsight']['PredictionTimeRange'],
        'anomalies': []
    }
    
    # Get related anomalies
    anomalies = devops_guru.list_anomalies_for_insight(
        InsightId=insight_id,
        Type='PROACTIVE'
    )
    
    for anomaly in anomalies['ProactiveAnomalies']:
        analysis['anomalies'].append({
            'id': anomaly['Id'],
            'metric': anomaly['SourceDetails']['CloudWatchMetrics'][0],
            'severity': anomaly['Severity'],
            'description': anomaly['Description']
        })
    
    return analysis

Recommendations

Types of Recommendations

DevOps Guru provides:

  1. Performance Recommendations

    • Instance right-sizing
    • Cache optimization
    • Query optimization
  2. Reliability Recommendations

    • Multi-AZ deployment
    • Auto Scaling configuration
    • Backup improvements
  3. Cost Recommendations

    • Reserved capacity
    • Spot instance usage
    • Resource cleanup

Implementing Recommendations

def process_recommendations(insight_id):
    recommendations = devops_guru.list_recommendations(
        InsightId=insight_id
    )
    
    for rec in recommendations['Recommendations']:
        print(f"Category: {rec['Category']}")
        print(f"Name: {rec['Name']}")
        print(f"Description: {rec['Description']}")
        print(f"Link: {rec['Link']}")
        
        # Log for tracking
        log_recommendation(rec)
        
        # Create ticket for team
        if rec['Category'] == 'PERFORMANCE':
            create_jira_ticket(rec)

EventBridge Integration

Event Patterns

DevOpsGuruEventRule:
  Type: AWS::Events::Rule
  Properties:
    Name: devops-guru-insights
    EventPattern:
      source:
        - aws.devops-guru
      detail-type:
        - DevOps Guru New Insight Open
        - DevOps Guru Insight Severity Upgraded
    Targets:
      - Arn: !Ref AlertSNSTopic
        Id: sns-alert
      - Arn: !GetAtt ProcessingLambda.Arn
        Id: process-insight

Processing Events

def lambda_handler(event, context):
    detail = event['detail']
    insight_type = detail['insightType']
    severity = detail['insightSeverity']
    
    if severity == 'high':
        # Page on-call
        pagerduty_alert(detail)
        
    elif severity == 'medium':
        # Slack notification
        slack_notify(detail)
        
    # Always create OpsItem
    ssm = boto3.client('ssm')
    ssm.create_ops_item(
        Title=f"DevOps Guru Insight: {detail['insightDescription']}",
        Description=json.dumps(detail, indent=2),
        Source='DevOpsGuru',
        Severity=map_severity(severity),
        OperationalData={
            '/aws/dedup': {
                'Value': detail['insightId'],
                'Type': 'SearchableString'
            }
        }
    )
    
    return {'statusCode': 200}

Dashboard and Monitoring

CloudWatch Dashboard

DevOpsGuruDashboard:
  Type: AWS::CloudWatch::Dashboard
  Properties:
    DashboardName: devops-guru-overview
    DashboardBody: !Sub |
      {
        "widgets": [
          {
            "type": "metric",
            "properties": {
              "title": "DevOps Guru Insights",
              "metrics": [
                ["AWS/DevOpsGuru", "InsightCount", "InsightType", "PROACTIVE"],
                ["...", "REACTIVE"]
              ],
              "period": 86400,
              "stat": "Sum"
            }
          },
          {
            "type": "metric",
            "properties": {
              "title": "Anomalies Detected",
              "metrics": [
                ["AWS/DevOpsGuru", "AnomalyCount"]
              ],
              "period": 3600,
              "stat": "Sum"
            }
          }
        ]
      }

Cost Optimization

Managing DevOps Guru Costs

Optimize coverage efficiently:

# Focus on critical resources
ResourceCollection:
  Tags:
    - AppBoundaryKey: Criticality
      TagValues:
        - "critical"
        - "high"

Cost Monitoring

DevOpsGuruCostAlarm:
  Type: AWS::CloudWatch::Alarm
  Properties:
    AlarmName: devops-guru-monthly-limit
    MetricName: EstimatedCharges
    Namespace: AWS/Billing
    Dimensions:
      - Name: ServiceName
        Value: AmazonDevOpsGuru
    Statistic: Maximum
    Period: 86400
    Threshold: 100
    ComparisonOperator: GreaterThanThreshold

Multi-Account Strategy

Centralized Monitoring

# Management account
resource "aws_devopsguru_event_sources_config" "main" {
  event_sources {
    amazon_code_guru_profiler {
      status = "ENABLED"
    }
  }
}

# Enable in member accounts
resource "aws_devopsguru_resource_collection" "member" {
  for_each = toset(var.member_account_ids)
  
  provider = aws.member[each.key]
  
  type = "TAGS"
  
  tags {
    app_boundary_key = "DevOpsGuru"
    tag_values       = ["enabled"]
  }
}

Working with Warqline

We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.

Talk to an engineer

Best Practices

Implementation Checklist

  1. Start Small: Begin with critical stacks
  2. Configure Notifications: Set up SNS and EventBridge
  3. Review Regularly: Check insights dashboard weekly
  4. Track Recommendations: Log and implement suggestions
  5. Tune Coverage: Adjust based on signal quality
  6. Integrate Tools: Connect with existing ITSM

Severity Response Guide

Severity Response Time Action
High 15 min Page on-call
Medium 1 hour Slack alert
Low Next business day Email/ticket

Conclusion

AWS DevOps Guru provides ML-powered operational insights that help teams identify and resolve issues proactively. Combined with a Well-Architected review, organizations can achieve operational excellence in their cloud environments.