AWS DevOps Guru: Proactive Operational Insights
Leverage AWS DevOps Guru's ML-powered insights to detect operational issues before they impact your applications.
AWS DevOps Guru uses machine learning to detect abnormal application behavior, identify operational issues, and recommend remediation actions. This guide covers implementing DevOps Guru for proactive operations.
Understanding DevOps Guru
Core Capabilities
DevOps Guru provides:
- Anomaly Detection: ML identifies unusual behavior patterns
- Root Cause Analysis: Correlates metrics to find issue sources
- Recommendations: Actionable guidance for remediation
- Integration: Works with CloudWatch, X-Ray, and CloudFormation
How It Works
- Data Collection: Gathers CloudWatch metrics and logs
- ML Analysis: Applies machine learning models
- Correlation: Links related anomalies
- Notification: Alerts through SNS or EventBridge
- Guidance: Provides specific recommendations
Enabling DevOps Guru
Basic Setup
Enable DevOps Guru via console or CLI:
aws devops-guru update-service-integration \
--service-integration '{"opsCenter":{"optInStatus":"ENABLED"}}'
CloudFormation Configuration
DevOpsGuruNotificationChannel:
Type: AWS::DevOpsGuru::NotificationChannel
Properties:
Config:
Sns:
TopicArn: !Ref OpsAlertsTopic
DevOpsGuruResourceCollection:
Type: AWS::DevOpsGuru::ResourceCollection
Properties:
ResourceCollectionFilter:
CloudFormation:
StackNames:
- "production-*"
- "shared-infrastructure"
Terraform Setup
resource "aws_devopsguru_notification_channel" "main" {
sns {
topic_arn = aws_sns_topic.devops_alerts.arn
}
}
resource "aws_devopsguru_resource_collection" "production" {
type = "CLOUD_FORMATION"
cloudformation {
stack_names = ["production-*", "api-*"]
}
}
resource "aws_devopsguru_service_integration" "main" {
logs_anomaly_detection {
opt_in_status = "ENABLED"
}
kms_server_side_encryption {
kms_key_id = aws_kms_key.devops_guru.arn
opt_in_status = "ENABLED"
type = "CUSTOMER_MANAGED_KEY"
}
}
Resource Coverage
Coverage Types
CloudFormation Stack Coverage
Monitor all resources in stacks:
ResourceCollection:
CloudFormation:
StackNames:
- production-api
- production-web
- shared-database
Tag-Based Coverage
Monitor resources by tags:
ResourceCollection:
Tags:
- AppBoundaryKey: DevOpsGuruCoverage
TagValues:
- "enabled"
- "critical"
Account-Wide Coverage
Monitor everything:
{
"ResourceCollectionFilter": {
"All": {}
}
}
Insight Types
Reactive Insights
Issues detected after they occur:
- Application latency spikes
- Error rate increases
- Resource exhaustion
- Service failures
Proactive Insights
Predictions before issues occur:
- Capacity limits approaching
- Configuration drift detection
- Performance degradation trends
- Cost optimization opportunities
Understanding Insight Details
import boto3
devops_guru = boto3.client('devops-guru')
def analyze_insight(insight_id):
insight = devops_guru.describe_insight(Id=insight_id)
analysis = {
'id': insight['ProactiveInsight']['Id'],
'severity': insight['ProactiveInsight']['Severity'],
'status': insight['ProactiveInsight']['Status'],
'prediction': insight['ProactiveInsight']['PredictionTimeRange'],
'anomalies': []
}
# Get related anomalies
anomalies = devops_guru.list_anomalies_for_insight(
InsightId=insight_id,
Type='PROACTIVE'
)
for anomaly in anomalies['ProactiveAnomalies']:
analysis['anomalies'].append({
'id': anomaly['Id'],
'metric': anomaly['SourceDetails']['CloudWatchMetrics'][0],
'severity': anomaly['Severity'],
'description': anomaly['Description']
})
return analysis
Recommendations
Types of Recommendations
DevOps Guru provides:
-
Performance Recommendations
- Instance right-sizing
- Cache optimization
- Query optimization
-
Reliability Recommendations
- Multi-AZ deployment
- Auto Scaling configuration
- Backup improvements
-
Cost Recommendations
- Reserved capacity
- Spot instance usage
- Resource cleanup
Implementing Recommendations
def process_recommendations(insight_id):
recommendations = devops_guru.list_recommendations(
InsightId=insight_id
)
for rec in recommendations['Recommendations']:
print(f"Category: {rec['Category']}")
print(f"Name: {rec['Name']}")
print(f"Description: {rec['Description']}")
print(f"Link: {rec['Link']}")
# Log for tracking
log_recommendation(rec)
# Create ticket for team
if rec['Category'] == 'PERFORMANCE':
create_jira_ticket(rec)
EventBridge Integration
Event Patterns
DevOpsGuruEventRule:
Type: AWS::Events::Rule
Properties:
Name: devops-guru-insights
EventPattern:
source:
- aws.devops-guru
detail-type:
- DevOps Guru New Insight Open
- DevOps Guru Insight Severity Upgraded
Targets:
- Arn: !Ref AlertSNSTopic
Id: sns-alert
- Arn: !GetAtt ProcessingLambda.Arn
Id: process-insight
Processing Events
def lambda_handler(event, context):
detail = event['detail']
insight_type = detail['insightType']
severity = detail['insightSeverity']
if severity == 'high':
# Page on-call
pagerduty_alert(detail)
elif severity == 'medium':
# Slack notification
slack_notify(detail)
# Always create OpsItem
ssm = boto3.client('ssm')
ssm.create_ops_item(
Title=f"DevOps Guru Insight: {detail['insightDescription']}",
Description=json.dumps(detail, indent=2),
Source='DevOpsGuru',
Severity=map_severity(severity),
OperationalData={
'/aws/dedup': {
'Value': detail['insightId'],
'Type': 'SearchableString'
}
}
)
return {'statusCode': 200}
Dashboard and Monitoring
CloudWatch Dashboard
DevOpsGuruDashboard:
Type: AWS::CloudWatch::Dashboard
Properties:
DashboardName: devops-guru-overview
DashboardBody: !Sub |
{
"widgets": [
{
"type": "metric",
"properties": {
"title": "DevOps Guru Insights",
"metrics": [
["AWS/DevOpsGuru", "InsightCount", "InsightType", "PROACTIVE"],
["...", "REACTIVE"]
],
"period": 86400,
"stat": "Sum"
}
},
{
"type": "metric",
"properties": {
"title": "Anomalies Detected",
"metrics": [
["AWS/DevOpsGuru", "AnomalyCount"]
],
"period": 3600,
"stat": "Sum"
}
}
]
}
Cost Optimization
Managing DevOps Guru Costs
Optimize coverage efficiently:
# Focus on critical resources
ResourceCollection:
Tags:
- AppBoundaryKey: Criticality
TagValues:
- "critical"
- "high"
Cost Monitoring
DevOpsGuruCostAlarm:
Type: AWS::CloudWatch::Alarm
Properties:
AlarmName: devops-guru-monthly-limit
MetricName: EstimatedCharges
Namespace: AWS/Billing
Dimensions:
- Name: ServiceName
Value: AmazonDevOpsGuru
Statistic: Maximum
Period: 86400
Threshold: 100
ComparisonOperator: GreaterThanThreshold
Multi-Account Strategy
Centralized Monitoring
# Management account
resource "aws_devopsguru_event_sources_config" "main" {
event_sources {
amazon_code_guru_profiler {
status = "ENABLED"
}
}
}
# Enable in member accounts
resource "aws_devopsguru_resource_collection" "member" {
for_each = toset(var.member_account_ids)
provider = aws.member[each.key]
type = "TAGS"
tags {
app_boundary_key = "DevOpsGuru"
tag_values = ["enabled"]
}
}
Working with Warqline
We are a cloud engineering consultancy and an official AWS and Google Cloud partner. If you are running this in production and want a second pair of eyes, we scope work in a free 45-minute technical call: you describe what you are running and what worries you, and we tell you what we would look at first.
Best Practices
Implementation Checklist
- Start Small: Begin with critical stacks
- Configure Notifications: Set up SNS and EventBridge
- Review Regularly: Check insights dashboard weekly
- Track Recommendations: Log and implement suggestions
- Tune Coverage: Adjust based on signal quality
- Integrate Tools: Connect with existing ITSM
Severity Response Guide
| Severity | Response Time | Action |
|---|---|---|
| High | 15 min | Page on-call |
| Medium | 1 hour | Slack alert |
| Low | Next business day | Email/ticket |
Conclusion
AWS DevOps Guru provides ML-powered operational insights that help teams identify and resolve issues proactively. Combined with a Well-Architected review, organizations can achieve operational excellence in their cloud environments.