SRE on Google Cloud: Implementing SLOs, Error Budgets, and Incident Response

Site Reliability Engineering isn't just a job title — it's a set of practices that make services more reliable without sacrificing development velocity. This guide shows how to implement SLOs, error budgets, and incident response workflows on Google Cloud using Cloud Monitoring, Cloud Alerting, and the SLO API.

Google invented Site Reliability Engineering as a practice, and Google Cloud has native tooling for implementing SRE concepts that most cloud providers don't have. Yet most teams using GCP skip the SLO API entirely and just set up basic uptime checks. This article covers the full SRE toolkit on GCP: defining meaningful SLIs, measuring them via Cloud Monitoring, setting error budgets, building burn rate alerts, and running effective incident response.

Whether you're starting an SRE practice from scratch or trying to make your existing on-call rotation less painful, these patterns are directly applicable to production GKE workloads, Cloud Run services, and stateful backends.

SLI → SLO → Error Budget: A Quick Refresher

If you want to skip the fundamentals, jump to the implementation sections. But it's worth aligning on terminology since these terms get used inconsistently.

Service Level Indicator (SLI): A quantitative measure of some aspect of service behavior. Good SLIs are measurable with real data, and they reflect whether users are happy. Common SLIs:

  • Availability: percentage of requests that succeed (HTTP 2xx/3xx)
  • Latency: percentage of requests completed under a threshold
  • Throughput: number of requests processed per second
  • Error rate: percentage of requests that fail

Service Level Objective (SLO): A target value for an SLI over a time window. "99.9% of requests return successfully over a 30-day rolling window." SLOs define what "good enough" looks like.

Error Budget: The amount of unreliability allowed by the SLO. A 99.9% SLO has a 0.1% error budget — that's ~43 minutes of downtime per month. The error budget is spent on outages, planned maintenance, and experiments. When you run out, development velocity slows until the budget recovers.

The crucial insight: error budgets turn reliability into a team objective rather than an operations burden. "We have 30 minutes of error budget left this month" is a concrete statement everyone can act on.

Cloud Monitoring SLO API

Cloud Monitoring has native SLO support that computes error budget consumption automatically. You define an SLO, it computes compliance, and you can alert on burn rate.

Creating an SLO for a GKE Service

First, create a monitoring service to represent your application:

# Create a custom service (if your app isn't using Istio/Anthos Service Mesh)
gcloud monitoring services create   --display-name="Payment API"   --project=my-project   --custom

# Note the service ID from the output: projects/my-project/services/custom:my-service-id

Then define an SLO on that service:

# Create a request-based availability SLO
gcloud monitoring slos create   --display-name="Payment API 99.9% Availability"   --service=projects/my-project/services/custom:payment-api   --project=my-project   --goal=0.999   --rolling-period-days=30   --request-based   --good-service-filter='metric.type="run.googleapis.com/request_count" AND resource.type="cloud_run_revision" AND metric.labels.response_code_class="2xx"'   --total-service-filter='metric.type="run.googleapis.com/request_count" AND resource.type="cloud_run_revision"'

Latency-Based SLO

For a latency SLO (e.g., 95% of requests under 500ms):

{
  "displayName": "Payment API P95 Latency < 500ms",
  "goal": 0.95,
  "rollingPeriod": "2592000s",
  "requestBased": {
    "goodTotalRatio": {
      "goodServiceFilter": "metric.type=\"run.googleapis.com/request_latencies\" AND metric.labels.response_code_class=\"2xx\" AND metric.labels.percentile=\"95\" AND value.distribution_value.mean < 500",
      "totalServiceFilter": "metric.type=\"run.googleapis.com/request_latencies\" AND metric.labels.response_code_class=\"2xx\""
    }
  }
}

For GKE workloads instrumented with Prometheus, you can use the prometheus-to-cloud-monitoring adapter to create SLOs based on custom metrics:

# prometheus-adapter-config.yaml
- seriesQuery: 'http_requests_total{namespace!="",pod!=""}'
  resources:
    overrides:
      namespace: {resource: "namespace"}
      pod: {resource: "pod"}
  name:
    matches: "http_requests_total"
    as: "http_requests_total"
  metricsQuery: 'sum(rate(<<.Series>>{<<.LabelMatchers>>}[5m])) by (<<.GroupBy>>)'

Building Error Budget Burn Rate Alerts

A burn rate alert fires when you're consuming your error budget faster than sustainable. The key insight from Google's SRE book: if you're burning your 30-day budget in 1 hour, that's 720x burn rate — you need to know immediately. Burning at 5x is still alarming but less urgent.

Effective burn rate alerting uses multiple windows to reduce false positives:

Alert Short Window Long Window Burn Rate Response
Page (critical) 5 min 1 hour 14.4x Wake up the on-call
Page (high) 30 min 6 hours 6x On-call responds within 30 min
Ticket (medium) 6 hours 3 days 3x Fix within 24 hours
Ticket (low) 3 days — 1x Fix this sprint

Creating Burn Rate Alerts via Cloud Monitoring

# Create an alerting policy for critical burn rate
cat > burn-rate-alert.json << 'EOF'
{
  "displayName": "Payment API - Critical Error Budget Burn",
  "conditions": [
    {
      "displayName": "Error budget burn rate > 14.4x (5-min window)",
      "conditionThreshold": {
        "filter": "select_slo_burn_rate(\"projects/my-project/services/payment-api/serviceLevelObjectives/slo-id\", 300)",
        "comparison": "COMPARISON_GT",
        "threshold_value": 14.4,
        "duration": "0s",
        "trigger": {
          "count": 1
        }
      }
    },
    {
      "displayName": "Error budget burn rate > 14.4x (1-hour window)",
      "conditionThreshold": {
        "filter": "select_slo_burn_rate(\"projects/my-project/services/payment-api/serviceLevelObjectives/slo-id\", 3600)",
        "comparison": "COMPARISON_GT",
        "threshold_value": 14.4,
        "duration": "0s",
        "trigger": {
          "count": 1
        }
      }
    }
  ],
  "combiner": "AND",
  "alertStrategy": {
    "notificationRateLimit": {
      "period": "300s"
    },
    "autoClose": "1800s"
  },
  "notificationChannels": [
    "projects/my-project/notificationChannels/pagerduty-channel"
  ],
  "severity": "CRITICAL"
}
EOF

gcloud monitoring policies create   --policy-from-file=burn-rate-alert.json   --project=my-project

Setting Up Notification Channels

# Create a PagerDuty notification channel
gcloud monitoring channels create   --display-name="PagerDuty On-Call"   --type=pagerduty   --channel-labels=service_key=YOUR_PAGERDUTY_KEY   --project=my-project

# Create a Slack notification channel
gcloud monitoring channels create   --display-name="Slack #incidents"   --type=slack   --channel-labels=channel_name=#incidents   --project=my-project

Cloud Monitoring Uptime Checks

For external SLI measurement (from outside your network), Cloud Monitoring uptime checks probe your service from multiple Google PoPs:

# Create an HTTP uptime check
gcloud monitoring uptime-checks create http   --display-name="Payment API Health"   --hostname=api.example.com   --path=/health   --port=443   --use-ssl   --period=60s   --timeout=10s   --regions=EUROPE,USA

# Create an alerting policy for uptime check failure
gcloud monitoring policies create   --display-name="Payment API Down"   --condition-filter='resource.type="uptime_url" AND metric.type="monitoring.googleapis.com/uptime_check/check_passed" AND metric.labels.check_id="check_id_here"'   --condition-comparison=COMPARISON_LT   --condition-threshold=1.0   --condition-duration=120s   --notification-channels=pagerduty-channel-id

GKE SLO Implementation with Prometheus

For GKE workloads, you likely want SLOs based on application-level metrics rather than infrastructure metrics. Instrument your service with Prometheus client libraries:

# Python example with prometheus_client
from prometheus_client import Counter, Histogram, start_http_server
import time

REQUEST_COUNT = Counter(
    'http_requests_total',
    'Total HTTP requests',
    ['method', 'endpoint', 'status_code']
)

REQUEST_LATENCY = Histogram(
    'http_request_duration_seconds',
    'HTTP request duration',
    ['method', 'endpoint'],
    buckets=[0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5]
)

def handle_request(method, endpoint):
    start = time.time()
    try:
        result = process_request(method, endpoint)
        REQUEST_COUNT.labels(method=method, endpoint=endpoint, status_code=200).inc()
        return result
    except Exception as e:
        REQUEST_COUNT.labels(method=method, endpoint=endpoint, status_code=500).inc()
        raise
    finally:
        REQUEST_LATENCY.labels(method=method, endpoint=endpoint).observe(time.time() - start)

Then use Prometheus rules to compute SLI ratios:

# prometheus-rules.yaml
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: payment-api-sli
  namespace: production
spec:
  groups:
  - name: sli.rules
    rules:
    # 5-minute error rate (for burn rate alerting)
    - record: job:http_requests:error_rate5m
      expr: |
        sum(rate(http_requests_total{status_code=~"5.."}[5m])) by (job)
        /
        sum(rate(http_requests_total[5m])) by (job)

    # 1-hour error rate (for SLO compliance)
    - record: job:http_requests:error_rate1h
      expr: |
        sum(rate(http_requests_total{status_code=~"5.."}[1h])) by (job)
        /
        sum(rate(http_requests_total[1h])) by (job)

    # P95 latency
    - record: job:http_request_duration_seconds:p95_5m
      expr: |
        histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[5m])) by (le, job))

Incident Response Playbooks

Good SLO alerting only helps if your on-call knows what to do when an alert fires. Document runbooks for each alert condition.

Incident Response Structure

1. ACKNOWLEDGE (< 5 min)
   - Acknowledge the alert in PagerDuty/Opsgenie
   - Post in #incidents: "I'm on this. Investigating."
   - Start the incident timer

2. ASSESS (< 15 min)
   - What SLI is degraded?
   - What is the burn rate?
   - How many users affected?
   - Has anything changed recently? (deployments, config changes)

3. MITIGATE (< 30 min for P1)
   - Can we roll back the last deployment?
   - Can we scale up the service?
   - Can we route traffic away from the bad instance?
   - If unclear: escalate to senior engineer

4. RESOLVE
   - Confirm SLI returns to normal
   - Post resolution in #incidents
   - Start post-mortem doc

5. REVIEW (within 48 hours)
   - Blameless post-mortem
   - Root cause
   - Action items with owners and deadlines

GKE Incident Investigation Commands

# Check for recent deployments that might have caused issues
kubectl rollout history deployment/payment-api -n production

# Roll back to previous version if needed
kubectl rollout undo deployment/payment-api -n production

# Check pod status across nodes
kubectl get pods -n production -o wide | grep payment-api

# Check recent pod events
kubectl describe pods -n production -l app=payment-api | grep -A 20 "Events:"

# Check logs from the last 30 minutes
kubectl logs -n production -l app=payment-api --since=30m --tail=100

# Check if HPA is throttling the service
kubectl describe hpa payment-api-hpa -n production

# Check node pressure
kubectl describe nodes | grep -A 5 "Conditions:"

Error Budget Policies

An error budget policy defines what happens when the error budget is exhausted. Without a policy, the error budget is just a number on a dashboard.

Example error budget policy:

Error Budget Policy — Payment API

If remaining error budget drops below 25%:
  - Engineering lead is notified daily
  - Feature work continues, but risky changes require extra review

If remaining error budget drops below 10%:
  - Development velocity is reduced
  - Only bug fixes and reliability improvements are deployed
  - No experiments or large feature launches

If error budget is exhausted (0% remaining):
  - Feature deployments are frozen
  - All engineering effort shifts to reliability work
  - Budget review with engineering and product leadership within 48 hours
  - Development resumes when budget is > 10% remaining or 30-day window resets

Document this policy, get product and engineering leadership to sign off, and reference it in your runbooks.

SLO Dashboards in Cloud Monitoring

Create a Cloud Monitoring dashboard to visualize error budget consumption:

# Create an SLO dashboard
gcloud monitoring dashboards create   --config-from-file=slo-dashboard.json   --project=my-project

Key widgets to include:

  • Current SLO compliance (gauge chart)
  • Error budget remaining % (gauge with red zone below 10%)
  • Error budget burn rate over time (line chart with 1x, 6x, 14.4x reference lines)
  • Request error rate (line chart)
  • P50/P95/P99 latency (line charts)
  • Active incidents table

For the observability stack that feeds into SLOs, see our GCP observability guide. For securing the GKE clusters running these services, see our GKE security hardening guide.