GCP Observability: Cloud Monitoring, Logging, and Tracing Guide

Observability isn't just about having logs and metrics — it's about being able to answer 'what's wrong and why' in under 10 minutes when an alert fires at 3am. This guide covers the complete GCP observability stack: Cloud Monitoring metrics, Cloud Logging structured logs, Cloud Trace distributed tracing, and Cloud Profiler.

When something breaks in production, you have one goal: figure out what's wrong and fix it as fast as possible. Observability is the property of your system that makes that possible. Not alerting — alerting just tells you something is wrong. Observability is what lets you understand the system state from outside, using logs, metrics, and traces.

Google Cloud's observability suite has improved dramatically over the past few years. Cloud Operations (formerly Stackdriver) now covers metrics, logs, traces, profiling, and error reporting with deep GKE integration. This guide covers the practical setup for each component, how they work together, and the alerting patterns that actually reduce mean time to resolution.

Cloud Monitoring: Metrics and Dashboards

Built-In GKE Metrics

When you create a GKE cluster with Cloud Monitoring enabled, you get a rich set of pre-built dashboards and metrics automatically:

# Enable Cloud Monitoring for an existing cluster
gcloud container clusters update my-cluster   --region=europe-west4   --enable-managed-prometheus  # Preferred for GKE workloads

Key automatically collected metrics:

  • container/cpu/core_usage_time — Container CPU usage
  • container/memory/used_bytes — Container memory usage
  • container/restart_count — Container restarts (critical for crash loop detection)
  • kubernetes.io/container/cpu/request_utilization — CPU request utilization %
  • kubernetes.io/container/memory/request_utilization — Memory request utilization %

Managed Prometheus on GKE

Google's Managed Service for Prometheus (GMP) lets you use PromQL and Grafana with metrics stored in Cloud Monitoring — no Prometheus cluster to operate:

# Enable scraping for your application pods
apiVersion: monitoring.googleapis.com/v1
kind: PodMonitoring
metadata:
  name: payment-api-monitoring
  namespace: production
spec:
  selector:
    matchLabels:
      app: payment-api
  endpoints:
  - port: metrics
    interval: 30s
    path: /metrics
# Verify metrics are being collected
gcloud monitoring time-series list   --filter='metric.type="prometheus.googleapis.com/http_requests_total/counter"'   --project=my-project   --freshness=PT5M

Custom Dashboards via Terraform

resource "google_monitoring_dashboard" "service_dashboard" {
  project        = var.project_id
  dashboard_json = jsonencode({
    displayName = "Payment API - Service Dashboard"
    gridLayout = {
      columns = 12
      widgets = [
        {
          title = "Request Rate"
          xyChart = {
            dataSets = [{
              timeSeriesQuery = {
                timeSeriesFilter = {
                  filter = "metric.type=\"run.googleapis.com/request_count\" AND resource.type=\"cloud_run_revision\""
                  aggregation = {
                    alignmentPeriod = "60s"
                    perSeriesAligner = "ALIGN_RATE"
                    crossSeriesReducer = "REDUCE_SUM"
                  }
                }
              }
            }]
          }
        },
        {
          title = "P95 Latency"
          xyChart = {
            dataSets = [{
              timeSeriesQuery = {
                timeSeriesFilter = {
                  filter = "metric.type=\"run.googleapis.com/request_latencies\" AND resource.type=\"cloud_run_revision\""
                  aggregation = {
                    alignmentPeriod = "60s"
                    perSeriesAligner = "ALIGN_PERCENTILE_95"
                  }
                }
              }
            }]
          }
        }
      ]
    }
  })
}

Alert Policies for Common GKE Problems

# Alert on container restart loops
gcloud monitoring policies create   --display-name="GKE Container Restart Loop"   --condition-filter='resource.type="k8s_container" AND metric.type="kubernetes.io/container/restart_count"'   --condition-comparison=COMPARISON_GT   --condition-threshold=5   --condition-duration=60s   --aggregation-aligner=ALIGN_RATE   --aggregation-period=60s   --notification-channels=projects/my-project/notificationChannels/slack-channel

# Alert on high memory utilization
gcloud monitoring policies create   --display-name="GKE Memory Request Utilization > 85%"   --condition-filter='resource.type="k8s_container" AND metric.type="kubernetes.io/container/memory/request_utilization"'   --condition-comparison=COMPARISON_GT   --condition-threshold=0.85   --condition-duration=300s

Cloud Logging: Structured Logs at Scale

Why Structured Logging Matters

Structured logs (JSON) enable fast filtering, metric extraction, and alerting based on log content. Unstructured text logs are searchable but not queryable efficiently.

Bad (unstructured):

2025-01-15 10:23:45 ERROR Payment failed for user 12345: insufficient funds

Good (structured):

{
  "severity": "ERROR",
  "message": "Payment failed",
  "user_id": "12345",
  "error_code": "INSUFFICIENT_FUNDS",
  "amount_cents": 4999,
  "currency": "EUR",
  "trace": "projects/my-project/traces/abc123",
  "httpRequest": {
    "requestMethod": "POST",
    "requestUrl": "/api/payments",
    "responseStatusCode": 402,
    "latency": "0.245s"
  }
}

Structured Logging in Python

import logging
import json
from google.cloud import logging as cloud_logging

# Initialize Cloud Logging client
client = cloud_logging.Client()
client.setup_logging()

class StructuredLogger:
    def __init__(self, name: str):
        self.logger = logging.getLogger(name)

    def info(self, message: str, **kwargs):
        self._log(logging.INFO, message, **kwargs)

    def error(self, message: str, **kwargs):
        self._log(logging.ERROR, message, **kwargs)

    def _log(self, level: int, message: str, **kwargs):
        extra = {
            "json_fields": {
                "message": message,
                **kwargs
            }
        }
        self.logger.log(level, message, extra=extra)

# Usage
logger = StructuredLogger("payment-api")
logger.info(
    "Payment processed",
    user_id="12345",
    amount_cents=4999,
    currency="EUR",
    payment_method="card",
    duration_ms=245
)

Log-Based Metrics

Log-based metrics let you create Cloud Monitoring metrics from log entries, turning log data into alertable signals:

# Create a counter metric for payment failures
gcloud logging metrics create payment_failures   --description="Count of payment failures"   --log-filter='resource.type="cloud_run_revision" AND jsonPayload.error_code!="" AND severity="ERROR"'   --project=my-project

# Create a distribution metric for payment latency
gcloud logging metrics create payment_latency   --description="Distribution of payment processing latency"   --log-filter='resource.type="cloud_run_revision" AND jsonPayload.duration_ms!=""'   --project=my-project   --value-extractor='EXTRACT(jsonPayload.duration_ms)'   --bucket-type=LINEAR   --num-buckets=20   --width=50   --offset=0

Log Sinks for Long-Term Retention and SIEM

Cloud Logging retains logs for 30 days by default. For compliance or security analysis, export to Cloud Storage or BigQuery:

# Export all application logs to BigQuery for 1-year retention
gcloud logging sinks create prod-logs-bq   bigquery.googleapis.com/projects/my-project/datasets/application_logs   --log-filter='resource.type=("cloud_run_revision" OR "k8s_container") AND severity>=WARNING'   --project=my-project

# Export security-relevant logs to Cloud Storage
gcloud logging sinks create security-audit-gcs   storage.googleapis.com/my-security-logs-bucket   --log-filter='protoPayload.@type="type.googleapis.com/google.cloud.audit.AuditLog"'   --project=my-project

Cloud Trace: Distributed Tracing

Distributed tracing shows you exactly how a request flows through your microservices — which service calls which, how long each step takes, and where errors originate.

Auto-Instrumentation for GKE

The easiest way to get traces is the OpenTelemetry operator on GKE, which can auto-instrument Java, Python, and Node.js applications without code changes:

# Install the OpenTelemetry Operator
# First, apply the cert-manager (prerequisite)
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.11.0/cert-manager.yaml

# Install OpenTelemetry Operator
kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml

# Configure auto-instrumentation
apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
  name: my-instrumentation
  namespace: production
spec:
  exporter:
    endpoint: http://otel-collector:4317
  propagators:
    - tracecontext
    - baggage
    - b3
  python:
    env:
    - name: OTEL_EXPORTER_OTLP_ENDPOINT
      value: http://otel-collector:4317
  java:
    env:
    - name: OTEL_EXPORTER_OTLP_ENDPOINT
      value: http://otel-collector:4317

Manual Tracing in Python

from opentelemetry import trace
from opentelemetry.exporter.cloud_trace import CloudTraceSpanExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor

# Configure Cloud Trace exporter
provider = TracerProvider()
provider.add_span_processor(
    BatchSpanProcessor(CloudTraceSpanExporter())
)
trace.set_tracer_provider(provider)

tracer = trace.get_tracer("payment-api")

def process_payment(user_id: str, amount_cents: int) -> dict:
    with tracer.start_as_current_span("process_payment") as span:
        span.set_attribute("user.id", user_id)
        span.set_attribute("payment.amount_cents", amount_cents)

        # Call to external service — creates a child span
        with tracer.start_as_current_span("validate_card"):
            result = validate_card(user_id)
            span.set_attribute("payment.card_valid", result.valid)

        with tracer.start_as_current_span("charge_card"):
            charge = charge_card(user_id, amount_cents)

        return {"status": "success", "transaction_id": charge.id}

Cloud Profiler

Cloud Profiler continuously profiles CPU and memory usage in production with very low overhead (~1% CPU). It's particularly valuable for finding performance regressions introduced by code changes.

# Enable Cloud Profiler in your Python application
import googlecloudprofiler

try:
    googlecloudprofiler.start(
        service="payment-api",
        service_version="1.2.3",
        verbose=3,
    )
except (ValueError, NotImplementedError) as exc:
    print(f"Profiler not enabled: {exc}")

View profiles in the console: Cloud Profiler → Select service → Compare profiles over time. Look for flame graph changes between versions that indicate performance regressions.

Correlating Signals: Logs + Metrics + Traces

The real power of GCP observability is the correlation between signals. In Cloud Monitoring, when you click on an alert, it shows you the metric spike. From the metric spike, you can drill to the logs using the "View Logs" link. From a log entry with a trace ID, you can jump to the distributed trace.

This correlation is automatic when you:

  1. Use structured logs with trace and spanId fields populated
  2. Instrument with OpenTelemetry or Cloud Trace SDK
  3. Enable Cloud Monitoring on GKE with the Cloud Logging agent

For SLO-based alerting that uses these observability signals, see our SRE practices guide. For the CI/CD pipeline that deploys changes to monitored services, see our GCP DevOps guide.