GCP Observability: Cloud Monitoring, Logging, and Tracing Guide
Observability isn't just about having logs and metrics — it's about being able to answer 'what's wrong and why' in under 10 minutes when an alert fires at 3am. This guide covers the complete GCP observability stack: Cloud Monitoring metrics, Cloud Logging structured logs, Cloud Trace distributed tracing, and Cloud Profiler.
When something breaks in production, you have one goal: figure out what's wrong and fix it as fast as possible. Observability is the property of your system that makes that possible. Not alerting — alerting just tells you something is wrong. Observability is what lets you understand the system state from outside, using logs, metrics, and traces.
Google Cloud's observability suite has improved dramatically over the past few years. Cloud Operations (formerly Stackdriver) now covers metrics, logs, traces, profiling, and error reporting with deep GKE integration. This guide covers the practical setup for each component, how they work together, and the alerting patterns that actually reduce mean time to resolution.
Cloud Monitoring: Metrics and Dashboards
Built-In GKE Metrics
When you create a GKE cluster with Cloud Monitoring enabled, you get a rich set of pre-built dashboards and metrics automatically:
# Enable Cloud Monitoring for an existing cluster
gcloud container clusters update my-cluster --region=europe-west4 --enable-managed-prometheus # Preferred for GKE workloads
Key automatically collected metrics:
container/cpu/core_usage_time— Container CPU usagecontainer/memory/used_bytes— Container memory usagecontainer/restart_count— Container restarts (critical for crash loop detection)kubernetes.io/container/cpu/request_utilization— CPU request utilization %kubernetes.io/container/memory/request_utilization— Memory request utilization %
Managed Prometheus on GKE
Google's Managed Service for Prometheus (GMP) lets you use PromQL and Grafana with metrics stored in Cloud Monitoring — no Prometheus cluster to operate:
# Enable scraping for your application pods
apiVersion: monitoring.googleapis.com/v1
kind: PodMonitoring
metadata:
name: payment-api-monitoring
namespace: production
spec:
selector:
matchLabels:
app: payment-api
endpoints:
- port: metrics
interval: 30s
path: /metrics
# Verify metrics are being collected
gcloud monitoring time-series list --filter='metric.type="prometheus.googleapis.com/http_requests_total/counter"' --project=my-project --freshness=PT5M
Custom Dashboards via Terraform
resource "google_monitoring_dashboard" "service_dashboard" {
project = var.project_id
dashboard_json = jsonencode({
displayName = "Payment API - Service Dashboard"
gridLayout = {
columns = 12
widgets = [
{
title = "Request Rate"
xyChart = {
dataSets = [{
timeSeriesQuery = {
timeSeriesFilter = {
filter = "metric.type=\"run.googleapis.com/request_count\" AND resource.type=\"cloud_run_revision\""
aggregation = {
alignmentPeriod = "60s"
perSeriesAligner = "ALIGN_RATE"
crossSeriesReducer = "REDUCE_SUM"
}
}
}
}]
}
},
{
title = "P95 Latency"
xyChart = {
dataSets = [{
timeSeriesQuery = {
timeSeriesFilter = {
filter = "metric.type=\"run.googleapis.com/request_latencies\" AND resource.type=\"cloud_run_revision\""
aggregation = {
alignmentPeriod = "60s"
perSeriesAligner = "ALIGN_PERCENTILE_95"
}
}
}
}]
}
}
]
}
})
}
Alert Policies for Common GKE Problems
# Alert on container restart loops
gcloud monitoring policies create --display-name="GKE Container Restart Loop" --condition-filter='resource.type="k8s_container" AND metric.type="kubernetes.io/container/restart_count"' --condition-comparison=COMPARISON_GT --condition-threshold=5 --condition-duration=60s --aggregation-aligner=ALIGN_RATE --aggregation-period=60s --notification-channels=projects/my-project/notificationChannels/slack-channel
# Alert on high memory utilization
gcloud monitoring policies create --display-name="GKE Memory Request Utilization > 85%" --condition-filter='resource.type="k8s_container" AND metric.type="kubernetes.io/container/memory/request_utilization"' --condition-comparison=COMPARISON_GT --condition-threshold=0.85 --condition-duration=300s
Cloud Logging: Structured Logs at Scale
Why Structured Logging Matters
Structured logs (JSON) enable fast filtering, metric extraction, and alerting based on log content. Unstructured text logs are searchable but not queryable efficiently.
Bad (unstructured):
2025-01-15 10:23:45 ERROR Payment failed for user 12345: insufficient funds
Good (structured):
{
"severity": "ERROR",
"message": "Payment failed",
"user_id": "12345",
"error_code": "INSUFFICIENT_FUNDS",
"amount_cents": 4999,
"currency": "EUR",
"trace": "projects/my-project/traces/abc123",
"httpRequest": {
"requestMethod": "POST",
"requestUrl": "/api/payments",
"responseStatusCode": 402,
"latency": "0.245s"
}
}
Structured Logging in Python
import logging
import json
from google.cloud import logging as cloud_logging
# Initialize Cloud Logging client
client = cloud_logging.Client()
client.setup_logging()
class StructuredLogger:
def __init__(self, name: str):
self.logger = logging.getLogger(name)
def info(self, message: str, **kwargs):
self._log(logging.INFO, message, **kwargs)
def error(self, message: str, **kwargs):
self._log(logging.ERROR, message, **kwargs)
def _log(self, level: int, message: str, **kwargs):
extra = {
"json_fields": {
"message": message,
**kwargs
}
}
self.logger.log(level, message, extra=extra)
# Usage
logger = StructuredLogger("payment-api")
logger.info(
"Payment processed",
user_id="12345",
amount_cents=4999,
currency="EUR",
payment_method="card",
duration_ms=245
)
Log-Based Metrics
Log-based metrics let you create Cloud Monitoring metrics from log entries, turning log data into alertable signals:
# Create a counter metric for payment failures
gcloud logging metrics create payment_failures --description="Count of payment failures" --log-filter='resource.type="cloud_run_revision" AND jsonPayload.error_code!="" AND severity="ERROR"' --project=my-project
# Create a distribution metric for payment latency
gcloud logging metrics create payment_latency --description="Distribution of payment processing latency" --log-filter='resource.type="cloud_run_revision" AND jsonPayload.duration_ms!=""' --project=my-project --value-extractor='EXTRACT(jsonPayload.duration_ms)' --bucket-type=LINEAR --num-buckets=20 --width=50 --offset=0
Log Sinks for Long-Term Retention and SIEM
Cloud Logging retains logs for 30 days by default. For compliance or security analysis, export to Cloud Storage or BigQuery:
# Export all application logs to BigQuery for 1-year retention
gcloud logging sinks create prod-logs-bq bigquery.googleapis.com/projects/my-project/datasets/application_logs --log-filter='resource.type=("cloud_run_revision" OR "k8s_container") AND severity>=WARNING' --project=my-project
# Export security-relevant logs to Cloud Storage
gcloud logging sinks create security-audit-gcs storage.googleapis.com/my-security-logs-bucket --log-filter='protoPayload.@type="type.googleapis.com/google.cloud.audit.AuditLog"' --project=my-project
Cloud Trace: Distributed Tracing
Distributed tracing shows you exactly how a request flows through your microservices — which service calls which, how long each step takes, and where errors originate.
Auto-Instrumentation for GKE
The easiest way to get traces is the OpenTelemetry operator on GKE, which can auto-instrument Java, Python, and Node.js applications without code changes:
# Install the OpenTelemetry Operator
# First, apply the cert-manager (prerequisite)
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.11.0/cert-manager.yaml
# Install OpenTelemetry Operator
kubectl apply -f https://github.com/open-telemetry/opentelemetry-operator/releases/latest/download/opentelemetry-operator.yaml
# Configure auto-instrumentation
apiVersion: opentelemetry.io/v1alpha1
kind: Instrumentation
metadata:
name: my-instrumentation
namespace: production
spec:
exporter:
endpoint: http://otel-collector:4317
propagators:
- tracecontext
- baggage
- b3
python:
env:
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: http://otel-collector:4317
java:
env:
- name: OTEL_EXPORTER_OTLP_ENDPOINT
value: http://otel-collector:4317
Manual Tracing in Python
from opentelemetry import trace
from opentelemetry.exporter.cloud_trace import CloudTraceSpanExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
# Configure Cloud Trace exporter
provider = TracerProvider()
provider.add_span_processor(
BatchSpanProcessor(CloudTraceSpanExporter())
)
trace.set_tracer_provider(provider)
tracer = trace.get_tracer("payment-api")
def process_payment(user_id: str, amount_cents: int) -> dict:
with tracer.start_as_current_span("process_payment") as span:
span.set_attribute("user.id", user_id)
span.set_attribute("payment.amount_cents", amount_cents)
# Call to external service — creates a child span
with tracer.start_as_current_span("validate_card"):
result = validate_card(user_id)
span.set_attribute("payment.card_valid", result.valid)
with tracer.start_as_current_span("charge_card"):
charge = charge_card(user_id, amount_cents)
return {"status": "success", "transaction_id": charge.id}
Cloud Profiler
Cloud Profiler continuously profiles CPU and memory usage in production with very low overhead (~1% CPU). It's particularly valuable for finding performance regressions introduced by code changes.
# Enable Cloud Profiler in your Python application
import googlecloudprofiler
try:
googlecloudprofiler.start(
service="payment-api",
service_version="1.2.3",
verbose=3,
)
except (ValueError, NotImplementedError) as exc:
print(f"Profiler not enabled: {exc}")
View profiles in the console: Cloud Profiler → Select service → Compare profiles over time. Look for flame graph changes between versions that indicate performance regressions.
Correlating Signals: Logs + Metrics + Traces
The real power of GCP observability is the correlation between signals. In Cloud Monitoring, when you click on an alert, it shows you the metric spike. From the metric spike, you can drill to the logs using the "View Logs" link. From a log entry with a trace ID, you can jump to the distributed trace.
This correlation is automatic when you:
- Use structured logs with
traceandspanIdfields populated - Instrument with OpenTelemetry or Cloud Trace SDK
- Enable Cloud Monitoring on GKE with the Cloud Logging agent
For SLO-based alerting that uses these observability signals, see our SRE practices guide. For the CI/CD pipeline that deploys changes to monitored services, see our GCP DevOps guide.