Google Kubernetes Engine: Production-Ready Cluster Configuration Guide

Setting up a GKE cluster for production requires more than clicking 'create' in the console. This guide covers VPC-native networking, private clusters, workload identity, cluster autoscaling, and every configuration decision that separates a production cluster from a development sandbox.

There's a meaningful gap between "a GKE cluster that runs workloads" and "a GKE cluster that's ready for production." The former takes about 10 minutes. The latter requires working through a series of decisions that most tutorials skip entirely: private networking configuration, workload identity for pod-level AWS and GCP permissions, release channel selection, node pool design, cluster autoscaling strategy, monitoring integration, and security baseline.

This guide walks through all of it. We'll use Terraform throughout because a GKE cluster with 30-plus configuration options is not something you want to manage through the console — drift will occur and it's nearly impossible to audit.

By the end, you'll have a template for a production GKE cluster that follows Google's recommended practices, is auditable through code, and can be reliably recreated in a new region.

Cluster Architecture Decisions

Before writing any Terraform, make four architecture decisions:

1. Private vs public nodes? Production clusters should use private nodes — worker nodes have no public IP addresses and all traffic routes through private Google APIs. This prevents direct internet exposure of your Kubernetes API server and nodes. Use a private cluster unless you have a specific reason not to.

2. Regional vs zonal? Regional clusters run the control plane and node pools across three zones, providing higher availability. Zonal clusters run in a single zone and are cheaper but go down if that zone experiences an outage. Use regional for production, zonal for development.

3. Standard vs Autopilot? Autopilot abstracts node management entirely — GKE handles node provisioning, scaling, and patching. Standard mode gives you full control over node configuration. For most production workloads without specialized hardware requirements (GPUs, TPUs, specific machine families), Autopilot is worth serious consideration. See our GKE Autopilot vs Standard comparison.

4. Release channel? GKE offers RAPID, REGULAR, and STABLE channels. Stable gets new features 2-3 months after Regular. For production, Regular is usually the right choice — it's tested enough without being too far behind the Kubernetes release schedule.

Terraform Foundation

Start with the VPC, then the cluster. Keeping them in separate Terraform modules avoids a monolithic configuration file that's hard to reason about.

# modules/gke-vpc/main.tf

resource "google_compute_network" "gke_vpc" {
  name                    = var.network_name
  auto_create_subnetworks = false  # Always use custom subnets in production
  project                 = var.project_id
}

resource "google_compute_subnetwork" "gke_nodes" {
  name          = "${var.network_name}-nodes"
  network       = google_compute_network.gke_vpc.id
  region        = var.region
  ip_cidr_range = var.nodes_cidr  # e.g., "10.0.0.0/20" — 4096 addresses

  private_ip_google_access = true  # Required for private clusters to reach GCP APIs

  secondary_ip_range {
    range_name    = "pods"
    ip_cidr_range = var.pods_cidr  # e.g., "10.16.0.0/14" — 262144 pod addresses
  }

  secondary_ip_range {
    range_name    = "services"
    ip_cidr_range = var.services_cidr  # e.g., "10.20.0.0/20" — 4096 service IPs
  }
}

# Cloud Router + NAT for private nodes to reach the internet
resource "google_compute_router" "nat_router" {
  name    = "${var.network_name}-router"
  network = google_compute_network.gke_vpc.id
  region  = var.region
}

resource "google_compute_router_nat" "nat_config" {
  name                               = "${var.network_name}-nat"
  router                             = google_compute_router.nat_router.name
  region                             = var.region
  nat_ip_allocate_option             = "AUTO_ONLY"
  source_subnetwork_ip_ranges_to_nat = "ALL_SUBNETWORKS_ALL_IP_RANGES"

  log_config {
    enable = true
    filter = "ERRORS_ONLY"
  }
}

Now the cluster itself:

# modules/gke-cluster/main.tf

resource "google_container_cluster" "primary" {
  name     = var.cluster_name
  location = var.region  # Regional cluster (use zone for zonal, e.g., "europe-west4-a")
  project  = var.project_id

  # Delete the default node pool and manage node pools separately
  # This gives you more control over node pool updates
  remove_default_node_pool = true
  initial_node_count       = 1

  # Networking
  network    = var.network_id
  subnetwork = var.subnetwork_id

  networking_mode = "VPC_NATIVE"  # Required for private clusters

  ip_allocation_policy {
    cluster_secondary_range_name  = "pods"
    services_secondary_range_name = "services"
  }

  # Private cluster configuration
  private_cluster_config {
    enable_private_nodes    = true
    enable_private_endpoint = false  # Keep API server accessible from authorized networks
    master_ipv4_cidr_block  = "172.16.0.0/28"  # Control plane CIDR
  }

  # Restrict API server access to your corporate IP ranges
  master_authorized_networks_config {
    cidr_blocks {
      cidr_block   = var.authorized_cidr
      display_name = "corporate-vpn"
    }
  }

  # Release channel for automatic upgrades
  release_channel {
    channel = "REGULAR"
  }

  # Workload Identity — binds GKE service accounts to GCP IAM
  workload_identity_config {
    workload_pool = "${var.project_id}.svc.id.goog"
  }

  # Binary Authorization — only allow signed container images
  binary_authorization {
    evaluation_mode = "PROJECT_SINGLETON_POLICY_ENFORCE"
  }

  # Enable GKE add-ons
  addons_config {
    horizontal_pod_autoscaling {
      disabled = false
    }
    http_load_balancing {
      disabled = false  # Required for GKE Ingress (L7 LB)
    }
    dns_cache_config {
      enabled = true  # NodeLocal DNSCache reduces DNS latency
    }
    gce_persistent_disk_csi_driver_config {
      enabled = true
    }
    gcs_fuse_csi_driver_config {
      enabled = true  # Mount GCS buckets as filesystems
    }
  }

  # Network policies
  network_policy {
    enabled  = true
    provider = "CALICO"
  }

  # Logging and monitoring
  logging_config {
    enable_components = ["SYSTEM_COMPONENTS", "WORKLOADS"]
  }

  monitoring_config {
    enable_components = ["SYSTEM_COMPONENTS", "WORKLOADS"]

    managed_prometheus {
      enabled = true  # Google-managed Prometheus (no infra to run)
    }
  }

  # Security
  enable_shielded_nodes = true

  database_encryption {
    state    = "ENCRYPTED"
    key_name = var.kms_key_name  # Customer-managed encryption key
  }
}

Node Pool Design

Node pools should be designed around workload requirements. Never run everything on a single node pool — you lose the ability to right-size nodes for different workload types.

# System node pool: for kube-system workloads
resource "google_container_node_pool" "system" {
  name     = "system-pool"
  cluster  = google_container_cluster.primary.id
  location = var.region

  # Distribute across zones
  node_count = 1  # Per zone — regional cluster has 3 zones = 3 nodes total

  autoscaling {
    min_node_count = 1
    max_node_count = 3
    location_policy = "BALANCED"
  }

  management {
    auto_repair  = true
    auto_upgrade = true  # Keeps nodes on latest GKE version in release channel
  }

  node_config {
    machine_type = "e2-standard-4"  # 4 vCPU, 16 GB — right-sized for system workloads

    # Use service account with minimal permissions
    service_account = google_service_account.gke_nodes.email
    oauth_scopes = [
      "https://www.googleapis.com/auth/cloud-platform",
    ]

    workload_metadata_config {
      mode = "GKE_METADATA"  # Required for Workload Identity
    }

    shielded_instance_config {
      enable_secure_boot          = true
      enable_integrity_monitoring = true
    }

    # Taint system pool to prevent user workloads
    taint {
      key    = "node-role"
      value  = "system"
      effect = "NO_SCHEDULE"
    }

    labels = {
      "node-role" = "system"
    }
  }

  upgrade_settings {
    max_surge       = 1
    max_unavailable = 0
  }
}

# Application node pool: for production workloads
resource "google_container_node_pool" "app" {
  name     = "app-pool"
  cluster  = google_container_cluster.primary.id
  location = var.region

  autoscaling {
    min_node_count = 2   # Per zone
    max_node_count = 20  # Per zone — allows up to 60 nodes regionally
    location_policy = "BALANCED"
  }

  management {
    auto_repair  = true
    auto_upgrade = true
  }

  node_config {
    machine_type = "n2-standard-8"  # 8 vCPU, 32 GB

    service_account = google_service_account.gke_nodes.email
    oauth_scopes    = ["https://www.googleapis.com/auth/cloud-platform"]

    workload_metadata_config {
      mode = "GKE_METADATA"
    }

    shielded_instance_config {
      enable_secure_boot          = true
      enable_integrity_monitoring = true
    }

    labels = {
      "node-role" = "app"
    }
  }

  upgrade_settings {
    max_surge       = 2   # Add 2 new nodes before removing old ones
    max_unavailable = 0   # Never remove a node without a replacement
  }
}

# Spot node pool: for batch and stateless workloads
resource "google_container_node_pool" "spot" {
  name     = "spot-pool"
  cluster  = google_container_cluster.primary.id
  location = var.region

  autoscaling {
    min_node_count = 0   # Can scale to zero when no batch jobs
    max_node_count = 50
    location_policy = "ANY"  # Spot doesn't guarantee zone balance
  }

  node_config {
    machine_type = "n2-standard-8"
    spot         = true  # 60-90% cheaper than on-demand

    service_account = google_service_account.gke_nodes.email
    oauth_scopes    = ["https://www.googleapis.com/auth/cloud-platform"]

    workload_metadata_config {
      mode = "GKE_METADATA"
    }

    taint {
      key    = "cloud.google.com/gke-spot"
      value  = "true"
      effect = "NO_SCHEDULE"
    }

    labels = {
      "node-role" = "spot"
    }
  }
}

Workload Identity Configuration

Workload Identity is GKE's mechanism for giving pods GCP IAM permissions without storing credentials. Each Kubernetes ServiceAccount maps to a GCP IAM service account.

# IAM service account for a specific application
resource "google_service_account" "app_sa" {
  account_id   = "my-app-sa"
  display_name = "My Application Service Account"
  project      = var.project_id
}

# Grant BigQuery access to the app
resource "google_project_iam_member" "app_bigquery" {
  project = var.project_id
  role    = "roles/bigquery.dataViewer"
  member  = "serviceAccount:${google_service_account.app_sa.email}"
}

# Allow the Kubernetes ServiceAccount to impersonate the GCP SA
resource "google_service_account_iam_binding" "workload_identity" {
  service_account_id = google_service_account.app_sa.name
  role               = "roles/iam.workloadIdentityUser"

  members = [
    "serviceAccount:${var.project_id}.svc.id.goog[my-namespace/my-service-account]",
  ]
}

Then annotate your Kubernetes ServiceAccount:

apiVersion: v1
kind: ServiceAccount
metadata:
  name: my-service-account
  namespace: my-namespace
  annotations:
    iam.gke.io/gcp-service-account: my-app-sa@my-project.iam.gserviceaccount.com

Pods using this ServiceAccount automatically get GCP credentials that respect the IAM permissions you've granted — no secret management, no key rotation.

Cluster Hardening Checklist

With the cluster built, verify these hardening measures are in place:

# Verify private nodes
gcloud container clusters describe my-cluster   --region=europe-west4   --format="value(privateClusterConfig.enablePrivateNodes)"
# Expected: True

# Verify workload identity
gcloud container clusters describe my-cluster   --region=europe-west4   --format="value(workloadIdentityConfig.workloadPool)"
# Expected: my-project.svc.id.goog

# Verify shielded nodes
gcloud container clusters describe my-cluster   --region=europe-west4   --format="value(shieldedNodes.enabled)"
# Expected: True

# Check for any publicly accessible services
kubectl get svc --all-namespaces | grep LoadBalancer

# Verify network policy is enabled
gcloud container clusters describe my-cluster   --region=europe-west4   --format="value(networkPolicy.enabled)"
# Expected: True

Setting Up Cloud Monitoring and Alerting

Google-managed Prometheus collects metrics automatically from GKE clusters. Create alerting policies for critical signals:

# Alert when node CPU utilization exceeds 80%
gcloud alpha monitoring policies create   --display-name="GKE Node CPU High"   --condition-display-name="node CPU > 80%"   --condition-filter='resource.type="k8s_node" AND metric.type="kubernetes.io/node/cpu/allocatable_utilization"'   --condition-threshold-value=0.80   --condition-threshold-comparison=COMPARISON_GT   --condition-duration=300s   --notification-channels=$CHANNEL_ID

# Alert when pods are failing to schedule (pending > 5 minutes)
gcloud alpha monitoring policies create   --display-name="GKE Pods Pending"   --condition-display-name="pending pods > 0 for 5 min"   --condition-filter='resource.type="k8s_pod" AND metric.type="kubernetes.io/pod/volume/request_bytes" AND metadata.user_labels.phase="Pending"'   --condition-duration=300s

The Result

Following this guide produces a GKE cluster that:

  • Runs in three availability zones with automatic failover
  • Has nodes with no public IP addresses
  • Uses Workload Identity instead of service account keys
  • Encrypts etcd with a customer-managed key
  • Enforces network policies between pods
  • Automatically repairs and upgrades nodes
  • Scales from 6 nodes to 180 nodes in response to load
  • Ships logs and metrics to Cloud Monitoring without any agent configuration

For securing this cluster further, see our GKE security hardening guide. For optimizing the cluster's cost profile, see our GKE cost optimization guide.