GKE Cost Optimization: Spot VMs, Node Auto-Provisioning, and Right-Sizing

Kubernetes clusters are expensive when sized for peak load and never revisited. This guide covers GKE-specific cost optimization: Spot VMs for up to 91% savings on interruptible workloads, Node Auto-Provisioning for automatic right-sizing, and committed use discounts for predictable baseline load.

Kubernetes clusters have a way of growing in cost without growing in value. You provision nodes for peak load, peak load doesn't materialize as often as expected, and suddenly your compute bill is supporting infrastructure that sits at 30% utilization. Add in the namespace entropy that happens over time — every team adding their own services with generous resource requests — and a cluster that started at a reasonable cost becomes difficult to justify.

GKE provides a set of cost optimization tools that, used together, can reduce cluster costs by 40-70% for most organizations. This guide covers the practical implementation: Spot VMs for interruptible workloads, Node Auto-Provisioning to eliminate over-provisioning, Horizontal Pod Autoscaler for scaling pods with demand, and committed use discounts for baseline predictable load.

We'll put actual savings estimates on each strategy so you can prioritize.

Understanding Your Current Spend

Before optimizing, measure. The first step is understanding where your cluster money is going.

# Export detailed billing data to BigQuery for analysis
# (Assumes billing export is already configured)

# Query: Cost by node pool for last 30 days
SELECT
  labels.value AS node_pool,
  SUM(cost) AS total_cost,
  SUM(cost) / 30 AS daily_average
FROM `my-project.billing.gcp_billing_export_v1_*`
WHERE service.description = 'Compute Engine'
  AND labels.key = 'goog-gke-node-pool'
  AND DATE(_PARTITIONTIME) >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
GROUP BY 1
ORDER BY 2 DESC;

# Query: CPU utilization vs provisioned (requires metrics export)
SELECT
  node_pool,
  AVG(cpu_requested_millicores) / AVG(cpu_allocatable_millicores) AS cpu_utilization_pct,
  AVG(memory_requested_bytes) / AVG(memory_allocatable_bytes) AS memory_utilization_pct
FROM `my-project.metrics.node_utilization`
GROUP BY 1
ORDER BY 2 ASC;

Typical findings: node pools running at 20-40% CPU utilization with much higher memory reservation because developers set conservative CPU requests. This is exactly the scenario that Spot VMs and right-sizing address.

Spot VMs: 60-91% Cost Reduction for Eligible Workloads

Google Cloud Spot VMs (formerly Preemptible VMs) are excess compute capacity available at a steep discount. Google can reclaim them with 30-second notice when capacity is needed. In practice, most Spot VMs run for hours to days without interruption — but you can't count on it.

Eligible workloads for Spot VMs:

  • Batch processing jobs: Data transformation, ML training, report generation
  • CI/CD runners: Build pods don't need to survive interruption — just retry the build
  • Stateless web services with horizontal scaling: If one pod is interrupted, others handle the traffic
  • Development environments: Dev namespaces can tolerate occasional disruption

Not eligible:

  • Stateful services (databases, message queues)
  • Jobs that can't resume from checkpoints and take longer than a few hours
  • Anything where pod interruption causes customer-visible errors

Setting Up a Spot Node Pool

resource "google_container_node_pool" "spot_workers" {
  name     = "spot-workers"
  cluster  = google_container_cluster.primary.id
  location = var.region

  autoscaling {
    min_node_count = 0
    max_node_count = 100
    location_policy = "ANY"  # Spot availability varies by zone — use ANY for more supply
  }

  node_config {
    machine_type = "n2-standard-8"  # 8 vCPU, 32 GB
    spot         = true

    # Spot nodes are tainted so only pods that tolerate it get scheduled here
    taint {
      key    = "cloud.google.com/gke-spot"
      value  = "true"
      effect = "NO_SCHEDULE"
    }

    labels = {
      "cloud.google.com/gke-spot" = "true"
    }

    service_account = google_service_account.gke_nodes.email
    oauth_scopes    = ["https://www.googleapis.com/auth/cloud-platform"]
  }
}

Configuring Workloads to Use Spot Nodes

Pods need to explicitly tolerate the Spot taint and optionally use a node affinity to prefer Spot nodes:

spec:
  tolerations:
  - key: "cloud.google.com/gke-spot"
    operator: "Equal"
    value: "true"
    effect: "NoSchedule"
  affinity:
    nodeAffinity:
      preferredDuringSchedulingIgnoredDuringExecution:
      - weight: 100
        preference:
          matchExpressions:
          - key: "cloud.google.com/gke-spot"
            operator: "In"
            values: ["true"]
  # Ensure pods terminate gracefully when Spot node is reclaimed
  terminationGracePeriodSeconds: 25  # Spot gives 30s warning
  containers:
  - name: batch-worker
    image: my-batch-worker:latest
    resources:
      requests:
        cpu: "2"
        memory: "4Gi"

For batch jobs using Kubernetes Job resources, configure restart behavior:

apiVersion: batch/v1
kind: Job
metadata:
  name: data-transform
spec:
  completions: 100
  parallelism: 20
  backoffLimit: 6  # Retry up to 6 times on failure (Spot interruption = failure)
  template:
    spec:
      restartPolicy: OnFailure
      tolerations:
      - key: "cloud.google.com/gke-spot"
        operator: "Equal"
        value: "true"
        effect: "NoSchedule"

Savings estimate: A 20-node n2-standard-8 pool costs approximately $8,000/month on-demand. The same pool using Spot VMs costs approximately $1,600-$2,400/month — savings of $5,600-$6,400/month.

Node Auto-Provisioning: Eliminating Over-Provisioned Node Pools

Standard cluster autoscaler scales nodes within a node pool. If your workloads need a machine type that isn't in any existing pool, they stay Pending. Node Auto-Provisioning (NAP) takes this further — it creates new node pools automatically when needed and deletes them when their workloads finish.

# Enable Node Auto-Provisioning
gcloud container clusters update my-cluster   --region=europe-west4   --enable-autoprovisioning   --max-cpu=500   --max-memory=2000   --autoprovisioning-scopes=https://www.googleapis.com/auth/cloud-platform

# Configure auto-upgrade and auto-repair for provisioned pools
gcloud container clusters update my-cluster   --region=europe-west4   --autoprovisioning-management=ENABLED   --autoprovisioning-upgrade-strategy=SURGE   --autoprovisioning-max-surge-upgrade=3   --autoprovisioning-max-unavailable-upgrade=0
# Alternatively in Terraform
resource "google_container_cluster" "primary" {
  # ... other config ...

  cluster_autoscaling {
    enabled = true  # Enables Node Auto-Provisioning

    resource_limits {
      resource_type = "cpu"
      minimum       = 4
      maximum       = 500
    }

    resource_limits {
      resource_type = "memory"
      minimum       = 8
      maximum       = 2000
    }

    auto_provisioning_defaults {
      service_account = google_service_account.gke_nodes.email
      oauth_scopes    = ["https://www.googleapis.com/auth/cloud-platform"]

      management {
        auto_upgrade = true
        auto_repair  = true
      }

      upgrade_settings {
        max_surge       = 2
        max_unavailable = 0
      }
    }
  }
}

With NAP enabled, GKE automatically creates the right machine type for each workload. A pod requesting 32 vCPU gets an n2-standard-32 node provisioned for it. A pod requesting 1 vCPU and 2 GB gets an e2-medium. This eliminates the common pattern of provisioning large nodes for occasional large workloads.

Horizontal Pod Autoscaler and Cluster Autoscaler Working Together

HPA and cluster autoscaler work in tandem — HPA adds pods, cluster autoscaler adds nodes to schedule those pods. Configure them correctly or they conflict.

# HPA for a web service
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: web-service-hpa
  namespace: my-app
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web-service
  minReplicas: 2
  maxReplicas: 50
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 60  # Scale up when avg CPU > 60%
  - type: Resource
    resource:
      name: memory
      target:
        type: Utilization
        averageUtilization: 70
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300  # Wait 5 min before scaling down
      policies:
      - type: Percent
        value: 10
        periodSeconds: 60  # Remove at most 10% of pods per minute when scaling down
    scaleUp:
      stabilizationWindowSeconds: 60
      policies:
      - type: Percent
        value: 100
        periodSeconds: 30  # Double pod count every 30 seconds when scaling up

The cluster autoscaler will scale nodes to accommodate new pods created by HPA. Set the stabilizationWindowSeconds for scale-down carefully — too short and you'll thrash nodes; too long and you'll pay for idle capacity.

Right-Sizing Resource Requests

This is the highest-impact, lowest-effort optimization most teams skip. If your pods are requesting 2x what they actually use, your nodes are half as full as they could be — you're paying for twice the infrastructure.

Use the Vertical Pod Autoscaler (VPA) in recommendation mode first:

# Install VPA
kubectl apply -f https://github.com/kubernetes/autoscaler/releases/download/vertical-pod-autoscaler-0.14.0/vpa-v1.crds.yaml
kubectl apply -f https://github.com/kubernetes/autoscaler/releases/download/vertical-pod-autoscaler-0.14.0/vpa-v0.14.0.yaml

# Create a VPA object in recommendation mode (no automatic changes)
cat << 'EOF' | kubectl apply -f -
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: web-service-vpa
  namespace: my-app
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: web-service
  updatePolicy:
    updateMode: "Off"  # Recommendation only — don't change pods automatically
EOF

# After a few days of data collection, view recommendations
kubectl describe vpa web-service-vpa -n my-app

The VPA recommendation output shows:

  • Lower bound: minimum safe requests
  • Target: recommended requests based on observed usage
  • Upper bound: maximum usage observed

Start by applying VPA target values as your resource requests. This typically reduces cluster size by 20-40%.

Committed Use Discounts for Baseline Load

Spot VMs handle bursty workloads cheaply. For your stable baseline — the minimum number of nodes that always run — committed use discounts (CUDs) provide 37-57% savings.

# Purchase 1-year commitment for your baseline node pool compute
gcloud compute commitments create gke-baseline-2025   --region=europe-west4   --plan=12-month   --resources=vcpu=80,memory=320GB  # Match your baseline node pool size

# View existing commitments
gcloud compute commitments list --region=europe-west4

A practical CUD strategy for GKE:

  1. Run your cluster for 2-4 weeks to understand stable baseline utilization
  2. Purchase CUDs for 70-80% of baseline (leave headroom for growth)
  3. Use cluster autoscaler (with on-demand nodes) for burst capacity beyond CUD coverage
  4. Use Spot VMs for batch/interruptible burst capacity

For a cluster with 20 baseline n2-standard-8 nodes at $0.416/node-hour:

  • On-demand cost: 20 × $0.416 × 730 = ~$6,073/month
  • 1-year CUD (37% discount): ~$3,826/month
  • 3-year CUD (57% discount): ~$2,611/month

Putting It Together: A Real Cost Reduction Story

Here's a representative example of a 100-node GKE cluster optimization:

Before:

  • 100 on-demand n2-standard-8 nodes, always running
  • Average utilization: 35% CPU, 55% memory
  • Monthly cost: ~$30,360

After (3-month optimization project):

  • 20 baseline nodes with 3-year CUDs: ~$5,222/month
  • 15 on-demand nodes for burst: ~$4,554/month
  • 30 Spot nodes for batch workloads: ~$2,160/month
  • 35 nodes eliminated via right-sizing (VPA + HPA)
  • Monthly cost: ~$11,936

Total savings: ~$18,424/month (61% reduction)

The optimization project required:

  1. 2 weeks of VPA data collection and right-sizing resource requests
  2. 1 week of NAP configuration and Spot node pool setup
  3. 2 weeks of HPA tuning and validating Spot workload behavior
  4. CUD purchase after establishing stable baseline

For the production cluster configuration that gives you the foundation to apply these optimizations, see our GKE production configuration guide. For securing the cluster through these changes, see our GKE security hardening guide.