GKE Cost Optimization: Spot VMs, Node Auto-Provisioning, and Right-Sizing
Kubernetes clusters are expensive when sized for peak load and never revisited. This guide covers GKE-specific cost optimization: Spot VMs for up to 91% savings on interruptible workloads, Node Auto-Provisioning for automatic right-sizing, and committed use discounts for predictable baseline load.
Kubernetes clusters have a way of growing in cost without growing in value. You provision nodes for peak load, peak load doesn't materialize as often as expected, and suddenly your compute bill is supporting infrastructure that sits at 30% utilization. Add in the namespace entropy that happens over time — every team adding their own services with generous resource requests — and a cluster that started at a reasonable cost becomes difficult to justify.
GKE provides a set of cost optimization tools that, used together, can reduce cluster costs by 40-70% for most organizations. This guide covers the practical implementation: Spot VMs for interruptible workloads, Node Auto-Provisioning to eliminate over-provisioning, Horizontal Pod Autoscaler for scaling pods with demand, and committed use discounts for baseline predictable load.
We'll put actual savings estimates on each strategy so you can prioritize.
Understanding Your Current Spend
Before optimizing, measure. The first step is understanding where your cluster money is going.
# Export detailed billing data to BigQuery for analysis
# (Assumes billing export is already configured)
# Query: Cost by node pool for last 30 days
SELECT
labels.value AS node_pool,
SUM(cost) AS total_cost,
SUM(cost) / 30 AS daily_average
FROM `my-project.billing.gcp_billing_export_v1_*`
WHERE service.description = 'Compute Engine'
AND labels.key = 'goog-gke-node-pool'
AND DATE(_PARTITIONTIME) >= DATE_SUB(CURRENT_DATE(), INTERVAL 30 DAY)
GROUP BY 1
ORDER BY 2 DESC;
# Query: CPU utilization vs provisioned (requires metrics export)
SELECT
node_pool,
AVG(cpu_requested_millicores) / AVG(cpu_allocatable_millicores) AS cpu_utilization_pct,
AVG(memory_requested_bytes) / AVG(memory_allocatable_bytes) AS memory_utilization_pct
FROM `my-project.metrics.node_utilization`
GROUP BY 1
ORDER BY 2 ASC;
Typical findings: node pools running at 20-40% CPU utilization with much higher memory reservation because developers set conservative CPU requests. This is exactly the scenario that Spot VMs and right-sizing address.
Spot VMs: 60-91% Cost Reduction for Eligible Workloads
Google Cloud Spot VMs (formerly Preemptible VMs) are excess compute capacity available at a steep discount. Google can reclaim them with 30-second notice when capacity is needed. In practice, most Spot VMs run for hours to days without interruption — but you can't count on it.
Eligible workloads for Spot VMs:
- Batch processing jobs: Data transformation, ML training, report generation
- CI/CD runners: Build pods don't need to survive interruption — just retry the build
- Stateless web services with horizontal scaling: If one pod is interrupted, others handle the traffic
- Development environments: Dev namespaces can tolerate occasional disruption
Not eligible:
- Stateful services (databases, message queues)
- Jobs that can't resume from checkpoints and take longer than a few hours
- Anything where pod interruption causes customer-visible errors
Setting Up a Spot Node Pool
resource "google_container_node_pool" "spot_workers" {
name = "spot-workers"
cluster = google_container_cluster.primary.id
location = var.region
autoscaling {
min_node_count = 0
max_node_count = 100
location_policy = "ANY" # Spot availability varies by zone — use ANY for more supply
}
node_config {
machine_type = "n2-standard-8" # 8 vCPU, 32 GB
spot = true
# Spot nodes are tainted so only pods that tolerate it get scheduled here
taint {
key = "cloud.google.com/gke-spot"
value = "true"
effect = "NO_SCHEDULE"
}
labels = {
"cloud.google.com/gke-spot" = "true"
}
service_account = google_service_account.gke_nodes.email
oauth_scopes = ["https://www.googleapis.com/auth/cloud-platform"]
}
}
Configuring Workloads to Use Spot Nodes
Pods need to explicitly tolerate the Spot taint and optionally use a node affinity to prefer Spot nodes:
spec:
tolerations:
- key: "cloud.google.com/gke-spot"
operator: "Equal"
value: "true"
effect: "NoSchedule"
affinity:
nodeAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
preference:
matchExpressions:
- key: "cloud.google.com/gke-spot"
operator: "In"
values: ["true"]
# Ensure pods terminate gracefully when Spot node is reclaimed
terminationGracePeriodSeconds: 25 # Spot gives 30s warning
containers:
- name: batch-worker
image: my-batch-worker:latest
resources:
requests:
cpu: "2"
memory: "4Gi"
For batch jobs using Kubernetes Job resources, configure restart behavior:
apiVersion: batch/v1
kind: Job
metadata:
name: data-transform
spec:
completions: 100
parallelism: 20
backoffLimit: 6 # Retry up to 6 times on failure (Spot interruption = failure)
template:
spec:
restartPolicy: OnFailure
tolerations:
- key: "cloud.google.com/gke-spot"
operator: "Equal"
value: "true"
effect: "NoSchedule"
Savings estimate: A 20-node n2-standard-8 pool costs approximately $8,000/month on-demand. The same pool using Spot VMs costs approximately $1,600-$2,400/month — savings of $5,600-$6,400/month.
Node Auto-Provisioning: Eliminating Over-Provisioned Node Pools
Standard cluster autoscaler scales nodes within a node pool. If your workloads need a machine type that isn't in any existing pool, they stay Pending. Node Auto-Provisioning (NAP) takes this further — it creates new node pools automatically when needed and deletes them when their workloads finish.
# Enable Node Auto-Provisioning
gcloud container clusters update my-cluster --region=europe-west4 --enable-autoprovisioning --max-cpu=500 --max-memory=2000 --autoprovisioning-scopes=https://www.googleapis.com/auth/cloud-platform
# Configure auto-upgrade and auto-repair for provisioned pools
gcloud container clusters update my-cluster --region=europe-west4 --autoprovisioning-management=ENABLED --autoprovisioning-upgrade-strategy=SURGE --autoprovisioning-max-surge-upgrade=3 --autoprovisioning-max-unavailable-upgrade=0
# Alternatively in Terraform
resource "google_container_cluster" "primary" {
# ... other config ...
cluster_autoscaling {
enabled = true # Enables Node Auto-Provisioning
resource_limits {
resource_type = "cpu"
minimum = 4
maximum = 500
}
resource_limits {
resource_type = "memory"
minimum = 8
maximum = 2000
}
auto_provisioning_defaults {
service_account = google_service_account.gke_nodes.email
oauth_scopes = ["https://www.googleapis.com/auth/cloud-platform"]
management {
auto_upgrade = true
auto_repair = true
}
upgrade_settings {
max_surge = 2
max_unavailable = 0
}
}
}
}
With NAP enabled, GKE automatically creates the right machine type for each workload. A pod requesting 32 vCPU gets an n2-standard-32 node provisioned for it. A pod requesting 1 vCPU and 2 GB gets an e2-medium. This eliminates the common pattern of provisioning large nodes for occasional large workloads.
Horizontal Pod Autoscaler and Cluster Autoscaler Working Together
HPA and cluster autoscaler work in tandem — HPA adds pods, cluster autoscaler adds nodes to schedule those pods. Configure them correctly or they conflict.
# HPA for a web service
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web-service-hpa
namespace: my-app
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web-service
minReplicas: 2
maxReplicas: 50
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 60 # Scale up when avg CPU > 60%
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # Wait 5 min before scaling down
policies:
- type: Percent
value: 10
periodSeconds: 60 # Remove at most 10% of pods per minute when scaling down
scaleUp:
stabilizationWindowSeconds: 60
policies:
- type: Percent
value: 100
periodSeconds: 30 # Double pod count every 30 seconds when scaling up
The cluster autoscaler will scale nodes to accommodate new pods created by HPA. Set the stabilizationWindowSeconds for scale-down carefully — too short and you'll thrash nodes; too long and you'll pay for idle capacity.
Right-Sizing Resource Requests
This is the highest-impact, lowest-effort optimization most teams skip. If your pods are requesting 2x what they actually use, your nodes are half as full as they could be — you're paying for twice the infrastructure.
Use the Vertical Pod Autoscaler (VPA) in recommendation mode first:
# Install VPA
kubectl apply -f https://github.com/kubernetes/autoscaler/releases/download/vertical-pod-autoscaler-0.14.0/vpa-v1.crds.yaml
kubectl apply -f https://github.com/kubernetes/autoscaler/releases/download/vertical-pod-autoscaler-0.14.0/vpa-v0.14.0.yaml
# Create a VPA object in recommendation mode (no automatic changes)
cat << 'EOF' | kubectl apply -f -
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: web-service-vpa
namespace: my-app
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: web-service
updatePolicy:
updateMode: "Off" # Recommendation only — don't change pods automatically
EOF
# After a few days of data collection, view recommendations
kubectl describe vpa web-service-vpa -n my-app
The VPA recommendation output shows:
- Lower bound: minimum safe requests
- Target: recommended requests based on observed usage
- Upper bound: maximum usage observed
Start by applying VPA target values as your resource requests. This typically reduces cluster size by 20-40%.
Committed Use Discounts for Baseline Load
Spot VMs handle bursty workloads cheaply. For your stable baseline — the minimum number of nodes that always run — committed use discounts (CUDs) provide 37-57% savings.
# Purchase 1-year commitment for your baseline node pool compute
gcloud compute commitments create gke-baseline-2025 --region=europe-west4 --plan=12-month --resources=vcpu=80,memory=320GB # Match your baseline node pool size
# View existing commitments
gcloud compute commitments list --region=europe-west4
A practical CUD strategy for GKE:
- Run your cluster for 2-4 weeks to understand stable baseline utilization
- Purchase CUDs for 70-80% of baseline (leave headroom for growth)
- Use cluster autoscaler (with on-demand nodes) for burst capacity beyond CUD coverage
- Use Spot VMs for batch/interruptible burst capacity
For a cluster with 20 baseline n2-standard-8 nodes at $0.416/node-hour:
- On-demand cost: 20 × $0.416 × 730 = ~$6,073/month
- 1-year CUD (37% discount): ~$3,826/month
- 3-year CUD (57% discount): ~$2,611/month
Putting It Together: A Real Cost Reduction Story
Here's a representative example of a 100-node GKE cluster optimization:
Before:
- 100 on-demand n2-standard-8 nodes, always running
- Average utilization: 35% CPU, 55% memory
- Monthly cost: ~$30,360
After (3-month optimization project):
- 20 baseline nodes with 3-year CUDs: ~$5,222/month
- 15 on-demand nodes for burst: ~$4,554/month
- 30 Spot nodes for batch workloads: ~$2,160/month
- 35 nodes eliminated via right-sizing (VPA + HPA)
- Monthly cost: ~$11,936
Total savings: ~$18,424/month (61% reduction)
The optimization project required:
- 2 weeks of VPA data collection and right-sizing resource requests
- 1 week of NAP configuration and Spot node pool setup
- 2 weeks of HPA tuning and validating Spot workload behavior
- CUD purchase after establishing stable baseline
For the production cluster configuration that gives you the foundation to apply these optimizations, see our GKE production configuration guide. For securing the cluster through these changes, see our GKE security hardening guide.