GCP Rightsizing Guide: Optimizing Compute Engine, GKE, and Cloud SQL
Rightsizing is the single highest-ROI cloud optimization activity. Most GCP workloads run on machines that are 2-3x larger than necessary, paying for CPU and memory that sits idle. This guide covers systematic rightsizing for Compute Engine VMs, GKE node pools, and Cloud SQL instances using real utilization data.
Here's a reality check most cloud engineers don't talk about: the average GCP workload uses about 30-40% of the CPU it's paying for. Memory utilization is often even lower. We size for peak, add a "safety margin," and then rarely revisit those decisions as workloads evolve.
Rightsizing — matching your machine types to actual utilization — consistently delivers 30-60% cost reductions on compute. It's not glamorous work, but it's the highest ROI optimization you can do before reaching for committed use discounts or other financial instruments.
This guide covers rightsizing for the three biggest GCP compute cost centers: Compute Engine VMs, GKE node pools, and Cloud SQL instances. For each, we'll look at how to measure utilization, what targets to aim for, and how to safely resize without causing downtime.
The Rightsizing Mindset
Before diving into specifics, two mental model shifts that make rightsizing less scary:
Target 60-70% utilization, not 100%: You're not trying to jam resources to the limit. You want headroom for traffic spikes and process bursts. A VM with 65% average CPU utilization with spikes to 80% is well-sized. A VM running at 8% average is dramatically oversized.
Use percentiles, not averages: Average CPU utilization is misleading. A VM that sits at 5% for 23 hours but spikes to 90% for one hour has an average of ~9% — but you absolutely need that capacity for that one hour. Use P95 or P99 utilization over a 2-4 week window for sizing decisions.
Rightsizing Compute Engine VMs
Finding Oversized VMs with Cloud Monitoring
The Recommendations Hub gives you machine-type recommendations automatically, but it's helpful to understand what data it's using.
# Get VM rightsizing recommendations for a project
gcloud recommender recommendations list --project=my-project --location=us-central1-a --recommender=google.compute.instance.MachineTypeRecommender --format="table(name,description,primaryImpact.costProjection.cost.units,stateInfo.state)"
For custom analysis using Cloud Monitoring metrics:
from google.cloud import monitoring_v3
from datetime import datetime, timedelta
import time
client = monitoring_v3.MetricServiceClient()
project_name = f"projects/my-project"
# Query P95 CPU utilization for all instances over 14 days
now = time.time()
interval = monitoring_v3.TimeInterval({
"end_time": {"seconds": int(now)},
"start_time": {"seconds": int(now - 14 * 86400)},
})
results = client.list_time_series(
request={
"name": project_name,
"filter": 'metric.type="compute.googleapis.com/instance/cpu/utilization"',
"interval": interval,
"aggregation": {
"alignment_period": {"seconds": 86400},
"per_series_aligner": monitoring_v3.Aggregation.Aligner.ALIGN_PERCENTILE_95,
"cross_series_reducer": monitoring_v3.Aggregation.Reducer.REDUCE_NONE,
},
}
)
for series in results:
instance_name = series.resource.labels["instance_id"]
p95_values = [point.value.double_value for point in series.points]
avg_p95 = sum(p95_values) / len(p95_values) if p95_values else 0
if avg_p95 < 0.20: # P95 CPU below 20% — likely oversized
print(f"OVERSIZED: {instance_name} | P95 CPU: {avg_p95*100:.1f}%")
Compute Engine Machine Type Selection
When you find an oversized VM, choose the right replacement machine series:
| Use Case | Machine Series | Notes |
|---|---|---|
| General workload | N2 or N2D | Best price/performance for most apps |
| Memory-intensive | M2 | DB caching, in-memory analytics |
| Compute-intensive | C2 | Video transcoding, batch compute |
| Cost-sensitive, interruptible | Spot VMs | 60-91% savings vs on-demand |
| ARM workloads | T2A | Smaller per-core cost, great for containerized apps |
# Resize a stopped VM to a smaller machine type
gcloud compute instances set-machine-type api-server-prod --zone=europe-west4-a --machine-type=n2-standard-4 # was n2-standard-16
# Check available machine types in a zone
gcloud compute machine-types list --filter="zone:europe-west4-a AND name~n2-standard" --format="table(name,guestCpus,memoryMb,description)"
Safe Resizing Procedure
Never resize a production VM blindly. Follow this sequence:
# 1. Snapshot the disk before resizing
gcloud compute disks snapshot api-server-prod-disk --zone=europe-west4-a --snapshot-names=pre-resize-snapshot-$(date +%Y%m%d)
# 2. Stop the instance
gcloud compute instances stop api-server-prod --zone=europe-west4-a
# 3. Change machine type
gcloud compute instances set-machine-type api-server-prod --zone=europe-west4-a --machine-type=n2-standard-4
# 4. Start the instance
gcloud compute instances start api-server-prod --zone=europe-west4-a
# 5. Monitor for 15 minutes — if issues, revert
# gcloud compute instances set-machine-type api-server-prod # --zone=europe-west4-a # --machine-type=n2-standard-16
For instances that can't tolerate downtime, use a blue/green approach: provision a new smaller VM, migrate traffic via load balancer, verify, then decommission the old VM.
Rightsizing GKE Node Pools
GKE adds a layer of complexity because you're sizing nodes to fit pods, not applications directly. The unit of analysis is the node pool, and the goal is maximizing bin packing efficiency — the percentage of allocated resources actually used.
Understanding GKE Resource Allocation
In GKE, every node has a portion of its CPU and memory reserved for system overhead. For an n2-standard-4 (4 vCPU, 16 GB):
- System reserved CPU: ~0.09 vCPU
- System reserved memory: ~1.1 GB
- Eviction threshold: ~100 MB memory
- Allocatable CPU: ~3.91 vCPU
- Allocatable memory: ~14.4 GB
Your pods' requests (not limits) consume allocatable capacity. Bin packing efficiency = (sum of pod requests) / (allocatable capacity).
# View node allocatable capacity and current allocations
kubectl describe nodes | grep -A 5 "Allocatable:"
kubectl top nodes
# View pod resource requests vs limits
kubectl get pods -A -o custom-columns="NAMESPACE:.metadata.namespace,NAME:.metadata.name,CPU_REQ:.spec.containers[*].resources.requests.cpu,MEM_REQ:.spec.containers[*].resources.requests.memory"
Identifying Underutilized Node Pools
# Enable cluster autoscaler metrics view
kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml
# Get per-node utilization
kubectl top nodes --sort-by=cpu
# Find nodes with low utilization (candidates for consolidation)
kubectl top nodes | awk 'NR>1 && $3+0 < 30 {print "LOW_CPU:", $0}'
For production clusters, use the GKE Cost Optimization Insights dashboard in the console (Kubernetes Engine → Clusters → Cost → Optimization Insights). It shows:
- Node pool CPU/memory utilization
- Over-provisioning percentage
- Estimated monthly savings from rightsizing
Node Pool Rightsizing Strategy
The most effective GKE rightsizing approach combines three techniques:
1. Enable Vertical Pod Autoscaler (VPA) in recommendation mode first
VPA analyzes pod utilization and recommends appropriate requests/limits without automatically changing anything:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: api-vpa
namespace: production
spec:
targetRef:
apiVersion: "apps/v1"
kind: Deployment
name: api
updatePolicy:
updateMode: "Off" # Recommendation only, no auto-updates
resourcePolicy:
containerPolicies:
- containerName: api
minAllowed:
cpu: 100m
memory: 128Mi
maxAllowed:
cpu: 2
memory: 2Gi
# View VPA recommendations
kubectl get vpa api-vpa -o yaml | grep -A 20 "recommendation:"
2. Right-size pod requests based on VPA recommendations
After a week of VPA recommendations, update your deployments to match recommended requests:
resources:
requests:
cpu: "250m" # was "1000m"
memory: "512Mi" # was "2Gi"
limits:
cpu: "1000m"
memory: "1Gi"
3. Scale node pool machine type down to match improved density
After right-sizing pod requests, you can use fewer or smaller nodes:
# Create new right-sized node pool
gcloud container node-pools create rightsized-pool --cluster=production --region=europe-west4 --machine-type=n2-standard-4 --num-nodes=3 --enable-autoscaling --min-nodes=2 --max-nodes=8
# Cordon and drain the old oversized pool
kubectl cordon -l cloud.google.com/gke-nodepool=large-pool
kubectl drain -l cloud.google.com/gke-nodepool=large-pool --ignore-daemonsets --delete-emptydir-data
# Delete the old pool
gcloud container node-pools delete large-pool --cluster=production --region=europe-west4
Node Auto-Provisioning for Automatic Rightsizing
If you're regularly adding new workloads with different resource profiles, enable Node Auto-Provisioning (NAP). NAP automatically creates node pools with machine types optimized for your pending pods:
gcloud container clusters update production --region=europe-west4 --enable-autoprovisioning --min-cpu=4 --max-cpu=200 --min-memory=16 --max-memory=800 --autoprovisioning-scopes=https://www.googleapis.com/auth/cloud-platform
Rightsizing Cloud SQL Instances
Cloud SQL is often the most oversized resource in a GCP environment because database sizing decisions are made conservatively and rarely revisited. A db-n1-standard-16 instance running at 8% CPU and 20% memory utilization is an extremely common sight.
Cloud SQL Metrics to Analyze
Key metrics for rightsizing Cloud SQL:
| Metric | Location | Rightsizing Trigger |
|---|---|---|
| CPU utilization | cloudsql.googleapis.com/database/cpu/utilization |
P95 < 40% → downsize |
| Memory usage | cloudsql.googleapis.com/database/memory/usage |
P95 < 50% of total → downsize |
| Connections | cloudsql.googleapis.com/database/postgresql/num_backends |
Compare to max_connections |
| Disk IOPS | cloudsql.googleapis.com/database/disk/read_ops_count |
Check if storage tier is overprovisioned |
# Check Cloud SQL instance metrics via gcloud
gcloud monitoring metrics list --filter="metric.type:cloudsql" --format="value(metric.type)" | head -20
# Get CPU utilization for a specific instance
gcloud monitoring read 'cloudsql.googleapis.com/database/cpu/utilization' --project=my-project --filter='resource.labels.database_id="my-project:my-instance"' --freshness=PT2W --align=ALIGN_PERCENTILE_95
Cloud SQL Rightsizing Procedure
# List current Cloud SQL instances and their tiers
gcloud sql instances list --format="table(name,databaseVersion,settings.tier,region,state)"
# For PostgreSQL: check pg_stat_activity for connection count
# Connect via Cloud SQL proxy or Cloud Shell
# Edit an instance to downsize (requires brief restart)
gcloud sql instances patch my-postgres-instance --tier=db-n1-standard-4 # was db-n1-standard-16
# For high-availability instances, zero-downtime failover first
gcloud sql instances failover my-postgres-ha-instance
Important: Cloud SQL tier changes cause a brief restart (typically 60-120 seconds). Schedule during a maintenance window. For HA instances, the restart is faster due to automatic failover, but plan for it.
Cloud SQL Storage Rightsizing
Cloud SQL also charges for provisioned storage even when unused. Enable automatic storage increases, but check that existing instances aren't dramatically overprovisioned:
# Check storage utilization
gcloud sql instances describe my-instance --format="table(name,settings.dataDiskSizeGb,diskUsedInMb)"
# Storage can only be increased, not decreased — be conservative with initial sizing
# For a new instance, start small and enable autoresize:
gcloud sql instances create new-instance --database-version=POSTGRES_15 --tier=db-n1-standard-4 --region=europe-west4 --storage-size=20GB --storage-auto-increase
Tracking Rightsizing Savings
After rightsizing, measure the actual savings in your BigQuery billing export. Compare the 4 weeks before and after the change for each resource:
-- Compare costs before and after rightsizing date
WITH before AS (
SELECT
resource.name,
ROUND(SUM(cost), 2) AS cost_before
FROM `my-project.billing_export.gcp_billing_export_v1_*`
WHERE DATE(usage_start_time) BETWEEN '2025-01-01' AND '2025-01-31'
AND resource.name LIKE '%api-server%'
GROUP BY resource.name
),
after AS (
SELECT
resource.name,
ROUND(SUM(cost), 2) AS cost_after
FROM `my-project.billing_export.gcp_billing_export_v1_*`
WHERE DATE(usage_start_time) BETWEEN '2025-02-01' AND '2025-02-28'
AND resource.name LIKE '%api-server%'
GROUP BY resource.name
)
SELECT
before.resource_name,
cost_before,
cost_after,
ROUND(cost_before - cost_after, 2) AS monthly_savings,
ROUND((cost_before - cost_after) / cost_before * 100, 1) AS savings_pct
FROM before
JOIN after USING (resource_name)
WHERE cost_before > cost_after;
Building a Rightsizing Practice
One-time rightsizing is better than nothing, but the real value comes from making it a quarterly practice. Cloud workloads evolve — a service that was CPU-heavy six months ago may now be memory-bound after a refactor. Set a quarterly calendar reminder to:
- Pull Recommendations Hub for all projects
- Query P95 utilization from Cloud Monitoring for top 20 cost items
- Run VPA in recommendation mode for 1 week on underanalyzed GKE workloads
- Review Cloud SQL instance utilization
- Implement changes and document savings
For the broader cost management picture including labeling and budgets, see our GCP cloud billing guide. For committed use discounts on your rightsized baseline, see our Google Cloud FinOps guide.