GCP Rightsizing Guide: Optimizing Compute Engine, GKE, and Cloud SQL

Rightsizing is the single highest-ROI cloud optimization activity. Most GCP workloads run on machines that are 2-3x larger than necessary, paying for CPU and memory that sits idle. This guide covers systematic rightsizing for Compute Engine VMs, GKE node pools, and Cloud SQL instances using real utilization data.

Here's a reality check most cloud engineers don't talk about: the average GCP workload uses about 30-40% of the CPU it's paying for. Memory utilization is often even lower. We size for peak, add a "safety margin," and then rarely revisit those decisions as workloads evolve.

Rightsizing — matching your machine types to actual utilization — consistently delivers 30-60% cost reductions on compute. It's not glamorous work, but it's the highest ROI optimization you can do before reaching for committed use discounts or other financial instruments.

This guide covers rightsizing for the three biggest GCP compute cost centers: Compute Engine VMs, GKE node pools, and Cloud SQL instances. For each, we'll look at how to measure utilization, what targets to aim for, and how to safely resize without causing downtime.

The Rightsizing Mindset

Before diving into specifics, two mental model shifts that make rightsizing less scary:

Target 60-70% utilization, not 100%: You're not trying to jam resources to the limit. You want headroom for traffic spikes and process bursts. A VM with 65% average CPU utilization with spikes to 80% is well-sized. A VM running at 8% average is dramatically oversized.

Use percentiles, not averages: Average CPU utilization is misleading. A VM that sits at 5% for 23 hours but spikes to 90% for one hour has an average of ~9% — but you absolutely need that capacity for that one hour. Use P95 or P99 utilization over a 2-4 week window for sizing decisions.

Rightsizing Compute Engine VMs

Finding Oversized VMs with Cloud Monitoring

The Recommendations Hub gives you machine-type recommendations automatically, but it's helpful to understand what data it's using.

# Get VM rightsizing recommendations for a project
gcloud recommender recommendations list   --project=my-project   --location=us-central1-a   --recommender=google.compute.instance.MachineTypeRecommender   --format="table(name,description,primaryImpact.costProjection.cost.units,stateInfo.state)"

For custom analysis using Cloud Monitoring metrics:

from google.cloud import monitoring_v3
from datetime import datetime, timedelta
import time

client = monitoring_v3.MetricServiceClient()
project_name = f"projects/my-project"

# Query P95 CPU utilization for all instances over 14 days
now = time.time()
interval = monitoring_v3.TimeInterval({
    "end_time": {"seconds": int(now)},
    "start_time": {"seconds": int(now - 14 * 86400)},
})

results = client.list_time_series(
    request={
        "name": project_name,
        "filter": 'metric.type="compute.googleapis.com/instance/cpu/utilization"',
        "interval": interval,
        "aggregation": {
            "alignment_period": {"seconds": 86400},
            "per_series_aligner": monitoring_v3.Aggregation.Aligner.ALIGN_PERCENTILE_95,
            "cross_series_reducer": monitoring_v3.Aggregation.Reducer.REDUCE_NONE,
        },
    }
)

for series in results:
    instance_name = series.resource.labels["instance_id"]
    p95_values = [point.value.double_value for point in series.points]
    avg_p95 = sum(p95_values) / len(p95_values) if p95_values else 0

    if avg_p95 < 0.20:  # P95 CPU below 20% — likely oversized
        print(f"OVERSIZED: {instance_name} | P95 CPU: {avg_p95*100:.1f}%")

Compute Engine Machine Type Selection

When you find an oversized VM, choose the right replacement machine series:

Use Case Machine Series Notes
General workload N2 or N2D Best price/performance for most apps
Memory-intensive M2 DB caching, in-memory analytics
Compute-intensive C2 Video transcoding, batch compute
Cost-sensitive, interruptible Spot VMs 60-91% savings vs on-demand
ARM workloads T2A Smaller per-core cost, great for containerized apps
# Resize a stopped VM to a smaller machine type
gcloud compute instances set-machine-type api-server-prod   --zone=europe-west4-a   --machine-type=n2-standard-4  # was n2-standard-16

# Check available machine types in a zone
gcloud compute machine-types list   --filter="zone:europe-west4-a AND name~n2-standard"   --format="table(name,guestCpus,memoryMb,description)"

Safe Resizing Procedure

Never resize a production VM blindly. Follow this sequence:

# 1. Snapshot the disk before resizing
gcloud compute disks snapshot api-server-prod-disk   --zone=europe-west4-a   --snapshot-names=pre-resize-snapshot-$(date +%Y%m%d)

# 2. Stop the instance
gcloud compute instances stop api-server-prod --zone=europe-west4-a

# 3. Change machine type
gcloud compute instances set-machine-type api-server-prod   --zone=europe-west4-a   --machine-type=n2-standard-4

# 4. Start the instance
gcloud compute instances start api-server-prod --zone=europe-west4-a

# 5. Monitor for 15 minutes — if issues, revert
# gcloud compute instances set-machine-type api-server-prod #   --zone=europe-west4-a #   --machine-type=n2-standard-16

For instances that can't tolerate downtime, use a blue/green approach: provision a new smaller VM, migrate traffic via load balancer, verify, then decommission the old VM.

Rightsizing GKE Node Pools

GKE adds a layer of complexity because you're sizing nodes to fit pods, not applications directly. The unit of analysis is the node pool, and the goal is maximizing bin packing efficiency — the percentage of allocated resources actually used.

Understanding GKE Resource Allocation

In GKE, every node has a portion of its CPU and memory reserved for system overhead. For an n2-standard-4 (4 vCPU, 16 GB):

  • System reserved CPU: ~0.09 vCPU
  • System reserved memory: ~1.1 GB
  • Eviction threshold: ~100 MB memory
  • Allocatable CPU: ~3.91 vCPU
  • Allocatable memory: ~14.4 GB

Your pods' requests (not limits) consume allocatable capacity. Bin packing efficiency = (sum of pod requests) / (allocatable capacity).

# View node allocatable capacity and current allocations
kubectl describe nodes | grep -A 5 "Allocatable:"
kubectl top nodes

# View pod resource requests vs limits
kubectl get pods -A -o custom-columns="NAMESPACE:.metadata.namespace,NAME:.metadata.name,CPU_REQ:.spec.containers[*].resources.requests.cpu,MEM_REQ:.spec.containers[*].resources.requests.memory"

Identifying Underutilized Node Pools

# Enable cluster autoscaler metrics view
kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml

# Get per-node utilization
kubectl top nodes --sort-by=cpu

# Find nodes with low utilization (candidates for consolidation)
kubectl top nodes | awk 'NR>1 && $3+0 < 30 {print "LOW_CPU:", $0}'

For production clusters, use the GKE Cost Optimization Insights dashboard in the console (Kubernetes Engine → Clusters → Cost → Optimization Insights). It shows:

  • Node pool CPU/memory utilization
  • Over-provisioning percentage
  • Estimated monthly savings from rightsizing

Node Pool Rightsizing Strategy

The most effective GKE rightsizing approach combines three techniques:

1. Enable Vertical Pod Autoscaler (VPA) in recommendation mode first

VPA analyzes pod utilization and recommends appropriate requests/limits without automatically changing anything:

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: api-vpa
  namespace: production
spec:
  targetRef:
    apiVersion: "apps/v1"
    kind: Deployment
    name: api
  updatePolicy:
    updateMode: "Off"  # Recommendation only, no auto-updates
  resourcePolicy:
    containerPolicies:
    - containerName: api
      minAllowed:
        cpu: 100m
        memory: 128Mi
      maxAllowed:
        cpu: 2
        memory: 2Gi
# View VPA recommendations
kubectl get vpa api-vpa -o yaml | grep -A 20 "recommendation:"

2. Right-size pod requests based on VPA recommendations

After a week of VPA recommendations, update your deployments to match recommended requests:

resources:
  requests:
    cpu: "250m"    # was "1000m"
    memory: "512Mi"  # was "2Gi"
  limits:
    cpu: "1000m"
    memory: "1Gi"

3. Scale node pool machine type down to match improved density

After right-sizing pod requests, you can use fewer or smaller nodes:

# Create new right-sized node pool
gcloud container node-pools create rightsized-pool   --cluster=production   --region=europe-west4   --machine-type=n2-standard-4   --num-nodes=3   --enable-autoscaling   --min-nodes=2   --max-nodes=8

# Cordon and drain the old oversized pool
kubectl cordon -l cloud.google.com/gke-nodepool=large-pool
kubectl drain -l cloud.google.com/gke-nodepool=large-pool --ignore-daemonsets --delete-emptydir-data

# Delete the old pool
gcloud container node-pools delete large-pool   --cluster=production   --region=europe-west4

Node Auto-Provisioning for Automatic Rightsizing

If you're regularly adding new workloads with different resource profiles, enable Node Auto-Provisioning (NAP). NAP automatically creates node pools with machine types optimized for your pending pods:

gcloud container clusters update production   --region=europe-west4   --enable-autoprovisioning   --min-cpu=4 --max-cpu=200   --min-memory=16 --max-memory=800   --autoprovisioning-scopes=https://www.googleapis.com/auth/cloud-platform

Rightsizing Cloud SQL Instances

Cloud SQL is often the most oversized resource in a GCP environment because database sizing decisions are made conservatively and rarely revisited. A db-n1-standard-16 instance running at 8% CPU and 20% memory utilization is an extremely common sight.

Cloud SQL Metrics to Analyze

Key metrics for rightsizing Cloud SQL:

Metric Location Rightsizing Trigger
CPU utilization cloudsql.googleapis.com/database/cpu/utilization P95 < 40% → downsize
Memory usage cloudsql.googleapis.com/database/memory/usage P95 < 50% of total → downsize
Connections cloudsql.googleapis.com/database/postgresql/num_backends Compare to max_connections
Disk IOPS cloudsql.googleapis.com/database/disk/read_ops_count Check if storage tier is overprovisioned
# Check Cloud SQL instance metrics via gcloud
gcloud monitoring metrics list   --filter="metric.type:cloudsql"   --format="value(metric.type)" | head -20

# Get CPU utilization for a specific instance
gcloud monitoring read   'cloudsql.googleapis.com/database/cpu/utilization'   --project=my-project   --filter='resource.labels.database_id="my-project:my-instance"'   --freshness=PT2W   --align=ALIGN_PERCENTILE_95

Cloud SQL Rightsizing Procedure

# List current Cloud SQL instances and their tiers
gcloud sql instances list   --format="table(name,databaseVersion,settings.tier,region,state)"

# For PostgreSQL: check pg_stat_activity for connection count
# Connect via Cloud SQL proxy or Cloud Shell

# Edit an instance to downsize (requires brief restart)
gcloud sql instances patch my-postgres-instance   --tier=db-n1-standard-4  # was db-n1-standard-16

# For high-availability instances, zero-downtime failover first
gcloud sql instances failover my-postgres-ha-instance

Important: Cloud SQL tier changes cause a brief restart (typically 60-120 seconds). Schedule during a maintenance window. For HA instances, the restart is faster due to automatic failover, but plan for it.

Cloud SQL Storage Rightsizing

Cloud SQL also charges for provisioned storage even when unused. Enable automatic storage increases, but check that existing instances aren't dramatically overprovisioned:

# Check storage utilization
gcloud sql instances describe my-instance   --format="table(name,settings.dataDiskSizeGb,diskUsedInMb)"

# Storage can only be increased, not decreased — be conservative with initial sizing
# For a new instance, start small and enable autoresize:
gcloud sql instances create new-instance   --database-version=POSTGRES_15   --tier=db-n1-standard-4   --region=europe-west4   --storage-size=20GB   --storage-auto-increase

Tracking Rightsizing Savings

After rightsizing, measure the actual savings in your BigQuery billing export. Compare the 4 weeks before and after the change for each resource:

-- Compare costs before and after rightsizing date
WITH before AS (
  SELECT
    resource.name,
    ROUND(SUM(cost), 2) AS cost_before
  FROM `my-project.billing_export.gcp_billing_export_v1_*`
  WHERE DATE(usage_start_time) BETWEEN '2025-01-01' AND '2025-01-31'
    AND resource.name LIKE '%api-server%'
  GROUP BY resource.name
),
after AS (
  SELECT
    resource.name,
    ROUND(SUM(cost), 2) AS cost_after
  FROM `my-project.billing_export.gcp_billing_export_v1_*`
  WHERE DATE(usage_start_time) BETWEEN '2025-02-01' AND '2025-02-28'
    AND resource.name LIKE '%api-server%'
  GROUP BY resource.name
)
SELECT
  before.resource_name,
  cost_before,
  cost_after,
  ROUND(cost_before - cost_after, 2) AS monthly_savings,
  ROUND((cost_before - cost_after) / cost_before * 100, 1) AS savings_pct
FROM before
JOIN after USING (resource_name)
WHERE cost_before > cost_after;

Building a Rightsizing Practice

One-time rightsizing is better than nothing, but the real value comes from making it a quarterly practice. Cloud workloads evolve — a service that was CPU-heavy six months ago may now be memory-bound after a refactor. Set a quarterly calendar reminder to:

  1. Pull Recommendations Hub for all projects
  2. Query P95 utilization from Cloud Monitoring for top 20 cost items
  3. Run VPA in recommendation mode for 1 week on underanalyzed GKE workloads
  4. Review Cloud SQL instance utilization
  5. Implement changes and document savings

For the broader cost management picture including labeling and budgets, see our GCP cloud billing guide. For committed use discounts on your rightsized baseline, see our Google Cloud FinOps guide.