Cloud Run Concurrency, Cold Starts, and Performance Tuning
Cloud Run's performance characteristics are often misunderstood. This guide explains how concurrency settings, CPU allocation modes, startup probes, and container optimization interact to determine your service's actual latency and scaling behavior.
Cloud Run's automatic scaling model hides infrastructure management from you but doesn't hide infrastructure behavior. A service that performs well at low traffic can degrade dramatically under load if you haven't configured concurrency correctly. A service that starts in 200ms in local testing can add 4 seconds of latency to first requests if the container initialization path isn't optimized.
Understanding how Cloud Run actually handles traffic — from the moment a request arrives to the moment a response is sent — lets you configure it for your specific performance requirements. This guide covers the mechanics and then the practical tuning knobs.
How Cloud Run Handles Requests
When a request arrives at a Cloud Run service, the following sequence occurs:
- Cloud Run's load balancer receives the request
- Load balancer routes to an existing container instance if one has capacity
- If no capacity exists, a new container instance is started (cold start)
- The container starts, runs initialization code, and signals readiness
- The request is routed to the container
- The container processes the request and returns a response
- After the configured idle timeout, instances with no traffic are terminated
Steps 3-4 are what we call a "cold start." Steps 5-6 are the actual request processing time. Your users feel the sum of all of these.
Concurrency: The Most Misunderstood Setting
Cloud Run's concurrency setting (default: 80) controls how many requests a single container instance can handle simultaneously. This is not the number of parallel workers — it's the number of concurrent HTTP connections the container accepts.
Getting this wrong is the most common source of Cloud Run performance problems.
Too high (e.g., 1000 for a CPU-bound service): Multiple concurrent requests compete for the same CPU. A service that takes 500ms under single-tenant load may take 5 seconds when 10 requests are sharing the same CPU core. The container appears available and Cloud Run sends more requests, making the problem worse.
Too low (e.g., 1 for an I/O-bound service): The container sits idle while waiting for network calls, but Cloud Run scales out more instances to handle additional requests. You end up paying for many idle container instances when a few containers with higher concurrency would handle the same load efficiently.
Choosing the Right Concurrency Setting
# For CPU-bound workloads (ML inference, image processing, report generation)
# Set concurrency = number of CPU cores allocated
gcloud run services update my-ml-service --region=europe-west4 --cpu=4 --concurrency=4 # 1 request per CPU core
# For I/O-bound workloads (database queries, API calls, file I/O)
# Set concurrency high — the container spends most time waiting, not computing
gcloud run services update my-api-service --region=europe-west4 --cpu=2 --concurrency=100 # Container can handle 100 concurrent I/O-bound requests
# For mixed workloads
# Benchmark to find the sweet spot — typically 10-50 for web services
gcloud run services update my-web-service --region=europe-west4 --cpu=2 --concurrency=20
Measuring the Right Concurrency Value
Load test your service to find the concurrency level where latency starts to degrade:
# Install hey (HTTP load testing tool)
go install github.com/rakyll/hey@latest
# Test with increasing concurrency
for concurrency in 1 5 10 20 50 100; do
echo "=== Concurrency: $concurrency ==="
hey -n 1000 -c $concurrency https://my-service-xxxxx-ew.a.run.app/api/endpoint
done
The output shows latency percentiles at each concurrency level. The inflection point where p99 latency starts growing faster than linearly is your concurrency ceiling.
Cold Start Reduction Strategies
Cold starts occur when Cloud Run needs to start a new container instance. The time includes:
- Pulling the container image from Artifact Registry
- Allocating compute resources
- Running your container's initialization code
- Waiting for the health check to pass
Strategy 1: Minimize Image Size
# Before: Full Python image — ~900MB, ~3.5s cold start
FROM python:3.11
WORKDIR /app
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
CMD ["python", "main.py"]
# After: Distroless base — ~150MB, ~1.2s cold start
FROM python:3.11-slim as builder
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir --prefix=/install -r requirements.txt
FROM gcr.io/distroless/python3-debian12
COPY --from=builder /install /usr/local
COPY . /app
WORKDIR /app
CMD ["main.py"]
Smaller images transfer faster from Artifact Registry to the Cloud Run instance. Distroless images also have a smaller attack surface — no shell, no package manager, only what your application needs.
Strategy 2: Move Heavy Initialization to Module Level
Cloud Run caches the container state between requests (when the instance stays warm). Move expensive initialization outside of request handlers:
# Bad: Loads the model on every cold start AND every request
from flask import Flask, request
app = Flask(__name__)
@app.route("/predict")
def predict():
import tensorflow as tf
model = tf.keras.models.load_model("model.h5") # 2-3 second load time
return model.predict(request.get_json())
# Good: Loads the model once when container starts
from flask import Flask, request
import tensorflow as tf
app = Flask(__name__)
model = tf.keras.models.load_model("model.h5") # Runs once on cold start
@app.route("/predict")
def predict():
return model.predict(request.get_json()) # Fast — model already loaded
The model load happens during cold start, but not during warm request processing. This adds time to cold starts (unavoidable) but keeps warm requests fast.
Strategy 3: Minimum Instances
The most direct cold start elimination is keeping minimum instances running:
# Keep 3 warm instances running at all times
gcloud run services update my-service --region=europe-west4 --min-instances=3 --max-instances=100
Cost implication: minimum instances are billed for CPU and memory allocation at the CPU-always-on rate, even when idle. For a 2 vCPU, 1 GB service:
- 3 idle instances × $0.000018/vCPU-second × 2 vCPU × 86400 seconds/day ≈ $9.33/day
For user-facing services where cold start latency is unacceptable, this is usually worth it.
Strategy 4: CPU Always Allocated
By default, Cloud Run throttles CPU to 0 when a container is not processing a request. This means background goroutines, connection pool keep-alives, and cache warming stop working between requests.
# Keep CPU active even between requests
gcloud run services update my-service --region=europe-west4 --cpu-throttling=false
With --cpu-throttling=false, your container can:
- Maintain warm database connection pools
- Run background cache warming
- Respond to keep-alive pings
The trade-off is higher cost — you're billed for CPU even when idle.
Startup Probes and Health Checks
Cloud Run supports startup probes (checks whether the container has finished initializing) and health checks (checks whether the container is healthy enough to receive traffic).
# Service YAML with startup probe
apiVersion: serving.knative.dev/v1
kind: Service
metadata:
name: my-service
spec:
template:
spec:
containers:
- image: europe-west4-docker.pkg.dev/my-project/repo/my-service:latest
startupProbe:
httpGet:
path: /health
port: 8080
initialDelaySeconds: 5
periodSeconds: 3
failureThreshold: 10 # Allow 30 seconds for slow startup (10 × 3s)
timeoutSeconds: 2
livenessProbe:
httpGet:
path: /health
port: 8080
periodSeconds: 30
failureThreshold: 3
Implement a health endpoint that reflects actual readiness:
from flask import Flask, jsonify
import time
app = Flask(__name__)
start_time = time.time()
WARMUP_DURATION = 10 # Seconds to complete initialization
@app.route("/health")
def health():
# Check if initialization is complete
if time.time() - start_time < WARMUP_DURATION:
return jsonify({"status": "starting"}), 503
# Check database connection
try:
db.ping()
except Exception as e:
return jsonify({"status": "unhealthy", "reason": str(e)}), 503
return jsonify({"status": "healthy"}), 200
CPU Throttling and Background Work
If your service uses background threads or goroutines (e.g., for connection pool management, cache population, or metric collection), CPU throttling can cause subtle bugs:
# This background thread behaves differently with CPU throttling
import threading
import time
def refresh_cache():
while True:
# With CPU throttling: this sleeps for wall time but CPU is 0
# The sleep completes correctly, but the refresh work may stall
data = fetch_from_database()
cache.set("key", data)
time.sleep(60)
thread = threading.Thread(target=refresh_cache, daemon=True)
thread.start()
For services with background work, either use --cpu-throttling=false or refactor to lazy initialization — fetch from database on first request, not on a timer.
Memory Configuration and OOM Prevention
Cloud Run terminates containers that exceed their memory limit with OOM (Out of Memory) kills. These appear in logs as container failures and trigger cold starts for the next request.
# Monitor memory usage in Cloud Monitoring
gcloud logging read 'resource.type="cloud_run_revision" AND textPayload:"Memory limit"' --limit=20 --format="table(timestamp,textPayload)"
# Increase memory if OOM kills are occurring
gcloud run services update my-service --region=europe-west4 --memory=2Gi # Increase from 512Mi
Set container memory limits in code for languages with garbage collectors:
# For JVM-based services: constrain heap to leave room for off-heap
# If Cloud Run memory limit is 2GB, set JVM max heap to ~1.5GB
ENV JAVA_OPTS="-Xmx1500m -XX:+UseG1GC -XX:MaxGCPauseMillis=200"
Execution Environments: Gen1 vs Gen2
Cloud Run offers two execution environments:
First generation (default): Uses gVisor for container isolation. Lower resource overhead but some syscall restrictions. Better cold start time.
Second generation: Full Linux environment, no syscall restrictions. Supports network file systems, faster local disk I/O, and compatibility with binaries that rely on unsupported syscalls. Slightly longer cold starts.
# Use Gen2 for services that need full Linux compatibility
gcloud run services update my-service --region=europe-west4 --execution-environment=gen2
Choose Gen2 if:
- Your container uses FUSE file systems (GCS FUSE, NFS)
- You're running compiled binaries that use advanced syscalls
- You need local SSD access
Choose Gen1 (default) otherwise — it has slightly faster cold starts and lower per-instance overhead.
Putting It Together: Performance Checklist
For a production Cloud Run service, work through this checklist:
# 1. Right-size CPU and memory
gcloud run services update my-service --cpu=2 --memory=1Gi
# 2. Set concurrency based on workload type
gcloud run services update my-service --concurrency=20 # Adjust based on benchmarking
# 3. Configure minimum instances for latency-sensitive paths
gcloud run services update my-service --min-instances=2
# 4. Set CPU always allocated if you have background work
gcloud run services update my-service --cpu-throttling=false
# 5. Set appropriate request timeout
gcloud run services update my-service --timeout=300 # Max 3600 for long-running operations
# 6. Enable execution environment gen2 if needed
gcloud run services update my-service --execution-environment=gen2
For the full Cloud Run architecture in event-driven systems, see our Cloud Run and Pub/Sub microservices guide. For securing Cloud Run services, see our GCP serverless security guide.