Cloud Run vs Cloud Functions vs App Engine: Choosing the Right GCP Serverless Platform
GCP offers three serverless compute platforms, each with a different operational model and set of trade-offs. This guide compares Cloud Run, Cloud Functions, and App Engine across cold start behavior, pricing, scaling, and workload fit to help you pick the right platform.
GCP has three serverless compute platforms that can all run application code without you managing servers. If you're new to GCP, the overlap is confusing: all three scale automatically, all three are billed based on usage, and all three have zero-infrastructure-management promises. So why does GCP offer all three?
The answer is that they target different operational models and workload types. Choosing the wrong one doesn't break anything immediately — you can deploy a web app to Cloud Functions and make it work — but you pay for the wrong abstractions in complexity, cost, or constraints down the line.
This guide gives you the mental model to choose correctly from the start. We'll compare each platform on architecture, cold start behavior, pricing, scaling, and workload fit, then give you a clear decision framework.
The Three Platforms, Briefly
Cloud Run runs any containerized workload. You bring a Docker image, Cloud Run handles the infrastructure. It scales from zero to thousands of containers in response to HTTP requests or Pub/Sub events, and charges only for the time requests are actively being processed. Cloud Run is the most flexible of the three.
Cloud Functions is a function-as-a-service platform (FaaS). You write a function in Python, Node.js, Go, Java, Ruby, PHP, or .NET and deploy source code. Cloud Functions manages the runtime, dependencies, and execution environment. It's designed for event handlers and lightweight API backends.
App Engine is the oldest of the three. It was GCP's original PaaS (Platform as a Service) offering and supports running web applications in Standard or Flexible environment. App Engine Standard uses proprietary sandboxed runtimes; App Engine Flexible runs Docker containers on managed VMs. App Engine is best for web applications that benefit from built-in features like Cron jobs, task queues, and the Datastore integration.
Architecture and Deployment Model
Cloud Run
Cloud Run deployments are container-based. You build an image, push it to Artifact Registry, and point Cloud Run at it:
# Build and push your container
gcloud builds submit --tag europe-west4-docker.pkg.dev/my-project/my-repo/my-service:latest
# Deploy to Cloud Run
gcloud run deploy my-service --image=europe-west4-docker.pkg.dev/my-project/my-repo/my-service:latest --region=europe-west4 --platform=managed --min-instances=0 --max-instances=100 --memory=512Mi --cpu=1 --allow-unauthenticated
Cloud Run containers must listen on the port defined by the PORT environment variable (default 8080) and respond to HTTP requests. Beyond that, the container can run any language, any framework, any dependencies.
Cloud Functions (2nd Gen)
Cloud Functions 2nd gen runs on Cloud Run under the hood. You write a function, deploy source code, and GCP handles containerization:
# main.py
import functions_framework
from flask import Request, jsonify
@functions_framework.http
def process_webhook(request: Request):
payload = request.get_json()
result = process_event(payload)
return jsonify(result), 200
gcloud functions deploy process-webhook --gen2 --runtime=python311 --region=europe-west4 --source=. --entry-point=process_webhook --trigger-http --allow-unauthenticated --memory=256Mi --max-instances=100
Cloud Functions is simpler to deploy but less flexible — you can't customize the container, add a web server middleware stack, or run background threads.
App Engine
App Engine deployments use app.yaml to describe the runtime and scaling configuration:
# app.yaml
runtime: python311
service: my-web-app
automatic_scaling:
min_instances: 1
max_instances: 20
target_cpu_utilization: 0.6
env_variables:
DATABASE_URL: "postgresql://..."
handlers:
- url: /static
static_dir: static
- url: /.*
script: auto
gcloud app deploy app.yaml --project=my-project
App Engine Standard has the most built-in features: Cron jobs (cron.yaml), task queues (queue.yaml), and dispatch routing (dispatch.yaml). But it constrains you to specific runtime versions and sandboxes.
Cold Start Performance
Cold starts are the latency added when your platform needs to initialize a new container or function instance to handle an incoming request.
| Platform | Typical Cold Start | Min Instances = 0 Impact |
|---|---|---|
| Cloud Run | 1-4 seconds | Every scaling event |
| Cloud Functions 2nd gen | 1-4 seconds | Every scaling event |
| Cloud Functions 1st gen | 0.5-2 seconds | Less severe (lighter) |
| App Engine Standard | 0.5-2 seconds | Less severe (lighter runtime) |
| App Engine Flexible | 30-60 seconds | Significant |
Cloud Functions 2nd gen and Cloud Run have similar cold start characteristics because 2nd gen runs on Cloud Run infrastructure.
Mitigating cold starts:
# Cloud Run: Set minimum instances to keep warm instances
gcloud run services update my-service --region=europe-west4 --min-instances=2 # Always keep 2 instances warm
# Cloud Run: CPU always allocated (even when not handling requests)
gcloud run services update my-service --region=europe-west4 --cpu-throttling=false # CPU stays active between requests
The --min-instances setting eliminates cold starts for sustained traffic at the cost of paying for idle instances. Use it for latency-sensitive user-facing services.
Pricing Model Comparison
This is where the platforms diverge most significantly.
Cloud Run Pricing
Cloud Run charges for:
- CPU: $0.00002400/vCPU-second (while processing requests)
- Memory: $0.00000250/GB-second (while processing requests)
- Requests: $0.40/million requests
If you set --cpu-throttling=false (CPU always allocated), you also pay for idle time:
- CPU (always allocated): $0.00001800/vCPU-second (slightly cheaper rate)
Example: A service receiving 1 million requests/day, each taking 200ms with 1 vCPU and 512MB:
- CPU: 1M × 0.2s × 1 vCPU × $0.000024 = $4.80/day
- Memory: 1M × 0.2s × 0.5 GB × $0.0000025 = $0.25/day
- Requests: 1M × $0.0000004 = $0.40/day
- Total: ~$5.45/day = ~$163/month
Cloud Functions Pricing
Cloud Functions 2nd gen uses the same Cloud Run pricing model (it's the same infrastructure). 1st gen pricing is slightly different but comparable for most workloads.
For lightweight functions (256MB, 0.3 GHz), 1st gen can be cheaper due to lower compute commitment per invocation.
App Engine Pricing
App Engine Standard pricing is based on instance hours — the time your instances are running, regardless of whether they're handling requests:
- B2 instance (600MHz, 256MB): $0.045/hour in europe-west4
- F2 instance (1.2GHz, 512MB): $0.055/hour
For a service with 2 minimum instances: 2 × $0.045 × 730 hours = $65.70/month just for idle capacity, regardless of request volume.
App Engine Flexible pricing is based on underlying Compute Engine VMs and is typically more expensive than Cloud Run for variable workloads.
Cost verdict: Cloud Run wins for variable, bursty, or low-traffic workloads where you want to pay only for actual processing. App Engine Standard can be competitive for high-traffic applications where instances are always busy.
Scaling Behavior
All three platforms support automatic scaling, but with different mechanics.
Cloud Run: Scales based on concurrent requests per container (default: 80 concurrent requests per instance). New instances spin up when all existing instances are at capacity. Scales from 0 to max-instances in roughly 1-3 seconds per instance.
# Configure concurrency - lower means more instances but better isolation
gcloud run services update my-service --region=europe-west4 --concurrency=10 # Good for CPU-intensive work
# --concurrency=1000 # Good for I/O-bound services
Cloud Functions: Each function invocation gets its own instance (concurrency=1 in 1st gen; configurable in 2nd gen). This makes functions naturally isolated — one invocation can't slow another — but means you pay more per unit of compute.
App Engine: Scales based on configurable metrics (CPU utilization, concurrent requests, request latency). Standard environment is faster to scale (5-10 seconds); Flexible takes 2-5 minutes.
Workload Fit: A Decision Guide
Use Cloud Run When:
- Your workload is already containerized or you're starting fresh
- You need custom runtime dependencies (specific OS packages, binary tools)
- You're building a long-running HTTP service or streaming response API
- You want WebSocket support or streaming
- You need to run any language not supported by Cloud Functions
- You're using Vertex AI models and want to wrap them in a service API
- You need precise control over concurrency model
# Cloud Run streaming example: stream LLM responses to browser
gcloud run deploy llm-api --image=my-llm-api:latest --region=europe-west4 --execution-environment=gen2 --cpu=4 --memory=8Gi --timeout=300 # Allow long-running streaming responses
Use Cloud Functions When:
- You want maximum deployment simplicity (deploy source code, not containers)
- You're building event handlers: Pub/Sub, Cloud Storage, Firestore triggers
- You're writing lightweight API endpoints or webhooks
- Your team prefers not to manage Docker images
- You need tight Eventarc integration for event-driven workflows
# Perfect Cloud Functions use case: handle Pub/Sub events
import functions_framework
import base64
import json
@functions_framework.cloud_event
def handle_order_event(cloud_event):
data = json.loads(base64.b64decode(cloud_event.data["message"]["data"]))
process_order(data["order_id"])
Use App Engine When:
- You're migrating an existing App Engine application and prefer not to rewrite
- You need built-in App Engine features: Cron, Task Queues, or dispatch routing
- Your application uses App Engine's built-in Datastore or memcache integration
- You have a monolithic web application with complex routing requirements
The Reality in 2025
Cloud Run has become the default choice for most new GCP serverless deployments. Cloud Functions 2nd gen runs on Cloud Run under the hood, which means the operational distinction between them has narrowed — Functions is essentially Cloud Run with source-based deployment. The main remaining differentiator for Cloud Functions is its native Eventarc trigger support and simpler developer experience for event handlers.
App Engine sees less new development activity. Google hasn't deprecated it, but the momentum and documentation investment has clearly shifted to Cloud Run. For new projects, start with Cloud Run.
For existing App Engine Standard applications, migration to Cloud Run is feasible but requires containerization work. For simple applications, the App Engine Buildpacks can generate a container automatically:
# Convert App Engine source to Cloud Run container using buildpacks
pack build my-app --builder=gcr.io/buildpacks/builder:v1 --env=GOOGLE_RUNTIME=python
gcloud run deploy my-app --image=my-app:latest --region=europe-west4
For production Cloud Run deployments, see our Cloud Run event-driven microservices guide and our Cloud Run performance tuning guide. For securing serverless workloads on GCP, see our GCP serverless security guide.