Vertex AI Platform: Complete MLOps Guide on Google Cloud

Vertex AI is Google's unified ML platform — it covers everything from data preparation to model serving and monitoring. This guide walks through a production MLOps workflow: training pipelines, model registry, endpoint deployment, prediction drift detection, and CI/CD automation with Cloud Build.

The gap between a model that works in a notebook and a model that reliably serves predictions in production is where most ML projects fail. Training a model is the easy part. Deploying it repeatably, versioning it properly, monitoring it for drift, retraining when performance degrades, and rolling back when something goes wrong — that's where teams struggle.

MLOps is the set of practices that bring software engineering discipline to machine learning workflows. Vertex AI is Google's managed platform for implementing those practices on GCP. It's not a single tool — it's a collection of managed services that cover each step of the ML lifecycle: Vertex AI Pipelines for workflow orchestration, Model Registry for versioning, Endpoints for serving, Feature Store for feature management, and Model Monitoring for drift detection.

This guide walks through a complete production MLOps workflow on Vertex AI, from pipeline definition to serving to monitoring. We'll use a customer churn prediction model as the example.

Vertex AI Architecture Overview

Before diving into code, here's the component map:

Vertex AI Workbench: Managed JupyterLab environments for exploration and experimentation Vertex AI Training: Managed compute for training jobs (custom containers or pre-built runtimes) Vertex AI Pipelines: Kubeflow Pipelines v2 for orchestrating multi-step ML workflows Vertex AI Model Registry: Versioned model storage with metadata and lineage Vertex AI Endpoints: Managed serving infrastructure for real-time predictions Vertex AI Batch Prediction: Asynchronous prediction on large datasets Vertex AI Feature Store: Managed feature serving with online and offline access Vertex AI Model Monitoring: Prediction distribution drift detection

Step 1: Define Your Training Pipeline

Vertex AI Pipelines uses Kubeflow Pipelines (KFP) SDK to define ML workflows as directed acyclic graphs (DAGs). Each component is a containerized step.

# pipeline.py
from kfp import dsl, compiler
from kfp.dsl import Dataset, Model, Input, Output, Metrics
from google.cloud import aiplatform

@dsl.component(
    base_image="python:3.11",
    packages_to_install=["pandas", "scikit-learn", "google-cloud-bigquery"],
)
def extract_training_data(
    project_id: str,
    dataset_id: str,
    start_date: str,
    end_date: str,
    output_dataset: Output[Dataset],
):
    """Extract training data from BigQuery."""
    from google.cloud import bigquery

    client = bigquery.Client(project=project_id)
    query = f"""
        SELECT * FROM `{project_id}.{dataset_id}.customer_features`
        WHERE snapshot_date BETWEEN '{start_date}' AND '{end_date}'
          AND churned IS NOT NULL
    """
    df = client.query(query).to_dataframe()
    df.to_parquet(output_dataset.path, index=False)
    print(f"Extracted {len(df)} rows for training")

@dsl.component(
    base_image="python:3.11",
    packages_to_install=["pandas", "scikit-learn", "xgboost"],
)
def train_model(
    input_dataset: Input[Dataset],
    test_size: float,
    n_estimators: int,
    output_model: Output[Model],
    metrics: Output[Metrics],
):
    """Train XGBoost churn model and save in XGBoost native format."""
    import pandas as pd
    from sklearn.model_selection import train_test_split
    from sklearn.metrics import roc_auc_score, f1_score
    from xgboost import XGBClassifier
    import json
    import os

    df = pd.read_parquet(input_dataset.path)
    feature_cols = [c for c in df.columns if c not in ["churned", "customer_id", "snapshot_date"]]

    X = df[feature_cols]
    y = df["churned"]

    X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=test_size, stratify=y)

    model = XGBClassifier(
        n_estimators=n_estimators,
        max_depth=6,
        learning_rate=0.1,
        use_label_encoder=False,
        eval_metric="logloss",
    )
    model.fit(X_train, y_train, eval_set=[(X_test, y_test)], early_stopping_rounds=20, verbose=False)

    y_pred_proba = model.predict_proba(X_test)[:, 1]
    auc = roc_auc_score(y_test, y_pred_proba)
    f1 = f1_score(y_test, (y_pred_proba > 0.5).astype(int))

    metrics.log_metric("auc", auc)
    metrics.log_metric("f1_score", f1)
    metrics.log_metric("n_training_samples", len(X_train))

    # Save in XGBoost native JSON format (not pickle)
    os.makedirs(output_model.path, exist_ok=True)
    model.save_model(f"{output_model.path}/model.json")

    with open(f"{output_model.path}/features.json", "w") as f:
        json.dump(feature_cols, f)

    output_model.metadata["framework"] = "xgboost"
    output_model.metadata["auc"] = auc

@dsl.component(
    base_image="python:3.11",
    packages_to_install=["google-cloud-aiplatform"],
)
def register_model(
    model: Input[Model],
    project_id: str,
    region: str,
    model_display_name: str,
    serving_container_image: str,
) -> str:
    """Register the model in Vertex AI Model Registry."""
    from google.cloud import aiplatform

    aiplatform.init(project=project_id, location=region)

    vertex_model = aiplatform.Model.upload(
        display_name=model_display_name,
        artifact_uri=model.uri,
        serving_container_image_uri=serving_container_image,
        serving_container_environment_variables={
            "MODEL_PATH": "/model/model.json",
            "FEATURES_PATH": "/model/features.json",
        },
        labels={
            "framework": "xgboost",
            "pipeline": "churn-prediction",
        },
    )

    return vertex_model.resource_name

@dsl.pipeline(
    name="churn-training-pipeline",
    description="Train and register customer churn prediction model",
)
def churn_pipeline(
    project_id: str,
    region: str,
    dataset_id: str = "analytics",
    start_date: str = "2024-01-01",
    end_date: str = "2024-09-30",
    test_size: float = 0.2,
    n_estimators: int = 200,
):
    extract_task = extract_training_data(
        project_id=project_id,
        dataset_id=dataset_id,
        start_date=start_date,
        end_date=end_date,
    )

    train_task = train_model(
        input_dataset=extract_task.outputs["output_dataset"],
        test_size=test_size,
        n_estimators=n_estimators,
    )

    register_task = register_model(
        model=train_task.outputs["output_model"],
        project_id=project_id,
        region=region,
        model_display_name="churn-predictor",
        serving_container_image=f"europe-west4-docker.pkg.dev/{project_id}/ml-models/churn-server:latest",
    )

    return register_task.output

compiler.Compiler().compile(churn_pipeline, "churn_pipeline.yaml")

Step 2: Submit the Pipeline

from google.cloud import aiplatform

aiplatform.init(project="my-project", location="europe-west4")

job = aiplatform.PipelineJob(
    display_name="churn-training-2025-01",
    template_path="churn_pipeline.yaml",
    pipeline_root="gs://my-project-pipelines/churn",
    parameter_values={
        "project_id": "my-project",
        "region": "europe-west4",
        "start_date": "2024-01-01",
        "end_date": "2024-09-30",
    },
)

job.submit(service_account="vertex-pipelines-sa@my-project.iam.gserviceaccount.com")

Step 3: Build the Serving Container

# serving/main.py
import os
import json
from xgboost import XGBClassifier
import numpy as np
from flask import Flask, request, jsonify

app = Flask(__name__)

MODEL_PATH = os.environ.get("MODEL_PATH", "/model/model.json")
FEATURES_PATH = os.environ.get("FEATURES_PATH", "/model/features.json")

# Load model in XGBoost native format (not pickle)
model = XGBClassifier()
model.load_model(MODEL_PATH)

with open(FEATURES_PATH) as f:
    feature_names = json.load(f)

@app.route("/predict", methods=["POST"])
def predict():
    body = request.get_json()
    instances = body.get("instances", [])

    if not instances:
        return jsonify({"error": "No instances provided"}), 400

    X = np.array([
        [instance.get(feat, 0) for feat in feature_names]
        for instance in instances
    ])

    probs = model.predict_proba(X)[:, 1]

    predictions = [
        {
            "churn_probability": float(prob),
            "churn_risk": "high" if prob > 0.7 else "medium" if prob > 0.4 else "low",
        }
        for prob in probs
    ]

    return jsonify({"predictions": predictions})

@app.route("/health", methods=["GET"])
def health():
    return jsonify({"status": "healthy"}), 200

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=int(os.environ.get("AIP_HTTP_PORT", 8080)))

Step 4: Deploy to Vertex AI Endpoint

from google.cloud import aiplatform

aiplatform.init(project="my-project", location="europe-west4")

endpoint = aiplatform.Endpoint.create(
    display_name="churn-prediction-endpoint",
    labels={"env": "production", "model": "churn"},
)

model = aiplatform.Model("projects/my-project/locations/europe-west4/models/MODEL_ID")

model.deploy(
    endpoint=endpoint,
    deployed_model_display_name="churn-v2",
    machine_type="n1-standard-4",
    min_replica_count=2,
    max_replica_count=10,
    traffic_percentage=10,
    sync=True,
)

Step 5: Configure Model Monitoring

Vertex AI Model Monitoring detects input distribution drift — an early warning that your model's predictions may be degrading.

from google.cloud.aiplatform import model_monitoring

monitoring_job = aiplatform.ModelDeploymentMonitoringJob.create(
    display_name="churn-model-monitoring",
    project="my-project",
    location="europe-west4",
    endpoint=endpoint.resource_name,
    logging_sampling_strategy=model_monitoring.RandomSampleConfig(sample_rate=0.2),
    schedule_config=model_monitoring.ScheduleConfig(monitor_interval=3600),
    alert_config=model_monitoring.EmailAlertConfig(
        user_emails=["ml-team@company.com"],
        enable_logging=True,
    ),
    drift_config=model_monitoring.TabularDriftMetricsConfig(
        feature_names=feature_names,
        default_drift_threshold=0.1,
    ),
)

Step 6: CI/CD Automation with Cloud Build

# cloudbuild.yaml
steps:
- name: python:3.11
  entrypoint: bash
  args:
  - -c
  - pip install -r requirements.txt && python -m pytest tests/ -v

- name: gcr.io/cloud-builders/docker
  args:
  - build
  - -t
  - europe-west4-docker.pkg.dev/$PROJECT_ID/ml-models/churn-server:$COMMIT_SHA
  - -f
  - serving/Dockerfile
  - serving/

- name: gcr.io/cloud-builders/docker
  args:
  - push
  - europe-west4-docker.pkg.dev/$PROJECT_ID/ml-models/churn-server:$COMMIT_SHA

- name: gcr.io/cloud-builders/docker
  args:
  - tag
  - europe-west4-docker.pkg.dev/$PROJECT_ID/ml-models/churn-server:$COMMIT_SHA
  - europe-west4-docker.pkg.dev/$PROJECT_ID/ml-models/churn-server:latest

- name: gcr.io/cloud-builders/docker
  args:
  - push
  - europe-west4-docker.pkg.dev/$PROJECT_ID/ml-models/churn-server:latest

- name: python:3.11
  entrypoint: bash
  args:
  - -c
  - |
    pip install google-cloud-aiplatform kfp
    python pipeline.py
    python submit_pipeline.py       --project=$PROJECT_ID       --region=europe-west4       --commit=$COMMIT_SHA

options:
  logging: CLOUD_LOGGING_ONLY

Vertex AI Experiments for Tracking

from google.cloud import aiplatform

aiplatform.init(project="my-project", location="europe-west4", experiment="churn-model-experiments")

with aiplatform.start_run("xgboost-tuned-v3"):
    aiplatform.log_params({
        "n_estimators": 200,
        "max_depth": 6,
        "learning_rate": 0.1,
    })
    aiplatform.log_metrics({
        "auc": 0.89,
        "f1_score": 0.72,
    })

All experiment runs appear in the Vertex AI console under Experiments, with comparison charts and lineage tracking across pipeline runs.

For building generative AI applications on top of Vertex AI, see our Vertex AI Gemini enterprise integration guide. For RAG architectures, see our Vertex AI RAG guide. For CI/CD integration details, see our MLOps CI/CD pipeline guide.