Anthos: Managing Hybrid and Multi-Cloud Workloads at Scale

Anthos extends GKE's Kubernetes management plane to on-premises and multi-cloud environments. This guide covers Anthos Config Management for GitOps, Anthos Service Mesh for service-to-service security, multi-cluster ingress, and the use cases where Anthos is worth the operational investment.

Running Kubernetes in a single cloud is complicated enough. Running it consistently across GCP, an on-premises data center, and AWS creates a whole new category of operational complexity: different node types, different networking models, different security configurations, different upgrade schedules.

Anthos is Google's answer to this: a managed Kubernetes platform that extends GKE's control plane to any environment. With Anthos, you manage clusters across environments through a unified interface, apply policies consistently through Config Management, and get service mesh networking through Anthos Service Mesh — regardless of where your clusters run.

The trade-off is complexity and cost. Anthos makes sense for organizations with genuine hybrid or multi-cloud requirements. If you're running everything on GKE and don't have regulatory or architectural reasons to go multi-cloud, Anthos adds cost without benefit.

Anthos Components and Architecture

Anthos Fleet: The management plane that groups clusters into a logical fleet. Fleet membership enables all other Anthos features.

Anthos Config Management (ACM): GitOps-based policy management. A Git repository is the source of truth for cluster configurations, namespace policies, and custom resource definitions. ACM syncs the desired state from Git to all fleet clusters automatically.

Anthos Service Mesh (ASM): Managed Istio for the fleet. Provides mTLS between services, traffic management, and observability across all clusters.

Multi-cluster Ingress: A single global load balancer that routes traffic to the appropriate cluster based on latency and availability.

Config Connector: Kubernetes CRDs for managing GCP resources (Cloud SQL, Pub/Sub, IAM) declaratively from Kubernetes YAML.

Setting Up Anthos Fleet

# Enable required APIs
gcloud services enable   anthos.googleapis.com   gkehub.googleapis.com   mesh.googleapis.com   multiclusteringress.googleapis.com

# Register an existing GKE cluster to the fleet
gcloud container fleet memberships register my-gke-cluster   --gke-cluster=europe-west4/my-gke-cluster   --enable-workload-identity

# Register an on-premises cluster (using Connect agent)
# First, download the Connect credentials
gcloud container fleet memberships register on-prem-cluster   --context=on-prem-cluster-context   --service-account-key-file=/path/to/service-account-key.json

# Verify fleet membership
gcloud container fleet memberships list

Anthos Config Management: GitOps for Multi-Cluster

ACM syncs configuration from a Git repository to all fleet clusters. The repository structure determines what gets applied where.

config-repo/
├── cluster/              # Applied to all clusters
│   ├── namespace.yaml    # Standard namespaces
│   ├── rbac.yaml         # Base RBAC roles
│   └── network-policy.yaml  # Baseline network policies
├── clusterregistry/      # Cluster-specific overrides
│   ├── gke-prod-eu.yaml  # Labels for production EU cluster
│   └── on-prem.yaml      # Labels for on-premises cluster
└── namespaces/           # Applied to specific namespaces
    ├── production/
    │   ├── namespace.yaml
    │   ├── resource-quota.yaml
    │   └── pod-disruption-budget.yaml
    └── monitoring/
        └── namespace.yaml
# Enable Config Management for the fleet
gcloud beta container fleet config-management enable

# Apply Config Management configuration
cat > config-management.yaml << 'EOF'
applySpecVersion: 1
spec:
  configSync:
    enabled: true
    sourceFormat: hierarchy
    syncRepo: https://github.com/my-org/config-repo
    syncBranch: main
    secretType: token
  policyController:
    enabled: true
    templateLibraryInstalled: true
    auditIntervalSeconds: 60
EOF

gcloud beta container fleet config-management apply   --config=config-management.yaml   --membership=my-gke-cluster

Policy Controller: Kubernetes Admission Control

Policy Controller (based on OPA Gatekeeper) enforces organizational policies at admission time:

# Constraint template: require resource limits on all containers
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
  name: requireresourcelimits
spec:
  crd:
    spec:
      names:
        kind: RequireResourceLimits
  targets:
  - target: admission.k8s.gatekeeper.sh
    rego: |
      package requireresourcelimits

      violation[{"msg": msg}] {
        container := input.review.object.spec.containers[_]
        not container.resources.limits.cpu
        msg := sprintf("Container '%v' must have CPU limits", [container.name])
      }

      violation[{"msg": msg}] {
        container := input.review.object.spec.containers[_]
        not container.resources.limits.memory
        msg := sprintf("Container '%v' must have memory limits", [container.name])
      }
---
# Activate the constraint
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: RequireResourceLimits
metadata:
  name: require-resource-limits
spec:
  match:
    kinds:
    - apiGroups: [""]
      kinds: ["Pod"]
    namespaces: ["production", "staging"]

When ACM syncs this to all fleet clusters, every Pod in production and staging namespaces across all clusters must have resource limits. Pods without limits are rejected at admission time.

Anthos Service Mesh: mTLS and Traffic Management

ASM provides managed Istio. Unlike self-managed Istio, you don't manage control plane upgrades.

# Install Anthos Service Mesh on a fleet cluster
gcloud container fleet mesh enable

gcloud container fleet mesh update   --management=automatic   --memberships=my-gke-cluster

# Verify ASM installation
gcloud container fleet mesh describe --memberships=my-gke-cluster

Enable mTLS for all workloads in the mesh:

# Enable STRICT mTLS for the production namespace
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: default
  namespace: production
spec:
  mtls:
    mode: STRICT  # All traffic between pods must use mTLS
---
# Traffic management: canary deployment with Istio
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: my-service
  namespace: production
spec:
  hosts:
  - my-service
  http:
  - match:
    - headers:
        x-canary:
          exact: "true"
    route:
    - destination:
        host: my-service
        subset: v2
  - route:
    - destination:
        host: my-service
        subset: v1
      weight: 90
    - destination:
        host: my-service
        subset: v2
      weight: 10
---
apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: my-service
  namespace: production
spec:
  host: my-service
  subsets:
  - name: v1
    labels:
      version: v1
  - name: v2
    labels:
      version: v2

Multi-Cluster Ingress

Multi-cluster ingress provides a single global Anycast IP that routes traffic to the closest healthy cluster:

# Enable Multi-cluster Ingress (requires a config cluster)
gcloud container fleet ingress enable   --config-membership=my-gke-cluster

# Deploy the multi-cluster ingress resource
cat << 'EOF' | kubectl apply -f -
apiVersion: networking.gke.io/v1
kind: MultiClusterIngress
metadata:
  name: my-service-ingress
  namespace: production
spec:
  template:
    spec:
      backend:
        serviceName: my-service
        servicePort: 80
---
apiVersion: networking.gke.io/v1
kind: MultiClusterService
metadata:
  name: my-service
  namespace: production
spec:
  template:
    spec:
      selector:
        app: my-service
      ports:
      - protocol: TCP
        port: 80
        targetPort: 8080
EOF

Multi-cluster ingress routes traffic to the cluster closest to the user (based on Anycast routing) and automatically removes clusters that fail health checks from rotation.

When Anthos Is Worth the Cost

Anthos licensing costs roughly $0.25-0.45/vCPU-hour on managed clusters. For a 20-node GKE cluster with 8 vCPU nodes, that's $280-$500/day in Anthos fees alone.

Anthos is worth the investment when you have:

  • Regulatory requirements: Data must stay on-premises for legal reasons, but you want GKE tooling
  • Multiple cloud environments: AWS or Azure clusters you need to manage consistently with GCP
  • Large fleet of clusters: 5+ clusters where consistent policy enforcement and GitOps become critical
  • Advanced traffic management: Multi-cluster ingress, blue/green deployments across regions

If you have a single GKE cluster or only GKE clusters, you get most of the same capabilities (Config Management, Service Mesh, binary authorization) without Anthos licensing by using the GKE-native equivalents.

For the Terraform configuration that provisions Anthos-registered clusters, see our GCP Terraform best practices guide. For GKE cluster security in the mesh, see our GKE security hardening guide.