Skip to content

Multi-Cloud Governance -- Explanation

How a multi-cloud governance model fits together, and why it is built that way. This page covers the governance control plane, how a policy travels from a pull request to runtime, the trade-offs between policy engines, CSPM/CNAPP and landing zones, networking and DNS topologies, IaC and GitOps, observability, FinOps, identity federation, and the threat model. The examples span AWS, GCP, Alibaba Cloud, and Tencent Cloud, with Azure where relevant.

Look-up tables (versions, service limits, CIS benchmarks, tool matrices) are in Reference. Commands and config recipes are in How-to Guides.

Governance Control Plane

Multi-cloud governance is not one product. It is a control plane assembled from four kinds of control, applied at each layer where a change can enter:

  • Identity: one IdP federates humans and workloads into every cloud, so access can be granted and revoked in one place.
  • Preventive controls: policy checks that block a bad change before it exists. Examples are CI checks on IaC, policy sets on Terraform runs, cloud-organization guardrails, and Kubernetes admission control.
  • Detective and corrective controls: scanners that find drift and misconfiguration in live resources and optionally fix them. Examples are Cloud Custodian, Prowler, and commercial CSPM/CNAPP.
  • Evidence and cost: audit logs, policy reports, and FOCUS-format billing data, which prove compliance and allocate spend.

The diagram shows how the pieces connect. Policies live in Git beside the infrastructure code. They are enforced at CI, at the Terraform run, at the cloud organization boundary, and at the Kubernetes API server. Everything that runs then feeds evidence and cost data back.

flowchart TB
    subgraph IDN["Identity plane"]
        IDP["Central IdP<br/>(Entra ID / Okta / Keycloak)"]
    end
    subgraph GITP["Policy and config as code (Git)"]
        PREPO["Policy repo<br/>(Rego, Kyverno CEL, Sentinel, c7n YAML)"]
        IREPO["IaC repo<br/>(Terraform / OpenTofu / Crossplane XRs)"]
    end
    subgraph PREV["Preventive controls"]
        CI["CI checks<br/>(Conftest, Checkov, Trivy)"]
        TFC["HCP Terraform / TFE<br/>(Sentinel or OPA policy sets)"]
        ADM["Kubernetes admission<br/>(Gatekeeper / Kyverno / VAP)"]
    end
    subgraph GUARD["Cloud organization guardrails"]
        AWSO["AWS Organizations<br/>(SCP / RCP, Control Tower)"]
        AZP["Azure Policy<br/>(ALZ management groups)"]
        GOP["GCP Organization Policy"]
        ALIRD["Alibaba Resource Directory<br/>control policies"]
        TENC["Tencent CAM and Organization"]
    end
    subgraph DET["Detective and corrective controls"]
        C7N["Cloud Custodian"]
        CSPM["CSPM / CNAPP<br/>(Prowler, Wiz, Cortex Cloud, Defender for Cloud)"]
    end
    subgraph EVD["Evidence and cost"]
        SIEM["SIEM and audit-log archive"]
        FOCUSX["FOCUS billing exports<br/>to FinOps platform"]
    end
    IDP -->|"SAML / OIDC"| AWSO
    IDP -->|"SAML / OIDC"| AZP
    IDP -->|"SAML / OIDC"| GOP
    IDP -->|"SAML / OIDC"| ALIRD
    IDP -->|"SAML / OIDC"| TENC
    PREPO --> CI
    PREPO --> TFC
    PREPO --> ADM
    PREPO --> C7N
    IREPO --> CI
    CI --> TFC
    TFC -->|"apply"| AWSO
    TFC -->|"apply"| GOP
    TFC -->|"apply"| ALIRD
    C7N -->|"scan and remediate"| AWSO
    CSPM -->|"read-only scan"| GOP
    ADM --> SIEM
    C7N --> SIEM
    CSPM --> SIEM
    AWSO -->|"billing data"| FOCUSX
    GOP -->|"billing data"| FOCUSX

Design choices behind this shape:

  1. Guardrails at the organization boundary are the backstop. A Terraform policy only sees changes made through Terraform. A console click, a CLI call, or a compromised CI token bypasses it. Organization-level controls (AWS SCPs and RCPs, Azure Policy, GCP Organization Policy, Alibaba control policies) apply to every API call, whatever the source.
  2. Shift-left checks exist for fast feedback, not for security. A policy failure in a pull request costs minutes. The same failure discovered by a CSPM scan after deployment costs a ticket, a rollback, or an incident.
  3. Detective controls cover what prevention cannot express. Examples are resource age, unused resources, cost anomalies, and posture across accounts. They also catch drift from before a guardrail existed.
  4. One policy source, several engines. No single engine runs everywhere. The practical target is one Git repository, one review process, and one exception process, even if the files inside are Rego, CEL, Sentinel, and Custodian YAML.

Policy Flow: From Pull Request to Runtime

The sequence below follows one change, for example a new storage bucket plus a Kubernetes deployment, through every enforcement point. Each hop can reject the change. Later hops exist because earlier ones can be bypassed.

sequenceDiagram
    autonumber
    participant Dev as Engineer
    participant PR as Git pull request
    participant CI as CI (Conftest, Checkov, Trivy)
    participant TFC as HCP Terraform (Sentinel / OPA)
    participant Org as Cloud org guardrails
    participant GitOps as Argo CD / Flux
    participant API as K8s API server + Gatekeeper / Kyverno
    participant Det as Cloud Custodian / CSPM
    Dev->>PR: Push IaC and manifests
    PR->>CI: Run static policy checks on code and plan JSON
    CI-->>PR: Block merge on violations
    PR->>TFC: Merge triggers plan
    TFC->>TFC: Evaluate policy sets against the plan
    TFC->>Org: Apply (cloud API calls)
    Org-->>TFC: AccessDenied if an SCP / Org Policy blocks the call
    PR->>GitOps: Manifests land on the deploy branch
    GitOps->>API: Apply manifests
    API->>API: Admission webhook validates or mutates
    API-->>GitOps: Reject non-compliant objects
    Det->>Org: Scheduled or event-driven scan of live resources
    Det-->>Dev: Finding, ticket, or auto-remediation

Two properties matter when you design this flow:

  • Exceptions must be first-class. Every engine has an exception mechanism (Gatekeeper excludedNamespaces and constraint match, Kyverno PolicyException, Sentinel enforcement levels, SCP conditions on principal tags). Route them all through the same review and expiry process, or they turn into permanent holes.
  • Audit mode before enforce mode. Gatekeeper's audit controller, Kyverno's Audit action, and Azure Policy's audit effect report violations without blocking. Roll a new policy out in audit mode, fix the backlog, then switch to deny.

Policy Engines and Their Trade-offs

Engine Where it runs Strength Cost
OPA (Rego) Library, sidecar, CLI (Conftest), HCP Terraform General-purpose: any JSON input, so one language for Terraform plans, Kubernetes, and API authorization Rego is a new language for most teams. OPA 1.0 (2024-12) made the if/contains keywords mandatory, so older policies need migration
OPA Gatekeeper Kubernetes admission webhook plus audit controller Reusable parameterized ConstraintTemplates. Audit of existing objects. CEL engine alongside Rego Kubernetes only. Two layers (template plus constraint) to learn
Kyverno Kubernetes admission webhook, background controller, CLI Policies are Kubernetes resources. Validate, mutate, generate, clean up, and verify images. CNCF Graduated 2026-03 Kubernetes-centric. The move from legacy ClusterPolicy to CEL-based ValidatingPolicy and related types (v1 since 1.17) means a policy rewrite; the 1.17 announcement targeted removal of ClusterPolicy for 2026-10
ValidatingAdmissionPolicy Inside the Kubernetes API server (GA 1.30) No webhook to run or fail. CEL Validation only. No external data. Limited reporting
Sentinel HCP Terraform / Terraform Enterprise Tight integration with the Terraform run and its enforcement levels Proprietary. Tied to HashiCorp's platform (HCP Terraform also accepts OPA policy sets)
Cloud Custodian CLI, scheduled jobs, or serverless (Lambda, Azure Functions, Cloud Functions) Detect and act on live resources across AWS, Azure, GCP, Kubernetes, OCI, and Tencent Cloud, including cost hygiene (stop, tag, delete) Detective, not preventive. Event modes differ per cloud. No Alibaba Cloud provider
Cloud-native guardrails Cloud control plane Cannot be bypassed by any client. No infrastructure to run Different language and semantics per cloud. SCPs cap permissions and do not grant them

The flowchart below is a starting point for choosing an engine by what is being governed.

flowchart TD
    Q1{"What is being governed?"}
    Q1 -->|"Kubernetes objects"| Q2{"Validation only,<br/>no external data?"}
    Q2 -->|"Yes"| VAP["ValidatingAdmissionPolicy<br/>(built in, Kubernetes 1.30+)"]
    Q2 -->|"Need mutate, generate,<br/>or image verification"| KYV["Kyverno"]
    Q2 -->|"Rego already used<br/>across the stack"| GK["OPA Gatekeeper"]
    Q1 -->|"IaC changes before apply"| Q3{"Runs on HCP Terraform or TFE?"}
    Q3 -->|"Yes"| SEN["Sentinel or OPA policy sets"]
    Q3 -->|"No"| CONF["Conftest (OPA), Checkov,<br/>or Trivy in CI"]
    Q1 -->|"Live cloud resources"| Q4{"Prevent or detect?"}
    Q4 -->|"Prevent"| ORG["Cloud-native guardrails<br/>(SCP / RCP, Azure Policy, Org Policy)"]
    Q4 -->|"Detect and remediate"| C7N["Cloud Custodian or CSPM / CNAPP"]

OPA governance after Styra

In 2025-08 OPA's founding maintainers joined Apple. The project stays a CNCF Graduated project with the same maintainer process, and Styra's commercial pieces (EOPA, OPA Control Plane, the Regal linter) were open-sourced. Releases continue: 1.21.0 shipped on 2026-09-24. Sources: OPA blog, CHANGELOG.

CSPM and CNAPP Landscape

Cloud Security Posture Management (CSPM) tools read cloud control-plane APIs and compare configuration against benchmarks such as CIS. Cloud-Native Application Protection Platforms (CNAPP) combine CSPM with workload scanning, identity risk (CIEM), IaC scanning, and runtime detection. In a multi-cloud estate the key question is coverage: native tools are strongest in their own cloud, while cross-cloud tools trade depth for one pane of glass.

The market consolidated in 2025-2026:

  • Wiz became part of Google Cloud on 2026-03-11 (a $32B deal). Google says Wiz keeps its brand and continues to support all major clouds (Google).
  • Prisma Cloud became Cortex Cloud (announced 2025-02), merging CNAPP into Palo Alto's Cortex SecOps platform (Palo Alto Networks).
  • AWS Security Hub was renamed Security Hub CSPM when AWS launched a new unified Security Hub. It supports CIS AWS Foundations v5.0.0 (added 2025-10), while CIS itself published v7.0.0 in 2026-04.
  • Open source: Prowler now covers AWS, Azure, GCP, Kubernetes, M365, OCI, Alibaba Cloud, and more. tfsec was folded into Trivy. ScoutSuite has not released since 2024-05.

For China-region clouds, cross-cloud coverage is thinnest. Alibaba Cloud is covered by Prowler, Cortex Cloud, Wiz, and Orca. Tencent Cloud coverage is mostly limited to Tencent's own security products and Cloud Custodian's c7n-tencentcloud. Tool-by-tool coverage is in Reference: Automated Compliance Scanning.

Landing Zones

A landing zone is the pre-built multi-account (or multi-project/subscription) foundation that every workload lands in: account hierarchy, identity, guardrail policies, central logging, network hub, and security tooling. Governance is far cheaper when it is built in from day one than when it is retrofitted onto existing accounts.

Each cloud has its own accelerator, and they do not interoperate:

  • AWS: Control Tower provides the account factory and managed guardrails. Landing Zone Accelerator on AWS (LZA) layers configuration-driven networking, security services, and compliance baselines on top.
  • Azure: Azure Landing Zones (ALZ) is the reference. The Terraform implementation moved to Azure Verified Modules pattern modules (avm-ptn-alz and related modules). The classic caf-enterprise-scale module is deprecated.
  • Google Cloud: The Enterprise Foundations Blueprint (terraform-example-foundation) or Cloud Foundation Fabric FAST. Both are fork-and-own code, not managed services.
  • Alibaba Cloud: Cloud Governance Center sets up Resource Directory, core accounts, and an Account Factory with account baselines.
  • Tencent Cloud: Control Center sets up the landing zone and account factory on top of Tencent Cloud Organization (Configuring a Landing Zone). The code-first path is the tencentcloud-landing-zone-booster Terraform modules.

The multi-cloud implication: standardize the intent (account taxonomy, mandatory tags, log retention, network CIDR plan, break-glass process) in one document, and let each cloud's accelerator implement it natively. Do not build a lowest-common-denominator abstraction over all of them. See Reference: Landing-Zone Accelerators.

Networking Patterns

Hub-and-Spoke

The most common multi-cloud topology. A central network hub (often an on-premises data center or a dedicated transit VPC/VNet) routes all inter-cloud traffic. Each cloud environment is a "spoke" connected via dedicated interconnect.

                    [ On-Prem DC / Transit Hub ]
                     /        |          \
                    /         |           \
            [ AWS TGW ]  [ GCP NCC ]  [ Alibaba CEN ]
                |            |              |
            [ VPCs ]     [ VPCs ]       [ VPCs ]

Key services per cloud:

Cloud Hub Service Interconnect Service
AWS Transit Gateway / Cloud WAN Direct Connect
GCP Network Connectivity Center (NCC) Dedicated / Partner / Cross-Cloud Interconnect
Alibaba Cloud Cloud Enterprise Network (CEN) Express Connect (physical dedicated line)
Tencent Cloud Cloud Connect Network (CCN) Direct Connect

Advantages: Centralized routing policy, single audit point, clear blast-radius boundary. Disadvantages: Hub is a single point of failure. All cross-cloud traffic incurs hub hairpinning latency.

Full Mesh

Every cloud peer connects directly to every other cloud peer, typically via dedicated point-to-point circuits or an SD-WAN overlay. No central transit hub.

            [ AWS TGW ] ------ [ GCP NCC ]
                |    \          /     |
                |     \        /      |
                |    [ Alibaba CEN ]   |
                |         |            |
                +---[ Tencent CCN ]----+

Advantages: Lowest inter-cloud latency (direct paths). No single point of failure. Disadvantages: O(n^2) circuit management. Complex routing tables. Higher cost at scale.

Transit (Backbone)

A provider-agnostic backbone -- often a third-party fabric like Equinix Fabric, Megaport, or PacketFabric -- acts as the Layer 2/3 transit layer. Each cloud connects to the nearest fabric node via its dedicated interconnect service. The backbone provides any-to-any reachability with per-circuit QoS and bandwidth policies.

            [ Equinix Fabric / Megaport Backbone ]
              /         |            |          \
        [ AWS DC ]  [ GCP DC ]  [ Alibaba PoP ] [ Tencent PoP ]
             |           |            |              |
          [ VPCs ]    [ VPCs ]     [ VPCs ]       [ VPCs ]

Advantages: Centralized bandwidth management. Sub-cloud provisioning. Any-to-any reachability without full mesh circuits. Disadvantages: Added cost of third-party fabric. Dependency on fabric provider SLA.

Providers now sell direct cloud-to-cloud links that remove the colocation and fabric step:

  • GCP Cross-Cloud Interconnect provisions dedicated 10 or 100 Gbps links from Google Cloud to AWS, Azure, OCI, and Alibaba Cloud.
  • AWS Interconnect -- multicloud reached GA in 2026-04 with Google Cloud as the first partner. AWS provisions redundant links, BGP, and MACsec, with 1-100 Gbps bandwidth and a 99.99% SLA. Azure and OCI support was announced for later in 2026. AWS published the interconnect specification under Apache-2.0 (InfoQ).

These links simplify a two-cloud mesh between Western hyperscalers. They do not yet cover Tencent Cloud, and Alibaba coverage exists only on the Google side. So APAC estates that include China-region clouds still rely on Express Connect / Direct Connect plus a fabric provider.

SD-WAN Overlay

Some enterprises cannot justify dedicated circuits at every edge. For these enterprises, SD-WAN solutions create encrypted IPsec/GRE tunnels over the public internet between cloud VPCs and branch offices. Common products are Cisco Catalyst SD-WAN (formerly Viptela), VMware VeloCloud SD-WAN, Palo Alto Prisma SD-WAN, and Fortinet Secure SD-WAN.

Advantages: Rapid provisioning. Internet-based (no colo requirement). Built-in WAN optimization. Disadvantages: Higher and variable latency. Bandwidth limited by internet path. Not suitable for latency-sensitive workloads.

Decision Matrix

Factor Hub-Spoke Full Mesh Transit Backbone SD-WAN Overlay
Latency Medium (hairpin) Low (direct) Low-Medium High-Variable
Complexity Low High Medium Low-Medium
Cost Medium High Medium-High Low
Blast Radius Hub is SPOF Isolated per link Fabric is SPOF Per-tunnel
Use Case Regulated enterprise Low-latency apps Global scale Branch / DR

Bandwidth and protocol details per service: Reference: Key Interconnect Services Reference.

DNS and Traffic Management

Global DNS is the first hop for multi-region, multi-cloud traffic steering. Per-provider routing policies are listed in Reference: Provider DNS Services.

Multi-Cloud DNS Strategy Patterns

1. Delegated Subdomain per Cloud

Each cloud owns a subdomain zone (for example, aws.example.com, gcp.example.com, cn.example.com). A global apex zone delegates NS records per subdomain to the respective cloud DNS service.

example.com (Route 53 / Cloudflare)
  ├── aws.example.com   NS → Route 53 hosted zone
  ├── gcp.example.com   NS → Cloud DNS zone
  ├── cn.example.com    NS → Alidns zone
  └── global.example.com → GSLB (weighted/latency routing across clouds)

2. GSLB Overlay with Health Checks

A global DNS layer (Route 53, Cloudflare, NS1) resolves global.example.com by evaluating health-check endpoints in each cloud. On failure, traffic shifts to the next healthy cloud within TTL convergence time.

  • Route 53 health checks: HTTP/HTTPS/TCP, 10 s or 30 s request intervals, configurable failure threshold.
  • Cloudflare Load Balancing: Health checks per pool, failover steering, session affinity.
  • NS1 Filter Chain: Programmatic DNS decisions based on telemetry feeds (Pulsar), availability, and geolocation.

3. China-Specific DNS Consideration

Serving a site from mainland China requires an ICP filing for the domain. Alidns and DNSPod both support ISP-line resolution (routing queries from China Telecom, China Unicom, and China Mobile to the nearest endpoint). For multi-cloud within China, DNSPod or Alidns serves as the apex. Global-with-China architectures typically use split-horizon DNS: international queries resolve via Route 53/Cloudflare, and China queries resolve via Alidns/DNSPod.

Service Mesh and L7 Traffic

DNS steers users to a cloud. Inside and between clusters, a service mesh (Istio multi-primary, Cilium Cluster Mesh) handles service-to-service routing, mTLS, and failover. Provider-specific meshes are poor multi-cloud choices: AWS App Mesh reaches end of support on 2026-09-30, and Google's Traffic Director is now part of Cloud Service Mesh. See Reference: Traffic-Management Tools Beyond DNS and the Istio topic.

Infrastructure-as-Code

Terraform / OpenTofu

The dominant multi-cloud IaC tool. HCL configurations declare resources across providers in a single state or split across workspaces. HashiCorp moved Terraform to BSL 1.1 in 2023 (1.6.0+), and HashiCorp became part of IBM in 2025-02. OpenTofu is the MPL-2.0 fork under the Linux Foundation, in the CNCF Sandbox since 2025-04. Both use the same providers: hashicorp/aws, hashicorp/google, aliyun/alicloud, and tencentcloudstack/tencentcloud. Current versions are in Reference.

State isolation strategy: Use separate workspaces or separate state backends per cloud (and per environment) to limit blast radius. Common backend choices are S3 (AWS), GCS (GCP), OSS (Alibaba), and COS (Tencent). One state per cloud also means an outage of one cloud's API does not block plans for the others.

Terraform Stacks (GA in HCP Terraform, 2025-09) deploy the same set of components into many deployments. Each deployment can get its own short-lived cloud credentials from an identity_token block (OIDC workload identity) instead of static keys. Recipes: How-to: Configure Multiple Cloud Providers.

Pulumi

General-purpose IaC in TypeScript, Python, Go, .NET, Java, or YAML. It uses bridged Terraform providers (including @pulumi/alicloud and @pulumi/tencentcloud) plus native providers (AWS Native, Azure Native, Google Native) generated from cloud API specs. Pulumi ESC (Environments, Secrets, and Configuration) issues dynamic OIDC credentials for AWS, Azure, and GCP, which removes static access keys. Stack references let one stack consume another's outputs, so dependencies across cloud boundaries stay explicit. See the Pulumi topic.

Crossplane

Kubernetes-native IaC. Infrastructure is declared as Kubernetes resources and reconciled by provider controllers running in a control-plane cluster. Platform teams publish their own APIs as composite resources (XRs) defined by CompositeResourceDefinitions (XRDs), and implement them with composition functions.

Crossplane v2 (v2.0.0 tagged 2025-08-08; current line v2.4, 2026-08) changed the model (What's New in v2):

  • XRs and managed resources (MRs) are namespaced, and claims are gone for v2-style XRs. Tenancy maps directly onto Kubernetes namespaces and RBAC.
  • Compositions can compose any Kubernetes resource, not only MRs. An XR can bundle a Deployment, a Service, and an RDSInstance.
  • Operations (Operation, CronOperation, WatchOperation, alpha) run function pipelines to completion for day-2 tasks.
  • Managed resource definitions activate only the MRs you need, which reduces CRD load on the API server.
  • Breaking removals: native patch-and-transform composition (use function-patch-and-transform), ControllerConfig, external secret stores, XR connection details, and the default package registry.

Crossplane graduated in the CNCF in late 2025. It releases quarterly with about nine months of support per minor release, and v1.20 reaches EOL in 2026-11.

For multi-cloud, the catch is provider depth. Namespaced MRs shipped first for AWS; GCP and Azure followed. Alibaba (provider-upjet-alibabacloud) and Tencent (provider-tencentcloud) providers exist in crossplane-contrib. The Alibaba provider is at v1.3.0 (2026-06-16) and still active; the Tencent provider is pre-1.0, with its last release v0.8.6 on 2025-04-24 (git tags, checked 2026-09-27). Crossplane is strongest when the organization already runs Kubernetes as its platform layer, because XRs, GitOps, and admission policy all share the same API machinery.

Decision Matrix

Factor Terraform/OpenTofu Pulumi Crossplane
Learning curve Low (HCL) Medium (code) High (Kubernetes CRDs and functions)
Multi-cloud maturity Highest High Growing (AWS strongest; Alibaba/Tencent community providers)
GitOps native Via Atlantis, HCP Terraform VCS runs Via Pulumi Deployments, Kubernetes operator Native (Kubernetes resources)
State management External backend Pulumi Cloud or self-managed Kubernetes API plus the cloud itself (continuous reconciliation)
Platform abstraction Modules, Stacks ComponentResource XRDs and composition functions
Policy integration Sentinel/OPA in HCP Terraform, Conftest in CI CrossGuard policy packs Admission policy (Gatekeeper/Kyverno) on XRs and MRs
Best fit General infra teams Developer-heavy teams Platform engineering / Kubernetes-first orgs

For a deeper tool-level comparison, see IaC Comparison.

CI/CD and GitOps

GitOps Patterns for Multi-Cloud

GitOps uses Git as the single source of truth for declarative infrastructure and application state. Automated controllers (Argo CD, Flux) continuously reconcile the live state with the desired state in Git.

Hub-and-Spoke GitOps:

A central management cluster runs Argo CD or Flux, which deploys to remote target clusters across clouds via ApplicationSets or Flux Kustomization resources.

[ Management Cluster (Argo CD/Flux) ]
        |                |               \
   [ AWS EKS ]     [ GCP GKE ]    [ Alibaba ACK ]
  • Argo CD ApplicationSets with cluster, git, or matrix generators create one Application per target cluster.
  • Flux Kustomization resources with spec.kubeConfig reference kubeconfigs for remote clusters.

Per-Cluster GitOps:

Each cluster runs its own Argo CD or Flux instance. The blast radius is smaller, but global policy is harder to enforce. This suits regulated environments where clusters must be autonomous.

Progressive Delivery:

  • Argo Rollouts: Canary / blue-green / experiment strategies with metric analysis.
  • Flagger (with Flux): Automated canary deployments with Prometheus metric analysis.
  • Argo CD Image Updater / Flux Image Automation: Update image tags in Git automatically when new images are published.

Multi-Cloud CI Pipeline Structure

[ Git Push ]
     |
[ CI Pipeline (GitHub Actions / GitLab CI) ]
  ├── Build container image
  ├── Run tests and policy checks (Conftest / Checkov / Trivy)
  ├── Push image to multi-cloud registries
  │   ├── ECR (AWS)
  │   ├── Artifact Registry (GCP)
  │   ├── ACR (Alibaba Container Registry)
  │   └── TCR (Tencent Container Registry)
  ├── Update GitOps manifest (image tag)
  └── GitOps controller detects change and rolls out

Secret management in GitOps: Sealed Secrets, SOPS (now a CNCF project) with KMS per cloud, or External Secrets Operator. External Secrets Operator syncs from HashiCorp Vault or cloud-native secret stores (AWS Secrets Manager, GCP Secret Manager, Alibaba KMS). Tool licenses and multi-cluster features: Reference: GitOps Tool Reference. See also the GitOps comparison.

Observability Architecture

The OpenTelemetry Standard

OpenTelemetry (OTel) is the CNCF vendor-neutral observability framework. It provides APIs, SDKs, the OTLP protocol, and the Collector pipeline for traces, metrics, and logs. A fourth signal, profiles, is still in development. In a multi-cloud environment, OTel normalizes telemetry regardless of which cloud produced it. See the OpenTelemetry and OpenTelemetry Collector topics.

Collector Topology

The recommended multi-cloud pattern is per-cloud collector deployment with centralized backend export.

[ AWS Cluster ]                 [ GCP Cluster ]
  OTel Collector (DaemonSet)      OTel Collector (DaemonSet)
       | OTLP                          | OTLP
       v                               v
[ AWS Gateway Collector ] -----> [ Central Backend ]
                                    (Grafana LGTM /
                                     Datadog / Dynatrace)
       ^                               ^
       | OTLP                          | OTLP
[ Alibaba Cluster ]             [ Tencent Cluster ]
  OTel Collector (DaemonSet)      OTel Collector (DaemonSet)

DaemonSet agents on each cluster send to a per-cloud gateway collector. The gateway batches, retries, enriches with cloud.* resource attributes, and does tail sampling before a single egress to the central backend. This keeps cross-cloud egress (which is billed) to one compressed OTLP stream per cloud. Deployment modes, per-cloud integrations, and semantic conventions are in Reference: Observability Reference. The collector config is in How-to: Deploy Per-Cloud OTel Collectors.

Logging Architecture

[ Applications with OTel SDK ]
         |
         v
[ OTel Collector -- log pipeline ]
         |
         +---> [ Central Backend (Loki / Datadog) ]
         |
         +---> [ Cloud-native Log Store (SLS / CLS / CloudWatch) ]
                  |
                  +---> [ Compliance archive (OSS / COS / S3 Glacier) ]

Key considerations:

  • Retention policies differ per cloud. Standardize retention at the central backend level.
  • Compliance-critical logs (audit trails) must also be stored in immutable, append-only storage in each cloud, in addition to centralized aggregation.
  • Log volume across clouds can be significant. Use the Collector's batch processor, filtering, and sampling to control ingestion and egress costs.
  • For Alibaba Cloud SLS and Tencent CLS, verify OTLP ingestion endpoints and limits per region before relying on them.

Distributed Tracing Across Clouds

Traces that span services on different clouds need consistent context propagation. W3C Trace Context and W3C Baggage are the standard headers.

  • Trace Context: traceparent and tracestate HTTP headers (or gRPC metadata) carry the trace ID across service boundaries.
  • Cross-cloud correlation: All services must use the same W3C propagator. OTel SDKs default to it. Load balancers and API gateways between clouds must pass the headers through.
  • Trace sampling: Tail-based sampling at the gateway collector bases decisions on the complete trace (for example, keep all error traces and 10% of successful ones). All spans of a trace must reach the same gateway replica, so front the gateways with the loadbalancing exporter keyed on trace ID.

FinOps Model

Phases and Maturity

The FinOps Foundation framework describes three iterative phases, which every team cycles through continuously:

Phase Focus Typical Activities
Inform Visibility and allocation Tagging, cost allocation, showback, forecasting
Optimize Rates and usage Rightsizing, commitment coverage (RIs, Savings Plans, CUDs), spot, anomaly response
Operate Continuous improvement and governance Unit economics, chargeback, policy, automation, executive alignment

Separately, each capability is assessed on a Crawl / Walk / Run maturity scale. A team can be at Run for allocation and at Crawl for commitment management. The two are not the same axis. The 2026 framework update extended FinOps beyond public cloud (AI, SaaS, licensing, data center), deepened Scopes (segments of spend such as public cloud or GenAI), and added an Executive Strategy Alignment capability (FinOps Foundation).

FOCUS as the Common Billing Schema

Every cloud has its own billing schema: AWS CUR, Azure cost exports, the GCP BigQuery export, Alibaba bills, and Tencent bills. Normalizing them used to be the first job of every multi-cloud FinOps team. FOCUS (FinOps Open Cost and Usage Specification) solves this with one column set (BilledCost, EffectiveCost, ServiceName, ChargePeriodStart, and so on). FOCUS 1.4 was ratified on 2026-06-04. AWS, Azure, Google Cloud, and OCI publish FOCUS exports, Alibaba Cloud has a FOCUS 1.0 export in invitational preview, and Tencent Cloud converts bills to FOCUS 1.0.

The flow below shows the resulting pipeline. Each provider's native FOCUS export lands in one store, Kubernetes allocation data is joined in, and all reporting runs on a single schema.

flowchart LR
    AWSX["AWS Data Exports<br/>(FOCUS)"] --> LAKE
    AZX["Azure Cost Management<br/>(FOCUS export)"] --> LAKE
    GCPX["GCP BigQuery billing<br/>(FOCUS view)"] --> LAKE
    ALIX["Alibaba FOCUS 1.0 export<br/>(OSS, preview)"] --> LAKE
    TENX["Tencent FOCUS 1.0 bills<br/>(COS)"] --> LAKE
    K8S["OpenCost / Kubecost<br/>(namespace and pod allocation)"] --> LAKE
    LAKE["Cost store<br/>(one FOCUS schema)"] --> SHOW["Showback and chargeback"]
    LAKE --> UNIT["Unit economics<br/>(joined with OTel business metrics)"]
    LAKE --> ANOM["Budgets and anomaly detection"]

FOCUS normalizes the schema, not the prices or the discount semantics. Commitments, credits, and marketplace charges still need per-provider interpretation. Check which FOCUS version each export emits before joining them.

Unit Economics

Shift from raw spend to cost per business metric:

  • cost-per-transaction = total cloud spend / number of transactions processed.
  • cost-per-active-user = total cloud spend / monthly active users.
  • cost-per-api-call = total cloud spend / API calls served.
  • revenue-per-dollar-of-cloud = revenue / cloud spend.

OTel metrics can count transactions and active users, which lets a dashboard (Grafana, Datadog) correlate them directly with spend. When one transaction spans two clouds, allocate by a measurable driver (for example, requests or CPU-seconds per cloud from OTel) rather than splitting evenly.

Green FinOps

Carbon-aware workload placement is an emerging practice:

  • Google Cloud Carbon Footprint: Carbon emissions data per project and region.
  • AWS Customer Carbon Footprint Tool: Estimated emissions per service and region.
  • Cloud Carbon Footprint (open source): Estimates energy and emissions across AWS, GCP, and Azure.
  • Strategy: Schedule batch workloads in regions with lower carbon intensity, prefer newer instance types with better performance per watt, and include carbon cost in placement decisions.

Reference Architecture

A vendor-neutral enterprise multi-cloud architecture built on CNCF and open-source components. The diagram groups components by layer. Every layer fans out to the four workload clouds.

graph TB
    subgraph "Identity Layer"
        IdP[IdP: Okta / Entra ID / Keycloak]
        SPIFFE[SPIFFE / SPIRE]
    end

    subgraph "Control Plane"
        Git[Git Repository]
        CI[CI Pipeline]
        ArgoCD[Argo CD / Flux]
        IaC[Terraform / Crossplane]
    end

    subgraph "Networking Layer"
        Fabric[Equinix Fabric / Megaport / Cross-Cloud Interconnect]
        DNS[Global DNS: Route 53 / Cloudflare]
    end

    subgraph "Workload Plane"
        AWS[AWS EKS]
        GCP[GCP GKE]
        ALI[Alibaba ACK]
        TENCENT[Tencent TKE]
    end

    subgraph "Observability Plane"
        OTel[OTel Collectors]
        Backend[Grafana LGTM / Datadog]
    end

    subgraph "Policy Layer"
        OPA[Gatekeeper / Kyverno]
        Sentinel[Sentinel / Cloud Org Policies]
    end

    IdP -->|SAML/OIDC| AWS
    IdP -->|SAML/OIDC| GCP
    IdP -->|SAML/OIDC| ALI
    IdP -->|SAML/OIDC| TENCENT

    Git --> CI --> ArgoCD
    ArgoCD --> AWS
    ArgoCD --> GCP
    ArgoCD --> ALI
    ArgoCD --> TENCENT

    IaC --> AWS
    IaC --> GCP
    IaC --> ALI
    IaC --> TENCENT

    Fabric --> AWS
    Fabric --> GCP
    Fabric --> ALI
    Fabric --> TENCENT

    DNS --> AWS
    DNS --> GCP
    DNS --> ALI
    DNS --> TENCENT

    OTel --> Backend
    AWS --> OTel
    GCP --> OTel
    ALI --> OTel
    TENCENT --> OTel

    OPA --> AWS
    OPA --> GCP
    OPA --> ALI
    OPA --> TENCENT
    Sentinel --> IaC

Layer Descriptions

Layer Components Purpose
Identity IdP (SAML/OIDC), SPIFFE/SPIRE Unified human + workload identity
Control Plane Git, CI, GitOps controller, IaC engine Declarative intent and reconciliation
Networking Interconnect fabric, global DNS Cross-cloud connectivity and traffic steering
Workload Plane Managed Kubernetes clusters per cloud Application runtime
Observability OTel Collectors, central backend Traces, metrics, logs (profiles emerging)
Policy Gatekeeper/Kyverno, Sentinel/OPA, org policies Governance, compliance, cost guardrails

How It Works

  1. Identity federation: The IdP issues SAML assertions (human SSO) and OIDC tokens (workload identity) to each cloud. SPIRE agents on each node provide mTLS workload identity independent of cloud IAM.
  2. Infrastructure provisioning: Terraform or Crossplane declares VPCs, databases, queues, and other cloud resources. Changes land in Git. CI validates them (terraform plan plus policy checks, or crossplane render for compositions) and applies after approval.
  3. Application delivery: CI builds images, pushes to per-cloud registries, and updates GitOps manifests. Argo CD or Flux reconciles manifests to clusters in all clouds.
  4. Networking: A fabric provider or provider-managed cross-cloud links provide private connectivity. Global DNS routes users to the nearest healthy cloud endpoint.
  5. Observability: OTel Collectors in each cluster collect traces, metrics, and logs. They export to a central backend via OTLP.
  6. Policy enforcement: Gatekeeper or Kyverno enforces Kubernetes-level policies (resource limits, allowed images, labels). Cloud-native org policies (AWS SCPs and RCPs, Azure Policy, GCP Org Policy, Alibaba control policies) enforce cloud-resource-level governance. Sentinel or OPA policy sets guard Terraform runs. See Policy Flow.

Identity Federation

Problem

Each cloud has its own IAM system (AWS IAM, GCP IAM, Alibaba RAM, Tencent CAM). Without federation, operators maintain separate credentials per cloud. This defeats centralized identity governance and creates credential-sprawl risk. Protocol and per-cloud mechanism tables are in Reference: Identity Reference.

How Each Cloud Federates

  • AWS: IAM Identity Center accepts SAML assertions (and SCIM provisioning) from an external IdP and maps groups to permission sets. IAM OIDC providers trust external issuers (GitHub Actions, GitLab, EKS) for sts:AssumeRoleWithWebIdentity. IAM SAML providers enable sts:AssumeRoleWithSAML for direct role federation.
  • GCP: Workforce Identity Federation handles humans from an external SAML/OIDC IdP. Workload Identity Federation exchanges an external OIDC token at Google STS for a federated token, which can call APIs directly or impersonate a service account.
  • Alibaba Cloud: RAM supports role-based SAML SSO and user-based SSO. For workloads, a RAM OIDC provider plus a role trust policy with a Federated principal lets a caller exchange a JWT via the STS AssumeRoleWithOIDC API. RRSA applies the same mechanism to ACK pods.
  • Tencent Cloud: CAM supports SAML role SSO and OIDC identity providers. OIDC-based roles are assumed via STS AssumeRoleWithWebIdentity.

The trust-policy and provider snippets are in How-to: Configure Identity Federation per Cloud.

Humans come through one IdP over SAML (or OIDC). Pipelines and workloads come through OIDC issuers such as GitHub Actions or the cluster's service-account issuer. No path uses long-lived keys.

flowchart TB
    IDP["Central IdP<br/>(Okta / Entra ID / Keycloak)"]
    CICD["OIDC issuers for workloads<br/>(GitHub Actions, GitLab, cluster SA issuer)"]
    IDP -->|"SAML + SCIM"| AWSIC["AWS IAM Identity Center<br/>(permission sets)"]
    IDP -->|"SAML / OIDC"| GWF["GCP Workforce<br/>Identity Federation"]
    IDP -->|"SAML"| RAMSSO["Alibaba RAM SSO"]
    IDP -->|"SAML"| CAMSSO["Tencent CAM SAML SSO"]
    CICD -->|"AssumeRoleWithWebIdentity"| AWSOIDC["AWS IAM OIDC provider"]
    CICD -->|"STS token exchange"| GWIF["GCP Workload<br/>Identity Federation"]
    CICD -->|"AssumeRoleWithOIDC"| RAMOIDC["Alibaba RAM OIDC provider"]
    CICD -->|"AssumeRoleWithWebIdentity"| CAMOIDC["Tencent CAM OIDC provider"]

A CI job that deploys to three clouds uses one ID token per audience and never stores a cloud key:

sequenceDiagram
    participant Job as GitHub Actions job
    participant GHO as GitHub OIDC issuer
    participant STS as AWS STS
    participant GSTS as Google STS
    participant IAMC as GCP IAM Credentials API
    participant ASTS as Alibaba Cloud STS
    Job->>GHO: Request ID token (audience per cloud)
    GHO-->>Job: Signed JWT (sub = repo and ref)
    Job->>STS: AssumeRoleWithWebIdentity (JWT, role ARN)
    STS-->>Job: Temporary credentials
    Job->>GSTS: Exchange JWT for federated token
    GSTS-->>Job: Federated access token
    Job->>IAMC: generateAccessToken (impersonate service account)
    IAMC-->>Job: Service account access token
    Job->>ASTS: AssumeRoleWithOIDC (JWT, OIDC provider ARN, role ARN)
    ASTS-->>Job: STS token

Best practices:

  1. Use SAML 2.0 (or OIDC) for human user SSO across all clouds.
  2. Use OIDC for workload-to-cloud and CI/CD identity federation.
  3. Map IdP groups to cloud roles with least-privilege permission sets. Do not map individual users.
  4. Enforce MFA at the IdP level. Require MFA in cloud role trust conditions where supported.
  5. Pin trust policies to specific claims (sub, aud, repository, branch), never just the issuer.
  6. Centralize audit logs (CloudTrail, Cloud Audit Logs, ActionTrail, CloudAudit) into a SIEM.
  7. Rotate IdP signing certificates on a regular cadence with overlap periods.
  8. Use conditional access policies (device compliance, IP ranges) at the IdP layer.

SPIFFE / SPIRE for Workload Identity

SPIFFE (Secure Production Identity Framework for Everyone) provides a cloud-agnostic workload identity layer. SPIRE is its reference implementation. Both are CNCF Graduated.

  • Each workload gets a SPIFFE ID (for example, spiffe://example.org/billing-service).
  • SPIRE agents on each node issue X.509 or JWT SVIDs (SPIFFE Verifiable Identity Documents) to workloads via the Workload API, after attesting the node (using cloud instance identity documents on AWS, GCP, and Azure) and the workload.
  • mTLS between workloads is bootstrapped by SVIDs, independent of cloud-specific IAM.
  • In a multi-cloud Kubernetes deployment, SPIRE federates trust domains across clusters. Workloads on AWS can then authenticate workloads on Alibaba Cloud without cloud-specific trust configuration.

Integration points:

Component Integration
Istio SPIRE as the identity provider for Istio mTLS
Envoy SDS (Secret Discovery Service) with SPIRE
Kubernetes SPIRE agent DaemonSet plus the SPIFFE CSI driver
Cloud KMS SPIRE server keys stored in AWS KMS / GCP KMS / Azure Key Vault

Security Baseline and Compliance

A multi-cloud security baseline has three layers:

  1. Benchmarks define what "secure" means per cloud. CIS publishes Foundations benchmarks for AWS (v7.0.0, 2026-04), Azure, GCP, and Alibaba Cloud. There is no CIS benchmark for Tencent Cloud, so Tencent estates usually map to China's MLPS 2.0 requirements.
  2. Frameworks (ISO 27001, SOC 2, NIST CSF 2.0, PCI DSS, GDPR/PIPL) define what auditors ask for. Most controls are met once at the organization level and inherited from each provider's own attestations for the physical layer.
  3. Enforcement combines the policy-as-code engines (preventive) with CSPM scanning (detective) described above. The same control, such as "storage must be encrypted", appears as an SCP or Org Policy, as a CI policy on Terraform plans, and as a CSPM check. That redundancy is deliberate.

Tables for benchmarks, frameworks, scanners, and policy tools, plus the baseline checklist, are in Reference: Security and Compliance Reference.

Threat Model

Multi-Cloud-Specific Threats

Threat Description Mitigation
Credential compromise in one cloud leading to lateral movement Attacker obtains credentials for one cloud and attempts to pivot to others via shared identity Conditional access policies, IP-based trust conditions, MFA requirement in trust policies, SPIFFE workload identity boundary
Misconfigured federation trust policy Overly broad SAML/OIDC trust allows any user or repository from the issuer to assume high-privilege roles Tight attribute conditions (sub, aud, repo, branch) in trust policies, group-to-role mapping with least privilege
Inconsistent security baseline across clouds Different teams manage different clouds with varying security posture CIS benchmark scanning, policy-as-code, org-level guardrails, cross-cloud security dashboards
Policy bypass outside the pipeline Console or CLI changes skip CI and Terraform policy checks Org-level guardrails (SCP/RCP, Azure Policy, Org Policy), admission control, drift detection via CSPM or Cloud Custodian
Data exfiltration via inter-cloud transfer Attacker moves data from a secured cloud to a less-secured cloud or external endpoint DLP policies, egress filtering, data-perimeter policies (RCPs, VPC Service Controls), flow-log anomaly detection
Supply chain attack on IaC modules Compromised Terraform module, provider, or Helm chart deploys malicious infrastructure Pin module and provider versions (lock files), scan IaC with Checkov/Trivy, sign commits, require PR approval
DNS hijacking across multi-cloud zones Attacker compromises DNS credentials and redirects traffic DNSSEC, MFA on the DNS registrar and providers, DNS change monitoring
Cloud provider supply chain Vulnerability in a managed service affects workloads Multi-cloud deployment for critical workloads, regular patch cycle, vulnerability scanning

Sources