Multi-Cloud Governance -- How-to Guides¶
Task recipes for running governance across AWS, GCP, Alibaba Cloud, and Tencent Cloud (with Azure where relevant): policy-as-code, posture scanning, identity federation, IaC, GitOps, observability, and FinOps. For why these patterns exist, see Explanation. For versions and tool matrices, see Reference.
Placeholders
Account IDs, project IDs, bucket names, and hostnames below are examples. Replace them with your own values. Tool versions were checked on 2026-09-25. Pin versions in real pipelines.
Policy as Code¶
Check Terraform Plans with Conftest (OPA)¶
Goal: fail a pull request when a planned resource lacks the mandatory cost-center tag, before anything is applied.
-
Write a Rego policy in
policy/tags.rego. Therego.v1import keeps it valid on both older and 1.x OPA-based tools:package main import rego.v1 taggable := {"aws_instance", "aws_s3_bucket", "alicloud_instance", "tencentcloud_instance"} deny contains msg if { some rc in input.resource_changes taggable[rc.type] rc.change.actions[_] in {"create", "update"} not rc.change.after.tags["cost-center"] msg := sprintf("%s is missing the cost-center tag", [rc.address]) } -
Produce the plan as JSON and test it in CI:
-
On HCP Terraform or Terraform Enterprise, attach the same Rego as an OPA policy set, or write a Sentinel policy, so that the check also runs on remote runs.
Scan IaC with Checkov and Trivy¶
# Checkov: Terraform, Kubernetes, Helm, and more (Apache-2.0)
pip install checkov
checkov -d ./infra --framework terraform --compact
# Trivy replaces tfsec for IaC misconfiguration scanning
trivy config ./infra
Enforce Kubernetes Policy with Gatekeeper¶
Install Gatekeeper (3.23.x as of 2026-09) and add a required-labels constraint from the Gatekeeper library:
helm repo add gatekeeper https://open-policy-agent.github.io/gatekeeper/charts
helm install gatekeeper/gatekeeper --name-template=gatekeeper \
--namespace gatekeeper-system --create-namespace
# ConstraintTemplate K8sRequiredLabels from the gatekeeper-library
kubectl apply -f https://raw.githubusercontent.com/open-policy-agent/gatekeeper-library/master/library/general/requiredlabels/template.yaml
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
name: namespaces-need-cost-labels
spec:
enforcementAction: dryrun # audit first, switch to deny once clean
match:
kinds:
- apiGroups: [""]
kinds: ["Namespace"]
parameters:
message: "Namespaces must carry team and cost-center labels"
labels:
- key: team
- key: cost-center
Check the audit results with kubectl get k8srequiredlabels namespaces-need-cost-labels -o yaml (see status.violations).
Enforce Kubernetes Policy with Kyverno¶
Kyverno 1.17+ uses CEL-based policy types at policies.kyverno.io/v1. The legacy ClusterPolicy is deprecated.
helm repo add kyverno https://kyverno.github.io/kyverno/
helm install kyverno kyverno/kyverno -n kyverno --create-namespace
apiVersion: policies.kyverno.io/v1
kind: ValidatingPolicy
metadata:
name: require-cost-labels
spec:
validationActions: [Audit] # change to [Deny] after the backlog is fixed
matchConstraints:
resourceRules:
- apiGroups: ["apps"]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["deployments"]
validations:
- expression: >-
has(object.metadata.labels) &&
'team' in object.metadata.labels &&
'cost-center' in object.metadata.labels
message: "Deployments must carry team and cost-center labels."
Kyverno writes results to PolicyReport objects: kubectl get policyreports -A.
Detect and Remediate Untagged Resources with Cloud Custodian¶
Cloud Custodian covers AWS, Azure, GCP, Kubernetes, OCI, and Tencent Cloud (c7n-tencentcloud). It does not cover Alibaba Cloud.
# policy.yml
policies:
- name: ec2-missing-cost-center
resource: aws.ec2
filters:
- "tag:cost-center": absent
actions:
- type: mark-for-op
tag: c7n_tag_compliance
op: stop
days: 3
pip install c7n c7n-gcp c7n-azure c7n-tencentcloud
custodian validate policy.yml
custodian run --dryrun -s out/ policy.yml # report only
custodian run -s out/ policy.yml # apply actions
For continuous enforcement, add a mode block (for example type: periodic or a CloudTrail event mode on AWS) so that Custodian deploys the policy as a serverless function.
Run CIS Benchmark Scans with Prowler¶
pip install prowler
prowler aws --list-compliance # shows the framework IDs available
prowler aws --compliance cis_7.0_aws
prowler gcp --compliance cis_5.0_gcp
prowler azure --compliance cis_6.0_azure
prowler alibabacloud --compliance cis_2.0_alibabacloud
Prowler has no Tencent Cloud provider. For Tencent, use Tencent Cloud's own security products or Cloud Custodian policies.
Configure Identity Federation per Cloud¶
AWS: SAML Trust Policy for a Federated Role¶
{
"Version": "2012-10-17",
"Statement": [{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::123456789012:saml-provider/Okta"
},
"Action": "sts:AssumeRoleWithSAML",
"Condition": {
"StringEquals": {
"SAML:aud": "https://signin.aws.amazon.com/saml"
}
}
}]
}
For workforce access, prefer IAM Identity Center with SCIM provisioning and permission sets. Use direct SAML roles like this one only for special cases.
GCP: Workload Identity Pool Provider for GitHub Actions (Terraform)¶
resource "google_iam_workload_identity_pool_provider" "github" {
workload_identity_pool_id = google_iam_workload_identity_pool.github.workload_identity_pool_id
workload_identity_pool_provider_id = "github-actions"
attribute_mapping = {
"google.subject" = "assertion.sub"
"attribute.repository" = "assertion.repository"
"attribute.ref" = "assertion.ref"
}
# Restrict which repositories can use the pool
attribute_condition = "assertion.repository_owner == 'example-org'"
oidc {
issuer_uri = "https://token.actions.githubusercontent.com"
}
}
Alibaba Cloud: RAM Role Trust Policy for an OIDC Provider¶
A RAM OIDC trust policy uses sts:AssumeRole with a Federated principal. The caller exchanges its JWT through the STS AssumeRoleWithOIDC API.
{
"Statement": [{
"Action": "sts:AssumeRole",
"Effect": "Allow",
"Principal": {
"Federated": ["acs:ram::1234567890123456:oidc-provider/github-actions"]
},
"Condition": {
"StringEquals": {
"oidc:aud": "sts.aliyuncs.com",
"oidc:iss": "https://token.actions.githubusercontent.com",
"oidc:sub": "repo:example-org/infra:ref:refs/heads/main"
}
}
}],
"Version": "1"
}
For ACK pods, RRSA creates the cluster's OIDC provider for you. The oidc:sub value is then system:serviceaccount:<namespace>:<serviceaccount> (Alibaba Cloud docs).
Tencent Cloud: OIDC Role¶
Create an OIDC identity provider in CAM, then a role whose trust policy names that provider and restricts the aud/sub claims. Workloads call STS AssumeRoleWithWebIdentity. The exact trust-policy JSON keys differ from AWS. Copy them from the CAM console's generated policy rather than reusing an AWS template.
Infrastructure-as-Code Recipes¶
Configure Multiple Cloud Providers in One Terraform Configuration¶
terraform {
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 6.0"
}
google = {
source = "hashicorp/google"
version = "~> 8.0"
}
alicloud = {
source = "aliyun/alicloud"
version = "~> 1.290"
}
tencentcloud = {
source = "tencentcloudstack/tencentcloud"
version = "~> 1.83"
}
}
}
provider "aws" {
region = "ap-southeast-1"
}
provider "google" {
project = "my-project"
region = "asia-southeast1"
}
provider "alicloud" {
region = "cn-hangzhou"
}
provider "tencentcloud" {
region = "ap-guangzhou"
}
In production, split state per cloud (one root module or workspace per cloud and environment) rather than using one state for all four providers.
Use OIDC Identity Tokens in Terraform Stacks¶
In a Stack's *.tfdeploy.hcl, identity_token blocks give each deployment short-lived workload-identity JWTs that the components pass to their providers:
identity_token "aws" {
audience = ["aws.workload.identity"]
}
identity_token "gcp" {
audience = ["//iam.googleapis.com/projects/123456789012/locations/global/workloadIdentityPools/hcp-terraform/providers/hcp-terraform"]
}
deployment "multi_cloud" {
inputs = {
aws_token = identity_token.aws.jwt
gcp_token = identity_token.gcp.jwt
aws_role = "arn:aws:iam::123456789012:role/terraform-role"
gcp_sa = "terraform@my-project.iam.gserviceaccount.com"
}
}
The AWS role and the GCP pool must trust HCP Terraform's OIDC issuer. See the Terraform topic for Stacks details.
Run a Multi-Cloud Terraform Plan¶
terraform init
terraform plan \
-var-file="aws-ap-southeast-1.tfvars" \
-var-file="alibaba-cn-hangzhou.tfvars" \
-out=multi-cloud.tfplan
terraform apply multi-cloud.tfplan
Configure Pulumi Providers per Cloud¶
import * as aws from "@pulumi/aws";
import * as gcp from "@pulumi/gcp";
import * as alicloud from "@pulumi/alicloud";
const awsProvider = new aws.Provider("aws-provider", {
region: "ap-southeast-1",
});
const gcpProvider = new gcp.Provider("gcp-provider", {
project: "my-project",
region: "asia-southeast1",
});
const aliProvider = new alicloud.Provider("ali-provider", {
region: "cn-hangzhou",
});
pulumi stack select prod-aws
pulumi preview --diff
pulumi up --stack prod-aws --yes
pulumi stack select prod-alibaba
pulumi up --stack prod-alibaba --yes
# Read an output from another stack
pulumi stack output --stack prod-aws vpc_id
Install Crossplane Providers and Define a Cloud-Agnostic API (v2)¶
Crossplane v2 removed the default package registry, so package references must be fully qualified:
apiVersion: pkg.crossplane.io/v1
kind: Provider
metadata:
name: crossplane-contrib-provider-aws-s3
spec:
package: xpkg.crossplane.io/crossplane-contrib/provider-aws-s3:v2.0.0
For Alibaba Cloud and Tencent Cloud, take the package path and current version from the provider-upjet-alibabacloud and provider-tencentcloud READMEs.
A namespaced, cloud-agnostic XRD (v2 API) that compositions can implement per cloud:
apiVersion: apiextensions.crossplane.io/v2
kind: CompositeResourceDefinition
metadata:
name: xdatastores.example.org
spec:
scope: Namespaced
group: example.org
names:
kind: XDataStore
plural: xdatastores
versions:
- name: v1alpha1
served: true
referenceable: true
schema:
openAPIV3Schema:
type: object
properties:
spec:
type: object
properties:
engine:
type: string
enum: ["mysql", "postgresql"]
size:
type: string
cloudProvider:
type: string
enum: ["aws", "gcp", "alibaba"]
kubectl apply -f provider-aws-s3.yaml -f xrd.yaml -f compositions/
kubectl get providers
kubectl get xdatastores -A
kubectl get managed -A
kubectl describe xdatastore my-prod-db -n team-a
The Crossplane CLI was restructured in 2026 (new crossplane/cli repository). Check crossplane --help for the current names of the render and trace commands in your CLI version.
GitOps Across Clusters¶
Register Clusters and Sync with Argo CD¶
# Register remote clusters (kubeconfig contexts)
argocd cluster add gke_my-project_asia-southeast1_prod --name prod-gke
argocd cluster add alibaba-ack-context --name prod-ack
# Inspect the ApplicationSet and the Applications it generated
argocd appset get my-appset
argocd app list --output wide
# Sync generated Applications explicitly (if auto-sync is off)
argocd app sync prod-gke-my-app prod-ack-my-app --prune
argocd appset only manages ApplicationSet objects (create, delete, generate, get, list, update). Syncing happens per Application, or automatically when the template sets syncPolicy.automated. See the Argo CD topic.
Observability¶
Deploy Per-Cloud OTel Collectors¶
Deploy DaemonSet agents in each cluster and one gateway collector per cloud. The gateway config below enriches telemetry with cloud attributes and exports to a central backend:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 75
spike_limit_percentage: 15
batch:
send_batch_size: 8192
timeout: 5s
resource:
attributes:
- key: cloud.provider
value: "aws" # per cloud: aws, gcp, azure, alibaba_cloud, tencent_cloud
action: upsert
- key: cloud.region
value: "ap-southeast-1"
action: upsert
exporters:
otlphttp/grafana:
endpoint: "https://otlp-gateway-prod-eu-west-0.grafana.net/otlp"
headers:
Authorization: "Basic ${env:GRAFANA_OTLP_AUTH}"
debug:
verbosity: basic # the old `logging` exporter was removed; use `debug`
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, resource, batch]
exporters: [otlphttp/grafana]
metrics:
receivers: [otlp]
processors: [memory_limiter, resource, batch]
exporters: [otlphttp/grafana]
logs:
receivers: [otlp]
processors: [memory_limiter, resource, batch]
exporters: [otlphttp/grafana]
To dual-export traces to AWS X-Ray, use an otlphttp exporter pointed at the X-Ray OTLP endpoint (https://xray.<region>.amazonaws.com) with the contrib sigv4auth extension, or use the ADOT Collector, which bundles it.
Check Collector Health¶
kubectl get pods -n otel-system -l app.kubernetes.io/name=opentelemetry-collector
# zpages (55679) and internal metrics (8888) bind to localhost by default in
# recent releases, so port-forward instead of curling the Service
kubectl port-forward -n otel-system deploy/otel-gateway 55679:55679 8888:8888 &
curl -s http://localhost:55679/debug/tracez | head
curl -s http://localhost:8888/metrics | grep otelcol_exporter_send_failed
Set Up Cross-Cloud Alerting¶
- Each cluster runs a Prometheus-compatible scraper (Prometheus, VictoriaMetrics agent, or the OTel Collector Prometheus receiver).
- Metrics flow to one central Prometheus-compatible store (Mimir, VictoriaMetrics, Thanos).
- Alertmanager or Grafana Alerting evaluates rules against the unified metrics.
- Alerts route to a single notification system (PagerDuty, Opsgenie, Slack).
Useful cross-cloud alert rules:
- Inter-cloud latency SLO breach (measured via synthetic probes between clouds).
- Error-budget burn rate across clouds (combined service-level indicator).
- Cost anomaly (spend spike in any cloud above a threshold).
- Security events (new admin role, root account usage, failed federation attempts).
- Certificate expiration across clouds (cert-manager, ACM, Certificate Authority Service).
FinOps and Cost Management¶
Enforce a Cross-Cloud Tagging Policy¶
- Agree on the mandatory keys (see Reference: Tagging Strategy). Use lowercase keys that are valid GCP labels.
- Bake the tags into shared IaC modules (
default_tagsin the AWS provider,labelsdefaults in Google modules). - Block untagged plans in CI (Conftest recipe).
- Enforce at the organization level: AWS tag policies or SCP conditions, Azure Policy, GCP custom org-policy constraints, Alibaba tag policies.
- Detect stragglers with Cloud Custodian or CSPM, and report unallocated spend weekly.
Query FOCUS Exports Across Clouds¶
Once each provider's FOCUS export lands in object storage (one prefix per provider), one query covers every cloud. Example with DuckDB over Parquet files:
duckdb -c "
SELECT regexp_extract(filename, 'focus/([^/]+)/', 1) AS source,
ServiceName,
round(sum(EffectiveCost), 2) AS effective_cost
FROM read_parquet('focus/*/*.parquet', filename = true, union_by_name = true)
WHERE ChargePeriodStart >= TIMESTAMP '2026-09-01'
GROUP BY ALL
ORDER BY effective_cost DESC
LIMIT 20;"
Use read_csv instead for CSV exports (for example Tencent's converted bills). Check the FOCUS version of each export, because column sets differ between 1.0 and 1.4.
Pull Native Cost Reports¶
# AWS -- daily cost by service
aws ce get-cost-and-usage \
--time-period Start=2026-09-01,End=2026-09-25 \
--granularity DAILY \
--metrics UnblendedCost \
--group-by Type=DIMENSION,Key=SERVICE
# GCP -- billing export query (BigQuery)
bq query --use_legacy_sql=false \
'SELECT service.description, SUM(cost) AS total_cost
FROM `my-project.billing_dataset.gcp_billing_export_v1_XXXXXX_XXXXXX_XXXXXX`
WHERE invoice.month = "202609"
GROUP BY service.description
ORDER BY total_cost DESC'
# Alibaba -- via OpenAPI
aliyun bssopenapi QueryBill \
--BillingCycle 2026-09 \
--PageNum 1 \
--PageSize 100
Plan Commitments per Cloud¶
- Analyze steady-state compute usage per cloud. Identify workloads running 24/7.
- Buy RIs, Savings Plans, or CUDs for the steady-state baseline in each cloud (discount levels: Reference: Commitment-Based Savings).
- Use spot/preemptible capacity for fault-tolerant, batch, or stateless workloads.
- Review coverage and utilization monthly. Underused commitments are wasted spend.
- For workloads that can run in any cloud, include spot price and commitment headroom in placement decisions.
Troubleshooting¶
Common Policy and Guardrail Issues¶
| Symptom | Likely Cause | Resolution |
|---|---|---|
| Cluster operations hang or fail after installing an admission controller | Webhook unreachable with failurePolicy: Fail |
Exempt system namespaces, run multiple webhook replicas, and check webhook Service endpoints |
Terraform apply fails with AccessDenied although the IAM role allows the action |
An SCP, RCP, or permission boundary denies it | Check the organization policies on the account. SCPs cap permissions and never grant them |
| Policy passes in CI but a violation appears in CSPM | Change made outside the pipeline (console/CLI) or pre-existing drift | Add an org-level guardrail for the rule, and remediate with Cloud Custodian |
| Kyverno results change after migrating a policy | A rewritten ValidatingPolicy does not match exactly what the legacy ClusterPolicy matched |
Test both versions with the Kyverno CLI (kyverno apply / kyverno test) in CI, and run the new policy in Audit first |
Common Multi-Cloud Observability Issues¶
| Symptom | Likely Cause | Resolution |
|---|---|---|
| Missing traces in one cloud | OTel Collector pod crashlooping or misconfigured exporter | Check collector pod logs. Verify OTLP endpoint and credentials |
| High collector memory | Insufficient memory_limiter or excessive log volume | Tune limit_percentage and spike_limit_percentage. Add sampling |
| Attribute conflicts in backend | Different cloud.provider values for same service |
Verify the resource processor sets consistent semantic-convention values per deployment |
| Cross-cloud trace breaks | Missing W3C Trace Context propagation in a service hop | Verify all services use OTel SDK with the W3C propagator. Check load-balancer pass-through of trace headers |
| Collector config rejected after upgrade | Removed components such as the logging exporter |
Replace with debug. Read the collector changelog before upgrading |
Common Multi-Cloud Networking Issues¶
| Symptom | Likely Cause | Resolution |
|---|---|---|
| Intermittent inter-cloud latency spikes | Traffic routing over public internet instead of dedicated interconnect | Verify route tables point to Direct Connect / Express Connect / Interconnect. Check BGP routes |
| DNS resolution failures between clouds | Split-horizon DNS misconfiguration or stale NS delegation | Verify NS records at the apex zone match the cloud DNS zone name servers. Check DNS propagation |
| Connection resets between clouds | MTU mismatch on interconnect circuits | Verify MTU settings on the VBR (Alibaba), Direct Connect virtual interface (AWS), and interconnect attachment |
Common FinOps Issues¶
| Symptom | Likely Cause | Resolution |
|---|---|---|
| Unattributed spend (>20% of total) | Missing tags on resources | Enforce tagging at provisioning time via IaC modules and cloud org policies |
| RI/Savings Plan underutilization | Over-purchased commitments or workload migration | Rightsize the commitment portfolio. Exchange convertible RIs. Adjust Savings Plan coverage targets |
| Cost anomaly not detected | Alert threshold too high or no anomaly detection configured | Configure AWS Cost Anomaly Detection and budget alerts in every cloud |
| Totals differ between FOCUS data and invoices | Mixed FOCUS versions, or EffectiveCost compared with BilledCost |
Compare like with like. Use the FOCUS 1.4 invoice dataset for reconciliation where available |