Skip to content

GCP Landing Zone: How-to Guides

Deployment recipes, Terraform configurations, CLI commands, and operational tasks for each architecture pattern. Placeholders are in UPPER_CASE. Look-up facts (constraint names, IP ranges, prices, versions) are in the Reference. Design reasoning is in the Explanation.

Test in a sandbox organization first

Organization policies, hierarchical firewall policies, and VPC Service Controls apply to everything below the node you set them on. A wrong rule at the organization level can break every project at once. Use dry-run modes where they exist.

  1. Configure Cloud Identity or Google Workspace, create admin groups, and let the Organization resource be created
  2. Create the folder hierarchy (for example Production, Non-Production, Shared-Infrastructure)
  3. Review the security baseline constraints already enforced on new orgs, then apply the remaining baseline organization policies
  4. Create shared infrastructure projects (networking, security, logging)
  5. Configure Shared VPC with subnets per environment and region
  6. Apply a hierarchical firewall policy at the organization or folder, and network firewall policies per VPC
  7. Configure IAM groups and roles at folder level. Add Privileged Access Manager entitlements for admin roles
  8. Enable Security Command Center. Create VPC Service Controls perimeters in dry-run mode first
  9. Configure Cloud Interconnect or Cloud VPN for hybrid connectivity
  10. Configure centralized logging sinks and monitoring dashboards
  11. Implement a project factory (Fabric FAST or CFT project-factory module) with guardrails
  12. Create CI/CD pipelines for infrastructure changes

For a full automated deployment, start from Cloud Foundation Fabric FAST (fast/stages/0-org-setup onward in recent releases) or the enterprise foundations blueprint (terraform-example-foundation, stages 0-bootstrap to 5-app-infra). Follow the README of the release you pin, because stage names and variables change between major versions. See Landing-Zone Blueprints and Tooling Versions.

Project and Hierarchy Management

Project Factory Pattern

Use Terraform with the Cloud Foundation Toolkit (CFT) project factory module to standardize project creation. Version 18.x needs Terraform 1.3+ and supports Google provider v7.

module "project" {
  source  = "terraform-google-modules/project-factory/google"
  version = "~> 18.3"

  name              = "prod-payments-api"
  random_project_id = true
  org_id            = var.org_id
  folder_id         = google_folder.production.name
  billing_account   = var.billing_account_id

  # Attach to the Shared VPC host and grant networkUser on these subnets
  svpc_host_project_id = "shared-net-prod"
  shared_vpc_subnets = [
    "projects/shared-net-prod/regions/us-central1/subnetworks/app-subnet",
  ]

  # Labels for cost attribution
  labels = {
    team        = "payments"
    environment = "prod"
    cost-center = "cc-1234"
    managed-by  = "terraform"
  }

  # APIs to enable
  activate_apis = [
    "compute.googleapis.com",
    "container.googleapis.com",
    "sqladmin.googleapis.com",
    "monitoring.googleapis.com",
  ]
}

Source: terraform-google-project-factory README.

Folder Creation

# Create top-level folders under the organization
gcloud resource-manager folders create \
  --display-name="Production" \
  --organization=ORGANIZATION_ID

gcloud resource-manager folders create \
  --display-name="Non-Production" \
  --organization=ORGANIZATION_ID

gcloud resource-manager folders create \
  --display-name="Shared-Infrastructure" \
  --organization=ORGANIZATION_ID

# Create sub-folders
gcloud resource-manager folders create \
  --display-name="Team-A" \
  --folder=PRODUCTION_FOLDER_ID

Baseline Organization Policies

Apply these at the organization level with Terraform. The example uses google_org_policy_policy (Organization Policy v2 API). The older google_organization_policy resource (v1 API) still works but lacks conditions and dry-run support.

# Enforce OS Login (boolean constraint)
resource "google_org_policy_policy" "require_os_login" {
  name   = "organizations/${var.org_id}/policies/compute.requireOsLogin"
  parent = "organizations/${var.org_id}"

  spec {
    rules {
      enforce = "TRUE"
    }
  }
}

# Restrict resource locations to US and EU value groups (list constraint)
resource "google_org_policy_policy" "resource_locations" {
  name   = "organizations/${var.org_id}/policies/gcp.resourceLocations"
  parent = "organizations/${var.org_id}"

  spec {
    rules {
      values {
        allowed_values = ["in:us-locations", "in:eu-locations"]
      }
    }
  }
}

# Disable service account key creation
resource "google_org_policy_policy" "disable_sa_keys" {
  name   = "organizations/${var.org_id}/policies/iam.disableServiceAccountKeyCreation"
  parent = "organizations/${var.org_id}"

  spec {
    rules {
      enforce = "TRUE"
    }
  }
}

# Deny external IPs on all VMs
resource "google_org_policy_policy" "vm_external_ip" {
  name   = "organizations/${var.org_id}/policies/compute.vmExternalIpAccess"
  parent = "organizations/${var.org_id}"

  spec {
    rules {
      deny_all = "TRUE"
    }
  }
}

# Enforce Shielded VMs
resource "google_org_policy_policy" "require_shielded_vm" {
  name   = "organizations/${var.org_id}/policies/compute.requireShieldedVm"
  parent = "organizations/${var.org_id}"

  spec {
    rules {
      enforce = "TRUE"
    }
  }
}

Location value groups use in:

Value groups are written in:us-locations, not in/us-locations. Source: Restrict resource locations.

Test an Organization Policy in Dry-Run Mode

Managed constraints support a dryRunSpec that logs violations without blocking. Write the policy to a file:

# policy.yaml
name: organizations/ORGANIZATION_ID/policies/compute.managed.requireOsLogin
dryRunSpec:
  rules:
  - enforce: true

Apply it, then look for would-be violations in Cloud Audit Logs before you move the rules into spec:

gcloud org-policies set-policy policy.yaml

# Inspect the effective policy on a project
gcloud org-policies describe compute.managed.requireOsLogin \
  --project=PROJECT_ID --effective

Policy changes can take up to 15 minutes to propagate. Source: Managed constraints.

Networking Operations

Shared VPC Setup

# 1. Enable the host project (needs roles/compute.xpnAdmin)
gcloud compute shared-vpc enable shared-net-prod

# 2. Attach service projects
gcloud compute shared-vpc associated-projects add payments-prod \
  --host-project=shared-net-prod

# 3. Grant subnet-level access to service project principals
gcloud compute networks subnets add-iam-policy-binding app-subnet \
  --project=shared-net-prod \
  --region=us-central1 \
  --member="group:payments-team@example.com" \
  --role="roles/compute.networkUser"

Also grant roles/compute.networkUser on the subnet to the service project's service agents that create resources in it (for example the GKE service agent service-PROJECT_NUMBER@container-engine-robot.iam.gserviceaccount.com and the Google APIs service agent PROJECT_NUMBER@cloudservices.gserviceaccount.com). Source: Shared VPC provisioning.

VPC Peering Setup

# Create peering from network-a to network-b
gcloud compute networks peerings create peer-a-to-b \
  --network=network-a \
  --peer-network=network-b \
  --peer-project=project-b \
  --export-custom-routes \
  --import-custom-routes

# Create reciprocal peering from network-b to network-a (run in project-b)
gcloud compute networks peerings create peer-b-to-a \
  --network=network-b \
  --peer-network=network-a \
  --peer-project=project-a \
  --export-custom-routes \
  --import-custom-routes

# Verify peering status (state should be ACTIVE on both sides)
gcloud compute networks peerings list --network=network-a

Cloud NAT Configuration

# Create Cloud Router (required for Cloud NAT)
gcloud compute routers create nat-router \
  --network=shared-vpc \
  --region=us-central1

# Create Cloud NAT on the router
gcloud compute routers nats create nat-gateway \
  --router=nat-router \
  --region=us-central1 \
  --nat-all-subnet-ip-ranges \
  --auto-allocate-nat-ips \
  --enable-logging

Private Service Connect for Google APIs

A PSC endpoint for Google APIs is a global internal address plus a global forwarding rule. The forwarding rule name must be 1-20 lowercase letters or digits and start with a letter.

# Reserve a global internal IP for the endpoint
gcloud compute addresses create psc-apis-ip \
  --global \
  --purpose=PRIVATE_SERVICE_CONNECT \
  --addresses=10.255.255.254 \
  --network=shared-vpc

# Create the endpoint: all-apis (like private.googleapis.com)
# or vpc-sc (like restricted.googleapis.com, for VPC SC perimeters)
gcloud compute forwarding-rules create pscapis \
  --global \
  --network=shared-vpc \
  --address=psc-apis-ip \
  --target-google-apis-bundle=all-apis

Then create Cloud DNS private zones (for example googleapis.com) that resolve API hostnames to the endpoint IP. Source: Access Google APIs through endpoints.

Hierarchical Firewall Policy (Organization Baseline)

# Create the policy under the organization
gcloud compute firewall-policies create \
  --organization=ORGANIZATION_ID \
  --short-name=org-baseline \
  --description="Org-wide baseline rules"

# Allow IAP TCP forwarding to SSH everywhere
gcloud compute firewall-policies rules create 1000 \
  --firewall-policy=org-baseline \
  --organization=ORGANIZATION_ID \
  --direction=INGRESS \
  --action=allow \
  --src-ip-ranges=35.235.240.0/20 \
  --layer4-configs=tcp:22

# Delegate everything else to folder and network policies
gcloud compute firewall-policies rules create 65000 \
  --firewall-policy=org-baseline \
  --organization=ORGANIZATION_ID \
  --direction=INGRESS \
  --action=goto_next \
  --src-ip-ranges=0.0.0.0/0

# Associate the policy with the organization
gcloud compute firewall-policies associations create \
  --firewall-policy=org-baseline \
  --organization=ORGANIZATION_ID \
  --name=org-baseline-assoc

Source: Create hierarchical firewall policies.

Global Network Firewall Policy with Secure Tags

# Create the policy and associate it with the VPC
gcloud compute network-firewall-policies create prod-fw-policy \
  --global \
  --description="Workload rules for shared-vpc"

gcloud compute network-firewall-policies associations create \
  --firewall-policy=prod-fw-policy \
  --network=shared-vpc \
  --name=prod-fw-assoc \
  --global-firewall-policy

# Allow LB health checks only to VMs bound to the web secure tag value
gcloud compute network-firewall-policies rules create 1000 \
  --firewall-policy=prod-fw-policy \
  --global-firewall-policy \
  --direction=INGRESS \
  --action=allow \
  --src-ip-ranges=35.191.0.0/16,130.211.0.0/22 \
  --layer4-configs=tcp:80,tcp:443 \
  --target-secure-tags=tagValues/TAG_VALUE_ID

Create the secure tag key with --purpose=GCE_FIREWALL first, and bind tag values to VM instances. Source: Create and manage secure tags.

Security Operations

VPC Service Controls Perimeter

Create an access level from a YAML list of conditions:

# access-level.yaml
- ipSubnetworks:
  - 203.0.113.0/24
# Create an access level (allowed network ranges)
gcloud access-context-manager levels create corporate_network \
  --title="Corporate Network" \
  --basic-level-spec=access-level.yaml \
  --policy=ACCESS_POLICY_ID

# Create a perimeter in dry-run mode first
gcloud access-context-manager perimeters dry-run create prod_perimeter \
  --perimeter-title="Production Perimeter" \
  --perimeter-type=regular \
  --perimeter-resources=projects/PROJECT_NUMBER \
  --perimeter-restricted-services=bigquery.googleapis.com,storage.googleapis.com \
  --perimeter-access-levels=corporate_network \
  --policy=ACCESS_POLICY_ID

# After reviewing dry-run violations in audit logs, enforce it
gcloud access-context-manager perimeters dry-run enforce prod_perimeter \
  --policy=ACCESS_POLICY_ID

To create an enforced perimeter directly, use gcloud access-context-manager perimeters create with --title, --resources, --restricted-services, and --access-levels. Access level and perimeter names allow letters, digits, and underscores. Source: Create a basic access level, VPC SC dry run.

Cloud Armor WAF Policy

# Create security policy
gcloud compute security-policies create prod-waf-policy \
  --description="Production WAF policy"

# Add preconfigured WAF rules (OWASP CRS 4.22 based)
gcloud compute security-policies rules create 1000 \
  --security-policy=prod-waf-policy \
  --expression="evaluatePreconfiguredWaf('sqli-v422-stable', {'sensitivity': 1})" \
  --action=deny-403 \
  --description="Block SQL injection"

gcloud compute security-policies rules create 1001 \
  --security-policy=prod-waf-policy \
  --expression="evaluatePreconfiguredWaf('xss-v422-stable', {'sensitivity': 1})" \
  --action=deny-403 \
  --description="Block XSS"

# Attach to backend service
gcloud compute backend-services update prod-backend \
  --security-policy=prod-waf-policy \
  --global

Start new rules with --preview to log matches without blocking, then remove the flag. Source: Set up preconfigured WAF rules.

Enforce TLS 1.2+ on an External Load Balancer

gcloud compute ssl-policies create tls12-modern \
  --profile=MODERN \
  --min-tls-version=1.2

gcloud compute target-https-proxies update prod-https-proxy \
  --ssl-policy=tls12-modern

IAP TCP Forwarding for SSH

# Enable OS Login project-wide (IAM-based SSH access)
gcloud compute project-info add-metadata \
  --metadata=enable-oslogin=TRUE

# Allow the IAP range to reach SSH (VPC firewall rule variant)
gcloud compute firewall-rules create allow-iap-ssh \
  --network=shared-vpc \
  --direction=INGRESS \
  --action=allow \
  --rules=tcp:22 \
  --source-ranges=35.235.240.0/20

# Grant the IAP-secured Tunnel User role on the instance
gcloud compute instances add-iam-policy-binding my-instance \
  --zone=us-central1-a \
  --member="user:admin@example.com" \
  --role="roles/iap.tunnelResourceAccessor"

# SSH through IAP (no external IP needed on the VM)
gcloud compute ssh my-instance --zone=us-central1-a --tunnel-through-iap

With OS Login, the user also needs roles/compute.osLogin (or roles/compute.osAdminLogin). Source: Using IAP for TCP forwarding.

Compute and Orchestration

Regional Managed Instance Group

# Create instance template in a Shared VPC subnet
gcloud compute instance-templates create app-template \
  --machine-type=e2-standard-4 \
  --image-family=debian-12 \
  --image-project=debian-cloud \
  --subnet=projects/shared-net-prod/regions/us-central1/subnetworks/app-subnet \
  --region=us-central1 \
  --no-address \
  --shielded-secure-boot \
  --tags=app-server

# Create regional MIG
gcloud compute instance-groups managed create app-mig \
  --template=app-template \
  --region=us-central1 \
  --size=3 \
  --health-check=http-health-check \
  --initial-delay=120

# Set autoscaling
gcloud compute instance-groups managed set-autoscaling app-mig \
  --region=us-central1 \
  --min-num-replicas=3 \
  --max-num-replicas=20 \
  --target-cpu-utilization=0.7

GKE Regional Cluster in Shared VPC

gcloud container clusters create prod-cluster \
  --project=prod-project \
  --region=us-central1 \
  --network=projects/shared-net-prod/global/networks/shared-vpc \
  --subnetwork=projects/shared-net-prod/regions/us-central1/subnetworks/gke-subnet \
  --cluster-secondary-range-name=gke-pods \
  --services-secondary-range-name=gke-services \
  --enable-ip-alias \
  --enable-private-nodes \
  --num-nodes=1 \
  --enable-autoscaling \
  --min-nodes=1 \
  --max-nodes=10 \
  --workload-pool=prod-project.svc.id.goog \
  --binauthz-evaluation-mode=PROJECT_SINGLETON_POLICY_ENFORCE

--num-nodes is per zone, so a regional cluster in three zones starts with three nodes. --enable-binary-authorization is deprecated in favor of --binauthz-evaluation-mode. Shielded GKE nodes are on by default. Source: Binary Authorization with GKE.

Data Layer Operations

Cloud SQL HA with Cross-Region Replica

Private IP in a Shared VPC needs private services access (an allocated range peered with servicenetworking.googleapis.com) configured once in the host project.

# Create HA instance (automatic zonal failover within the region)
gcloud sql instances create prod-db \
  --database-version=POSTGRES_16 \
  --edition=ENTERPRISE_PLUS \
  --tier=db-perf-optimized-N-4 \
  --region=us-central1 \
  --availability-type=REGIONAL \
  --network=projects/shared-net-prod/global/networks/shared-vpc \
  --no-assign-ip

# Create cross-region read replica for DR
gcloud sql instances create prod-db-replica \
  --master-instance-name=prod-db \
  --region=us-east1 \
  --tier=db-perf-optimized-N-4 \
  --network=projects/shared-net-prod/global/networks/shared-vpc \
  --no-assign-ip

For Enterprise edition, use --edition=ENTERPRISE with a db-custom-CPU-MEMORY tier (for example db-custom-4-16384). PostgreSQL 16 and later defaults to Enterprise Plus. Source: Create instances.

Spanner Multi-Region

# Multi-region configurations need the Enterprise Plus edition
gcloud spanner instances create prod-spanner \
  --config=nam6 \
  --edition=ENTERPRISE_PLUS \
  --description="Production Spanner" \
  --nodes=3

# nam6: read-write us-central1 + us-east1, read-only us-west1 + us-west2, witness us-central2
# List every available configuration:
gcloud spanner instance-configs list

Source: Spanner instance configurations, Spanner editions.

Dual-Region Cloud Storage

# Predefined dual-region NAM4 (us-central1 + us-east1) with turbo replication (15-minute RPO)
gcloud storage buckets create gs://prod-assets-bucket \
  --location=NAM4 \
  --default-storage-class=STANDARD \
  --rpo=ASYNC_TURBO

# Configurable dual-region: continent code plus placement
gcloud storage buckets create gs://prod-assets-bucket-2 \
  --location=US \
  --placement=us-central1,us-east1

Source: Bucket locations, Create a bucket.

DR Operations

Failover Runbook (Warm Standby)

  1. Detect: Cloud Monitoring alert fires (primary region health check failures)
  2. Verify: On-call SRE confirms the primary region is degraded
  3. Scale the DR region:
    # Scale up compute in DR region
    gcloud compute instance-groups managed resize app-mig-dr \
      --region=us-east1 \
      --size=10
    
    # Promote Cloud SQL replica to a standalone primary (breaks replication)
    gcloud sql instances promote-replica prod-db-replica
    
  4. Shift traffic: The global LB routes to the healthy DR region based on health checks
  5. Verify: Confirm the application is serving traffic from the DR region
  6. Communicate: Notify stakeholders of failover
  7. Post-incident: After the primary recovers, re-establish replication and plan failback

For Cloud SQL Enterprise Plus, a designated DR replica supports switchover (planned role reversal without rebuilding the old primary) instead of a one-way promote. For MySQL and PostgreSQL this needs a cross-region Enterprise Plus primary and replica pair, with the replica named in the primary's replication_cluster.failover_dr_replica_name. SQL Server uses a cascadable replica in another region instead (Terraform google_sql_database_instance docs, Switchover section, checked 2026-09-27).

Automated Failover (Active-Active)

For active-active, failover is automatic. The global load balancer stops sending traffic to an unhealthy region based on health check results. The key operational task is testing:

# Game day: drain the primary region's backend, watch traffic move, then restore
gcloud compute backend-services update-backend web-backend \
  --global \
  --instance-group=app-mig \
  --instance-group-region=us-central1 \
  --capacity-scaler=0.0

# ... verify traffic shifts to the secondary region in Cloud Monitoring ...

gcloud compute backend-services update-backend web-backend \
  --global \
  --instance-group=app-mig \
  --instance-group-region=us-central1 \
  --capacity-scaler=1.0

Monitoring and Observability

Centralized Logging Sink

# Create aggregated log sink at organization level
gcloud logging sinks create org-logs-sink \
  bigquery.googleapis.com/projects/shared-logging/datasets/org_logs \
  --organization=ORGANIZATION_ID \
  --include-children \
  --log-filter='severity>=WARNING OR logName:"cloudaudit.googleapis.com"'

# Grant the sink's writer identity access to the destination
WRITER=$(gcloud logging sinks describe org-logs-sink \
  --organization=ORGANIZATION_ID --format='value(writerIdentity)')
gcloud projects add-iam-policy-binding shared-logging \
  --member="$WRITER" \
  --role="roles/bigquery.dataEditor"
Alert Condition Severity
Region health check failures >50% backends unhealthy for 2 min Critical
Cross-region replication lag Lag > 60 seconds for 5 min Warning
Cloud NAT port exhaustion Allocated ports >80% for 5 min, or dropped-packet count rising Warning
IAM policy changes Any change to org-level IAM Info
VPC SC violations Any VPC_SERVICE_CONTROLS denial in audit logs for enforced perimeters Warning
Budget threshold Spend >80% of monthly budget Warning

Troubleshooting

Shared VPC: Service Project Cannot Use Subnet

Check that the principal (and the relevant service agents) has roles/compute.networkUser on the specific subnet:

gcloud compute networks subnets get-iam-policy app-subnet \
  --project=shared-net-prod \
  --region=us-central1

VPC Peering: Routes Not Propagating

Verify both sides have matching import/export settings:

gcloud compute networks peerings list --network=network-a
# Check: exportCustomRoutes=true, importCustomRoutes=true on both sides

gcloud compute networks peerings list-routes peer-a-to-b \
  --network=network-a --region=us-central1 --direction=INCOMING

Cloud NAT: Connection Failures

Check NAT port allocation per VM:

gcloud compute routers get-nat-mapping-info nat-router --region=us-central1

If ports run out, raise --min-ports-per-vm or enable dynamic port allocation (--enable-dynamic-port-allocation) on the NAT gateway.

Firewall: Traffic Blocked Unexpectedly

List the effective rules (hierarchical, VPC, and network policy) that apply to one VM NIC:

gcloud compute instances network-interfaces get-effective-firewalls my-instance \
  --zone=us-central1-a

VPC Service Controls: 403 "Request is prohibited by organization's policy"

Find the denial in audit logs by the unique ID in the error message, then add an ingress or egress rule (or fix the access level):

gcloud logging read \
  'protoPayload.metadata.@type="type.googleapis.com/google.cloud.audit.VpcServiceControlAuditMetadata"' \
  --project=PROJECT_ID --limit=10 --freshness=1d

DR: Cloud SQL Replica Promote Fails

Make sure the replica is healthy before promoting:

gcloud sql instances describe prod-db-replica --format="value(state)"
# Must be RUNNABLE before promote

Sources