Skip to content

GCP Landing Zone: Architecture and Design

This page explains how each major Google Cloud project setup pattern works, why you would pick it, and how the security layers fit together. Look-up tables (constraints, IP ranges, prices, versions, checklists) live in the Reference. Commands live in the How-to Guides.

Landing Zone Architecture Overview

A landing zone combines four control planes: the resource hierarchy (where policy and IAM inherit), identity (Cloud Identity groups, service accounts, federation), network (Shared VPC, firewall policies, hybrid links), and data perimeter (VPC Service Controls). The diagram shows how they connect in a typical FAST-style or enterprise-foundations layout.

graph TD
    subgraph Identity["Cloud Identity / Workspace"]
        Groups["Google Groups<br/>gcp-org-admins, gcp-network-admins"]
        IdP["External IdP via<br/>Workforce Identity Federation"]
    end

    subgraph Hierarchy["Resource Manager hierarchy"]
        Org["Organization node<br/>org policies + hierarchical firewall policy"]
        FCommon["Folder: Common / Shared"]
        FProd["Folder: Production"]
        FNonProd["Folder: Non-Production"]
        Org --> FCommon
        Org --> FProd
        Org --> FNonProd
    end

    subgraph Shared["Shared projects"]
        NetHost["Shared VPC host project<br/>VPC, Cloud Router, Cloud NAT, Interconnect"]
        SecProj["Security project<br/>Cloud KMS, SCC, Secret Manager"]
        LogProj["Logging project<br/>org log sink, BigQuery, log buckets"]
    end

    FCommon --> NetHost
    FCommon --> SecProj
    FCommon --> LogProj
    FProd --> SvcProj["Service projects<br/>GKE, Cloud Run, Cloud SQL"]
    SvcProj -->|"compute.networkUser on subnets"| NetHost
    Groups -->|"IAM bindings at folder level"| Hierarchy
    IdP --> Groups
    Perimeter["VPC Service Controls perimeter"] -.->|"protects APIs of"| SvcProj
    Org -->|"aggregated sink"| LogProj

1. Single Project with Single VPC

Architecture Summary

The simplest GCP deployment: all resources live in one project with one VPC network. The VPC is global; subnets are regional (each subnet has a primary range and optional secondary ranges, for example for GKE Pods and Services). All internal communication uses internal IP addresses.

The diagram shows the single-project layout with private egress and private API access.

graph TD
    Internet["Internet"] --> LB["External Application Load Balancer"]
    LB --> Sub1

    subgraph VPC["VPC network (custom mode)"]
        subgraph Region1["Region: us-central1"]
            Sub1["Subnet 10.0.1.0/24<br/>App tier"]
            Sub2["Subnet 10.0.2.0/24<br/>Data tier"]
        end
        Sub1 --> Sub2
    end

    Sub1 --> NAT["Cloud NAT"]
    Sub1 --> PGA["Private Google Access"]
    PGA --> APIs["Google APIs"]

Key Services

  • VPC (custom mode) -- define your own subnet CIDR ranges
  • Cloud NAT -- outbound internet for private VMs
  • Private Google Access -- reach Google APIs without external IPs
  • Cloud NGFW -- VPC firewall rules or a global network firewall policy with secure tags
  • Cloud Load Balancing -- distribute incoming traffic
  • Use custom-mode VPC (not auto-mode) for explicit subnet control
  • Separate subnets per tier (app, data, management) within each region
  • Enable VPC Flow Logs for all subnets
  • Use Regional Managed Instance Groups for cross-zone HA
  • Target firewall rules by identity (secure tags or service accounts), not IP ranges. Secure tags are IAM-governed. Network tags are not

Real-World Example

A startup running a Django web application with a Cloud SQL PostgreSQL backend. Everything lives in myapp-prod. Two subnets in us-central1: one for Compute Engine MIGs serving traffic, one reserved for management. Cloud SQL uses a private IP through private services access. Cloud NAT handles package updates. A global external Application Load Balancer serves traffic from the MIG backends.

2. Multi-VPC Architecture

Choosing a Connectivity Model

Shared VPC, VPC Network Peering, Network Connectivity Center (NCC), and Private Service Connect (PSC) solve different problems. The decision flowchart shows which one fits.

flowchart TD
    Start["Need connectivity across projects or VPCs"] --> SameOrg{"Same organization and<br/>central network team?"}
    SameOrg -->|Yes| SVPC["Shared VPC<br/>host project owns network,<br/>service projects consume subnets"]
    SameOrg -->|No| Scope{"Expose one service<br/>or a whole network?"}
    Scope -->|"One service"| PSC["Private Service Connect<br/>endpoint to service attachment"]
    Scope -->|"Whole network"| Count{"More than a few VPCs<br/>or need transitivity?"}
    Count -->|No| Peer["VPC Network Peering<br/>non-transitive, no overlap"]
    Count -->|Yes| NCC["NCC hub with VPC spokes<br/>mesh or star topology"]
    SVPC --> Multi{"Several Shared VPCs<br/>per environment?"}
    Multi -->|Yes| NCC

2A. Shared VPC

Architecture Summary

A centralized networking model where a host project owns the VPC networks, subnets, firewall policies, and routes. Multiple service projects attach to the host and deploy compute resources (VMs, GKE clusters, Cloud Run with Direct VPC egress) into shared subnets. Network administration is separated from service administration.

The diagram shows one host project serving two service projects with subnet-level IAM.

graph TD
    Org["Organization"] --> HostProj["Host project: shared-net-prod"]
    Org --> SP1["Service project: payments-team"]
    Org --> SP2["Service project: orders-team"]

    HostProj --> VPC["Shared VPC network"]
    VPC --> Sub1["Subnet 10.0.1.0/24<br/>us-central1"]
    VPC --> Sub2["Subnet 10.0.2.0/24<br/>us-east1"]

    SP1 -->|"roles/compute.networkUser"| Sub1
    SP2 -->|"roles/compute.networkUser"| Sub2

    VPC --> VPN["Cloud VPN / Cloud Interconnect"]
    VPN --> OnPrem["On-premises network"]

Key Services

  • Shared VPC (roles/compute.xpnAdmin required to enable hosts and attach projects)
  • IAM -- roles/compute.networkUser for subnet-level delegation
  • Cloud DNS -- private zones in the host project, visible to service projects
  • Cloud Interconnect / Cloud VPN -- centralized hybrid connectivity in the host project
  • One host project per environment (prod, non-prod) to isolate network policies
  • Grant compute.networkUser at subnet level (not project level) for least privilege
  • Use organization policy constraints:
    • constraints/compute.restrictSharedVpcHostProjects -- limit which projects can be hosts
    • constraints/compute.restrictSharedVpcSubnetworks -- limit which subnets service projects can use
  • Keep Cloud VPN / Interconnect attachments in the host project for centralized routing
  • Use Cloud Router with custom advertisement mode to advertise subnet ranges to on-premises
  • Firewall data-processing charges for Cloud NGFW Standard/Enterprise rules on a Shared VPC are billed to the host project (pricing)

Real-World Example

A financial services company with 8 application teams. The networking-host-prod project owns the production VPC. Each team has its own service project (payments-prod, trading-prod, and others). The networking team controls firewall policies and routing centrally. Compute billing is attributed to each team's service project.

2B. VPC Network Peering

Architecture Summary

A peer-to-peer connectivity model where two VPC networks exchange subnet routes to enable internal IP communication. Peered networks remain administratively separate. Peering is not transitive -- if A peers with B and B peers with C, A cannot reach C through B.

The diagram shows why a "transit" VPC does not give spoke-to-spoke reachability.

graph LR
    VPCA["VPC A: 10.0.0.0/16"] <-->|Peering| VPCB["VPC B: 10.1.0.0/16"]
    VPCB <-->|Peering| VPCC["VPC C: 10.2.0.0/16"]
    VPCB -->|"Export custom routes"| OnPrem["On-premises via Cloud VPN in B"]

    noteA["VPC A cannot reach VPC C<br/>peering is not transitive"]

Key Services

  • VPC Network Peering -- managed by roles/compute.networkAdmin on each side
  • Cloud Router -- for exporting/importing custom (dynamic) routes through peering
  • Plan CIDR ranges carefully -- subnet ranges cannot overlap across peered VPCs
  • Enable --export-custom-routes and --import-custom-routes for on-premises transit scenarios (spokes reach on-prem through the hub's VPN/Interconnect)
  • A central "transit" VPC peered with every spoke gives spoke-to-hub and spoke-to-on-prem reachability, but not spoke-to-spoke. For that, use an NCC hub with VPC spokes (mesh or star topology), which provides transitive connectivity between spokes (NCC VPC spokes)
  • Firewall rules, network tags, and service accounts do not cross peering boundaries
  • Use Cloud DNS peering zones, or authorize the same private zone to all peered VPCs, for DNS resolution

Real-World Example

A SaaS provider offers a managed analytics platform. Their VPC (saas-prod) peers with each customer's VPC. Customers can access the analytics service using internal IP addresses. Each customer's VPC remains isolated from other customers. Many providers now publish through PSC instead, because peering needs non-overlapping CIDRs and counts against per-network peering limits.

2C. Private Service Connect

Architecture Summary

A service-oriented connectivity model where consumers access managed services (Google APIs, third-party SaaS, or internal services) using internal IP addresses in their own VPC. Unlike peering, PSC provides unidirectional, service-level access without exposing the full VPC. Producer traffic is NATed, so consumer and producer CIDRs may overlap.

The diagram shows a PSC endpoint to a published service and a global PSC endpoint for Google APIs.

graph TD
    ConsumerVPC["Consumer VPC"] -->|"PSC endpoint<br/>regional internal IP"| SA["Service attachment"]
    SA --> ProducerLB["Producer internal load balancer"]
    ProducerLB --> ProducerVPC["Producer VPC: managed service"]

    ConsumerVPC -->|"PSC endpoint, global<br/>bundle all-apis or vpc-sc"| GoogleAPIs["Google APIs<br/>Cloud Storage, BigQuery"]

Key Services

  • Private Service Connect endpoints -- forwarding rules mapping internal IPs to service attachments or to Google API bundles (all-apis or vpc-sc)
  • Private Service Connect backends -- NEGs behind consumer load balancers for advanced traffic management
  • Private Service Connect interfaces -- producer-initiated connections into a consumer network
  • Use PSC endpoints for Google API access instead of the default Private Google Access VIPs when you need a custom internal IP or per-service DNS control
  • Use the vpc-sc bundle (equivalent to restricted.googleapis.com) inside VPC Service Controls perimeters
  • Use PSC backends when you need custom URLs, TLS certificates, or failover between regional service endpoints
  • Service producers use consumer accept lists on the service attachment to control which projects can connect
  • Combine PSC with VPC Service Controls for defense in depth

Real-World Example

A company consumes a third-party data enrichment API hosted on GCP. Instead of routing traffic over the internet, they create a PSC endpoint in their VPC pointing to the vendor's service attachment. All traffic stays within Google's network. They also use a PSC endpoint for Google APIs to reach Cloud Storage and BigQuery via an internal IP.

3. Multi-Project Strategy

Architecture Summary

The GCP resource hierarchy (Organization > Folders > Projects > resources) provides policy inheritance and delegation. IAM allow policies are additive down the tree. Organization policies inherit by default, and a child can merge with or override the parent policy unless the parent prevents it. A well-structured hierarchy is the foundation of enterprise cloud governance.

The diagram shows an environment-first hierarchy with a shared-infrastructure folder.

graph TD
    Org["Organization: example.com"] --> FProd["Folder: Production"]
    Org --> FNonProd["Folder: Non-Production"]
    Org --> FShared["Folder: Shared-Infrastructure"]
    Org --> FSandbox["Folder: Sandboxes"]

    FProd --> FTeamA["Folder: Team-A"]
    FProd --> FTeamB["Folder: Team-B"]
    FTeamA --> ProdProjA["Project: prod-team-a-app"]
    FTeamB --> ProdProjB["Project: prod-team-b-app"]

    FNonProd --> FDev["Folder: Development"]
    FNonProd --> FStaging["Folder: Staging"]
    FDev --> DevProj["Project: dev-team-a"]

    FShared --> NetProj["Project: shared-networking"]
    FShared --> SecProj["Project: shared-security"]
    FShared --> LogProj["Project: shared-logging"]

Key Services

  • Resource Manager -- manages the org/folder/project hierarchy and tags
  • Organization Policy Service -- constraints applied at org, folder, or project level
  • IAM -- roles granted at any hierarchy level, inherited downward. Deny policies and principal access boundary policies add guardrails
  • Cloud Billing -- billing accounts linked to projects. Labels for cost attribution

How Organization Policies Are Evaluated

  • A policy set on a node applies to that node and all descendants unless a descendant sets its own policy.
  • For list constraints, a child can merge with the parent (inheritFromParent: true) or replace it. Boolean constraints are simply enforced or not.
  • Conditions based on tags let one policy enforce differently for tagged folders or projects (for example, allow external IPs only on projects tagged public-edge).
  • Dry-run specs and Policy Simulator show what a policy would block before you enforce it. Managed constraints (compute.managed.*, iam.managed.*) are designed around this safe-rollout flow.
  • Organizations created on or after 2024-05-03 start with security baseline constraints already enforced. See Baseline Organization Policy Constraints.

Hierarchy Design: - Top-level folders by environment (Production, Non-Production, Shared-Infrastructure) or by business unit, then environment - Second-level folders by team or business unit - Keep depth to 3-4 levels in practice (the hard limit is 10) so inheritance stays understandable - Manage the hierarchy with Terraform (Cloud Foundation Fabric or the Cloud Foundation Toolkit)

Essential Organization Policies: the baseline constraint list, including managed equivalents, is in Baseline Organization Policy Constraints.

IAM Best Practices: - Grant roles at the highest appropriate level (org > folder > project) - Use Google Groups (not individual users) for role bindings - Use custom roles for least privilege when predefined roles are too broad - Separate duties: project creator is not project owner is not billing admin - Use Privileged Access Manager for just-in-time elevation instead of standing admin grants

Real-World Example

A global retailer adopts GCP. They create top-level folders for Production, Non-Production, and Platform. The Platform folder contains shared networking (Shared VPC host project), security tooling (SCC, Cloud KMS), and centralized logging. Each business unit (ecommerce, supply chain, retail analytics) has a sub-folder under Production with its own projects. Org policies enforce gcp.resourceLocations to US/EU value groups and disable service account key creation.

4. Multi-Zone and Multi-Region Deployment

Architecture Summary

A region has three or more zones (most have 3, some have 4). Zonal resources (Compute Engine VMs, zonal Persistent Disks) live in one zone. Regional resources (Cloud SQL HA, regional MIGs, regional GKE clusters) span zones automatically. Multi-region resources (Spanner multi-region configurations, multi-region Cloud Storage) span regions.

The diagram shows a two-region deployment behind a global Application Load Balancer.

graph TD
    GLB["Global external Application Load Balancer"] --> MIG1
    GLB --> MIG2

    subgraph Region1["Region: us-central1"]
        MIG1["Regional MIG: 3 zones"]
        SQL1["Cloud SQL HA: regional"]
    end

    subgraph Region2["Region: europe-west1"]
        MIG2["Regional MIG: 3 zones"]
        SQL2["Cloud SQL cross-region replica"]
    end

    SQL1 -->|"Async cross-region replica"| SQL2
    GCS["Cloud Storage multi-region or<br/>dual-region bucket"] -.-> MIG1
    GCS -.-> MIG2

Key Services

  • Regional MIGs -- auto-distribute VMs across zones with autoscaling
  • Regional Persistent Disk / Hyperdisk Balanced High Availability -- synchronously replicate block data across two zones
  • GKE regional clusters -- control plane and nodes spread across zones
  • Global external Application Load Balancer -- L7 routing with health-check-driven failover
  • Global external proxy Network Load Balancer -- L4 TCP/SSL proxy for non-HTTP workloads. External passthrough Network Load Balancers are regional
  • Cloud DNS -- failover and geolocation routing policies with health checks

Data Layer Strategy: the per-service multi-zone and multi-region options are in Data Layer Strategy.

Compute Strategy: - Use regional MIGs, not zonal MIGs, for production workloads - GKE: use regional clusters. Fleet features (multi-cluster management, Config Sync, Policy Controller) are now included in GKE at no extra cost - Cloud Run: deploy to multiple regions. Use a global load balancer with serverless NEGs for routing - Configure health checks with aggressive thresholds for fast failover

Real-World Example

A media streaming service deploys in us-central1 (primary) and europe-west1 (secondary). Regional MIGs in each region run the API tier. A global external Application Load Balancer routes users to the nearest healthy region. Spanner in a multi-region configuration (for example nam6 in North America, Enterprise Plus edition) handles the global data layer. Cloud Storage uses dual-region buckets for media assets.

5. Disaster Recovery Patterns

5A. Pilot Light

Architecture Summary

A minimal footprint of critical infrastructure runs continuously in a secondary region. Only core components (database replicas, IaC definitions) are active. The application tier is provisioned on demand during a disaster.

The diagram shows a replicating database and an empty MIG waiting in the DR region.

graph LR
    subgraph Primary["Primary region: us-central1"]
        MIG_P["Regional MIG: 10 instances"]
        SQL_P["Cloud SQL primary"]
    end

    subgraph DR["DR region: us-east1"]
        SQL_R["Cloud SQL read replica - active"]
        IaC["Terraform config - stored"]
        MIG_D["MIG: 0 instances<br/>scaled up on failover"]
    end

    SQL_P -->|"Async replication"| SQL_R
    GLB["Global LB"] -->|Healthy| MIG_P
    GLB -.->|Failover| MIG_D

Key Services

  • Cloud SQL cross-region read replicas -- async replication, manual promote
  • Persistent Disk snapshots -- scheduled snapshots stored multi-regionally
  • Terraform / Cloud Build -- automated provisioning of the compute tier on failover
  • Global external Application Load Balancer -- health-check-driven traffic rerouting

Configuration

  • RPO: minutes to hours (depends on replication lag and snapshot schedule)
  • RTO: hours (includes provisioning time for the compute tier)
  • Cost: low -- only database replication and storage snapshots are continuously billed
  • Test failover at least quarterly

5B. Warm Standby

Architecture Summary

A fully running but scaled-down copy of production operates in the secondary region. During failover, resources scale to production capacity.

The diagram shows the scaled-down standby receiving no traffic until failover.

graph LR
    subgraph Primary["Primary region: us-central1"]
        MIG_P["MIG: 10 instances"]
        SQL_P["Cloud SQL primary"]
    end

    subgraph Warm["Warm standby: us-east1"]
        MIG_W["MIG: 2 instances - scaled down"]
        SQL_W["Cloud SQL read replica - active"]
    end

    SQL_P -->|"Async replication"| SQL_W
    GLB["Global LB"] -->|"100% traffic"| MIG_P
    GLB -.->|Failover| MIG_W

Key Services

  • Regional MIGs -- running at reduced capacity. Scale on failover trigger
  • Cloud SQL cross-region replicas -- continuously replicating (Enterprise Plus adds DR-replica switchover)
  • Global external Application Load Balancer -- automatic failover with health checks
  • Cloud Monitoring + Cloud Run functions -- automated scale-up triggers

Configuration

  • RPO: minutes (async replication lag)
  • RTO: minutes (resources already running, only need scale-up)
  • Cost: moderate -- reduced compute continuously plus full database replication
  • Automate scale-up via Cloud Monitoring alerting + Cloud Run functions (formerly Cloud Functions)

5C. Active-Active Multi-Region

Architecture Summary

Full production workloads run simultaneously in two or more regions. A global load balancer sends traffic to the nearest healthy region. Data is replicated synchronously or asynchronously depending on the service.

The diagram shows two live regions sharing one Spanner multi-region instance.

graph TD
    Users["Users worldwide"] --> GLB["Global external Application LB"]
    GLB -->|"Nearest healthy region"| MIG1
    GLB -->|"Nearest healthy region"| MIG2

    subgraph Region1["us-central1"]
        MIG1["MIG: 10 instances"]
        Sp1["Spanner read-write replica"]
    end

    subgraph Region2["us-east1"]
        MIG2["MIG: 10 instances"]
        Sp2["Spanner read-write replica"]
    end

    Sp1 <-->|"Synchronous Paxos quorum<br/>witness in us-central2"| Sp2
    MIG1 --> Sp1
    MIG2 --> Sp2

Key Services

  • Global external Application Load Balancer -- routes to the nearest healthy region
  • Spanner -- synchronous multi-region replication (Enterprise Plus edition, 99.999% SLA)
  • Bigtable -- multi-cluster replication, eventually consistent across clusters
  • Cloud Storage -- dual-region or multi-region buckets
  • GKE fleets -- multi-cluster management and multi-cluster Gateway
  • Cloud DNS -- geolocation or failover routing policies

Configuration

  • RPO: near-zero for data in synchronously replicated stores (Spanner). Async stores still have lag
  • RTO: near-zero (traffic rerouted automatically by the global LB)
  • Cost: highest -- full duplicate infrastructure in multiple regions
  • Use Spanner for relational data that needs synchronous cross-region replication
  • Use committed use discounts for base capacity. Use Spot VMs for non-critical batch workloads

DR Pattern Comparison

The RPO/RTO/cost matrix for the three patterns is in DR Comparison Matrix. The core trade-off: each step up (pilot light, warm standby, active-active) buys lower RTO by paying for more idle or duplicated capacity, and active-active also needs a data layer that tolerates multi-region writes.

6. DMZ / Network Perimeter Patterns

Architecture Summary

GCP does not need a traditional DMZ subnet. External Application and proxy Network Load Balancers are Google-managed front ends that terminate outside your VPC. Backends live in private subnets with no external IPs. This "DMZ-less" design is the Google-recommended pattern.

The diagram shows the DMZ-less layout with managed ingress, IAP admin access, and managed egress.

graph TD
    Internet["Internet"] --> Armor["Cloud Armor policy"]
    Armor --> ExtLB["External Application LB<br/>Google-managed, outside VPC"]
    Internet --> IAP["Identity-Aware Proxy"]

    ExtLB --> FW1["Firewall: allow health-check ranges"]
    IAP --> FW2["Firewall: allow IAP range<br/>35.235.240.0/20"]

    subgraph VPC["VPC network"]
        FW1 --> AppSubnet["Private subnet: app tier<br/>no external IPs"]
        FW2 --> AppSubnet
        AppSubnet --> FW3["Firewall: allow from app tier only"]
        FW3 --> DataSubnet["Private subnet: data tier<br/>no external IPs"]
    end

    AppSubnet --> NAT["Cloud NAT<br/>outbound egress"]
    AppSubnet --> PGA["Private Google Access<br/>Google API access"]

Key Services

  • External load balancers (Application LB, proxy Network LB, passthrough Network LB) -- distribute to private backends
  • Cloud NAT -- outbound internet for private VMs (no inbound)
  • Identity-Aware Proxy (IAP) -- zero-trust SSH/RDP and web app access, replaces bastion hosts
  • Private Google Access -- reach Google APIs without external IPs
  • Cloud NGFW -- tier-based access control with secure tags
  • Cloud Armor -- WAF and DDoS protection on external LBs
  • Place all workloads in private subnets (no external IPs)
  • Use external LBs for all inbound traffic (no DMZ subnet needed)
  • Use Cloud NAT for all outbound traffic (updates, external API calls)
  • Use IAP tunneling for SSH/RDP access (no bastion host required)
  • Use Private Google Access to reach Google APIs from private subnets
  • Allow only LB health-check ranges and the IAP range (35.235.240.0/20) through the firewall
  • Attach Cloud Armor policies to external LBs for WAF protection

7. GCP Landing Zone

Architecture Summary

A landing zone is the foundation that sets secure, scalable, well-governed baselines for an entire GCP organization. It covers the resource hierarchy, IAM, networking, security, logging, and billing. See Landing Zone Architecture Overview for the component diagram.

The diagram summarizes the landing-zone building blocks.

graph TD
    Org["Organization"] --> LZ["Landing zone components"]

    LZ --> Hierarchy["Resource hierarchy<br/>Org, Folders, Projects"]
    LZ --> IAM["IAM and groups<br/>least privilege via groups"]
    LZ --> Network["Networking<br/>Shared VPC, Cloud NAT, hybrid"]
    LZ --> Security["Security<br/>SCC, VPC SC, Binary Authorization"]
    LZ --> Logging["Logging and monitoring<br/>aggregated sinks, dashboards"]
    LZ --> Billing["Billing<br/>budgets, labels, BigQuery export"]
    LZ --> IaC["Infrastructure as code<br/>Terraform: Fabric FAST or CFT"]

Key Components

Component Service Purpose
Resource Hierarchy Resource Manager Org > Folders > Projects for policy inheritance
IAM Cloud IAM + Groups Role-based access via groups, custom roles for least privilege
Networking Shared VPC + Cloud NAT Centralized networking with private egress
Hybrid Connectivity Cloud Interconnect / Cloud VPN / NCC On-premises and multi-VPC connectivity
Firewalling Cloud NGFW hierarchical and network firewall policies Org-wide baseline rules plus per-network rules with secure tags
Security Posture Security Command Center (Standard or Premium) Misconfiguration and threat detection
Data Protection VPC Service Controls API perimeter security, data exfiltration prevention
Container Security Binary Authorization Enforce signed container images in GKE and Cloud Run
Web Security Cloud Armor DDoS and WAF for external-facing services
Logging Cloud Logging Aggregated log sinks to a central project
Monitoring Cloud Monitoring Dashboards, alerting policies
Cost Management Cloud Billing + BigQuery Billing export, budgets, label-based attribution
Automation Terraform + Cloud Build (or GitHub/GitLab CI) IaC for all infrastructure, CI/CD pipelines
Access Identity-Aware Proxy Zero-trust access without VPN/bastion

Blueprint Options

Google offers three paths to a landing zone. They share the same concepts but differ in opinion and ownership. Versions are in Landing-Zone Blueprints and Tooling Versions.

Option Style Fit
Google Cloud Setup (console) Guided click-through, can export Terraform Small and mid-size orgs that want a sane default quickly
Enterprise foundations blueprint (terraform-example-foundation, CFT modules) Opinionated, staged (0-bootstrap to 5-app-infra), documented security controls Regulated enterprises that want a documented reference architecture
Cloud Foundation Fabric FAST Staged, YAML-driven factories (projects, folders, firewall, VPC SC), frequent releases Platform teams that want to fork and own their landing zone. Google PSO uses it widely

Community comparisons (for example meshcloud) describe FAST as the more modular, factory-driven option and the example foundation as the more prescriptive one. That is a characterization, not an official Google statement. Both remain maintained as of 2026-09.

The step-by-step order is in the how-to guide: Recommended Implementation Order.

Real-World Example

A regulated financial institution adopts GCP. They deploy the landing zone from open-source Terraform (Fabric FAST or the enterprise foundations blueprint). The hierarchy separates Production and Non-Production at the top level. Shared VPC host projects provide networking. VPC Service Controls protect BigQuery and Cloud Storage perimeters. Security Command Center Premium detects misconfigurations. All infrastructure is managed through Terraform in a GitOps pipeline. Project provisioning is self-service through a project factory that enforces naming, labeling, budgets, and org policy defaults.

Security Model Overview

Identity flow, threat model, access control, DMZ patterns, VPC Service Controls, and security practices for GCP project architectures. GCP security operates in layers:

  1. Identity and Access Management (IAM) -- who can do what on which resource (allow, deny, and principal access boundary policies)
  2. Organization Policies -- constraints on what configurations are allowed
  3. Network Security -- Cloud NGFW firewall policies, Cloud Armor, private connectivity
  4. Data Protection -- encryption at rest and in transit, VPC Service Controls perimeters
  5. Zero-Trust Access -- Identity-Aware Proxy (IAP), context-aware access (BeyondCorp model)

IAM Architecture

Role Hierarchy and Delegation

The diagram shows how Shared VPC administration is delegated from the organization down to service projects.

graph TD
    OrgAdmin["Organization Admin<br/>resourcemanager.organizationAdmin"] --> SharedVPCAdmin["Shared VPC Admin<br/>compute.xpnAdmin + projectIamAdmin"]
    OrgAdmin --> FolderAdmin["Folder Admin<br/>resourcemanager.folderAdmin"]

    SharedVPCAdmin --> NetAdmin["Network Admin<br/>compute.networkAdmin, host project"]
    SharedVPCAdmin --> SecAdmin["Security Admin<br/>compute.securityAdmin, host project"]
    SharedVPCAdmin --> SvcProjAdmin["Service Project Admin<br/>compute.networkUser on subnets"]

    SvcProjAdmin --> InstanceAdmin["Instance Admin<br/>compute.instanceAdmin, service project"]

IAM Best Practices

Practice Implementation
Grant roles to groups, not individuals Create Google Groups per team/function (for example, gcp-network-admins@example.com)
Grant at highest appropriate level Org-level for cross-cutting roles. Folder-level for team-scoped roles
Use custom roles for least privilege Narrow predefined roles when they grant too many permissions
Use deny policies for guardrails IAM deny policies are evaluated before allow policies (set at org/folder level)
Enforce MFA Required for all human principals via Google Workspace / Cloud Identity
Use Workload Identity Federation For external workloads (AWS, Azure, GitHub Actions, on-prem) instead of service account keys
Disable SA key creation Org policy: iam.managed.disableServiceAccountKeyCreation (enforced by default on orgs created since 2024-05-03)
Short-lived credentials Service account impersonation (IAM Service Account Credentials API) instead of keys
Just-in-time elevation Privileged Access Manager entitlements with approval and maximum grant duration (PAM)

Service Account Strategy

  • One service account per application/component (not shared)
  • Use Workload Identity Federation for GKE for pods (not service account keys)
  • Use constraints/iam.allowedPolicyMemberDomains to restrict IAM membership to your Cloud Identity customer IDs
  • Audit service account usage with Policy Analyzer and service account insights

Network Security

GCP's recommended security architecture removes the need for a traditional DMZ:

Traditional DMZ GCP DMZ-Less Equivalent
Public-facing subnet with bastion host Identity-Aware Proxy (IAP) for admin access
NAT gateway VMs in DMZ Cloud NAT (managed service, no VMs)
WAF appliances in DMZ Cloud Armor policies on external LB
API gateway in DMZ External Application LB + API Gateway / Apigee
Public subnet for load balancers External LBs are Google-managed (outside VPC)
Perimeter IPS appliances Cloud NGFW Enterprise (IDPS on zonal firewall endpoints)

VPC Firewall Strategy

Cloud NGFW evaluates several policy types in a fixed order: hierarchical policies first (organization, then folders), then VPC firewall rules and network firewall policies, then the implied rules. The flowchart shows the default AFTER_CLASSIC_FIREWALL order. Details are in Firewall Policy Evaluation Order.

flowchart TD
    Pkt["Packet to or from a VM NIC"] --> HOrg["Hierarchical policy at organization"]
    HOrg -->|"allow or deny"| Verdict["Verdict applied"]
    HOrg -->|"goto_next or no match"| HFolder["Hierarchical policies on folders<br/>top folder to bottom folder"]
    HFolder -->|"allow or deny"| Verdict
    HFolder -->|"goto_next or no match"| Classic["VPC firewall rules"]
    Classic -->|"match"| Verdict
    Classic -->|"no match"| GNet["Global network firewall policy"]
    GNet -->|"match"| Verdict
    GNet -->|"no match"| RNet["Regional network firewall policy"]
    RNet -->|"match"| Verdict
    RNet -->|"no match"| Implied["Implied rules<br/>deny ingress, allow egress"]
    Implied --> Verdict

The layers of a typical rule set, from broadest to narrowest:

graph TD
    subgraph Layers["Firewall layers"]
        FW1["1. Default deny all ingress<br/>implied rule, always present"]
        FW2["2. Allow health-check ranges<br/>35.191.0.0/16, 130.211.0.0/22"]
        FW3["3. Allow IAP range<br/>35.235.240.0/20"]
        FW4["4. Allow inter-tier traffic<br/>app tier to data tier only"]
        FW5["5. Deny all egress to internet<br/>explicit deny overrides implied allow"]
    end

Put organization-wide rules (for example, allow health checks and IAP, block known-bad ranges) in a hierarchical firewall policy so project owners cannot remove them. Put workload rules in a global network firewall policy that targets secure tags. Secure tags are IAM-governed, while any user who can edit an instance can change its network tags. The recommended rule set is in Recommended Firewall Rules.

VPC Service Controls

VPC Service Controls create security perimeters around Google-managed services (BigQuery, Cloud Storage, Pub/Sub, and others) to prevent data exfiltration. They work independently of, and in addition to, IAM: a request must pass both.

The diagram shows allowed and denied flows around one perimeter.

graph TD
    subgraph Perimeter["Service perimeter"]
        ProjA["Project A<br/>BigQuery dataset"]
        ProjB["Project B<br/>Cloud Storage bucket"]
        VPC1["VPC network<br/>VMs"]
    end

    VPC1 -->|Allowed| ProjA
    VPC1 -->|Allowed| ProjB
    ProjA -->|Allowed| ProjB

    ExtUser["External user<br/>outside perimeter"] -->|Denied| ProjA
    VPC1 -->|Denied| ExtBucket["External bucket<br/>outside perimeter"]

The sequence diagram shows, in simplified form, how one API call is evaluated when VPC Service Controls and IAM both apply. Both checks must pass. The internal order is not a documented contract.

sequenceDiagram
    participant C as Caller (VM or user)
    participant GFE as Google API front end
    participant VPCSC as VPC Service Controls
    participant ACM as Access Context Manager
    participant IAM as IAM
    participant SVC as BigQuery API

    C->>GFE: bigquery.jobs.insert on project in perimeter
    GFE->>VPCSC: Is caller inside perimeter, or matched by an ingress rule?
    VPCSC->>ACM: Evaluate access levels (IP, device, identity)
    ACM-->>VPCSC: Access level satisfied or not
    alt Perimeter check fails
        VPCSC-->>C: 403 Request is prohibited by organization's policy
    else Perimeter check passes
        VPCSC->>IAM: Check allow and deny policies
        IAM-->>SVC: Authorized
        SVC-->>C: Job created
    end

The concept glossary (perimeter, ingress/egress rules, access levels, bridges, restricted VIP, dry run) is in VPC Service Controls Key Concepts. Setup commands are in the How-to Guides.

Cloud Armor

Cloud Armor attaches security policies to the backend services of external load balancers. Preconfigured WAF rules are built from the OWASP ModSecurity Core Rule Set. Google recommends the CRS 4.22 rule sets (sqli-v422-stable, xss-v422-stable, and others) over the older 3.3 sets (sqli-v33-stable) (WAF rules). Minimum TLS version is not a Cloud Armor setting: it is set with an SSL policy on the target proxy. Commands are in Cloud Armor WAF Policy.

Identity-Aware Proxy (IAP)

IAP for SSH Access (Replaces Bastion Hosts)

IAP TCP forwarding wraps SSH/RDP in an HTTPS tunnel from the client to Google, authenticated by IAM. Traffic enters the VPC from 35.235.240.0/20, so the VM needs no external IP and only one firewall rule. Commands are in IAP TCP Forwarding for SSH.

IAP for Web Application Authentication

The diagram shows the IAP web flow: IAP authenticates at the load balancer, then the backend verifies the signed header.

graph LR
    User["User browser"] --> IAP["Identity-Aware Proxy on the LB"]
    IAP -->|"Verify identity + access level"| IAM["Cloud IAM"]
    IAM -->|Authorized| App["Application backend"]
    IAM -->|Denied| Reject["403 Forbidden"]

    App -->|"Validate IAP JWT"| Verify["Verify signed header<br/>X-Goog-Iap-Jwt-Assertion"]

The application backend validates the IAP-signed JWT header (X-Goog-Iap-Jwt-Assertion) to confirm the request passed through IAP. This gives zero-trust authentication without managing an identity provider in the application.

Threat Model

Threat Categories and Mitigations

Threat GCP Mitigation Configuration
Stolen credentials MFA, IAP, VPC Service Controls Enforce MFA org-wide. Restrict API access to authorized networks via access levels
Data exfiltration via misconfigured IAM VPC Service Controls Block data copy operations that cross perimeter boundaries
Data exfiltration via compromised VM VPC Service Controls + egress firewall + Cloud NAT logging VMs cannot reach APIs outside the perimeter. Deny internet egress by default
External IP on private workload Org policy + firewall compute.managed.vmExternalIpAccess or constraints/compute.vmExternalIpAccess denies external IPs
SQL injection / XSS Cloud Armor WAF rules on external LBs
Insider threat IAM deny policies, PAM, audit logs Centralized logging. Just-in-time elevation. Deny policies for sensitive operations
Service account key leakage Org policy iam.managed.disableServiceAccountKeyCreation. Use Workload Identity Federation
Container image tampering Binary Authorization Enforce attested images in GKE and Cloud Run
On-premises network breach VPC Service Controls + context-aware access Extend perimeter via VPN/Interconnect. Restrict by source IP
Lateral movement inside a VPC Cloud NGFW with secure tags, Cloud NGFW Enterprise IDPS Micro-segment by tag. Inspect east-west traffic where needed
DNS exfiltration Cloud DNS logging + monitoring Log all DNS queries. Alert on suspicious patterns

Encryption

Layer GCP Default Recommended Enhancement
At rest Google-managed encryption (AES-256) Customer-managed encryption keys (CMEK) via Cloud KMS for regulated data
In transit within Google Encryption between data centers Use Private Service Connect or private connectivity to avoid internet transit
In transit to/from internet TLS on external LBs Enforce TLS 1.2+ with an SSL policy on the LB target proxy
Application-level N/A Application-layer encryption for highly sensitive fields

Security Command Center

Security Command Center aggregates misconfiguration findings (Security Health Analytics), threat detections (Event, Container, and VM Threat Detection), and asset inventory. The Standard tier is free. Premium adds threat detection and compliance reporting. The Enterprise tier is deprecated and shuts down on or after 2027-05-21. The feature table and tier notes are in Security Command Center. The hardening checklist is in Security Baseline Checklist.

Sources