Skip to content

Alibaba Cloud (Aliyun) Architecture Patterns

What this page explains

How Alibaba Cloud's building blocks fit together and why you would pick one pattern over another. It covers single-VPC, multi-VPC (CEN/Transit Router, VPC peering, Express Connect), multi-account (Resource Directory), multi-zone and multi-region, DR, DMZ, and landing zone patterns. The last section covers the security model (RAM, STS, SSO, KMS, Cloud Firewall, China compliance). Region tables, service mappings, and limits are in Reference. Runnable commands are in How-to Guides.

Platform Architecture Overview

Alibaba Cloud has three layers. Resource Directory and Agentic Cloud Governance Center govern accounts at the top. Regions contain zones, and each zone contains VPCs and vSwitches. Cloud Enterprise Network (CEN) and Express Connect stitch regions and on-premises sites together. The diagram shows how these layers relate in a typical enterprise tenancy.

flowchart TB
    subgraph GOV["Governance plane (management account)"]
        RD["Resource Directory<br/>Root, folders, members"]
        CGC["Agentic Cloud Governance Center<br/>landing zone, account factory"]
        CP["Control Policies"]
        SSO["CloudSSO / RAM SAML"]
        CGC --> RD
        RD --> CP
        SSO -.->|"role access"| RD
    end
    subgraph NET["Network account"]
        CEN["CEN instance"]
        TRA["Transit Router<br/>ap-southeast-5 (Jakarta)"]
        TRB["Transit Router<br/>ap-southeast-1 (Singapore)"]
        CEN --- TRA
        CEN --- TRB
        TRA <-->|"inter-region attachment"| TRB
    end
    subgraph WL["Workload member account"]
        VPC["VPC<br/>vSwitch zone A + zone B"]
        ECS["ECS / ACK node pools"]
        DB["PolarDB / RDS"]
        VPC --> ECS
        VPC --> DB
    end
    subgraph ONP["On-premises DC"]
        CPE["Customer router"]
    end
    RD -->|"member of folder"| WL
    RD -->|"member of folder"| NET
    TRA -->|"VPC attachment"| VPC
    CPE -->|"Express Connect circuit"| VBR["VBR"]
    VBR --> ECR["Express Connect Router"]
    ECR --> TRA

1. Single Project with Single VPC

Architecture Summary

A single Virtual Private Cloud (VPC) is the basic building block on Alibaba Cloud. A VPC is an isolated virtual network. You define a private CIDR block, create vSwitches (zonal subnets) across availability zones, and attach route tables and gateways. For one project, a single VPC with multi-AZ vSwitches gives isolation, high availability, and simplicity.

The flowchart shows the canonical three-tier layout inside one VPC.

flowchart TB
    INET["Internet"] --> WAF["WAF (CNAME / cloud-native mode)"]
    WAF --> ALB["ALB in public vSwitches<br/>zone A + zone B"]
    subgraph VPC["VPC 10.0.0.0/16"]
        ALB --> ECSA["ECS app<br/>private vSwitch, zone A"]
        ALB --> ECSB["ECS app<br/>private vSwitch, zone B"]
        ECSA --> RDSP["RDS / PolarDB primary<br/>data vSwitch, zone A"]
        ECSB --> RDSP
        RDSP -.->|"sync replica"| RDSS["Standby<br/>zone B"]
        ECSA --> NAT["Enhanced NAT Gateway + EIP"]
        ECSB --> NAT
    end
    NAT --> INET

Key Services

Service Role
VPC Isolated virtual network with custom CIDR, vSwitches, route tables
vSwitch Subnet within a VPC, bound to a single AZ. Resources attach here
Route Table System route table (auto-created) + custom route tables. Controls traffic forwarding
ALB / NLB / CLB The Server Load Balancer (SLB) family. ALB is Layer 7, NLB is Layer 4, CLB is the legacy Layer 4/7 product. All distribute traffic across AZs
NAT Gateway Outbound Internet (SNAT) and inbound port forwarding (DNAT) for private resources. New gateways are always the Enhanced type
Security Group Stateful per-instance firewall, like AWS security groups
Network ACL Stateless subnet-level packet filter, like AWS NACLs
EIP (Elastic IP) Public IP that can be bound to a NAT Gateway, load balancer, or individual ECS instance
  • CIDR planning: Use RFC 1918 ranges (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16). Reserve headroom. A VPC CIDR must not overlap with any VPC you plan to peer or attach to CEN later. Start with /16 for the VPC and /24 for vSwitches.
  • Multi-AZ vSwitches: Create at least two vSwitches in different AZs. Place load balancer, ECS, and database replicas across both.
  • Subnet tiers: Split into public (DMZ), private (application), and data (database) vSwitches. Use route tables to enforce tiered traffic flow.
  • No public IPs on app/DB: Application and database instances must have no EIPs. Use NAT Gateway (SNAT) for outbound Internet and a load balancer for inbound.
  • Security Groups: One SG per tier (web, app, db), with least-privilege rules. Allow web SG -> app SG -> db SG only on required ports.

Real-World Example

A SaaS startup deploys a three-tier web application in the Indonesia (Jakarta, ap-southeast-5) region. One VPC (10.0.0.0/16) has vSwitches in two zones. Public vSwitches host an ALB instance. Private vSwitches host ECS auto-scaling groups. A managed RDS MySQL instance runs in high-availability (multi-zone) mode. An Enhanced NAT Gateway provides egress. Cloud Firewall is enabled on the Internet border.


2. Multi-VPC Architecture

Workloads often span several VPCs: for environment isolation (dev/staging/prod), business-unit separation, or regulatory compliance. Alibaba Cloud offers three primary inter-VPC connectivity options.

2a. Cloud Enterprise Network (CEN) -- Hub-and-Spoke

Architecture Summary

CEN is Alibaba Cloud's managed WAN service. A Transit Router (TR) in each region acts as the hub. VPCs, VPN connections, VBRs (Express Connect), Express Connect Routers (ECR), and Cloud Connect Network (CCN) instances attach as spokes. Traffic flows through the TR's route tables, which gives fine-grained routing policy, traffic isolation, and centralized egress.

TRs in different regions connect through inter-region (peer) attachments over Alibaba's private backbone. Each inter-region connection draws on a purchased bandwidth plan (BandwidthType = BandwidthPackage) or is billed pay-by-data-transfer (DataTransfer). The pre-2026-09 version of this note said an Enterprise Edition TR can connect up to 1,000 VPCs per region. Alibaba confirms 1,000 VPCs per Enterprise Edition transit router, up from 200 (product news, checked 2026-09-27). Check the CEN SLA for current backbone commitments.

The diagram shows a two-region hub-and-spoke design with a shared-services hub VPC and an on-premises link.

flowchart LR
    subgraph CEN["CEN instance (network account)"]
        subgraph RA["Region: ap-southeast-5 (Jakarta)"]
            TRA["Transit Router A<br/>route tables: spoke-rt, hub-rt"]
            HUB["Hub VPC<br/>Cloud Firewall, NAT, DNS"]
            P1["Prod VPC"]
            D1["Dev VPC"]
            HUB --- TRA
            P1 --- TRA
            D1 --- TRA
        end
        subgraph RB["Region: ap-southeast-1 (Singapore)"]
            TRB["Transit Router B"]
            P2["Prod VPC (DR)"]
            P2 --- TRB
        end
        TRA <-->|"inter-region attachment<br/>bandwidth plan or pay-by-data-transfer"| TRB
    end
    DC["On-prem DC"] -->|"Express Connect"| VBR["VBR"]
    VBR --- TRA

Key Services

Service Role
CEN Instance Global container for Transit Routers. One CEN holds TRs in many regions
Transit Router (TR) Regional hub. The Enterprise Edition supports custom route tables, route maps, multicast, and traffic steering
TR Route Table System or custom. Controls inter-VPC routing, isolation, and traffic steering through association and propagation
Bandwidth Plan Pre-purchased cross-region bandwidth (the alternative is pay-by-data-transfer)
Network Instance Connection Attachment types: VPC, VBR, VPN, CCN, ECR, and inter-region (peer) connections
  • Use Enterprise Edition TRs. Custom route tables, traffic isolation, and service chaining need them.
  • Hub VPC pattern: Deploy a centralized shared-services VPC (DNS, NTP, security appliances, NAT egress) attached to the TR. Use custom route tables to let spoke VPCs reach the hub but not each other.
  • Centralized DMZ egress: Route all Internet-bound traffic from spoke VPCs through a hub VPC with Cloud Firewall and NAT Gateway for unified inspection.
  • Bandwidth planning: Choose bandwidth plans for steady, predictable inter-region traffic and pay-by-data-transfer for bursty or low-volume traffic. Set per-connection bandwidth limits so one spoke cannot saturate the link.
  • Route map policies: Use TR route maps for path selection, route manipulation, and traffic engineering.

Real-World Example

A financial services company has separate VPCs for trading, risk analytics, and back-office in China (Shanghai) and China (Beijing). A CEN with Enterprise TRs in both regions connects them. A hub security VPC in Shanghai runs Cloud Firewall and IDS. Custom route tables send trading VPC traffic through the security VPC before it reaches the Internet, while back-office VPCs stay isolated from trading. The Shanghai-Beijing connection uses a 2 Gbit/s bandwidth plan.


2b. VPC Peering Connection

Architecture Summary

VPC Peering is a direct, one-to-one connection between two VPCs. It supports same-region and cross-region peering, and same-account and cross-account peering. Unlike CEN, it is a point-to-point link without a central hub. You add routes that point at the peering connection to each VPC's route table.

flowchart LR
    A["VPC-A 10.0.0.0/16<br/>route 10.1.0.0/16 -> pcc-xxx"] <-->|"VPC peering connection pcc-xxx"| B["VPC-B 10.1.0.0/16<br/>route 10.0.0.0/16 -> pcc-xxx"]

Key Services

Service Role
VPC Peering Connection Direct link between two VPCs. Inter-region connections have a bandwidth setting and a Gold (default) or Platinum link type
VPC Route Table Route entries whose next hop is the peering connection, added on both sides
  • Make sure CIDR blocks do not overlap between peered VPCs.
  • Same-account peering is simplest. Cross-account peering needs the accepter account's ID and acceptance by that account.
  • For more than about 5-10 VPCs, prefer CEN over individual peering links to avoid a combinatorial explosion of peering connections and route entries. Peering is not transitive.
  • Inter-region peering is billed by data transfer. Check current pricing for same-region peering before relying on it being free.

When to Choose VPC Peering over CEN

  • Small number of VPCs (2-4).
  • Simple, static connectivity without complex routing policy.
  • Lower cost for low-volume traffic, since there are no TR attachment or processing fees.

2c. Express Connect (Leased Line)

Architecture Summary

Express Connect provides physically dedicated, private circuits between your data center and Alibaba Cloud. A Virtual Border Router (VBR) terminates the circuit at the Alibaba Cloud edge. An Express Connect Router (ECR) is a newer global router that exchanges routes with VBRs through BGP and forwards traffic to VPCs and TRs in many regions. It replaces a mesh of VBR-to-VPC connections.

The sequence shows how routes are learned and traffic flows from an on-premises router to a VPC through ECR.

sequenceDiagram
    participant CPE as On-prem router
    participant VBR as VBR (access point)
    participant ECR as Express Connect Router
    participant TR as Transit Router
    participant VPC as Workload VPC
    CPE->>VBR: BGP session over dedicated circuit, advertise 192.168.0.0/16
    VBR->>ECR: Propagate on-prem prefixes
    ECR->>TR: Associate and propagate routes
    TR->>VPC: Install route 192.168.0.0/16 via TR attachment
    VPC-->>CPE: Return traffic follows the same path, never the Internet

Dedicated circuits run at up to 100 Gbit/s per port. The Terraform provider's port_type values range from 100Base-T and 1000Base-LX through 10GBase-LR to 40GBase-LR and 100GBase-LR (the last two since provider 1.185.0, alicloud_express_connect_physical_connection, checked 2026-09-27). See also the interconnect table on the Multi-Cloud Governance reference.

Key Services

Service Role
Express Connect Circuit Dedicated physical line, or a hosted connection from a partner
VBR (Virtual Border Router) Router at the Alibaba Cloud edge for circuit termination and BGP
ECR (Express Connect Router) Global hybrid-cloud router with dynamic BGP (Terraform support since provider v1.224.0)
Hosted Connection Shared circuit provided by an Express Connect partner
  • Use it for hybrid cloud (DC-to-VPC) when you need guaranteed bandwidth.
  • For production, provision redundant circuits to different access points and use ECMP or BGP path preference for failover.
  • Enable BGP dynamic routing through VBR/ECR for automatic route exchange.
  • Express Connect vs. VPN: Express Connect offers higher bandwidth, lower and steadier latency, and physical isolation. Use IPsec-VPN as a backup path or for non-critical sites.
  • Do not use Express Connect for VPC-to-VPC links inside Alibaba Cloud. Use VPC peering or CEN instead.

Choosing a Connectivity Option

The full side-by-side matrix is in Reference: Connectivity Comparison Matrix. The decision usually reduces to the flow below.

flowchart TD
    Q1{"Connecting an on-prem<br/>site?"} -->|"yes, need SLA bandwidth"| EC["Express Connect + VBR<br/>(+ ECR for many regions)"]
    Q1 -->|"yes, best-effort"| VPN["IPsec-VPN to VPN Gateway or TR"]
    Q1 -->|"no, VPC to VPC"| Q2{"More than ~4 VPCs, multiple<br/>accounts or need isolation?"}
    Q2 -->|"yes"| CEN["CEN + Enterprise Transit Router"]
    Q2 -->|"no"| PEER["VPC Peering"]
    EC --> Q3{"Also many VPCs?"}
    VPN --> Q3
    Q3 -->|"yes"| CEN

3. Multi-Account Strategy

Architecture Summary

Alibaba Cloud's multi-account strategy centers on Resource Directory (RD), a hierarchical account-management service similar to AWS Organizations. The management account creates a Root folder and up to five levels of sub-folders, and places members (Alibaba Cloud accounts) in them. Control Policies attached to folders or members cap what identities in those members can do. Agentic Cloud Governance Center (renamed from Cloud Governance Center on 2026-06-24) builds landing zones on top of RD. RAM and CloudSSO provide access.

The hierarchy below is the recommended starting shape. Folder names are conventions, not product requirements.

flowchart TB
    MA["Management account<br/>Resource Directory owner, billing"] --> ROOT["Root folder"]
    ROOT --> CORE["Core folder"]
    ROOT --> WLF["Workloads folder"]
    ROOT --> SBX["Sandbox folder"]
    CORE --> LOG["Log Archive account<br/>ActionTrail, SLS, OSS"]
    CORE --> SEC["Security account<br/>Security Center, Cloud Firewall, Config aggregator"]
    CORE --> NETA["Network account<br/>CEN, Transit Routers, shared VPCs"]
    CORE --> SHR["Shared Services account<br/>ACR images, KMS, DNS"]
    WLF --> BUA["BU-A folder"]
    WLF --> BUB["BU-B folder"]
    BUA --> AD["A-Dev"]
    BUA --> AS["A-Staging"]
    BUA --> AP["A-Prod"]
    BUB --> BP["B-Prod"]
    SBX --> EXP["Experiment accounts"]
    CPOL["Control Policies"] -.->|"attached to folders"| WLF
    CPOL -.-> SBX

Key Services

Service Role
Resource Directory Multi-account hierarchy: Root, folders (max 5 levels), members. Supports trusted access and delegated administrators for services
Agentic Cloud Governance Center Landing zone setup, account factory with baselines, guardrails, governance maturity checks (including AI workloads since 2026)
RAM (Resource Access Management) Access control within each account: users, groups, roles, policies. SAML/OIDC federation
CloudSSO Workforce SSO across all RD members, like AWS IAM Identity Center. Access configurations map to RAM roles in members
Resource Groups Logical grouping of resources inside an account for access and management
Tags Key-value labels for cost allocation, access control, and automation
ActionTrail Audit logging of API calls. Organization trails cover every RD member
Cloud Config Compliance rules and configuration tracking. Aggregators span accounts
Control Policies Organization-level guardrails like AWS SCPs. Effect scope is All (includes the member's root identity) or RAM (RAM users and roles only)
  • Separate accounts by lifecycle stage: dev, staging, and prod go in different accounts, not different VPCs in one account.
  • Centralized networking account: Owns the CEN instance, Transit Routers, and shared VPCs. Workload VPCs attach to the TR across accounts (the attachment records a VpcOwnerId).
  • Centralized logging: All accounts send ActionTrail and SLS (Simple Log Service) logs to the Log Archive account.
  • Centralized security: Security Center and Cloud Firewall are managed from the Security account as a delegated administrator.
  • Account factory: Use Agentic Cloud Governance Center's account factory, or Terraform (alicloud_governance_account / alicloud_resource_manager_account) or ROS, to create accounts with baselines (RAM roles, guardrails, logging, networking).
  • SSO: Use CloudSSO, or federate an external IdP (Okta, Microsoft Entra ID, and others) through SAML 2.0. Map IdP groups to RAM roles in each account.
  • Cost management: Use the management account as the payer for members and use tags for cost allocation.

Cross-Account VPC Attachment Flow

The sequence shows how a workload VPC in a member account joins the network account's Transit Router. It is the most common cross-account operation in an RD tenancy.

sequenceDiagram
    participant WL as Workload account (VPC owner)
    participant RAM as RAM / CEN authorization
    participant NET as Network account (CEN owner)
    participant TR as Transit Router
    WL->>RAM: Grant CEN instance cen-xxx permission to attach vpc-xxx
    NET->>TR: CreateTransitRouterVpcAttachment with VpcOwnerId and zone vSwitches
    TR->>WL: Create TR ENIs in the chosen vSwitches
    NET->>TR: Associate the attachment with spoke route table
    NET->>TR: Enable propagation into hub route table
    TR-->>WL: Routes to hub and on-prem become reachable

Real-World Example

A multinational enterprise organizes 50+ Alibaba Cloud accounts under Resource Directory. The Core folder holds the log-archive, security, and shared-network accounts. Each business unit has its own folder with dev/staging/prod accounts. The shared-network account owns a CEN instance with Enterprise TRs in Jakarta, Singapore, and Frankfurt. Workload accounts create VPCs and attach them to the TR after granting cross-account authorization. Governance Center enforces baseline policies: no public ECS instances, ActionTrail enabled, and an owner tag on every resource.


4. Multi-Zone and Multi-Region Deployment Patterns

Multi-Zone (Intra-Region)

Alibaba Cloud regions contain multiple availability zones (AZs). An AZ is a physically isolated data center with low-latency links to the others in its region. Deploying across at least two AZs protects against a single-datacenter failure. Some newer regions launched with one AZ (Mexico at launch). Check the zone count before you design for multi-AZ in a new region.

Key Services and Features

Service Multi-AZ Feature
ECS Deploy instances across AZs. Use an Auto Scaling group with multi-AZ policy
ALB / NLB / CLB Distribute traffic across AZs. Health checks remove unhealthy instances
RDS High-availability edition with a standby in another zone and automatic failover
PolarDB Cluster with one primary and multiple read-only nodes (up to 15 per the 2026-04 note) across AZs
OSS LRS (locally redundant, the default) or ZRS (zone-redundant across AZs in supported regions). Choose ZRS explicitly for AZ-level resilience
Tair (Redis OSS-compatible) Standard (dual-replica) or cluster edition with multi-zone deployment
  • Use at least two AZs for any production workload.
  • Load balancer listeners should include backend servers in both AZs.
  • RDS: use the high-availability edition with the standby in a second zone, and test the failover window for your engine and edition before you rely on it.
  • ECS Auto Scaling group: set MultiAZPolicy to BALANCE or COST_OPTIMIZED (the default is PRIORITY).

Multi-Region

Deploy across geographically separated regions for disaster recovery, compliance, or latency. Examples: Jakarta + Singapore for Southeast Asia, or Shanghai + Singapore for China plus APAC.

Key Services and Features

Service Multi-Region Feature
CEN Connects VPCs across regions through Transit Router inter-region connections
Global Accelerator (GA) Anycast/accelerated IPs that carry user traffic over Alibaba's backbone to the nearest healthy endpoint
Alibaba Cloud DNS Geo- and ISP-line-based DNS routing (similar to Route 53)
GTM (Global Traffic Manager) Health-check-based DNS failover between regions
DTS (Data Transmission Service) Real-time one-way or two-way data synchronization (RDS, PolarDB, MongoDB, Redis, and others) between regions
OSS Cross-Region Replication Asynchronous object replication between buckets in different regions
PolarDB GDN Global Database Network: physical replication from a primary cluster to secondary clusters in other regions
  • Use CEN for private inter-region connectivity. Choose a bandwidth plan or pay-by-data-transfer.
  • Use GTM or Alibaba Cloud DNS for DNS-based failover with health checks. Keep DNS TTLs low.
  • Use DTS for database replication. Two-way sync (sync_architecture = bidirectional) needs a conflict-handling policy.
  • Use OSS CRR for object storage replication (async, eventual consistency).
  • Deploy stateless application layers to simplify failover. Store session state in Tair (Redis OSS-compatible), which offers a global distributed cache option.

Real-World Example

An e-commerce platform runs in China (Shanghai) as primary and China (Beijing) as secondary. CEN connects the two regions with a 5 Gbit/s bandwidth plan. DTS replicates RDS MySQL from Shanghai to Beijing with near-real-time lag. OSS CRR replicates product images. GTM monitors health endpoints in both regions. If Shanghai becomes unhealthy, GTM switches DNS to Beijing, and clients follow within the DNS TTL. The application tier is stateless (ECS + Auto Scaling), so Beijing can scale from a minimal warm standby to full capacity in minutes.


5. Disaster Recovery (DR)

DR Tiers

Tier RPO RTO Pattern Cost
Level 1: Data backup only Hours-Days Days OSS backup, Cloud Backup (formerly HBR) Lowest
Level 2: Cold standby Hours Hours Infra defined in IaC. Data replicated. No running compute Low
Level 3: Warm standby Minutes Minutes Minimal compute in DR region. Scale on failover Medium
Level 4: Active-passive Seconds Seconds-minutes Full stack in both regions. Only primary serves traffic High
Level 5: Active-active ~0 Seconds Both regions serve traffic simultaneously. Bidirectional replication Highest

Active-Passive DR

The diagram shows the steady state: Region A serves all traffic while DTS and OSS CRR keep Region B current.

flowchart LR
    GTM["GTM health checks<br/>DNS -> Region A"]
    subgraph A["Region A (primary)"]
        LBA["ALB"] --> ECSA["ECS (full)"]
        ECSA --> RDSA["RDS primary"]
        ECSA --> TA["Tair primary"]
        OSSA["OSS bucket"]
    end
    subgraph B["Region B (standby)"]
        LBB["ALB"] --> ECSB["ECS (full or scaled down)"]
        ECSB --> RDSB["RDS (DTS target, read-only)"]
        ECSB --> TB["Tair replica"]
        OSSB["OSS bucket"]
    end
    GTM --> LBA
    GTM -.->|"on failure"| LBB
    RDSA -->|"DTS sync"| RDSB
    OSSA -->|"CRR"| OSSB
    TA -->|"replication"| TB

The failover sequence shows who acts and in what order when Region A fails.

sequenceDiagram
    participant GTM as GTM
    participant A as Region A endpoint
    participant DTS as DTS task
    participant B as Region B stack
    participant Ops as Operator or runbook
    GTM->>A: Health check fails N times
    GTM->>GTM: Switch address pool to Region B
    Ops->>DTS: Stop sync task A to B
    Ops->>B: Promote RDS target to read-write, scale ECS
    GTM-->>B: New client sessions land on Region B after TTL
    Ops->>DTS: Later create reverse sync B to A for failback
  • DTS replicates data from Region A to Region B in near-real time.
  • Region B runs the full stack (or a scaled-down copy) but receives no user traffic while GTM points at Region A.
  • On failure, GTM detects the unhealthy Region A endpoint and updates DNS to Region B. You promote the database in Region B. DTS does not do this automatically.

Active-Active (Multi-Region)

The diagram shows both regions taking writes, with replication in both directions.

flowchart LR
    GTM["GTM geo routing<br/>per-region health checks"]
    subgraph A["Region A (active)"]
        LBA["ALB"] --> APPA["ECS / ACK"]
        APPA --> DBA["Database (writes)"]
        APPA --> CA["Tair"]
    end
    subgraph B["Region B (active)"]
        LBB["ALB"] --> APPB["ECS / ACK"]
        APPB --> DBB["Database (writes)"]
        APPB --> CB["Tair"]
    end
    GTM --> LBA
    GTM --> LBB
    DBA <-->|"DTS two-way sync<br/>conflict policy"| DBB
    CA <-->|"global distributed cache"| CB
  • Two-way logical replication: Allows writes in both regions. It needs a conflict-handling policy and ideally data partitioning (each user or tenant "homed" in one region) so real conflicts are rare. DTS provides this for supported engines. The 2026-04 note named PolarDB-X "DRC" as the engine for this. DRC (Data Replication Center) is Alibaba's internal replication platform behind its own unitized active-active setup. It is not sold as a separate cloud product: DTS is the public service that grew out of it (VLDB 2024 paper on DTS, checked 2026-09-27). DTS two-way sync is therefore the documented option.
  • PolarDB GDN: For active-passive with near-zero RPO. It replicates at the storage/redo layer (physical replication), with cross-region lag Alibaba describes as typically under 2 seconds. Only the primary cluster accepts writes. Secondary clusters can forward writes to the primary.
  • Tair global distributed cache: Bidirectional data synchronization between cache instances in different regions.
  • GTM geo-based routing: Splits traffic by geography, for example APAC users to Singapore and US users to Virginia. Per-region health checks enable automatic failover.

Hybrid DR (Cloud + On-Premises)

  • Cloud Backup (formerly Hybrid Backup Recovery, HBR): Backs up on-premises and cloud data (files, databases, VMs, NAS) to Alibaba Cloud.
  • Express Connect / VPN Gateway: Private connectivity between the on-premises DC and Alibaba Cloud VPCs.
  • Smart Access Gateway (SAG): Alibaba's SD-WAN appliance for branch offices. Parts of the family are being retired: SAG App, Cloud Intelligent Branch and SmartAG Maintenance premium reached end of marketing on 2026-03-17, and end of full support on 2026-09-13. End of service is 2027-03-12 (notice). Confirm with Alibaba before designing new branch connectivity around SAG.
  • DTS: Can replicate between on-premises databases and cloud RDS.
  • Define RTO/RPO targets first. They decide the DR tier and cost.
  • Use Infrastructure as Code (Terraform or ROS) to define the full stack. In a cold or warm standby, you can recreate the DR region from code.
  • Test failover regularly (quarterly at minimum). Rehearse GTM switchover and database promotion in a non-prod environment or a controlled window.
  • Stateless app tier: Makes failover much simpler. Keep all state in managed services (Tair, RDS, OSS).
  • Separate DR automation from production: Use separate Terraform state / ROS stacks for the DR region so they do not share a single point of failure.

6. DMZ Patterns

Architecture Summary

The DMZ (demilitarized zone) pattern on Alibaba Cloud uses subnet tiering inside a VPC to build layered defense. Public-facing resources sit in a public vSwitch (the DMZ). Application and data tiers sit in private vSwitches with no direct Internet exposure.

The diagram shows the three tiers and which controls sit at each boundary.

flowchart TB
    INET["Internet"] --> CFW["Cloud Firewall<br/>Internet border, IPS block mode"]
    subgraph VPC["VPC 10.0.0.0/16"]
        subgraph DMZ["Public vSwitch - DMZ tier"]
            WAF["WAF"]
            ALB["ALB"]
            NATG["NAT Gateway (SNAT)"]
            BAS["Bastionhost / VPN Gateway"]
        end
        subgraph APP["Private vSwitch - app tier"]
            ASG["ECS Auto Scaling group"]
            ECI["ECI / ACK pods"]
        end
        subgraph DATA["Private vSwitch - data tier"]
            DB["RDS / PolarDB, no public IP"]
            CACHE["Tair, no public IP"]
            OSSEP["OSS via VPC endpoint"]
        end
        WAF --> ALB
        ALB -->|"SG: 8080 only"| ASG
        ALB --> ECI
        ASG -->|"SG: 3306 only"| DB
        ASG --> CACHE
        ASG --> OSSEP
        ASG -->|"egress"| NATG
        BAS -->|"SSH/RDP audited"| ASG
    end
    CFW --> WAF
    NATG --> CFW

Key Services

Service Role in DMZ
Cloud Firewall Boundaries: Internet border (N/S on EIP/SLB), NAT border (egress), VPC border (E/W between VPCs), internal firewall (E/W between ECS). Includes IPS with threat intelligence
NAT Gateway Internet NAT Gateway provides SNAT (outbound) and DNAT (inbound port forwarding) for private resources. VPC NAT Gateway translates between overlapping private CIDRs
ALB / NLB / CLB Sits in the DMZ and distributes traffic to private ECS. ALB is Layer 7, NLB Layer 4, CLB legacy L4/L7
WAF (Web Application Firewall) Inspects HTTP/HTTPS traffic for OWASP Top 10, bots, and custom rules. Placed in front of ALB/CLB
Bastionhost Managed jump server for audited SSH/RDP access to private instances. Integrates with RAM
VPN Gateway IPsec-VPN or SSL-VPN for site-to-site or client-to-site private access
Security Center Unified threat detection, vulnerability scanning, compliance checking
  1. Network segmentation: Three tiers of vSwitches: public (DMZ), private (app), private (data). Route tables and security groups enforce DMZ -> app and app -> data. There is no direct DMZ -> data path.

  2. Cloud Firewall Internet border: Enable on all EIPs, public load balancers, and NAT Gateways. Configure allow-list rules. Set IPS to block mode, not just alert. Cloud Firewall inspects layers 3/4 and layer 7.

  3. Cloud Firewall VPC border: If you use CEN or VPC peering, enable the VPC border firewall. Default-deny, with explicit allow rules between VPCs.

  4. Cloud Firewall internal firewall: For micro-segmentation between ECS instances within a VPC. Group instances by role (web, app, db) and write policies per group.

  5. NAT Gateway placement: Deploy in the DMZ vSwitch. Configure SNAT entries for each private vSwitch that needs Internet egress. Avoid DNAT for inbound traffic where possible and prefer a load balancer.

  6. Load balancer placement: Internet-facing ALB/NLB in the DMZ vSwitches, backend servers in private vSwitches. Configure health checks and send access logs to SLS.

  7. Bastionhost / VPN Gateway: Place in the DMZ vSwitch. Never assign public IPs to app or DB instances. Use RAM roles and MFA for admin access.

  8. Defense-in-depth layers:

Layer Control
Edge Anti-DDoS + Cloud Firewall (Internet border): DDoS, IPS, access control
Perimeter WAF: OWASP, bot management, custom rules for HTTP/HTTPS
Network Security Groups (stateful, per-instance) + Network ACLs (stateless, per-subnet)
Host Security Center: vulnerability scanning, baseline checks
Application Application-level auth (RAM, OAuth, JWT)

Access Control Policy Matrix (Example)

Source Destination Protocol Ports Action
Internet ALB (DMZ) TCP 443 Allow
Internet Any Any Any Deny
ALB (DMZ) ECS App (private) TCP 8080 Allow
ECS App RDS (data) TCP 3306 Allow
Bastionhost All ECS TCP 22 Allow
Private subnet NAT GW (DMZ) TCP 80, 443 Allow (SNAT egress)
All others All others Any Any Deny

Real-World Example

A fintech company's production VPC in Singapore uses three tiers. The DMZ vSwitch hosts WAF, an ALB, a NAT Gateway, and a VPN Gateway. The app vSwitch runs ECS instances in an Auto Scaling group behind the ALB. The data vSwitch hosts PolarDB for MySQL and Tair. Cloud Firewall Internet border is enabled on the ALB EIP and NAT Gateway with IPS in block mode. The internal firewall enforces web->app on port 8080 and app->db on port 3306, and denies all other cross-tier traffic. All admin access goes through the VPN Gateway to Bastionhost.


7. Landing Zone Best Practices

Architecture Summary

A landing zone is a pre-configured, governed multi-account environment that gives every workload a secure baseline. Alibaba Cloud's implementation is Agentic Cloud Governance Center (renamed from Cloud Governance Center on 2026-06-24). It uses an IaC service to orchestrate Resource Directory, RAM, CloudSSO, ActionTrail, and Cloud Config. It then provides a setup wizard, an account factory with reusable baselines, and guardrails. The 2026 "agentic" release added AI Governance Maturity Checks for AI workloads (agent platforms, agent runtimes, AI gateways) across five pillars: security, reliability, cost, efficiency, and performance.

The diagram shows the landing zone components and which account owns each.

flowchart TB
    subgraph MGMT["Management account"]
        CGC["Agentic Cloud Governance Center<br/>setup wizard, account factory, guardrails"]
        RD["Resource Directory"]
        CSSO["CloudSSO"]
        BILL["Consolidated billing"]
        CGC --> RD
        CGC --> CSSO
    end
    subgraph CORE["Core folder"]
        LOGA["Log Archive<br/>ActionTrail org trail -> SLS + OSS"]
        SECA["Security<br/>Security Center, Cloud Firewall, Config aggregator"]
        NETA["Network<br/>CEN, TRs, shared VPCs, PrivateZone DNS"]
        SHRA["Shared Services<br/>ACR, KMS, images, certificates"]
    end
    subgraph WORK["Workload folders"]
        W1["BU-A dev / staging / prod"]
        W2["BU-B dev / staging / prod"]
        SB["Sandbox (time-limited)"]
    end
    RD --> CORE
    RD --> WORK
    CGC -->|"baseline: RAM roles, trail,<br/>Config rules, VPC attach"| WORK
    W1 -->|"TR attachment"| NETA
    W2 -->|"TR attachment"| NETA
    W1 -.->|"logs"| LOGA
    W2 -.->|"logs"| LOGA
    SECA -.->|"delegated admin"| WORK

Account Factory Flow

When the account factory vends a new account, the steps run in the order below. Terraform can drive the same flow with alicloud_governance_baseline and alicloud_governance_account (provider v1.228.0+).

sequenceDiagram
    participant Req as Requester or CI pipeline
    participant CGC as Governance Center account factory
    participant RD as Resource Directory
    participant NEW as New member account
    participant NET as Network account
    Req->>CGC: Enroll account with baseline_id and folder_id
    CGC->>RD: Create resource account in target folder
    RD-->>CGC: Member created (account ID)
    CGC->>NEW: Apply baseline items: password policy, RAM security, services
    CGC->>NEW: Create CloudSSO access roles and ActionTrail delivery
    CGC->>NEW: Enable Cloud Config rules and guardrails
    CGC->>NET: Create VPC and attach to Transit Router if baseline includes networking
    CGC-->>Req: Account ready, drift checked continuously

Key Services

Service Role in Landing Zone
Resource Directory Account hierarchy: Root, folders, members. Trusted access and delegated administration for services
Agentic Cloud Governance Center Landing zone setup wizard, account factory, baselines, guardrails, compliance dashboards, AI governance checks
CloudSSO / RAM Centralized identity: SSO to all members, external IdP federation (SAML 2.0, SCIM provisioning)
ActionTrail API-level audit logging for all accounts. Deliver the organization trail to the Log Archive account's SLS project or OSS bucket
Cloud Config Compliance rules (for example, "no public ECS," "all resources tagged," "encryption required"). Evaluates resources continuously
Control Policies Organization-level restrictions on what member accounts can do (like AWS SCPs). Applied at folder or member level
CEN (in Network Account) Centralized network connectivity for all workload accounts. Transit Router with custom route tables
Cloud Firewall (in Security Account) Centralized firewall management across accounts: Internet, NAT, VPC border, and internal firewall
ROS / Terraform IaC for reproducible landing zone deployment. Governance Center itself orchestrates through an IaC service
KMS (Key Management Service) Centralized key management. Keep keys per workload account or share from a security account per policy
  1. Start with Governance Center: Use the built-in landing zone setup wizard to create the core accounts (management, log archive, security, shared services) with recommended baselines. It is faster and less error-prone than manual setup.

  2. Account factory: Configure baselines so new accounts are provisioned with:

  3. RAM roles for admin and operator access (through CloudSSO)
  4. ActionTrail delivery to the Log Archive account
  5. Cloud Config rules (baseline compliance)
  6. Network attachment to the central CEN Transit Router
  7. Resource tags (owner, environment, cost-center)

  8. Guardrails (Control Policies + Config rules):

  9. Deny public ECS instances in prod folders
  10. Require encryption for OSS and RDS
  11. Deny creation of resources outside approved regions
  12. Require owner and environment tags on all resources
  13. Protect the landing zone itself: deny changes to ResourceDirectoryAccountAccessRole and the trail

  14. Centralized networking:

  15. The Network account owns the CEN instance and Transit Routers
  16. Shared VPCs (through Resource Sharing) for common services
  17. Workload VPCs attach to the TR from their own accounts
  18. Cloud Firewall VPC border enabled on TR attachments

  19. Centralized logging and monitoring:

  20. All accounts covered by one ActionTrail organization trail
  21. SLS centralized in the Log Archive account
  22. Cloud Config aggregation in the Security account

  23. Identity and access:

  24. Federate an external IdP (Okta, Microsoft Entra ID) into CloudSSO through SAML, with SCIM for user sync
  25. Map IdP groups to access configurations (RAM roles) in each account
  26. Enforce MFA for all human users
  27. Use RAM roles and STS (not long-lived AccessKeys) for cross-account and CI access

  28. IaC for the landing zone itself:

  29. Define the landing zone in Terraform or ROS
  30. Version-control the configuration
  31. Apply changes through CI/CD with a plan-and-apply pipeline

  32. Regular drift detection:

  33. Governance Center evaluates the environment and records evaluation history
  34. Configure alerts for non-compliant resources
  35. Review Control Policies periodically

Real-World Example

A large retail enterprise builds its landing zone with Governance Center. The management account holds Resource Directory with three top-level folders: Core, BusinessUnits, and Sandbox. Core contains the log-archive, security, and network accounts. BusinessUnits has a folder per department (Online, Stores, SupplyChain), each with dev/staging/prod accounts. The account factory provisions new accounts with baseline RAM roles, ActionTrail delivery, Cloud Config rules, and a CEN TR attachment in minutes. Control Policies enforce no public ECS in prod, encryption for all storage, and resources only in approved regions (Jakarta, Singapore, Frankfurt). The security account manages Cloud Firewall centrally across all accounts.


8. Security and Identity

Identity and access management, network security, data protection, and compliance considerations for Alibaba Cloud. The matching commands are in How-to Guides. The product fact tables are in Reference.

Identity & Access

RAM (Resource Access Management)

RAM is Alibaba Cloud's IAM service. It supports users, groups, roles, and fine-grained policies inside one Alibaba Cloud account. Resource Directory and CloudSSO extend it across accounts.

Concept Description
RAM User Long-lived identity with console login and/or AccessKey pair. The per-account limit is a quota: read the current default on the RAM quotas page or in Quota Center (see Reference)
RAM Group Collection of RAM users. Policies attached to a group apply to all members
RAM Role Virtual identity assumed by trusted entities (RAM users, Alibaba Cloud services, OIDC/SAML IdPs). Issues temporary STS credentials
RAM Policy JSON document defining allowed/denied actions, resources, and conditions. System policies are managed by Alibaba Cloud. Custom policies are yours

STS (Security Token Service)

STS issues temporary credentials (AccessKeyId, AccessKeySecret, SecurityToken). The duration runs from 900 seconds up to the role's maximum session duration: default 3,600 s, configurable up to 43,200 s. Prefer STS over long-lived AccessKey pairs for workloads. On ECS use instance RAM roles. On ACK use RRSA (OIDC). In CI use OIDC federation.

SSO Federation

Alibaba Cloud supports three SSO models:

  • User-based SSO (RAM): Maps external IdP users to RAM users through SAML 2.0. Each IdP user needs a matching RAM user.
  • Role-based SSO (RAM): Maps IdP users to RAM roles, so no RAM users are needed. The SAML assertion names the role to assume.
  • CloudSSO: The multi-account option. A directory at the RD management account, synced from the IdP (SAML + SCIM). Access configurations become RAM roles in each member. This is the recommended enterprise approach for Resource Directory tenancies.

The sequence shows role-based federation into a member account, the path most enterprise sign-ins take.

sequenceDiagram
    participant U as Engineer
    participant IdP as External IdP (Okta or Entra ID)
    participant SSO as CloudSSO portal
    participant STS as STS
    participant M as Member account console/API
    U->>IdP: Sign in with MFA
    IdP->>SSO: SAML assertion with group claims
    SSO->>SSO: Match groups to access configurations
    U->>SSO: Pick account A-Prod and role Operator
    SSO->>STS: AssumeRole for the provisioned RAM role
    STS-->>U: Temporary credentials (max session duration)
    U->>M: Call APIs, audited by ActionTrail

Multi-account SSO

For organizations on Resource Directory, configure SSO once at the management account (CloudSSO) and assume roles in member accounts. This avoids creating RAM users in every member account.

Network Security

Security Groups

Security groups are stateful, instance-level firewalls. Rules have a priority from 1 to 100, where a lower number means higher priority and the default is 1. Check Quota Center for the security-group quotas that apply to your account.

Best practices:

  • One security group per application tier (web, app, db)
  • Reference security group IDs in rules instead of IP ranges where possible
  • Default-deny: basic security groups deny all inbound traffic by default
  • Limit outbound traffic to known destinations for sensitive workloads

Network ACLs

Network ACLs are stateless, subnet-level packet filters. They supplement security groups for defense-in-depth. Rules are evaluated in order and the first match wins.

Cloud Firewall

Cloud Firewall inspects traffic at the Internet border, NAT border, VPC border, and inside the VPC (internal firewall). See Reference: Cloud Firewall Boundaries. It integrates with Alibaba Cloud threat intelligence for IP reputation scoring and automatic blocking of known malicious sources.

Anti-DDoS

Basic protection is on by default for every public IP. Paid tiers split by where the protected resource sits: mainland China (Anti-DDoS Pro) or outside it (Anti-DDoS Premium). See Reference: Anti-DDoS Family.

WAF (Web Application Firewall)

WAF sits in front of ALB/CLB to inspect HTTP/HTTPS traffic. It provides OWASP Top 10 protection, bot management, CC (HTTP flood) attack defense, and custom rules. In CNAME (reverse-proxy) mode, DNS points to the WAF CNAME, which forwards clean traffic to the origin. Cloud-native (load-balancer-integrated) mode avoids the DNS change.

Data Protection

KMS (Key Management Service)

KMS centralizes key management for encryption at rest. Keys are either customer-managed or Alibaba Cloud-managed service keys. Feature facts are in Reference: KMS Features.

TDE (Transparent Data Encryption)

RDS and PolarDB support TDE, which encrypts data files at rest without application changes. With a customer-managed KMS key, RDS also needs a RAM role (RoleARN) that lets it use the key.

ActionTrail Audit Logging

ActionTrail records API calls made to Alibaba Cloud services. Logs include the caller identity, timestamp, source IP, request parameters, and response. For multi-account setups, create an organization trail from the management account. It captures API calls from every member in Resource Directory. A new trail does not log until you start it.

Compliance

China Cybersecurity Law

The Cybersecurity Law of the People's Republic of China (in force since 2017-06-01) requires:

  • Data localization: Operators of critical information infrastructure (CII) must store personal information and important data collected in China domestically. The Data Security Law (2021) and PIPL (2021) extend cross-border controls to other data handlers. Cross-border transfer requires a CAC security assessment, a standard contract, or certification, depending on volume and data type.
  • Network security obligations: Network operators must implement security protections, retain network logs for at least 6 months, and maintain incident response plans.
  • Critical Information Infrastructure (CII): CII operators face additional requirements, including annual security reviews and procurement security reviews.
  • 2026 amendment: An amendment raising penalties took effect on 2026-01-01. Check its details with counsel.

MLPS 2.0 (Multi-Level Protection Scheme)

MLPS 2.0 (GB/T 22239-2019) is China's mandatory information-security grading system, with levels 1-5. Alibaba Cloud provides compliance packages mainly for levels 2 and 3:

Level Scope Alibaba Cloud Support
Level 2 General business systems Security Center baseline checks, Cloud Firewall, ActionTrail
Level 3 Important business systems All Level 2 + HSM-backed KMS, WAF, Anti-DDoS Pro, Bastionhost
Level 4 Very important systems No Alibaba Level 4 package found in the public sources checked (2026-09-27). Discuss with Alibaba; finance-cloud and government-cloud regions are the usual starting point

Data Residency

Alibaba Cloud runs separate sites, accounts, and legal entities for mainland China (aliyun.com) and international regions (alibabacloud.com). Data in Chinese regions stays in China unless you configure a transfer. Cross-border replication (for example DTS or OSS CRR from cn-* to international regions) must be set up explicitly and is subject to the rules above for personal and important data.

Available Certifications

See Reference: Compliance Facts for the certification list. Verify current scope per region in the Alibaba Cloud Trust Center before relying on it.

Sources