Alibaba Cloud (Aliyun) Architecture Patterns¶
What this page explains
How Alibaba Cloud's building blocks fit together and why you would pick one pattern over another. It covers single-VPC, multi-VPC (CEN/Transit Router, VPC peering, Express Connect), multi-account (Resource Directory), multi-zone and multi-region, DR, DMZ, and landing zone patterns. The last section covers the security model (RAM, STS, SSO, KMS, Cloud Firewall, China compliance). Region tables, service mappings, and limits are in Reference. Runnable commands are in How-to Guides.
Platform Architecture Overview¶
Alibaba Cloud has three layers. Resource Directory and Agentic Cloud Governance Center govern accounts at the top. Regions contain zones, and each zone contains VPCs and vSwitches. Cloud Enterprise Network (CEN) and Express Connect stitch regions and on-premises sites together. The diagram shows how these layers relate in a typical enterprise tenancy.
flowchart TB
subgraph GOV["Governance plane (management account)"]
RD["Resource Directory<br/>Root, folders, members"]
CGC["Agentic Cloud Governance Center<br/>landing zone, account factory"]
CP["Control Policies"]
SSO["CloudSSO / RAM SAML"]
CGC --> RD
RD --> CP
SSO -.->|"role access"| RD
end
subgraph NET["Network account"]
CEN["CEN instance"]
TRA["Transit Router<br/>ap-southeast-5 (Jakarta)"]
TRB["Transit Router<br/>ap-southeast-1 (Singapore)"]
CEN --- TRA
CEN --- TRB
TRA <-->|"inter-region attachment"| TRB
end
subgraph WL["Workload member account"]
VPC["VPC<br/>vSwitch zone A + zone B"]
ECS["ECS / ACK node pools"]
DB["PolarDB / RDS"]
VPC --> ECS
VPC --> DB
end
subgraph ONP["On-premises DC"]
CPE["Customer router"]
end
RD -->|"member of folder"| WL
RD -->|"member of folder"| NET
TRA -->|"VPC attachment"| VPC
CPE -->|"Express Connect circuit"| VBR["VBR"]
VBR --> ECR["Express Connect Router"]
ECR --> TRA
1. Single Project with Single VPC¶
Architecture Summary¶
A single Virtual Private Cloud (VPC) is the basic building block on Alibaba Cloud. A VPC is an isolated virtual network. You define a private CIDR block, create vSwitches (zonal subnets) across availability zones, and attach route tables and gateways. For one project, a single VPC with multi-AZ vSwitches gives isolation, high availability, and simplicity.
The flowchart shows the canonical three-tier layout inside one VPC.
flowchart TB
INET["Internet"] --> WAF["WAF (CNAME / cloud-native mode)"]
WAF --> ALB["ALB in public vSwitches<br/>zone A + zone B"]
subgraph VPC["VPC 10.0.0.0/16"]
ALB --> ECSA["ECS app<br/>private vSwitch, zone A"]
ALB --> ECSB["ECS app<br/>private vSwitch, zone B"]
ECSA --> RDSP["RDS / PolarDB primary<br/>data vSwitch, zone A"]
ECSB --> RDSP
RDSP -.->|"sync replica"| RDSS["Standby<br/>zone B"]
ECSA --> NAT["Enhanced NAT Gateway + EIP"]
ECSB --> NAT
end
NAT --> INET
Key Services¶
| Service | Role |
|---|---|
| VPC | Isolated virtual network with custom CIDR, vSwitches, route tables |
| vSwitch | Subnet within a VPC, bound to a single AZ. Resources attach here |
| Route Table | System route table (auto-created) + custom route tables. Controls traffic forwarding |
| ALB / NLB / CLB | The Server Load Balancer (SLB) family. ALB is Layer 7, NLB is Layer 4, CLB is the legacy Layer 4/7 product. All distribute traffic across AZs |
| NAT Gateway | Outbound Internet (SNAT) and inbound port forwarding (DNAT) for private resources. New gateways are always the Enhanced type |
| Security Group | Stateful per-instance firewall, like AWS security groups |
| Network ACL | Stateless subnet-level packet filter, like AWS NACLs |
| EIP (Elastic IP) | Public IP that can be bound to a NAT Gateway, load balancer, or individual ECS instance |
Recommended Configurations¶
- CIDR planning: Use RFC 1918 ranges (
10.0.0.0/8,172.16.0.0/12,192.168.0.0/16). Reserve headroom. A VPC CIDR must not overlap with any VPC you plan to peer or attach to CEN later. Start with/16for the VPC and/24for vSwitches. - Multi-AZ vSwitches: Create at least two vSwitches in different AZs. Place load balancer, ECS, and database replicas across both.
- Subnet tiers: Split into public (DMZ), private (application), and data (database) vSwitches. Use route tables to enforce tiered traffic flow.
- No public IPs on app/DB: Application and database instances must have no EIPs. Use NAT Gateway (SNAT) for outbound Internet and a load balancer for inbound.
- Security Groups: One SG per tier (web, app, db), with least-privilege rules. Allow web SG -> app SG -> db SG only on required ports.
Real-World Example¶
A SaaS startup deploys a three-tier web application in the Indonesia (Jakarta, ap-southeast-5) region. One VPC
(10.0.0.0/16) has vSwitches in two zones. Public vSwitches host an ALB instance. Private vSwitches host ECS
auto-scaling groups. A managed RDS MySQL instance runs in high-availability (multi-zone) mode. An Enhanced NAT
Gateway provides egress. Cloud Firewall is enabled on the Internet border.
2. Multi-VPC Architecture¶
Workloads often span several VPCs: for environment isolation (dev/staging/prod), business-unit separation, or regulatory compliance. Alibaba Cloud offers three primary inter-VPC connectivity options.
2a. Cloud Enterprise Network (CEN) -- Hub-and-Spoke¶
Architecture Summary¶
CEN is Alibaba Cloud's managed WAN service. A Transit Router (TR) in each region acts as the hub. VPCs, VPN connections, VBRs (Express Connect), Express Connect Routers (ECR), and Cloud Connect Network (CCN) instances attach as spokes. Traffic flows through the TR's route tables, which gives fine-grained routing policy, traffic isolation, and centralized egress.
TRs in different regions connect through inter-region (peer) attachments over Alibaba's private backbone. Each
inter-region connection draws on a purchased bandwidth plan (BandwidthType = BandwidthPackage) or is billed
pay-by-data-transfer (DataTransfer). The pre-2026-09 version of this note said an Enterprise Edition TR can
connect up to 1,000 VPCs per region. Alibaba confirms 1,000 VPCs per Enterprise Edition transit router, up from 200
(product news,
checked 2026-09-27). Check the CEN SLA for current backbone commitments.
The diagram shows a two-region hub-and-spoke design with a shared-services hub VPC and an on-premises link.
flowchart LR
subgraph CEN["CEN instance (network account)"]
subgraph RA["Region: ap-southeast-5 (Jakarta)"]
TRA["Transit Router A<br/>route tables: spoke-rt, hub-rt"]
HUB["Hub VPC<br/>Cloud Firewall, NAT, DNS"]
P1["Prod VPC"]
D1["Dev VPC"]
HUB --- TRA
P1 --- TRA
D1 --- TRA
end
subgraph RB["Region: ap-southeast-1 (Singapore)"]
TRB["Transit Router B"]
P2["Prod VPC (DR)"]
P2 --- TRB
end
TRA <-->|"inter-region attachment<br/>bandwidth plan or pay-by-data-transfer"| TRB
end
DC["On-prem DC"] -->|"Express Connect"| VBR["VBR"]
VBR --- TRA
Key Services¶
| Service | Role |
|---|---|
| CEN Instance | Global container for Transit Routers. One CEN holds TRs in many regions |
| Transit Router (TR) | Regional hub. The Enterprise Edition supports custom route tables, route maps, multicast, and traffic steering |
| TR Route Table | System or custom. Controls inter-VPC routing, isolation, and traffic steering through association and propagation |
| Bandwidth Plan | Pre-purchased cross-region bandwidth (the alternative is pay-by-data-transfer) |
| Network Instance Connection | Attachment types: VPC, VBR, VPN, CCN, ECR, and inter-region (peer) connections |
Recommended Configurations¶
- Use Enterprise Edition TRs. Custom route tables, traffic isolation, and service chaining need them.
- Hub VPC pattern: Deploy a centralized shared-services VPC (DNS, NTP, security appliances, NAT egress) attached to the TR. Use custom route tables to let spoke VPCs reach the hub but not each other.
- Centralized DMZ egress: Route all Internet-bound traffic from spoke VPCs through a hub VPC with Cloud Firewall and NAT Gateway for unified inspection.
- Bandwidth planning: Choose bandwidth plans for steady, predictable inter-region traffic and pay-by-data-transfer for bursty or low-volume traffic. Set per-connection bandwidth limits so one spoke cannot saturate the link.
- Route map policies: Use TR route maps for path selection, route manipulation, and traffic engineering.
Real-World Example¶
A financial services company has separate VPCs for trading, risk analytics, and back-office in China (Shanghai) and China (Beijing). A CEN with Enterprise TRs in both regions connects them. A hub security VPC in Shanghai runs Cloud Firewall and IDS. Custom route tables send trading VPC traffic through the security VPC before it reaches the Internet, while back-office VPCs stay isolated from trading. The Shanghai-Beijing connection uses a 2 Gbit/s bandwidth plan.
2b. VPC Peering Connection¶
Architecture Summary¶
VPC Peering is a direct, one-to-one connection between two VPCs. It supports same-region and cross-region peering, and same-account and cross-account peering. Unlike CEN, it is a point-to-point link without a central hub. You add routes that point at the peering connection to each VPC's route table.
flowchart LR
A["VPC-A 10.0.0.0/16<br/>route 10.1.0.0/16 -> pcc-xxx"] <-->|"VPC peering connection pcc-xxx"| B["VPC-B 10.1.0.0/16<br/>route 10.0.0.0/16 -> pcc-xxx"]
Key Services¶
| Service | Role |
|---|---|
| VPC Peering Connection | Direct link between two VPCs. Inter-region connections have a bandwidth setting and a Gold (default) or Platinum link type |
| VPC Route Table | Route entries whose next hop is the peering connection, added on both sides |
Recommended Configurations¶
- Make sure CIDR blocks do not overlap between peered VPCs.
- Same-account peering is simplest. Cross-account peering needs the accepter account's ID and acceptance by that account.
- For more than about 5-10 VPCs, prefer CEN over individual peering links to avoid a combinatorial explosion of peering connections and route entries. Peering is not transitive.
- Inter-region peering is billed by data transfer. Check current pricing for same-region peering before relying on it being free.
When to Choose VPC Peering over CEN¶
- Small number of VPCs (2-4).
- Simple, static connectivity without complex routing policy.
- Lower cost for low-volume traffic, since there are no TR attachment or processing fees.
2c. Express Connect (Leased Line)¶
Architecture Summary¶
Express Connect provides physically dedicated, private circuits between your data center and Alibaba Cloud. A Virtual Border Router (VBR) terminates the circuit at the Alibaba Cloud edge. An Express Connect Router (ECR) is a newer global router that exchanges routes with VBRs through BGP and forwards traffic to VPCs and TRs in many regions. It replaces a mesh of VBR-to-VPC connections.
The sequence shows how routes are learned and traffic flows from an on-premises router to a VPC through ECR.
sequenceDiagram
participant CPE as On-prem router
participant VBR as VBR (access point)
participant ECR as Express Connect Router
participant TR as Transit Router
participant VPC as Workload VPC
CPE->>VBR: BGP session over dedicated circuit, advertise 192.168.0.0/16
VBR->>ECR: Propagate on-prem prefixes
ECR->>TR: Associate and propagate routes
TR->>VPC: Install route 192.168.0.0/16 via TR attachment
VPC-->>CPE: Return traffic follows the same path, never the Internet
Dedicated circuits run at up to 100 Gbit/s per port. The Terraform provider's port_type values range from
100Base-T and 1000Base-LX through 10GBase-LR to 40GBase-LR and 100GBase-LR (the last two since provider
1.185.0, alicloud_express_connect_physical_connection,
checked 2026-09-27). See also the interconnect table on the
Multi-Cloud Governance reference.
Key Services¶
| Service | Role |
|---|---|
| Express Connect Circuit | Dedicated physical line, or a hosted connection from a partner |
| VBR (Virtual Border Router) | Router at the Alibaba Cloud edge for circuit termination and BGP |
| ECR (Express Connect Router) | Global hybrid-cloud router with dynamic BGP (Terraform support since provider v1.224.0) |
| Hosted Connection | Shared circuit provided by an Express Connect partner |
Recommended Configurations¶
- Use it for hybrid cloud (DC-to-VPC) when you need guaranteed bandwidth.
- For production, provision redundant circuits to different access points and use ECMP or BGP path preference for failover.
- Enable BGP dynamic routing through VBR/ECR for automatic route exchange.
- Express Connect vs. VPN: Express Connect offers higher bandwidth, lower and steadier latency, and physical isolation. Use IPsec-VPN as a backup path or for non-critical sites.
- Do not use Express Connect for VPC-to-VPC links inside Alibaba Cloud. Use VPC peering or CEN instead.
Choosing a Connectivity Option¶
The full side-by-side matrix is in Reference: Connectivity Comparison Matrix. The decision usually reduces to the flow below.
flowchart TD
Q1{"Connecting an on-prem<br/>site?"} -->|"yes, need SLA bandwidth"| EC["Express Connect + VBR<br/>(+ ECR for many regions)"]
Q1 -->|"yes, best-effort"| VPN["IPsec-VPN to VPN Gateway or TR"]
Q1 -->|"no, VPC to VPC"| Q2{"More than ~4 VPCs, multiple<br/>accounts or need isolation?"}
Q2 -->|"yes"| CEN["CEN + Enterprise Transit Router"]
Q2 -->|"no"| PEER["VPC Peering"]
EC --> Q3{"Also many VPCs?"}
VPN --> Q3
Q3 -->|"yes"| CEN
3. Multi-Account Strategy¶
Architecture Summary¶
Alibaba Cloud's multi-account strategy centers on Resource Directory (RD), a hierarchical account-management service similar to AWS Organizations. The management account creates a Root folder and up to five levels of sub-folders, and places members (Alibaba Cloud accounts) in them. Control Policies attached to folders or members cap what identities in those members can do. Agentic Cloud Governance Center (renamed from Cloud Governance Center on 2026-06-24) builds landing zones on top of RD. RAM and CloudSSO provide access.
The hierarchy below is the recommended starting shape. Folder names are conventions, not product requirements.
flowchart TB
MA["Management account<br/>Resource Directory owner, billing"] --> ROOT["Root folder"]
ROOT --> CORE["Core folder"]
ROOT --> WLF["Workloads folder"]
ROOT --> SBX["Sandbox folder"]
CORE --> LOG["Log Archive account<br/>ActionTrail, SLS, OSS"]
CORE --> SEC["Security account<br/>Security Center, Cloud Firewall, Config aggregator"]
CORE --> NETA["Network account<br/>CEN, Transit Routers, shared VPCs"]
CORE --> SHR["Shared Services account<br/>ACR images, KMS, DNS"]
WLF --> BUA["BU-A folder"]
WLF --> BUB["BU-B folder"]
BUA --> AD["A-Dev"]
BUA --> AS["A-Staging"]
BUA --> AP["A-Prod"]
BUB --> BP["B-Prod"]
SBX --> EXP["Experiment accounts"]
CPOL["Control Policies"] -.->|"attached to folders"| WLF
CPOL -.-> SBX
Key Services¶
| Service | Role |
|---|---|
| Resource Directory | Multi-account hierarchy: Root, folders (max 5 levels), members. Supports trusted access and delegated administrators for services |
| Agentic Cloud Governance Center | Landing zone setup, account factory with baselines, guardrails, governance maturity checks (including AI workloads since 2026) |
| RAM (Resource Access Management) | Access control within each account: users, groups, roles, policies. SAML/OIDC federation |
| CloudSSO | Workforce SSO across all RD members, like AWS IAM Identity Center. Access configurations map to RAM roles in members |
| Resource Groups | Logical grouping of resources inside an account for access and management |
| Tags | Key-value labels for cost allocation, access control, and automation |
| ActionTrail | Audit logging of API calls. Organization trails cover every RD member |
| Cloud Config | Compliance rules and configuration tracking. Aggregators span accounts |
| Control Policies | Organization-level guardrails like AWS SCPs. Effect scope is All (includes the member's root identity) or RAM (RAM users and roles only) |
Recommended Configurations¶
- Separate accounts by lifecycle stage: dev, staging, and prod go in different accounts, not different VPCs in one account.
- Centralized networking account: Owns the CEN instance, Transit Routers, and shared VPCs. Workload VPCs attach
to the TR across accounts (the attachment records a
VpcOwnerId). - Centralized logging: All accounts send ActionTrail and SLS (Simple Log Service) logs to the Log Archive account.
- Centralized security: Security Center and Cloud Firewall are managed from the Security account as a delegated administrator.
- Account factory: Use Agentic Cloud Governance Center's account factory, or Terraform
(
alicloud_governance_account/alicloud_resource_manager_account) or ROS, to create accounts with baselines (RAM roles, guardrails, logging, networking). - SSO: Use CloudSSO, or federate an external IdP (Okta, Microsoft Entra ID, and others) through SAML 2.0. Map IdP groups to RAM roles in each account.
- Cost management: Use the management account as the payer for members and use tags for cost allocation.
Cross-Account VPC Attachment Flow¶
The sequence shows how a workload VPC in a member account joins the network account's Transit Router. It is the most common cross-account operation in an RD tenancy.
sequenceDiagram
participant WL as Workload account (VPC owner)
participant RAM as RAM / CEN authorization
participant NET as Network account (CEN owner)
participant TR as Transit Router
WL->>RAM: Grant CEN instance cen-xxx permission to attach vpc-xxx
NET->>TR: CreateTransitRouterVpcAttachment with VpcOwnerId and zone vSwitches
TR->>WL: Create TR ENIs in the chosen vSwitches
NET->>TR: Associate the attachment with spoke route table
NET->>TR: Enable propagation into hub route table
TR-->>WL: Routes to hub and on-prem become reachable
Real-World Example¶
A multinational enterprise organizes 50+ Alibaba Cloud accounts under Resource Directory. The Core folder holds
the log-archive, security, and shared-network accounts. Each business unit has its own folder with
dev/staging/prod accounts. The shared-network account owns a CEN instance with Enterprise TRs in Jakarta,
Singapore, and Frankfurt. Workload accounts create VPCs and attach them to the TR after granting cross-account
authorization. Governance Center enforces baseline policies: no public ECS instances, ActionTrail enabled, and an
owner tag on every resource.
4. Multi-Zone and Multi-Region Deployment Patterns¶
Multi-Zone (Intra-Region)¶
Alibaba Cloud regions contain multiple availability zones (AZs). An AZ is a physically isolated data center with low-latency links to the others in its region. Deploying across at least two AZs protects against a single-datacenter failure. Some newer regions launched with one AZ (Mexico at launch). Check the zone count before you design for multi-AZ in a new region.
Key Services and Features¶
| Service | Multi-AZ Feature |
|---|---|
| ECS | Deploy instances across AZs. Use an Auto Scaling group with multi-AZ policy |
| ALB / NLB / CLB | Distribute traffic across AZs. Health checks remove unhealthy instances |
| RDS | High-availability edition with a standby in another zone and automatic failover |
| PolarDB | Cluster with one primary and multiple read-only nodes (up to 15 per the 2026-04 note) across AZs |
| OSS | LRS (locally redundant, the default) or ZRS (zone-redundant across AZs in supported regions). Choose ZRS explicitly for AZ-level resilience |
| Tair (Redis OSS-compatible) | Standard (dual-replica) or cluster edition with multi-zone deployment |
Recommended Configurations¶
- Use at least two AZs for any production workload.
- Load balancer listeners should include backend servers in both AZs.
- RDS: use the high-availability edition with the standby in a second zone, and test the failover window for your engine and edition before you rely on it.
- ECS Auto Scaling group: set
MultiAZPolicytoBALANCEorCOST_OPTIMIZED(the default isPRIORITY).
Multi-Region¶
Deploy across geographically separated regions for disaster recovery, compliance, or latency. Examples: Jakarta + Singapore for Southeast Asia, or Shanghai + Singapore for China plus APAC.
Key Services and Features¶
| Service | Multi-Region Feature |
|---|---|
| CEN | Connects VPCs across regions through Transit Router inter-region connections |
| Global Accelerator (GA) | Anycast/accelerated IPs that carry user traffic over Alibaba's backbone to the nearest healthy endpoint |
| Alibaba Cloud DNS | Geo- and ISP-line-based DNS routing (similar to Route 53) |
| GTM (Global Traffic Manager) | Health-check-based DNS failover between regions |
| DTS (Data Transmission Service) | Real-time one-way or two-way data synchronization (RDS, PolarDB, MongoDB, Redis, and others) between regions |
| OSS Cross-Region Replication | Asynchronous object replication between buckets in different regions |
| PolarDB GDN | Global Database Network: physical replication from a primary cluster to secondary clusters in other regions |
Recommended Configurations¶
- Use CEN for private inter-region connectivity. Choose a bandwidth plan or pay-by-data-transfer.
- Use GTM or Alibaba Cloud DNS for DNS-based failover with health checks. Keep DNS TTLs low.
- Use DTS for database replication. Two-way sync (
sync_architecture = bidirectional) needs a conflict-handling policy. - Use OSS CRR for object storage replication (async, eventual consistency).
- Deploy stateless application layers to simplify failover. Store session state in Tair (Redis OSS-compatible), which offers a global distributed cache option.
Real-World Example¶
An e-commerce platform runs in China (Shanghai) as primary and China (Beijing) as secondary. CEN connects the two regions with a 5 Gbit/s bandwidth plan. DTS replicates RDS MySQL from Shanghai to Beijing with near-real-time lag. OSS CRR replicates product images. GTM monitors health endpoints in both regions. If Shanghai becomes unhealthy, GTM switches DNS to Beijing, and clients follow within the DNS TTL. The application tier is stateless (ECS + Auto Scaling), so Beijing can scale from a minimal warm standby to full capacity in minutes.
5. Disaster Recovery (DR)¶
DR Tiers¶
| Tier | RPO | RTO | Pattern | Cost |
|---|---|---|---|---|
| Level 1: Data backup only | Hours-Days | Days | OSS backup, Cloud Backup (formerly HBR) | Lowest |
| Level 2: Cold standby | Hours | Hours | Infra defined in IaC. Data replicated. No running compute | Low |
| Level 3: Warm standby | Minutes | Minutes | Minimal compute in DR region. Scale on failover | Medium |
| Level 4: Active-passive | Seconds | Seconds-minutes | Full stack in both regions. Only primary serves traffic | High |
| Level 5: Active-active | ~0 | Seconds | Both regions serve traffic simultaneously. Bidirectional replication | Highest |
Active-Passive DR¶
The diagram shows the steady state: Region A serves all traffic while DTS and OSS CRR keep Region B current.
flowchart LR
GTM["GTM health checks<br/>DNS -> Region A"]
subgraph A["Region A (primary)"]
LBA["ALB"] --> ECSA["ECS (full)"]
ECSA --> RDSA["RDS primary"]
ECSA --> TA["Tair primary"]
OSSA["OSS bucket"]
end
subgraph B["Region B (standby)"]
LBB["ALB"] --> ECSB["ECS (full or scaled down)"]
ECSB --> RDSB["RDS (DTS target, read-only)"]
ECSB --> TB["Tair replica"]
OSSB["OSS bucket"]
end
GTM --> LBA
GTM -.->|"on failure"| LBB
RDSA -->|"DTS sync"| RDSB
OSSA -->|"CRR"| OSSB
TA -->|"replication"| TB
The failover sequence shows who acts and in what order when Region A fails.
sequenceDiagram
participant GTM as GTM
participant A as Region A endpoint
participant DTS as DTS task
participant B as Region B stack
participant Ops as Operator or runbook
GTM->>A: Health check fails N times
GTM->>GTM: Switch address pool to Region B
Ops->>DTS: Stop sync task A to B
Ops->>B: Promote RDS target to read-write, scale ECS
GTM-->>B: New client sessions land on Region B after TTL
Ops->>DTS: Later create reverse sync B to A for failback
- DTS replicates data from Region A to Region B in near-real time.
- Region B runs the full stack (or a scaled-down copy) but receives no user traffic while GTM points at Region A.
- On failure, GTM detects the unhealthy Region A endpoint and updates DNS to Region B. You promote the database in Region B. DTS does not do this automatically.
Active-Active (Multi-Region)¶
The diagram shows both regions taking writes, with replication in both directions.
flowchart LR
GTM["GTM geo routing<br/>per-region health checks"]
subgraph A["Region A (active)"]
LBA["ALB"] --> APPA["ECS / ACK"]
APPA --> DBA["Database (writes)"]
APPA --> CA["Tair"]
end
subgraph B["Region B (active)"]
LBB["ALB"] --> APPB["ECS / ACK"]
APPB --> DBB["Database (writes)"]
APPB --> CB["Tair"]
end
GTM --> LBA
GTM --> LBB
DBA <-->|"DTS two-way sync<br/>conflict policy"| DBB
CA <-->|"global distributed cache"| CB
- Two-way logical replication: Allows writes in both regions. It needs a conflict-handling policy and ideally data partitioning (each user or tenant "homed" in one region) so real conflicts are rare. DTS provides this for supported engines. The 2026-04 note named PolarDB-X "DRC" as the engine for this. DRC (Data Replication Center) is Alibaba's internal replication platform behind its own unitized active-active setup. It is not sold as a separate cloud product: DTS is the public service that grew out of it (VLDB 2024 paper on DTS, checked 2026-09-27). DTS two-way sync is therefore the documented option.
- PolarDB GDN: For active-passive with near-zero RPO. It replicates at the storage/redo layer (physical replication), with cross-region lag Alibaba describes as typically under 2 seconds. Only the primary cluster accepts writes. Secondary clusters can forward writes to the primary.
- Tair global distributed cache: Bidirectional data synchronization between cache instances in different regions.
- GTM geo-based routing: Splits traffic by geography, for example APAC users to Singapore and US users to Virginia. Per-region health checks enable automatic failover.
Hybrid DR (Cloud + On-Premises)¶
- Cloud Backup (formerly Hybrid Backup Recovery, HBR): Backs up on-premises and cloud data (files, databases, VMs, NAS) to Alibaba Cloud.
- Express Connect / VPN Gateway: Private connectivity between the on-premises DC and Alibaba Cloud VPCs.
- Smart Access Gateway (SAG): Alibaba's SD-WAN appliance for branch offices. Parts of the family are being retired: SAG App, Cloud Intelligent Branch and SmartAG Maintenance premium reached end of marketing on 2026-03-17, and end of full support on 2026-09-13. End of service is 2027-03-12 (notice). Confirm with Alibaba before designing new branch connectivity around SAG.
- DTS: Can replicate between on-premises databases and cloud RDS.
Recommended Configurations¶
- Define RTO/RPO targets first. They decide the DR tier and cost.
- Use Infrastructure as Code (Terraform or ROS) to define the full stack. In a cold or warm standby, you can recreate the DR region from code.
- Test failover regularly (quarterly at minimum). Rehearse GTM switchover and database promotion in a non-prod environment or a controlled window.
- Stateless app tier: Makes failover much simpler. Keep all state in managed services (Tair, RDS, OSS).
- Separate DR automation from production: Use separate Terraform state / ROS stacks for the DR region so they do not share a single point of failure.
6. DMZ Patterns¶
Architecture Summary¶
The DMZ (demilitarized zone) pattern on Alibaba Cloud uses subnet tiering inside a VPC to build layered defense. Public-facing resources sit in a public vSwitch (the DMZ). Application and data tiers sit in private vSwitches with no direct Internet exposure.
The diagram shows the three tiers and which controls sit at each boundary.
flowchart TB
INET["Internet"] --> CFW["Cloud Firewall<br/>Internet border, IPS block mode"]
subgraph VPC["VPC 10.0.0.0/16"]
subgraph DMZ["Public vSwitch - DMZ tier"]
WAF["WAF"]
ALB["ALB"]
NATG["NAT Gateway (SNAT)"]
BAS["Bastionhost / VPN Gateway"]
end
subgraph APP["Private vSwitch - app tier"]
ASG["ECS Auto Scaling group"]
ECI["ECI / ACK pods"]
end
subgraph DATA["Private vSwitch - data tier"]
DB["RDS / PolarDB, no public IP"]
CACHE["Tair, no public IP"]
OSSEP["OSS via VPC endpoint"]
end
WAF --> ALB
ALB -->|"SG: 8080 only"| ASG
ALB --> ECI
ASG -->|"SG: 3306 only"| DB
ASG --> CACHE
ASG --> OSSEP
ASG -->|"egress"| NATG
BAS -->|"SSH/RDP audited"| ASG
end
CFW --> WAF
NATG --> CFW
Key Services¶
| Service | Role in DMZ |
|---|---|
| Cloud Firewall | Boundaries: Internet border (N/S on EIP/SLB), NAT border (egress), VPC border (E/W between VPCs), internal firewall (E/W between ECS). Includes IPS with threat intelligence |
| NAT Gateway | Internet NAT Gateway provides SNAT (outbound) and DNAT (inbound port forwarding) for private resources. VPC NAT Gateway translates between overlapping private CIDRs |
| ALB / NLB / CLB | Sits in the DMZ and distributes traffic to private ECS. ALB is Layer 7, NLB Layer 4, CLB legacy L4/L7 |
| WAF (Web Application Firewall) | Inspects HTTP/HTTPS traffic for OWASP Top 10, bots, and custom rules. Placed in front of ALB/CLB |
| Bastionhost | Managed jump server for audited SSH/RDP access to private instances. Integrates with RAM |
| VPN Gateway | IPsec-VPN or SSL-VPN for site-to-site or client-to-site private access |
| Security Center | Unified threat detection, vulnerability scanning, compliance checking |
Recommended Configurations¶
-
Network segmentation: Three tiers of vSwitches: public (DMZ), private (app), private (data). Route tables and security groups enforce DMZ -> app and app -> data. There is no direct DMZ -> data path.
-
Cloud Firewall Internet border: Enable on all EIPs, public load balancers, and NAT Gateways. Configure allow-list rules. Set IPS to block mode, not just alert. Cloud Firewall inspects layers 3/4 and layer 7.
-
Cloud Firewall VPC border: If you use CEN or VPC peering, enable the VPC border firewall. Default-deny, with explicit allow rules between VPCs.
-
Cloud Firewall internal firewall: For micro-segmentation between ECS instances within a VPC. Group instances by role (
web,app,db) and write policies per group. -
NAT Gateway placement: Deploy in the DMZ vSwitch. Configure SNAT entries for each private vSwitch that needs Internet egress. Avoid DNAT for inbound traffic where possible and prefer a load balancer.
-
Load balancer placement: Internet-facing ALB/NLB in the DMZ vSwitches, backend servers in private vSwitches. Configure health checks and send access logs to SLS.
-
Bastionhost / VPN Gateway: Place in the DMZ vSwitch. Never assign public IPs to app or DB instances. Use RAM roles and MFA for admin access.
-
Defense-in-depth layers:
| Layer | Control |
|---|---|
| Edge | Anti-DDoS + Cloud Firewall (Internet border): DDoS, IPS, access control |
| Perimeter | WAF: OWASP, bot management, custom rules for HTTP/HTTPS |
| Network | Security Groups (stateful, per-instance) + Network ACLs (stateless, per-subnet) |
| Host | Security Center: vulnerability scanning, baseline checks |
| Application | Application-level auth (RAM, OAuth, JWT) |
Access Control Policy Matrix (Example)¶
| Source | Destination | Protocol | Ports | Action |
|---|---|---|---|---|
| Internet | ALB (DMZ) | TCP | 443 | Allow |
| Internet | Any | Any | Any | Deny |
| ALB (DMZ) | ECS App (private) | TCP | 8080 | Allow |
| ECS App | RDS (data) | TCP | 3306 | Allow |
| Bastionhost | All ECS | TCP | 22 | Allow |
| Private subnet | NAT GW (DMZ) | TCP | 80, 443 | Allow (SNAT egress) |
| All others | All others | Any | Any | Deny |
Real-World Example¶
A fintech company's production VPC in Singapore uses three tiers. The DMZ vSwitch hosts WAF, an ALB, a NAT Gateway, and a VPN Gateway. The app vSwitch runs ECS instances in an Auto Scaling group behind the ALB. The data vSwitch hosts PolarDB for MySQL and Tair. Cloud Firewall Internet border is enabled on the ALB EIP and NAT Gateway with IPS in block mode. The internal firewall enforces web->app on port 8080 and app->db on port 3306, and denies all other cross-tier traffic. All admin access goes through the VPN Gateway to Bastionhost.
7. Landing Zone Best Practices¶
Architecture Summary¶
A landing zone is a pre-configured, governed multi-account environment that gives every workload a secure baseline. Alibaba Cloud's implementation is Agentic Cloud Governance Center (renamed from Cloud Governance Center on 2026-06-24). It uses an IaC service to orchestrate Resource Directory, RAM, CloudSSO, ActionTrail, and Cloud Config. It then provides a setup wizard, an account factory with reusable baselines, and guardrails. The 2026 "agentic" release added AI Governance Maturity Checks for AI workloads (agent platforms, agent runtimes, AI gateways) across five pillars: security, reliability, cost, efficiency, and performance.
The diagram shows the landing zone components and which account owns each.
flowchart TB
subgraph MGMT["Management account"]
CGC["Agentic Cloud Governance Center<br/>setup wizard, account factory, guardrails"]
RD["Resource Directory"]
CSSO["CloudSSO"]
BILL["Consolidated billing"]
CGC --> RD
CGC --> CSSO
end
subgraph CORE["Core folder"]
LOGA["Log Archive<br/>ActionTrail org trail -> SLS + OSS"]
SECA["Security<br/>Security Center, Cloud Firewall, Config aggregator"]
NETA["Network<br/>CEN, TRs, shared VPCs, PrivateZone DNS"]
SHRA["Shared Services<br/>ACR, KMS, images, certificates"]
end
subgraph WORK["Workload folders"]
W1["BU-A dev / staging / prod"]
W2["BU-B dev / staging / prod"]
SB["Sandbox (time-limited)"]
end
RD --> CORE
RD --> WORK
CGC -->|"baseline: RAM roles, trail,<br/>Config rules, VPC attach"| WORK
W1 -->|"TR attachment"| NETA
W2 -->|"TR attachment"| NETA
W1 -.->|"logs"| LOGA
W2 -.->|"logs"| LOGA
SECA -.->|"delegated admin"| WORK
Account Factory Flow¶
When the account factory vends a new account, the steps run in the order below. Terraform can drive the same flow
with alicloud_governance_baseline and alicloud_governance_account (provider v1.228.0+).
sequenceDiagram
participant Req as Requester or CI pipeline
participant CGC as Governance Center account factory
participant RD as Resource Directory
participant NEW as New member account
participant NET as Network account
Req->>CGC: Enroll account with baseline_id and folder_id
CGC->>RD: Create resource account in target folder
RD-->>CGC: Member created (account ID)
CGC->>NEW: Apply baseline items: password policy, RAM security, services
CGC->>NEW: Create CloudSSO access roles and ActionTrail delivery
CGC->>NEW: Enable Cloud Config rules and guardrails
CGC->>NET: Create VPC and attach to Transit Router if baseline includes networking
CGC-->>Req: Account ready, drift checked continuously
Key Services¶
| Service | Role in Landing Zone |
|---|---|
| Resource Directory | Account hierarchy: Root, folders, members. Trusted access and delegated administration for services |
| Agentic Cloud Governance Center | Landing zone setup wizard, account factory, baselines, guardrails, compliance dashboards, AI governance checks |
| CloudSSO / RAM | Centralized identity: SSO to all members, external IdP federation (SAML 2.0, SCIM provisioning) |
| ActionTrail | API-level audit logging for all accounts. Deliver the organization trail to the Log Archive account's SLS project or OSS bucket |
| Cloud Config | Compliance rules (for example, "no public ECS," "all resources tagged," "encryption required"). Evaluates resources continuously |
| Control Policies | Organization-level restrictions on what member accounts can do (like AWS SCPs). Applied at folder or member level |
| CEN (in Network Account) | Centralized network connectivity for all workload accounts. Transit Router with custom route tables |
| Cloud Firewall (in Security Account) | Centralized firewall management across accounts: Internet, NAT, VPC border, and internal firewall |
| ROS / Terraform | IaC for reproducible landing zone deployment. Governance Center itself orchestrates through an IaC service |
| KMS (Key Management Service) | Centralized key management. Keep keys per workload account or share from a security account per policy |
Recommended Configurations¶
-
Start with Governance Center: Use the built-in landing zone setup wizard to create the core accounts (management, log archive, security, shared services) with recommended baselines. It is faster and less error-prone than manual setup.
-
Account factory: Configure baselines so new accounts are provisioned with:
- RAM roles for admin and operator access (through CloudSSO)
- ActionTrail delivery to the Log Archive account
- Cloud Config rules (baseline compliance)
- Network attachment to the central CEN Transit Router
-
Resource tags (owner, environment, cost-center)
-
Guardrails (Control Policies + Config rules):
- Deny public ECS instances in prod folders
- Require encryption for OSS and RDS
- Deny creation of resources outside approved regions
- Require
ownerandenvironmenttags on all resources -
Protect the landing zone itself: deny changes to
ResourceDirectoryAccountAccessRoleand the trail -
Centralized networking:
- The Network account owns the CEN instance and Transit Routers
- Shared VPCs (through Resource Sharing) for common services
- Workload VPCs attach to the TR from their own accounts
-
Cloud Firewall VPC border enabled on TR attachments
-
Centralized logging and monitoring:
- All accounts covered by one ActionTrail organization trail
- SLS centralized in the Log Archive account
-
Cloud Config aggregation in the Security account
-
Identity and access:
- Federate an external IdP (Okta, Microsoft Entra ID) into CloudSSO through SAML, with SCIM for user sync
- Map IdP groups to access configurations (RAM roles) in each account
- Enforce MFA for all human users
-
Use RAM roles and STS (not long-lived AccessKeys) for cross-account and CI access
-
IaC for the landing zone itself:
- Define the landing zone in Terraform or ROS
- Version-control the configuration
-
Apply changes through CI/CD with a plan-and-apply pipeline
-
Regular drift detection:
- Governance Center evaluates the environment and records evaluation history
- Configure alerts for non-compliant resources
- Review Control Policies periodically
Real-World Example¶
A large retail enterprise builds its landing zone with Governance Center. The management account holds Resource Directory with three top-level folders: Core, BusinessUnits, and Sandbox. Core contains the log-archive, security, and network accounts. BusinessUnits has a folder per department (Online, Stores, SupplyChain), each with dev/staging/prod accounts. The account factory provisions new accounts with baseline RAM roles, ActionTrail delivery, Cloud Config rules, and a CEN TR attachment in minutes. Control Policies enforce no public ECS in prod, encryption for all storage, and resources only in approved regions (Jakarta, Singapore, Frankfurt). The security account manages Cloud Firewall centrally across all accounts.
8. Security and Identity¶
Identity and access management, network security, data protection, and compliance considerations for Alibaba Cloud. The matching commands are in How-to Guides. The product fact tables are in Reference.
Identity & Access¶
RAM (Resource Access Management)¶
RAM is Alibaba Cloud's IAM service. It supports users, groups, roles, and fine-grained policies inside one Alibaba Cloud account. Resource Directory and CloudSSO extend it across accounts.
| Concept | Description |
|---|---|
| RAM User | Long-lived identity with console login and/or AccessKey pair. The per-account limit is a quota: read the current default on the RAM quotas page or in Quota Center (see Reference) |
| RAM Group | Collection of RAM users. Policies attached to a group apply to all members |
| RAM Role | Virtual identity assumed by trusted entities (RAM users, Alibaba Cloud services, OIDC/SAML IdPs). Issues temporary STS credentials |
| RAM Policy | JSON document defining allowed/denied actions, resources, and conditions. System policies are managed by Alibaba Cloud. Custom policies are yours |
STS (Security Token Service)¶
STS issues temporary credentials (AccessKeyId, AccessKeySecret, SecurityToken). The duration runs from 900 seconds up to the role's maximum session duration: default 3,600 s, configurable up to 43,200 s. Prefer STS over long-lived AccessKey pairs for workloads. On ECS use instance RAM roles. On ACK use RRSA (OIDC). In CI use OIDC federation.
SSO Federation¶
Alibaba Cloud supports three SSO models:
- User-based SSO (RAM): Maps external IdP users to RAM users through SAML 2.0. Each IdP user needs a matching RAM user.
- Role-based SSO (RAM): Maps IdP users to RAM roles, so no RAM users are needed. The SAML assertion names the role to assume.
- CloudSSO: The multi-account option. A directory at the RD management account, synced from the IdP (SAML + SCIM). Access configurations become RAM roles in each member. This is the recommended enterprise approach for Resource Directory tenancies.
The sequence shows role-based federation into a member account, the path most enterprise sign-ins take.
sequenceDiagram
participant U as Engineer
participant IdP as External IdP (Okta or Entra ID)
participant SSO as CloudSSO portal
participant STS as STS
participant M as Member account console/API
U->>IdP: Sign in with MFA
IdP->>SSO: SAML assertion with group claims
SSO->>SSO: Match groups to access configurations
U->>SSO: Pick account A-Prod and role Operator
SSO->>STS: AssumeRole for the provisioned RAM role
STS-->>U: Temporary credentials (max session duration)
U->>M: Call APIs, audited by ActionTrail
Multi-account SSO
For organizations on Resource Directory, configure SSO once at the management account (CloudSSO) and assume roles in member accounts. This avoids creating RAM users in every member account.
Network Security¶
Security Groups¶
Security groups are stateful, instance-level firewalls. Rules have a priority from 1 to 100, where a lower number means higher priority and the default is 1. Check Quota Center for the security-group quotas that apply to your account.
Best practices:
- One security group per application tier (web, app, db)
- Reference security group IDs in rules instead of IP ranges where possible
- Default-deny: basic security groups deny all inbound traffic by default
- Limit outbound traffic to known destinations for sensitive workloads
Network ACLs¶
Network ACLs are stateless, subnet-level packet filters. They supplement security groups for defense-in-depth. Rules are evaluated in order and the first match wins.
Cloud Firewall¶
Cloud Firewall inspects traffic at the Internet border, NAT border, VPC border, and inside the VPC (internal firewall). See Reference: Cloud Firewall Boundaries. It integrates with Alibaba Cloud threat intelligence for IP reputation scoring and automatic blocking of known malicious sources.
Anti-DDoS¶
Basic protection is on by default for every public IP. Paid tiers split by where the protected resource sits: mainland China (Anti-DDoS Pro) or outside it (Anti-DDoS Premium). See Reference: Anti-DDoS Family.
WAF (Web Application Firewall)¶
WAF sits in front of ALB/CLB to inspect HTTP/HTTPS traffic. It provides OWASP Top 10 protection, bot management, CC (HTTP flood) attack defense, and custom rules. In CNAME (reverse-proxy) mode, DNS points to the WAF CNAME, which forwards clean traffic to the origin. Cloud-native (load-balancer-integrated) mode avoids the DNS change.
Data Protection¶
KMS (Key Management Service)¶
KMS centralizes key management for encryption at rest. Keys are either customer-managed or Alibaba Cloud-managed service keys. Feature facts are in Reference: KMS Features.
TDE (Transparent Data Encryption)¶
RDS and PolarDB support TDE, which encrypts data files at rest without application changes. With a
customer-managed KMS key, RDS also needs a RAM role (RoleARN) that lets it use the key.
ActionTrail Audit Logging¶
ActionTrail records API calls made to Alibaba Cloud services. Logs include the caller identity, timestamp, source IP, request parameters, and response. For multi-account setups, create an organization trail from the management account. It captures API calls from every member in Resource Directory. A new trail does not log until you start it.
Compliance¶
China Cybersecurity Law¶
The Cybersecurity Law of the People's Republic of China (in force since 2017-06-01) requires:
- Data localization: Operators of critical information infrastructure (CII) must store personal information and important data collected in China domestically. The Data Security Law (2021) and PIPL (2021) extend cross-border controls to other data handlers. Cross-border transfer requires a CAC security assessment, a standard contract, or certification, depending on volume and data type.
- Network security obligations: Network operators must implement security protections, retain network logs for at least 6 months, and maintain incident response plans.
- Critical Information Infrastructure (CII): CII operators face additional requirements, including annual security reviews and procurement security reviews.
- 2026 amendment: An amendment raising penalties took effect on 2026-01-01. Check its details with counsel.
MLPS 2.0 (Multi-Level Protection Scheme)¶
MLPS 2.0 (GB/T 22239-2019) is China's mandatory information-security grading system, with levels 1-5. Alibaba Cloud provides compliance packages mainly for levels 2 and 3:
| Level | Scope | Alibaba Cloud Support |
|---|---|---|
| Level 2 | General business systems | Security Center baseline checks, Cloud Firewall, ActionTrail |
| Level 3 | Important business systems | All Level 2 + HSM-backed KMS, WAF, Anti-DDoS Pro, Bastionhost |
| Level 4 | Very important systems | No Alibaba Level 4 package found in the public sources checked (2026-09-27). Discuss with Alibaba; finance-cloud and government-cloud regions are the usual starting point |
Data Residency¶
Alibaba Cloud runs separate sites, accounts, and legal entities for mainland China (aliyun.com) and international
regions (alibabacloud.com). Data in Chinese regions stays in China unless you configure a transfer. Cross-border
replication (for example DTS or OSS CRR from cn-* to international regions) must be set up explicitly and is
subject to the rules above for personal and important data.
Available Certifications¶
See Reference: Compliance Facts for the certification list. Verify current scope per region in the Alibaba Cloud Trust Center before relying on it.
Sources¶
- Resource Directory overview
- Control Policy overview
- Agentic Cloud Governance Center introduction and Set up a landing zone
- Product renamed to Agentic Cloud Governance Center
- What is CEN
- What is Express Connect
- What is NAT Gateway
- Terraform provider resource docs:
cen_transit_router_peer_attachment,vpc_peer_connection,nat_gateway,ram_role,governance_account,governance_baseline,resource_manager_control_policy,oss_bucket,dts_synchronization_instance