0%
Cloud Architecture Patterns
Compute Patterns
Storage and Databases
Application Patterns
Infrastructure as Code
Reliability and Operations
Advanced Patterns
Regions, Availability Zones, and Edge
Every major cloud provider organizes its infrastructure around a simple hierarchy: regions contain availability zones, and availability zones contain the actual servers running your workloads. Understanding this hierarchy is not optional. It determines your application's latency, fault tolerance, compliance posture, and monthly bill. Every architectural decision you make, from database placement to cache strategy, flows from where you put your infrastructure on the map.
This lesson covers the physical infrastructure of the cloud: what regions and availability zones are, how they provide fault isolation, why data residency laws constrain your region choices, and how edge locations bring content closer to your users. By the end, you will understand the standard single-region multi-AZ architecture that serves as the starting point for most production systems, and know exactly when (and when not) to evolve beyond it.
What Is a Region?
A region is a geographic area where a cloud provider operates a cluster of data centers. AWS has regions like us-east-1 (Northern Virginia), eu-west-1 (Ireland), and ap-southeast-1 (Singapore). Google Cloud and Azure follow the same model with different naming conventions, but the concept is identical.
Regions are fully independent from each other. Each region has its own power grid connections, networking backbone, and operational staff. A catastrophic failure in us-east-1, a fire, a power grid collapse, a severed fiber bundle, does not affect eu-west-1. This independence is the foundation of disaster recovery. If your application runs in a single region and that region goes down, your application goes down with it. If your application runs in two regions, one region's failure becomes a traffic routing problem instead of an outage.
The number of regions continues to grow. AWS launched with a single US region in 2006. By 2026, it operates 30+ regions worldwide. Azure and Google Cloud follow similar trajectories. New regions open primarily in response to demand and regulation: when a country passes data sovereignty laws requiring local storage, cloud providers build a region there. This expansion means the geographic options available to you today are significantly broader than even five years ago.
Each cloud provider also offers specialized region types. AWS has GovCloud regions (isolated for US government workloads), Local Zones (single-AZ extensions into metro areas for ultra-low latency), and Wavelength Zones (embedded in 5G networks for mobile edge computing). These specialized offerings blur the line between the clean region/AZ/edge hierarchy, but the core concepts remain the same: proximity reduces latency, isolation provides fault tolerance, and geography determines compliance.
The distance between regions creates latency. A request from us-east-1 to eu-west-1 travels roughly 5,500 km across the Atlantic. At the speed of light in fiber, that round trip takes approximately 55ms. Add routing hops, protocol overhead, and serialization, and you land at 70-120ms per cross-region call. This is physics, not something you can optimize away with faster servers.
Choosing a Region
Region selection is one of the first infrastructure decisions you make, and changing it later is painful. Three factors drive the choice:
Proximity to users: If 80% of your users are on the US East Coast, us-east-1 puts your servers within 20-30ms of them. Choosing ap-southeast-1 instead would add 200ms to every request. For latency-sensitive applications (real-time collaboration, gaming, trading platforms), this difference is unacceptable.
Service availability: Not all cloud services are available in every region. Newer services often launch in us-east-1 and eu-west-1 first, then roll out to other regions over months. If you need a specific managed service (SageMaker endpoints, a particular instance type, a specific database engine version), verify it exists in your target region before committing.
Pricing: Identical resources cost different amounts in different regions. An m5.large instance in us-east-1 might cost $0.096/hr while the same instance in ap-northeast-1 (Tokyo) costs $0.124/hr, a 29% premium. Data transfer costs also vary by region. For cost-sensitive workloads without strict geographic requirements, us-east-1 or us-west-2 are typically the cheapest options.
A fourth factor is sometimes overlooked: disaster proximity. If your primary and disaster recovery regions are both on the same tectonic fault line or in the same hurricane zone, a natural disaster could hit both. Pairing us-east-1 (Virginia) with us-west-2 (Oregon) provides geographic diversity. Pairing us-east-1 with us-east-2 (Ohio), while better than nothing, keeps both regions in the eastern US where a large-scale grid failure could theoretically affect both. Think about what you are protecting against when choosing your DR region.
What Is an Availability Zone?
An availability zone (AZ) is one or more discrete data centers within a region, each with independent power, cooling, and networking. A typical region has three AZs, though some have as many as six. Each AZ sits in a separate physical facility, often kilometers apart, connected to the other AZs in the region through dedicated, high-bandwidth, low-latency fiber links.
The key design constraint: AZs are close enough for low-latency synchronous replication (under 2ms round trip) but far enough apart that a localized disaster, a flood, a fire, or a power substation failure, cannot take out more than one AZ simultaneously. This is the sweet spot that makes multi-AZ architectures practical: you get fault isolation without paying the latency cost of cross-region replication.
AZ Naming and Mapping
One subtle but important detail: AZ names (us-east-1a, us-east-1b, us-east-1c) do not map to the same physical facilities across AWS accounts. AWS randomizes the mapping to prevent everyone from defaulting to "AZ a" and overloading one facility. Your us-east-1a might be the same physical datacenter as another account's us-east-1c. If you need to coordinate AZ placement across accounts (for example, ensuring low-latency communication between services in different accounts), use AZ IDs (use1-az1, use1-az2) which are consistent across accounts.
This randomization also means that benchmarks or latency measurements from one account's AZ layout do not directly apply to another account. Always measure your own inter-AZ latency in your own account.
How Regions Connect to Each Other
Cloud providers operate private backbone networks that connect their regions. AWS has its own global fiber network (AWS Global Accelerator uses this backbone). Google Cloud has its own submarine cables. This private backbone is separate from the public internet and provides more consistent latency, lower packet loss, and higher throughput than traffic routed through public internet exchange points.
When your application in us-east-1 communicates with your database replica in eu-west-1, the traffic travels over this private backbone, not the public internet. This matters because public internet routing is unpredictable: traffic between two US cities might route through a European exchange point due to peering agreements, adding 100ms of unexpected latency. Private backbone routing is deterministic and optimized for the shortest path.
Services like AWS Global Accelerator and Google Cloud's Premium Tier networking let your end users' traffic enter the cloud provider's backbone at the nearest edge location rather than traversing the public internet to the destination region. This reduces the "last mile" variability and provides more consistent performance for latency-sensitive applications.
The practical difference is measurable. A user in Australia connecting to us-east-1 over the public internet might see latencies ranging from 180-350ms depending on the time of day and internet congestion. The same user connecting via Global Accelerator sees a consistent 160-180ms because the traffic enters the AWS backbone at the Sydney edge and travels the optimized private path.
For most applications, the public internet path is acceptable. Global Accelerator matters for latency-sensitive workloads like gaming, trading, or video conferencing where consistency matters as much as raw speed. The cost of Global Accelerator ($0.025/hr plus data transfer premiums) needs to be justified by a measurable improvement in user experience or SLA compliance.
Global Accelerator also provides static anycast IP addresses, which simplifies DNS configuration and supports instant failover between regions without waiting for DNS TTL propagation.
Latency at Each Level
These latency numbers should be internalized. They drive every placement decision:
Same AZ: Under 1ms round trip. Your application server and its database replica in the same AZ communicate essentially instantly. This is where you want your hot data path.
Cross-AZ (same region): 1-2ms round trip. Synchronous replication between AZs is feasible at this latency. A database write that synchronously replicates to a second AZ adds only 1-2ms to write latency, a small price for surviving a datacenter failure.
Cross-region: 50-200ms round trip depending on geographic distance. Same continent (us-east-1 to us-west-2) is around 60-80ms. Cross-ocean (us-east-1 to ap-southeast-1) is 150-200ms. Synchronous replication at these latencies is impractical for most workloads. This is why cross-region replication is almost always asynchronous.
These numbers become critical when you consider request chains. If a single user request triggers 3 sequential database queries, same-AZ placement costs 3ms total. Cross-AZ placement costs 6ms. Cross-region placement at 100ms per query costs 300ms, which is the difference between a snappy interface and a sluggish one. This is why you co-locate your application servers with their primary database in the same AZ whenever possible, and accept the cross-AZ cost only for replicas used for fault tolerance.
How to Measure Latency Yourself
Do not rely on published latency numbers. Measure from your own infrastructure. Run a simple ping or TCP latency test between instances in different AZs and regions in your account.
Same-AZ latency is straightforward: launch two instances in the same AZ and measure round trip time. You should see sub-millisecond results.
Cross-AZ latency: launch instances in different AZs within the same region. Expect 1-2ms. If you see 5ms+, investigate whether the instances are in the same VPC and subnet configuration, as routing through NAT or transit gateways adds overhead.
Cross-region latency: launch instances in two different regions and measure. The results depend on geographic distance. Keep a reference table of your measured latencies for architecture decisions. These measurements should be part of your infrastructure documentation and updated whenever you add a new region to your deployment.
Here are typical measured values for common region pairs:
us-east-1 to us-west-2 (Virginia to Oregon): 60-80ms
us-east-1 to eu-west-1 (Virginia to Ireland): 70-90ms
us-east-1 to ap-southeast-1 (Virginia to Singapore): 180-220ms
eu-west-1 to ap-southeast-1 (Ireland to Singapore): 150-180ms
These values assume the AWS private backbone. Public internet routing would add 20-50% more latency with higher variance.
In system design interviews, when someone asks how you would handle a datacenter failure, start with multi-AZ in a single region. Cross-region failover is expensive and complex. Most production systems achieve 99.99% availability with three AZs in one region. Only add a second region when regulatory requirements or global user distribution demands it.