Virtual Private Clouds

Topics Covered

VPC Design Principles

Why Isolation Matters

Choosing Your CIDR Block

Multi-AZ Design

The Three-Tier Pattern

Default VPC vs Custom VPC

DNS Resolution in a VPC

CIDR Notation Explained

VPC Limits and Quotas

Infrastructure as Code for VPCs

Subnets and Routing

Public vs Private Subnets

Route Tables

Subnet Sizing

Routing Between Subnets

Elastic Network Interfaces

Subnet Design for Microservices

Route Propagation

Gateway Route Tables

Troubleshooting Connectivity

IPv4 Exhaustion and IPv6 in VPCs

NAT Gateways and Internet Access

How NAT Works

Internet Gateway vs NAT Gateway

Cost Considerations

Security Groups vs Network ACLs

High Availability for NAT Gateways

Egress-Only Internet Gateways

Flow Logs

Security Group Best Practices

NAT Gateway vs NAT Instance

Understanding Ephemeral Ports

Elastic IP Addresses and NAT

Connection Draining and NAT Timeouts

VPC Peering and Connectivity

VPC Peering

Transit Gateway

VPN and Direct Connect

VPC Endpoints

PrivateLink

Choosing the Right Connectivity Option

Cross-Region Connectivity

Shared VPCs

DNS Resolution Across VPCs

Network Firewall

Bandwidth and Throughput Considerations

Security Considerations for Peering and Transit Gateway

A Virtual Private Cloud is your own isolated section of a cloud provider's network. Think of it as renting a floor in a massive office building: you share the physical infrastructure with thousands of other tenants, but your floor is logically separate. No one else's traffic touches yours. You control the layout, the doors, and who gets keys.

The foundation of every VPC is a CIDR block. CIDR (Classless Inter-Domain Routing) defines the range of private IP addresses available inside your VPC. When you create a VPC with the block 10.0.0.0/16, you are claiming 65,536 IP addresses (10.0.0.0 through 10.0.255.255). Every resource you launch inside this VPC, whether an EC2 instance, a Lambda function endpoint, or an RDS database, gets an IP address from this range.

One /16 cut into public and private subnets across zones, with the six layers that together decide reachability.

Why Isolation Matters

Without VPCs, every cloud resource would sit on a shared flat network. One customer's misconfigured database could be reachable by another customer's compromised web server. VPCs eliminate this risk by providing network-level isolation that is enforced by the cloud provider's hypervisor. Traffic between two VPCs does not flow unless you explicitly create a connection between them.

This isolation also gives you control over IP addressing. You pick the CIDR block, which means you can align your cloud IP ranges with your on-premises network to avoid conflicts when you connect them later. A company running 10.1.0.0/16 on-premises would choose 10.2.0.0/16 for their VPC so the two networks can be linked without overlapping addresses.

Choosing Your CIDR Block

The CIDR block size is a permanent decision for the life of the VPC. You cannot shrink it later, and expanding requires adding secondary CIDR blocks, which complicates routing. The most common choice is /16 (65,536 addresses), which gives you enough room for thousands of resources across multiple subnets and availability zones.

Smaller blocks like /24 (256 addresses) work for isolated test environments but run out fast in production. Larger blocks like /8 are wasteful and reduce your ability to create non-overlapping VPCs. The sweet spot for most production workloads is /16 or /20.

Interview Tip

Always plan your CIDR blocks on paper before creating VPCs. If you have three environments (dev, staging, production) and plan to peer them, each needs a non-overlapping range. A common scheme is 10.0.0.0/16 for production, 10.1.0.0/16 for staging, 10.2.0.0/16 for development. Getting this wrong means you cannot peer VPCs later without painful re-architecting.

Multi-AZ Design

Cloud providers organize their regions into availability zones (AZs), which are physically separate data centers within the same city. A well-designed VPC spreads resources across at least two AZs so that a single data center failure does not take down your entire application.

This means your VPC's CIDR block must be large enough to subdivide into subnets across multiple AZs. A /16 block can easily be split into subnets like 10.0.1.0/24 in AZ-A and 10.0.2.0/24 in AZ-B, each with 256 addresses. The architecture mirrors how you would design a physical data center with redundant network segments, except the cloud provider handles the physical redundancy for you.

The Three-Tier Pattern

The most common VPC architecture follows a three-tier model: a public subnet for load balancers and bastion hosts, a private subnet for application servers, and another private subnet for databases. Traffic flows inward: the internet reaches the load balancer, the load balancer forwards to app servers, and app servers query databases. No tier is reachable from a tier that should not talk to it.

This pattern enforces the principle of least privilege at the network level. Your database instances never have public IP addresses. Your application servers are not directly reachable from the internet. The only entry point is through the load balancer, which gives you a single chokepoint for security monitoring, rate limiting, and TLS termination.

Default VPC vs Custom VPC

Every AWS account comes with a default VPC in each region. The default VPC has public subnets in every AZ, an internet gateway already attached, and instances launch with public IPs by default. This is convenient for learning and prototyping but dangerous for production. Resources in the default VPC are internet-facing unless you actively lock them down.

Custom VPCs start with nothing: no subnets, no internet gateway, no routes to the outside world. You build exactly what you need. This forces intentional decisions about which subnets are public, which are private, and how traffic flows. For production workloads, always create a custom VPC. Treat the default VPC as a sandbox that should eventually be deleted or at least emptied of resources.

DNS Resolution in a VPC

Every VPC has a built-in DNS server (called the Amazon DNS server or Route 53 Resolver) running at the base of the VPC CIDR block plus two. For a VPC with CIDR 10.0.0.0/16, the DNS server is at 10.0.0.2. Two VPC settings control DNS behavior: enableDnsSupport (instances can use the VPC DNS server) and enableDnsHostnames (instances with public IPs get public DNS names).

Both settings must be enabled for VPC endpoints to work via private DNS, for RDS instances to be resolvable by hostname, and for services like ECS to discover each other. These are on by default in the default VPC but off by default in custom VPCs, which catches teams who create a custom VPC and then wonder why their RDS connection strings do not resolve.

CIDR Notation Explained

Understanding CIDR notation is essential for VPC design because every subnet, route, and security rule uses it. The number after the slash indicates how many bits of the 32-bit IP address are fixed (the network portion), and the remaining bits are available for host addresses.

/16 means the first 16 bits are fixed: 10.0.x.x where x can be anything. That gives you 2^16 = 65,536 addresses. /24 means the first 24 bits are fixed: 10.0.1.x where only the last octet varies. That gives you 2^8 = 256 addresses. /28 fixes 28 bits, leaving only 4 bits for hosts: 2^4 = 16 addresses.

The practical skill is calculating how many addresses a CIDR block provides: take 32 minus the prefix length, then raise 2 to that power. A /20 gives 2^12 = 4,096 addresses. A /22 gives 2^10 = 1,024 addresses. For subnets, remember to subtract the 5 reserved addresses that AWS claims, so a /24 has 251 usable addresses, not 256.

When designing a VPC, you also need to verify that subnet CIDR blocks do not overlap and that they fit within the VPC's CIDR block. A VPC with 10.0.0.0/16 can contain 10.0.1.0/24 and 10.0.2.0/24, but cannot contain 10.1.0.0/24 because that falls outside the 10.0.x.x range.

VPC Limits and Quotas

AWS enforces default limits on VPC resources that affect architectural decisions. The default limit is 5 VPCs per region (can be raised to hundreds). Each VPC supports up to 200 subnets. Each route table supports up to 50 static routes. Security groups are limited to 60 inbound and 60 outbound rules per group, and each ENI can be assigned up to 5 security groups.

The security group rules limit is the one that bites teams most often. A microservices architecture with 30 services, each needing its own security group reference, can exceed the 60-rule limit. The workaround is to use managed prefix lists (group multiple CIDR blocks into a single reference that counts as one rule) or to restructure security groups so services that share identical access patterns share a security group.

Route table limits matter when connecting to complex on-premises networks. If the on-premises network has 100 subnets, route propagation from the VPN would try to add 100 routes. With a 50-route default limit, you need to request a quota increase or summarize on-premises routes into fewer, larger CIDR blocks before advertising them to the VPC.

Infrastructure as Code for VPCs

VPC configuration is complex and interdependent: subnets reference route tables, route tables reference gateways, security groups reference other security groups. Manual creation through the console is error-prone and not repeatable. Infrastructure as Code tools (Terraform, CloudFormation, Pulumi) are essential for VPC management.

A Terraform module for a production VPC typically defines the VPC, subnets across AZs, route tables for public and private subnets, an internet gateway, NAT gateways, security groups for each tier, NACLs, and VPC flow logs. This module can be parameterized by environment (dev, staging, prod) with different CIDR blocks and instance sizes but identical architecture. When you need to create a new environment, you instantiate the module with new parameters and get an identical network topology in minutes.

Version-controlling your VPC configuration also creates an audit trail. Every change is reviewed in a pull request before being applied. If a network change causes connectivity issues, you can diff the current configuration against the last known good state and identify exactly what changed. This is vastly better than trying to reconstruct what happened from console click history.

A key pattern is to separate the VPC infrastructure module from the application infrastructure modules. The VPC module outputs subnet IDs, security group IDs, and VPC ID. Application modules consume these outputs as inputs. This separation means the networking team can update VPC configuration without touching application code, and application teams can deploy services without modifying network infrastructure. The interface between the two is a stable set of output values that both teams agree on.

Tagging is equally important. Every VPC resource should be tagged with the environment (prod, staging, dev), the team that owns it, and the purpose (public-web, private-app, private-db). Tags enable cost allocation (which team is spending how much on NAT gateway processing), operational visibility (which subnets belong to which environment), and automated compliance checks (flag any production subnet missing a flow log configuration). Without consistent tagging, a VPC with 30+ subnets across multiple AZs becomes an opaque maze that only the person who created it can navigate.