Load Balancing and Traffic Management

Topics Covered

Layer 4 vs Layer 7 Load Balancing

Why L4 Is Fast

L4 Routing Algorithms

Why L7 Exists

Content-Aware Routing

Sticky Sessions

Direct Server Return (DSR)

Multi-Tier Load Balancing

Application Load Balancers

TLS Termination

Routing Rules in Practice

Connection Draining and Deregistration

Header Manipulation

Request Rate Limiting and WAF Integration

Target Groups and Weighted Routing

ALB vs Nginx vs Envoy

Global Load Balancing

DNS-Based Global Load Balancing

Anycast

Multi-Region Active-Active vs Active-Passive

Global Server Load Balancing (GSLB) Beyond DNS

DNS TTL and Failover Speed

Health Checks at the DNS Level

Traffic Splitting Strategies

Canary Deployments

Blue-Green Deployments

Traffic Mirroring (Shadow Traffic)

Percentage-Based vs Header-Based Splitting

Progressive Delivery Pipeline

Feature Flags vs Traffic Splitting

Cross-Zone Load Balancing

Load Shedding

Weighted Least-Connections Algorithm

Health Checks and Failover

Health Check Types

Configuration Parameters

Shallow vs Deep Health Checks

Graceful Shutdown and Connection Draining

Detecting Failures Before Users Do

Multi-Tier Health Architecture

Circuit Breakers and Health Checks

Load balancers sit between clients and servers. Their job is deceptively simple: distribute incoming traffic across multiple backend instances so no single server drowns while others sit idle. But how a load balancer makes routing decisions depends on which layer of the network stack it operates at, and that choice has profound consequences for what you can and cannot do with your traffic.

Layer 4 (L4) load balancers operate at the transport layer. They see TCP or UDP packets. They know the source IP, destination IP, source port, and destination port. That is all. They cannot read HTTP headers, inspect URLs, or examine cookies. They forward raw TCP connections to a backend server using algorithms like round-robin or least-connections, and every packet in that connection goes to the same backend.

How far up the stack each balancer reads, with the routing that unlocks and the latency it adds per request.

Why L4 Is Fast

Because L4 load balancers never parse the application payload, they are extremely fast. They operate at near wire speed. An AWS Network Load Balancer (NLB) handles millions of requests per second with single-digit microsecond latency overhead. The load balancer does not terminate the TCP connection. It rewrites packet headers (destination IP/port) and forwards them, so the backend server sees something close to a direct connection.

This makes L4 ideal for non-HTTP protocols: database connections (MySQL on port 3306), message queues (Kafka on port 9092), gaming servers using UDP, or any custom TCP protocol. If the load balancer does not need to understand the payload, L4 is always the faster choice.

L4 Routing Algorithms

Since L4 cannot inspect request content, it relies on connection-level algorithms to distribute traffic:

Round-robin sends each new connection to the next server in sequence. Server A, then B, then C, then back to A. Simple and effective when all servers have equal capacity and all requests are roughly equal in cost.

Least-connections sends new connections to whichever server has the fewest active connections. This naturally adjusts for servers processing slow requests: a server stuck handling a 10-second database query accumulates connections, so the load balancer sends new traffic elsewhere. This is the better default for most workloads.

Source IP hash computes a hash of the client's IP address and maps it consistently to the same server. This provides a form of persistence without cookies. The downside is that a large corporate network behind one NAT IP gets all its traffic sent to one server. Adding or removing servers also changes the hash mapping for existing clients.

Weighted round-robin assigns different weights to servers based on their capacity. A server with 8 CPUs gets weight 4, a server with 2 CPUs gets weight 1. The load balancer sends 4 connections to the large server for every 1 connection to the small server. This is useful during mixed-fleet deployments where instance types differ, or when gradually introducing new hardware alongside older machines.

Why L7 Exists

Layer 7 (L7) load balancers operate at the application layer. They fully terminate the client's TCP connection, parse the HTTP request (method, URL, headers, sometimes the body), make a routing decision, and then open a separate connection to the chosen backend. This is slower than L4 because parsing HTTP adds latency, but it unlocks capabilities that L4 simply cannot offer.

Interview Tip

In interviews, default to L7 load balancing for web applications. L4 is the right answer only when you are load balancing non-HTTP traffic (databases, message brokers, game servers) or when you need absolute maximum throughput and do not need content-aware routing.

Content-Aware Routing

L7 load balancers can route based on any part of the HTTP request. This is the killer feature.

Path-based routing sends /api/* requests to your API servers, /uploads/* to your media service, and /static/* to a CDN origin. One load balancer fronts your entire application, routing traffic to specialized backend pools based on URL path.

One address in front of four backend pools, with the path rules quietly choosing between them per request.

Host-based routing inspects the Host header. api.example.com routes to your API cluster, dashboard.example.com to your frontend cluster, ws.example.com to your WebSocket servers. A single load balancer and IP address serves multiple subdomains without separate infrastructure for each.

Header-based routing enables patterns like routing requests with X-Version: 2 to a canary deployment, or sending mobile user-agents to a mobile-optimized backend. This is the foundation of A/B testing and gradual rollouts at the infrastructure level.

Query parameter routing sends requests with ?version=beta to a test backend. Combined with feature flags, this lets individual developers or QA teams test against specific backends without affecting production users.

All of these routing decisions happen at the load balancer, not in your application code. Your application does not need to know which path it was accessed from or which host header was used. The load balancer absorbs this routing complexity so your services can focus on business logic.

Sticky Sessions

L7 load balancers can pin a user to a specific backend by inspecting or injecting cookies. The first request gets routed to Server A, the load balancer sets a cookie like SERVERID=A, and subsequent requests with that cookie always go to Server A.

Sticky sessions exist because some applications store session state in server memory. But they undermine the core benefit of load balancing: even distribution. If a user with a sticky session generates heavy load, their assigned server bears it all. If that server crashes, the session is lost. Sticky sessions also prevent autoscaling from being effective, because new instances do not receive existing sessions.

The better approach is to store session state externally (Redis, DynamoDB) and make every server stateless. Then any server can handle any request, and sticky sessions become unnecessary.

Direct Server Return (DSR)

In a standard L4 setup, both the request and the response pass through the load balancer. For response-heavy workloads (video streaming, file downloads), the load balancer becomes a bandwidth bottleneck because it must relay large response bodies.

Direct Server Return solves this. The load balancer receives the client request and forwards it to the backend, but the backend responds directly to the client, bypassing the load balancer entirely. The load balancer only handles the small incoming request, not the large outgoing response. This dramatically reduces load balancer bandwidth requirements.

DSR requires the backend to be configured with the load balancer's IP as a loopback address (so it can respond to packets addressed to the VIP). It only works with L4 because the load balancer cannot modify the HTTP response if it never sees it. DSR is common in on-premises deployments but rare in cloud environments where managed load balancers handle this transparently.

Multi-Tier Load Balancing

Large-scale architectures often combine L4 and L7 in layers. An L4 load balancer (NLB) sits at the edge, handling raw TCP traffic at massive scale. Behind it, multiple L7 load balancers (ALBs or Envoy proxies) handle content-aware routing for specific services. The L4 tier distributes traffic across the L7 tier, and the L7 tier distributes traffic across application servers.

This separation is intentional. The L4 tier handles millions of connections with microsecond overhead and provides DDoS absorption (volumetric attacks get distributed across the L7 tier rather than hitting a single load balancer). The L7 tier handles thousands of connections with millisecond overhead but provides intelligent routing. Google's production load balancing architecture follows exactly this pattern: Google Front End (GFE) at L4, then Envoy proxies at L7.

In AWS terms, this looks like an NLB in front of multiple ALBs. The NLB provides a static IP address (or Elastic IP) that clients connect to, handles millions of TCP connections, and distributes them across ALB nodes in multiple AZs. Each ALB then performs content-aware routing to the appropriate target group.

This pattern is also used when you need a static IP for firewall allowlisting: ALBs do not have static IPs, but NLBs do, so placing an NLB in front gives you a fixed entry point with L7 routing behind it.

Level Expectations

Mid-level engineers know the difference between L4 and L7 and can pick the right one for a given use case. Senior engineers design systems that avoid sticky sessions entirely by externalizing state. Staff engineers evaluate the latency cost of L7 inspection at scale and decide where to place L4 vs L7 in a multi-tier load balancing architecture (e.g., L4 at the edge for raw throughput, L7 internally for content routing).