CDN and Edge Computing

Topics Covered

CDN Architecture and Cache Behavior

How CDN Caching Works

Cache-Control Headers and TTL

Cache Keys

Cache Hit Ratio

Vary Header and Cache Fragmentation

Push vs. Pull CDNs

Content Negotiation and CDN Behavior

Origin Configuration

Origin Shield

Origin Failover

SSL/TLS Termination at the Edge

Origin Connection Optimization

Request Collapsing

Cache Invalidation Strategies

Edge Functions

What Edge Functions Can Do

Platform Comparison

Cold Starts and Edge Limitations

Edge State with KV Stores

Multi-CDN Strategies

Why Multi-CDN

Implementation Approaches

Challenges of Multi-CDN

CDN for Dynamic Content

DDoS Protection at Edge

CDN for WebSocket and Real-Time Traffic

CDN and Compliance Considerations

Measuring CDN Effectiveness

Key Metrics

When a CDN is Not Worth It

CDN Observability

CDN Cost Model

Cost Optimization Strategies

A Content Delivery Network is a geographically distributed system of servers that caches content closer to users. The core promise is simple: instead of every request traveling across the world to your origin server, most requests are served from an edge location a few milliseconds away.

The five caches sitting between a reader and your origin, with what each one of them is allowed to hold.

How CDN Caching Works

When a user in Tokyo requests an image from your application hosted in Virginia, the first request travels to the nearest CDN edge location. The edge checks its cache. On a miss, it pulls the content from the origin server, stores a copy, and serves it to the user. The next request from anyone in the Tokyo region hits the cached copy directly. This is the pull-and-cache model: no content is pushed to edges preemptively. The CDN only caches what users actually request.

The latency difference is dramatic. A cache hit at a nearby edge location returns in 5-15ms. The same request routed to an origin server across the ocean takes 150-300ms due to TCP handshake, TLS negotiation, and physical distance. For a page loading 40 assets, that difference compounds into seconds of perceived load time.

The first request for an object and then the second, with an accounting of where the time goes in each.

Cache-Control Headers and TTL

The origin server controls caching behavior through HTTP headers. The Cache-Control header is the primary mechanism:

 
Cache-Control: public, max-age=86400

This tells the CDN (and browsers) to cache the response for 86,400 seconds (24 hours). After the TTL expires, the next request triggers a revalidation with the origin. The CDN sends a conditional request with If-None-Match using the stored ETag. If the content has not changed, the origin responds with 304 Not Modified and no body, saving bandwidth and origin compute.

Key directives you need to understand:

  • public allows any cache (CDN, browser, proxy) to store the response
  • private restricts caching to the user's browser only, excluding CDNs
  • no-cache forces revalidation on every request but still allows storage
  • no-store prevents caching entirely, used for sensitive data
  • s-maxage overrides max-age specifically for shared caches like CDNs

Cache Keys

A cache key determines whether two requests map to the same cached response. By default, the cache key is the full URL including query parameters. example.com/api/products?page=1 and example.com/api/products?page=2 are different cache entries.

This matters because poorly designed cache keys destroy your hit ratio. If your URLs include unnecessary query parameters (tracking IDs, session tokens, timestamps), each unique parameter combination creates a separate cache entry for identical content. A URL like example.com/image.jpg?session=abc123&utm_source=twitter caches separately from example.com/image.jpg?session=def456&utm_source=google even though the image is identical.

The fix is configuring your CDN to strip irrelevant query parameters from the cache key. Most CDNs let you define which parameters to include, which to ignore, and whether to sort the remaining parameters (so ?a=1&b=2 matches ?b=2&a=1).

Interview Tip

In interviews, when discussing CDN caching, always mention cache keys and hit ratios together. A beautifully configured TTL means nothing if your cache keys are so specific that every request is unique. The goal is maximizing the ratio of cache hits to total requests, and cache key design is the lever most teams overlook.

Cache Hit Ratio

Cache hit ratio is the percentage of requests served from the cache versus total requests. This is the single most important metric for evaluating your CDN's effectiveness.

The same traffic at two different hit ratios, with what the origin carries and what a reader actually feels.

A 90%+ cache hit ratio is the benchmark for a well-configured CDN. At this level, only 10% of requests reach your origin, reducing origin load by 10x. A 95% hit ratio means 20x reduction. Below 80%, you are paying for CDN infrastructure while still hammering your origin with most of the traffic. At that point, the CDN is an expensive proxy, not a performance optimization.

Common reasons for low hit ratios include overly specific cache keys, TTLs that are too short, large content catalogs with long-tail access patterns (items requested once and never again), and excessive cache invalidation.

Vary Header and Cache Fragmentation

The Vary header tells the CDN to maintain separate cache entries based on specific request headers. Vary: Accept-Encoding means a gzip-compressed response and a Brotli-compressed response are cached separately, which is correct behavior. But Vary: User-Agent creates a separate cache entry for every unique user agent string. Since there are thousands of distinct user agent strings in the wild, this effectively disables caching.

A common mistake is setting Vary: * which tells the CDN that every request header matters for caching. This makes every request a cache miss because no two requests have identical headers. If you see Vary: * in your response headers, remove it immediately. It is almost never intentional and it silently destroys your cache hit ratio.

The correct approach is to vary only on headers that actually change the response content. Vary: Accept-Encoding is standard. Vary: Accept-Language makes sense if you serve different translations from the same URL. Anything beyond that should be questioned carefully.

A diagnostic technique for Vary-related cache problems: add a custom response header (e.g., X-Cache-Key) that shows the actual cache key the CDN used. When debugging why two seemingly identical requests produce cache misses, comparing their X-Cache-Key values immediately reveals which header variation caused the split. Most CDNs let you add custom headers in their configuration panel without code changes.

Push vs. Pull CDNs

Most modern CDNs use the pull model: content is fetched from the origin on the first request and cached at the edge. The alternative is a push model where you upload content directly to the CDN before any user requests it. Push CDNs give you full control over what is cached and when, but they require explicit management of every asset.

In practice, the distinction has blurred. Pull CDNs with origin shield behave almost like push CDNs for popular content because the shield maintains a warm cache. True push CDNs (like uploading to S3 with CloudFront in front) are used primarily for static assets where you want guaranteed availability from the first request, with no cold-start penalty.

Some CDNs support a hybrid approach called cache warming or pre-fetching. After a deployment, you send a list of URLs to the CDN, and it proactively fetches them from the origin before any user requests them. This eliminates the cold-cache penalty for new content without requiring you to manage a full push workflow. The CDN handles it as a batch origin fetch and populates edges ahead of traffic. Cache warming is especially valuable for product launches where you expect sudden traffic spikes to new URLs that have never been cached. Without warming, the first wave of users all experience cache misses simultaneously, creating an origin load spike exactly when traffic is highest.

Content Negotiation and CDN Behavior

Content negotiation adds complexity to CDN caching. When a browser sends Accept: image/avif, image/webp, image/* and the origin responds with the best format, the CDN must cache each format variant separately. This is handled through the Vary: Accept header, which tells the CDN that different Accept headers may produce different responses for the same URL.

The problem is that Accept headers vary widely across browsers. Chrome, Safari, and Firefox each send different Accept values. Without careful configuration, you end up with dozens of cache entries for the same image. The solution is to normalize at the edge: the CDN groups similar Accept headers into a small number of buckets (supports AVIF, supports WebP, neither) and uses the bucket as the cache key instead of the raw header value. Most modern CDNs handle this automatically for image content types.