0%
Cloud Architecture Patterns
Cloud Foundations
Compute Patterns
Storage and Databases
Infrastructure as Code
Reliability and Operations
Advanced Patterns
API Gateway Patterns
An API gateway is a single entry point that sits between your clients and your backend services. Every request from a mobile app, web browser, or third-party integration flows through the gateway before reaching any internal service. This centralization is the point. Without a gateway, each client must know the address of every backend service, handle authentication independently, and implement its own retry and timeout logic. The gateway absorbs all of that cross-cutting complexity into one place.
What the Gateway Does
The gateway handles responsibilities that do not belong in any single backend service:
Routing: The gateway examines the incoming request path, method, and headers, then forwards it to the correct backend service. A request to /api/users/123 goes to the user service. A request to /api/orders goes to the order service. Clients never see internal service addresses.
Authentication: The gateway validates credentials (JWT tokens, API keys, OAuth tokens) before the request reaches any service. Invalid requests are rejected at the edge, saving backend resources. Backend services receive pre-validated identity information in headers, never raw tokens.
Rate limiting: The gateway tracks request counts per user, per API key, or globally. When a client exceeds its quota, the gateway returns 429 Too Many Requests without forwarding anything to the backend. This protects backend services from traffic spikes and abusive clients.
Request/response transformation: The gateway can rewrite requests before forwarding and responses before returning. This includes adding headers, converting between data formats (JSON to Protobuf), filtering sensitive fields from responses, and aggregating responses from multiple backend services into one client response.
Observability: The gateway logs every request, records latency metrics, and traces request flow across services. Because all traffic passes through a single point, you get a complete picture of your API's behavior without instrumenting each service individually.
Caching: The gateway can cache responses from read-heavy endpoints, serving them directly without forwarding to the backend. A product catalog that changes once per hour does not need every request hitting the database. The gateway stores the response, serves it for the cache TTL, and only forwards a request when the cache expires or is explicitly invalidated.
Logging and auditing: Every request that passes through the gateway can be logged with a correlation ID, timestamp, client identity, and response status. This creates a centralized audit trail for security compliance (who accessed what, when) and debugging (tracing a failed request through the system). Without a gateway, each service logs independently, and correlating a request across five services requires matching timestamps and building custom log aggregation.
In system design interviews, introduce the API gateway early in your architecture. It signals that you understand cross-cutting concerns like authentication, rate limiting, and routing belong in infrastructure, not in business logic. Interviewers look for this separation because it reflects how production systems actually work at scale.
Managed Gateway Services
You rarely build a gateway from scratch. Managed services handle the operational burden:
AWS API Gateway integrates tightly with Lambda, IAM, and CloudWatch. It handles scaling automatically and charges per request. Best for AWS-native architectures where Lambda functions are the backend.
Google Apigee focuses on API management and developer portals. It provides analytics, monetization, and developer onboarding features. Best for companies that sell APIs as products.
Kong is open-source and runs anywhere: cloud, on-premise, or hybrid. It uses a plugin architecture for extensibility. Best for teams that need control over gateway behavior and run multi-cloud or on-premise infrastructure.
Azure API Management provides a full lifecycle management solution with policies for transformation, validation, and caching. It integrates with Azure Active Directory for identity management. Best for Azure-native architectures.
The choice between managed and self-hosted depends on three factors: cloud lock-in tolerance (managed services tie you to one vendor), customization needs (self-hosted gives you full control), and operational capacity (managed services eliminate infrastructure management). A team of five engineers running on a single cloud provider should almost always choose the managed option. A platform team at a large company running multi-cloud infrastructure has the capacity and motivation to operate Kong or Envoy.
Gateway vs. Load Balancer
A load balancer distributes traffic across instances of the same service. A gateway routes traffic to different services based on the request. They operate at different levels: the load balancer asks "which instance should handle this?" while the gateway asks "which service should handle this?" In production, traffic flows through the gateway first (service routing), then through a load balancer (instance selection). They complement each other, they do not replace each other.
Single Point of Failure Concerns
The gateway sits on the critical path for every request. If it goes down, nothing works. This means the gateway itself must be highly available. In practice, you deploy multiple gateway instances behind a network load balancer (layer 4). The network load balancer is simple (TCP/UDP forwarding), stateless, and extremely reliable. Each gateway instance is identical and stateless: no session data, no in-memory caches that would be lost on restart. If one instance fails, the network load balancer routes traffic to the remaining instances.
Health checks monitor each gateway instance. The network load balancer pings a /health endpoint every few seconds and removes unresponsive instances from rotation. Auto-scaling adds instances when CPU or request count exceeds thresholds. This setup gives you the benefits of centralization (one place for auth, rate limiting, routing) without the fragility of a single machine.
Response Caching at the Gateway
The gateway can cache responses from backend services, serving repeated requests without forwarding them. For read-heavy endpoints like product catalogs or configuration data, this dramatically reduces backend load. The gateway respects Cache-Control headers from backend responses, stores the response with the appropriate TTL, and serves it directly for subsequent identical requests.
Cache keys typically combine the URL path, query parameters, and relevant headers (like Accept-Language for localized content). The gateway must be careful not to cache responses that contain user-specific data unless the cache key includes the user identity. A cached response for user A served to user B is a security incident.
Cache invalidation strategies at the gateway mirror those in any caching system. TTL-based: cache entries expire after a fixed time. Simple but serves stale data until expiry. Event-based: the backend publishes an event when data changes, and the gateway evicts the affected cache entry. More complex but ensures freshness. Hybrid: short TTL combined with event-based purging for critical data. The TTL catches cases where the event is lost.
Request Transformation in Practice
Request transformation is one of the gateway's most powerful capabilities. Consider a mobile client that sends REST JSON, but the backend order service uses gRPC with Protobuf. The gateway translates on the fly: it receives the JSON body, maps the fields to the Protobuf schema, serializes to binary, and forwards the gRPC call. The response follows the reverse path. Neither the client nor the service knows about the other's format.
Header injection is another common transformation. After validating a JWT, the gateway extracts claims (user ID, roles, tenant ID) and injects them as internal headers like X-User-Id and X-Tenant-Id. Backend services read these headers directly instead of re-parsing the JWT. This removes token validation logic from every service and ensures consistent identity extraction.
Body mapping handles differences between client expectations and backend requirements. A client might send {"firstName": "Alice"} while the backend expects {"first_name": "Alice"}. The gateway transforms field names (camelCase to snake_case), flattens nested objects, or enriches requests with data from other sources. This decouples the client-facing API contract from the internal service contract, letting each evolve independently.
Response transformation works the same way in reverse. The gateway can strip internal fields (database IDs, debug metadata) from responses before returning them to clients. It can also add pagination wrappers, HATEOAS links, or consistent error formatting that clients expect. A backend service returns a raw list; the gateway wraps it in {"data": [...], "pagination": {"next": "/api/users?page=3"}}.
Circuit Breaking at the Gateway
When a backend service starts failing, the gateway should stop sending requests to it rather than letting failures cascade. Circuit breaking implements this: the gateway tracks the error rate for each backend. When errors exceed a threshold (e.g., 50% of requests fail over 30 seconds), the circuit "opens" and the gateway returns 503 Service Unavailable immediately without forwarding requests. After a cool-down period, the circuit enters a "half-open" state where a small number of requests are forwarded to test if the backend has recovered. If those succeed, the circuit closes and normal traffic resumes.
Without circuit breaking, a failing backend accumulates requests in its queue, each timing out after 30 seconds. The gateway's connection pool fills up. Clients wait for timeout after timeout. The failure spreads from one backend to every client. Circuit breaking cuts this chain: the gateway fails fast (milliseconds, not seconds), clients get an immediate error they can handle, and the failing backend gets breathing room to recover.
Circuit breaker configuration requires tuning three parameters: the failure threshold (what percentage of requests must fail to open the circuit), the cool-down period (how long to wait before testing recovery), and the probe count (how many test requests to send in the half-open state). Too aggressive thresholds (circuit opens after one failure) cause false positives during normal transient errors. Too lenient thresholds (circuit opens after 90% failure) let cascade failures propagate too long. A good starting point is 50% failure rate over a 30-second window, with a 60-second cool-down and 3 probe requests.
Backend for Frontend (BFF) Pattern
A common extension of the gateway is the Backend for Frontend pattern. Different clients (mobile app, web browser, smart TV) need different views of the same data. A mobile app needs a compact response with small image URLs. A web dashboard needs detailed analytics with charts. Instead of building one API that serves all clients poorly, you create client-specific gateway layers.
Each BFF is a thin service that sits behind the main gateway and in front of the backend services. It aggregates data from multiple backends, formats responses for its specific client, and handles client-specific concerns (pagination format, field selection, image resizing). The main gateway handles cross-cutting concerns (auth, rate limiting), and each BFF handles client-specific orchestration.
This avoids overloading the main gateway with business logic while giving each client team autonomy over their API shape. The mobile team can add a new BFF endpoint without affecting the web team's BFF or the main gateway.
The BFF pattern is particularly valuable in microservice architectures where a single screen in the mobile app might require data from five different services (user profile, recent orders, recommendations, notifications, account balance). Without a BFF, the mobile app makes five separate API calls over the cellular network, each with its own latency. With a BFF, the mobile app makes one call to the BFF, which fans out to five services over the fast internal network and returns a single merged response. This reduces round trips from five to one, cutting load time significantly on slow networks.
Logging and Correlation IDs
The gateway is the best place to generate a correlation ID for every incoming request. This is a unique identifier (UUID) that travels with the request through every backend service. The gateway adds it as a header (X-Correlation-Id), and every service includes it in its logs, metrics, and error responses.
When a user reports that their request failed, you search logs for the correlation ID and instantly see the entire request path: which services it touched, where it failed, what each service returned. Without correlation IDs, debugging a request that traversed five services means searching logs by timestamp and hoping the clocks are synchronized.
The gateway also logs the overall request duration (client to gateway and back), which captures the total latency the user experiences. This is different from what individual services measure (just their processing time). The gateway's latency metric includes network hops, serialization, and queue time, giving you the true end-to-end number that matters for user experience.
IP Whitelisting and Geo-Blocking
The gateway can enforce network-level access control before any authentication logic runs. IP whitelisting restricts access to known IP ranges (corporate VPNs, partner data centers), blocking all other traffic at the gateway. Geo-blocking uses IP geolocation to restrict access by country, which may be required for regulatory compliance (data sovereignty laws) or business rules (service availability by region).
These checks happen before authentication, rate limiting, or routing, making them the first line of defense. They are computationally cheap (IP lookup in a hash set or trie) and eliminate unwanted traffic before it consumes any gateway resources. For internal-only APIs, IP whitelisting combined with mTLS provides strong defense-in-depth: even if an attacker obtains a valid certificate, they cannot reach the gateway from an unauthorized network.
Geo-blocking must account for VPNs and cloud provider IP ranges. A user in a blocked country using a VPN exits through an allowed country and bypasses the geo-block. Cloud provider IPs (AWS, GCP, Azure) change frequently, so static blocklists go stale. For strict compliance, combine geo-blocking with additional identity verification rather than relying on IP-based controls alone.
Request Validation at the Gateway
The gateway can validate incoming requests against the API's schema before forwarding to the backend. If the OpenAPI specification says POST /api/users requires a name field of type string, the gateway rejects requests missing that field or sending an integer, returning 400 Bad Request immediately. This offloads input validation from every backend service, ensuring consistent error formats and preventing malformed requests from consuming backend resources.
Schema validation at the gateway works best for structural validation (required fields, data types, string patterns). Business validation (is this email already registered? does this product exist?) must still happen in the service that owns the data. The gateway validates the shape of the request; the service validates the meaning.
Performance-sensitive gateways compile the OpenAPI schema into a validation function at startup rather than interpreting the schema at runtime. Compiled validators check a request in microseconds, adding negligible latency to the request path. The gateway reloads the compiled validator when the schema changes, which happens during deployment, not during request processing.
For APIs with large request bodies (file uploads, batch operations), the gateway should validate headers and URL parameters but skip body validation for performance. The backend service validates the body after accepting the request, possibly asynchronously for very large payloads.