Microservices Communication

Topics Covered

Synchronous vs Asynchronous Communication

Synchronous Communication

Asynchronous Communication

When Async Becomes Essential

Service Discovery

DNS-Based Discovery

Dedicated Service Registries

Client-Side vs Server-Side Discovery

Health Checking and Instance Management

Circuit Breakers and Retries

The Circuit Breaker Pattern

Why Circuit Breakers Prevent Cascading Failures

Retry Strategies

Implementation Considerations

The Saga Pattern

How Compensating Transactions Work

Choreography vs Orchestration

Saga Failure Scenarios

Idempotency in Distributed Systems

Idempotency Keys

Designing Idempotent Operations

At-Least-Once Plus Idempotent Equals Effectively Exactly-Once

Idempotency at Scale

Every time one microservice needs data or action from another, you face a fundamental choice: should the caller wait for a response, or should it fire a message and move on? This is the synchronous versus asynchronous decision, and it shapes your system's latency profile, failure modes, and scalability ceiling.

The same call made synchronously and then through a queue, with what happens in each when the callee is down.

Synchronous Communication

In synchronous communication, Service A sends a request to Service B and blocks until it gets a response. HTTP REST and gRPC are the dominant protocols here. The caller knows immediately whether the operation succeeded or failed, which makes control flow straightforward: call, check status, proceed or handle error.

The problem is temporal coupling. Both services must be running and reachable at the exact same moment. If Service B is down, overloaded, or slow, Service A is stuck waiting. In a chain of five synchronous calls (A calls B calls C calls D calls E), one slow service causes all upstream services to hold threads open, consuming resources while producing nothing. This is how a single slow database query in Service E takes down your entire platform.

 
1Service A  ──HTTP──>  Service B  ──gRPC──>  Service C
2   │ (blocked)           │ (blocked)            │
3   │ waiting...          │ waiting...            │ processing...
4   │                     │ <─── response ────────│
5   │ <─── response ──────│
6   │ continues

Synchronous calls work well when you need an immediate answer and the called service is fast and reliable. Reading a user profile from an auth service during login is a good example: you cannot proceed without the result, and the call should complete in single-digit milliseconds.

Asynchronous Communication

In asynchronous communication, Service A publishes a message to a broker (SQS, Kafka, RabbitMQ, Google Pub/Sub) and immediately continues its work. Service B picks up the message later and processes it independently. The two services never interact directly.

This achieves temporal decoupling. Service A does not know or care whether Service B is running right now. The broker stores the message until Service B is ready. If Service B goes down for maintenance, messages queue up and get processed when it returns. If Service B is slow, the queue absorbs the backpressure instead of propagating it upstream.

 
Service A  ──publish──>  [Message Queue]  ──consume──>  Service B
   │ (returns immediately)      │ (stores message)          │
   │ continues work             │                           │ processes later

The tradeoff is complexity. You lose the simple request-response model. Service A does not know if Service B succeeded or failed unless you build a feedback mechanism (a response queue, a status polling endpoint, or a callback webhook). Debugging is harder because there is no single request trace: you must correlate messages across multiple systems using distributed tracing.

Interview Tip

A practical rule of thumb: use synchronous calls when the caller needs the result to continue its own operation (fetching a price to calculate a total). Use asynchronous messaging when the caller does not need the result immediately (sending a notification email after an order is placed). Most production systems use both patterns, choosing per-interaction based on the coupling and latency requirements.

When Async Becomes Essential

Three scenarios make asynchronous communication the clear winner:

Spike absorption: A flash sale sends 100x normal traffic to the order service. With synchronous calls, the payment service gets 100x load simultaneously and likely crashes. With a message queue, the payment service processes messages at its own pace. The queue acts as a buffer, smoothing the spike over minutes instead of crashing in seconds.

Fan-out to multiple consumers: When an order is placed, you need to notify inventory, shipping, analytics, and email services. With synchronous calls, the order service makes four HTTP calls and waits for all four responses. With a Pub/Sub topic, the order service publishes once and each consumer subscribes independently. Adding a new consumer (say, a fraud detection service) requires zero changes to the order service.

Cross-region communication: Services in different regions have high network latency. Synchronous calls across regions add 100-300ms per hop. Asynchronous replication through a distributed queue (Kafka with cross-region mirroring) lets each region work with local data and synchronize in the background.