0%
Cloud Architecture Patterns
Cloud Foundations
Compute Patterns
Storage and Databases
Infrastructure as Code
Reliability and Operations
Advanced Patterns
Event-Driven Architecture
In a traditional request-response system, Service A calls Service B and waits for a reply. The caller knows exactly who it is talking to, what it expects back, and when the work is done. This creates tight coupling: Service A cannot function if Service B is down, slow, or has changed its interface.
Event-driven architecture flips this relationship. Instead of telling another service what to do, a service announces what just happened. "Order 456 was placed." "Payment for order 456 succeeded." "Inventory for SKU-789 dropped below threshold." These announcements are events. The producing service does not know or care who listens. It publishes the event and moves on.
Events vs Commands
This distinction is fundamental and frequently confused. An event describes something that already happened. A command instructs someone to do something. "OrderPlaced" is an event. "ProcessPayment" is a command. The difference matters because events are facts that cannot be rejected. The order was placed, that is a historical record. A command can fail, be refused, or be rerouted. When you model your system around events, you build a log of immutable facts. When you model around commands, you build a chain of requests that couples every link together.
Events use past tense: OrderPlaced, PaymentProcessed, InventoryReserved. Commands use imperative: ProcessPayment, ReserveInventory, SendNotification. If you find yourself naming events in imperative form, you are accidentally building a command bus, which recreates the tight coupling you were trying to escape.
There is a third concept that gets confused with both: a query. "GetOrderStatus" is neither an event nor a command. It is a request for data that does not change state. Events are state changes that already occurred. Commands are requests to change state. Queries are requests to read state. Keeping these three concepts separate (CQRS and event-driven architecture share this principle) leads to cleaner system boundaries.
Why Event-Driven Systems Scale
The producing service does not wait for consumers. It publishes the event and returns immediately. If the notification service takes 3 seconds to send an email, that delay does not slow down the order API response. This temporal decoupling is the primary scaling advantage. The order service can handle 10,000 orders per second regardless of how fast or slow downstream consumers are.
Consumers also scale independently. If the analytics pipeline falls behind, you add more consumer instances. If the notification service is idle, you scale it down. Each consumer has its own throughput bottleneck, and you address each one independently rather than scaling the entire system together.
Eventual Consistency: The Price of Decoupling
Event-driven systems are eventually consistent by design. When the order service publishes OrderPlaced, the inventory service does not reduce stock instantly. There is a delay: the event sits in the broker, the consumer picks it up, processes it, and updates its local state. During that window, the inventory count is stale.
This is not a bug. It is a deliberate trade-off for decoupling and scalability. The alternative (synchronous call to inventory, waiting for confirmation) gives you instant consistency but reintroduces coupling. In most business domains, eventual consistency is acceptable. A user seeing "in stock" for 500ms after the last item is purchased is a minor UX issue. A payment service being unavailable because the inventory service is down is a major business issue.
The key is knowing where eventual consistency is acceptable and where it is not. Payment amounts must be consistent (you cannot charge a stale price). Stock counts can be eventually consistent (overselling is handled by backorders). Design your event boundaries around these consistency requirements.
To manage eventual consistency in practice, track propagation delay. Measure the time between when a producer publishes an event and when each consumer finishes processing it. If the payment service consistently processes events within 200ms but the analytics service lags by 30 seconds, that tells you which consumers need scaling and which consistency windows your UI must communicate to users.
In interviews, when asked why microservices use event-driven communication, the answer is not just performance. It is about deployment independence. When Service A publishes events instead of calling Service B directly, you can deploy, scale, and restart Service B without touching Service A. This operational independence is often more valuable than the latency improvement.
Fan-Out: One Event, Many Reactions
When an order is placed, the payment service needs to charge the customer, the inventory service needs to reserve stock, and the notification service needs to send a confirmation email. In a synchronous system, the order service would call all three sequentially or in parallel, knowing about each one and handling each failure individually.
With fan-out, the order service publishes a single OrderPlaced event. The message broker delivers a copy to every subscriber. Each consumer processes its copy independently at its own pace. The payment service might process in 200ms while the analytics service takes 5 seconds. Neither blocks the other. Adding a new consumer (say, a fraud detection service) requires zero changes to the order service. You simply subscribe the new service to the topic. Removing a consumer is equally simple: unsubscribe and the order service never notices.
Fan-out is different from load balancing. In fan-out, every subscriber gets every event (one-to-many). In load balancing, each event goes to exactly one consumer in a group (one-to-one within the group). Kafka consumer groups implement load balancing: events within a partition go to one consumer in the group. But different consumer groups each receive all events, which is fan-out. Understanding this distinction prevents a common design mistake: using consumer groups when you need fan-out, or using separate topics when you need load balancing.
This is the architectural superpower of event-driven systems. The number of consumers can grow from 3 to 30 without modifying the producer. Each new business requirement (analytics, compliance logging, ML feature extraction) adds a subscriber, not a code change to the core service.
Event Delivery Guarantees
Message brokers offer different delivery guarantees, and the choice affects your entire system design.
At-most-once: The event is delivered zero or one time. If the consumer crashes before acknowledgment, the event is lost. Fast and simple, but only acceptable for non-critical events (analytics pings, debug logging).
At-least-once: The event is delivered one or more times. The broker retries until acknowledged. This means duplicates are possible. Most production systems use at-least-once because losing events is worse than processing duplicates, but consumers must be idempotent to handle the duplicates correctly.
Exactly-once: The event is delivered and processed exactly one time. This is the hardest guarantee to achieve and usually requires coordination between the broker and the consumer's storage (e.g., Kafka transactions with consumer offsets committed atomically with processing results). True exactly-once is expensive and often unnecessary if consumers are idempotent.
In practice, at-least-once delivery with idempotent consumers is the standard pattern. It gives you the reliability of guaranteed delivery without the complexity and performance cost of exactly-once coordination. The producer publishes, the broker guarantees delivery (retrying if needed), and the consumer handles duplicates gracefully. This pattern is used by nearly every large-scale event-driven system in production, from Uber's trip processing to LinkedIn's activity feeds.