0%
Data-Intensive Applications
Foundations of Data Systems
Distributed Data
Encoding and Evolution
Batch Processing
Stream Processing
Data Quality and Governance
The Architecture of Real-Time Analytics
A batch report that refreshes every hour tells you what happened. A real-time dashboard tells you what is happening right now. The difference between these two approaches is not just latency. It is the difference between reactive and proactive operations. When a payment service starts failing at 2:47 PM, a real-time dashboard lets your team detect and respond before customers flood the support queue. A batch report surfaces the same failure at 3:00 PM, by which time the damage is done.
Real-time analytics processes events as they arrive and surfaces metrics within seconds. The goal is not just speed for its own sake. It is about closing the gap between an event occurring and a human (or automated system) acting on it. A fraud detection system that flags suspicious transactions 45 minutes after they happen is useless. A live operations dashboard that shows server error rates from an hour ago cannot help you during an ongoing incident.
The business case for real-time analytics varies by domain. In e-commerce, real-time dashboards showing conversion rates and cart abandonment let marketing teams react to campaigns within minutes of launch. In infrastructure monitoring, sub-second alerting on error rate spikes can trigger auto-scaling or circuit breakers before users notice degradation. In financial services, real-time risk dashboards track exposure limits that regulators mandate be updated continuously. In each case, the value is not the data itself but the speed at which decisions are made from that data.
Real-Time vs Near-Real-Time
These two terms sound interchangeable but describe fundamentally different architectures. Real-time means sub-second latency through continuous streaming. Each event is processed individually as it arrives. Apache Flink, Apache Kafka Streams, and Apache Storm operate in this mode. A click event enters the pipeline and updates a counter within milliseconds.
Near-real-time means sub-minute latency through micro-batching. Events accumulate in small batches (typically 1-30 seconds) and are processed together. Apache Spark Structured Streaming operates this way by default. The latency is higher, but throughput is often better because batch processing amortizes overhead across many events. Micro-batching also simplifies exactly-once semantics: each batch is an atomic unit that either fully succeeds or fully retries, avoiding the complex checkpointing that per-event streaming requires.
The choice depends on your use case. A stock trading dashboard needs true real-time because a 10-second delay means stale prices. An e-commerce analytics dashboard showing page views per minute works perfectly with near-real-time micro-batches every 5 seconds. Choosing true real-time when near-real-time suffices wastes engineering effort and infrastructure cost for latency nobody notices.
A useful heuristic: if your dashboard's smallest time granularity is N seconds, any processing latency under N/2 seconds is invisible to the user. A per-minute dashboard (granularity 60 seconds) with 5-second micro-batch latency delivers results 55 seconds before the display interval changes. The user cannot tell whether the underlying pipeline is streaming or micro-batching.
The operational cost difference between real-time and near-real-time is significant. True streaming requires managing stateful operators, exactly-once checkpointing, and recovery from failures without data loss. Micro-batching handles most of this automatically because each batch is a discrete unit that either succeeds or retries. For teams without deep streaming expertise, near-real-time is the safer starting point.
Push-Based vs Pull-Based Dashboard Updates
Once your pipeline computes a metric, how does it reach the dashboard?
Pull-based (polling): The dashboard sends HTTP requests at regular intervals (every 1-5 seconds) asking for the latest data. Simple to implement. Works well for a handful of dashboards. Wasteful at scale because most polls return no new data, and the polling interval creates a floor on perceived latency. If you poll every 3 seconds, users see data that is 0-3 seconds stale.
Push-based (WebSocket or Server-Sent Events): The server pushes new data to the dashboard the moment it is available. Eliminates wasted polls and reduces latency to the speed of the pipeline itself. WebSockets provide full-duplex communication (the dashboard can send messages back, useful for interactive filters). Server-Sent Events (SSE) are simpler and unidirectional, suitable when the dashboard only receives updates.
Start with Server-Sent Events for read-only dashboards. SSE uses plain HTTP, works through proxies and load balancers without special configuration, reconnects automatically, and supports event IDs for resuming after disconnection. Switch to WebSockets only when the dashboard needs to send messages to the server, such as interactive filtering or drill-down queries.
For production dashboards at scale, push-based delivery is standard. A Grafana dashboard backed by Prometheus uses a pull model (Prometheus scrapes targets), but modern real-time dashboards backed by streaming pipelines use WebSocket connections to push metric updates to hundreds or thousands of browser sessions simultaneously. The server maintains a subscription for each connected dashboard and fans out updates as they arrive.
The fan-out challenge is real. If 500 operations engineers have the same dashboard open and your pipeline produces 100 metric updates per second, the server must deliver 50,000 messages per second. A pub/sub layer (Redis Pub/Sub, NATS, or a dedicated WebSocket gateway) between the pipeline and the dashboards decouples the computation from the delivery and handles fan-out efficiently.
Backpressure and Degradation
What happens when the dashboard cannot keep up with the update rate? A WebSocket connection to a browser tab that is in the background (or on a slow network) falls behind. If the server buffers updates for slow clients, memory grows unbounded. If the server drops updates, the dashboard shows stale data.
The standard approach is a combination of buffering with limits and snapshot recovery. The server buffers up to N messages per client (typically 100-1,000 messages). If the client falls behind and the buffer fills, the server drops intermediate updates and sends only the latest value for each metric. When the client catches up, it receives the most recent state rather than a replay of every update it missed. This is acceptable because dashboards care about the current value, not the history of every intermediate value. The history lives in the OLAP engine for anyone who needs it.
Some systems take this further with "last-value caching" at the gateway level. The gateway maintains the most recent value for every metric in memory. When a new client connects (or reconnects after a drop), it immediately receives the full current state as a snapshot rather than waiting for the next update cycle. This eliminates the blank-screen period that users experience when opening a dashboard for the first time.