Observability in the Cloud

Topics Covered

The Three Pillars

Metrics: Quantitative Signals

Logs: Qualitative Context

Traces: Causal Connections

How the Pillars Connect

From Monitoring to Observability

Metrics Collection and Dashboards

Metric Types

The RED Method

The USE Method

Dashboard Design

Collection Infrastructure

Metric Naming and Labels

Distributed Tracing

Traces and Spans

Context Propagation

Span Attributes and Events

Tracing Tools and Backends

Reading a Trace Waterfall

Sampling Strategies for Traces

Log Aggregation and Analysis

Structured Logging

Log Levels

Correlation IDs

Log Aggregation Systems

Log Retention and Cost

Alerting and On-Call

Alert on Symptoms, Not Causes

SLI/SLO-Based Alerting

Avoiding Alert Fatigue

On-Call Practices

Observability Cost Management

Observability is not monitoring. Monitoring tells you when something breaks. Observability lets you ask arbitrary questions about your system's internal state using its external outputs. In a distributed cloud environment where requests traverse dozens of services, you cannot predict every failure mode in advance. You need the ability to investigate novel problems without deploying new instrumentation. That ability rests on three pillars: metrics, logs, and traces.

Metrics, logs and traces, with the question each of the three answers and the one field that joins them.

Metrics: Quantitative Signals

Metrics are numeric measurements collected at regular intervals. They answer "how much" and "how fast." A counter tracking total HTTP requests tells you throughput. A gauge showing current memory usage tells you resource consumption. A histogram capturing request latency tells you the distribution of response times, not just the average.

Metrics are cheap to collect, cheap to store, and fast to query. A time-series database like Prometheus stores millions of data points per second and evaluates queries across weeks of data in milliseconds. This makes metrics ideal for dashboards and alerting: you can watch 500 metrics in real time without significant infrastructure cost.

But metrics are aggregates. They tell you that error rate jumped to 5% at 14:32. They do not tell you which users were affected, what the error messages said, or which downstream service caused the failures. For that, you need the other two pillars.

Logs: Qualitative Context

Logs are timestamped text records of discrete events. Where metrics tell you "error rate is 5%," logs tell you "at 14:32:07, user 4821 received a 500 error on POST /api/checkout because the payment service returned a connection timeout after 30 seconds." Logs carry the detail that metrics compress away.

The challenge with logs is volume. A single service processing 1,000 requests per second at four log lines per request produces 4,000 log entries per second. Across 50 services, that is 200,000 log lines per second, roughly 17 billion per day. Without structure and aggregation, this volume is unusable. You need structured logging (JSON with consistent fields) and a centralized log aggregation system to make logs queryable at scale.

Logs also carry risk. Accidentally logging sensitive data (credit card numbers, passwords, API keys, personal health information) creates compliance violations and security exposure. Build log sanitization into your logging framework: automatically redact known sensitive field patterns before the log entry is emitted. This is cheaper than scanning billions of stored log entries after the fact.

Traces: Causal Connections

Traces follow a single request as it moves through multiple services. A trace is a tree of spans, where each span represents one operation (an HTTP call, a database query, a cache lookup). The trace connects these spans with parent-child relationships, showing you the exact path a request took and how long each step consumed.

Traces answer the question metrics and logs cannot: "why is this specific request slow?" A metric shows p99 latency is 2 seconds. A log shows timeout errors from the payment service. A trace shows that the slow request hit the cache (2ms), called the user service (15ms), called the inventory service (18ms), then called the payment service which waited on a database query that took 1,800ms because it triggered a full table scan.

Interview Tip

In interviews, when discussing observability, start with the three pillars and explain how they complement each other. Metrics detect problems (alerting), logs explain problems (context), and traces locate problems (causality). No single pillar is sufficient. A mature observability stack integrates all three, typically by correlating them through a shared trace ID.

How the Pillars Connect

The real power comes from correlation. When a metric alert fires (error rate spike), you filter logs by the same time window and service to read the error messages. Then you pick a specific error log entry, extract its trace ID, and pull up the distributed trace to see the full request path. This workflow, metric to log to trace, is the standard investigation pattern.

The glue that connects them is the trace ID: a unique identifier generated at the edge when a request enters your system, propagated through every service call, and attached to every log entry and metric label along the way. Without this correlation ID, each pillar is an isolated data silo that you must manually cross-reference by timestamp, which is error-prone and slow.

From Monitoring to Observability

Traditional monitoring assumes you know what can go wrong. You define dashboards for known failure modes, set thresholds, and wait for alerts. This works for monoliths where the failure space is bounded: the database is slow, the disk is full, the process crashed.

In distributed cloud systems, the failure space is combinatorial. Twenty services with 10 possible states each create trillions of system configurations. You cannot pre-define dashboards for every combination. Observability shifts the model from "define what to watch" to "collect everything and query on demand." You instrument exhaustively (metrics, logs, traces), store cheaply (columnar storage, object storage, sampling), and explore interactively when something goes wrong. The investment is in instrumentation breadth and query tooling, not in predicting every failure mode.

OpenTelemetry has emerged as the standard instrumentation framework, providing vendor-neutral SDKs that emit metrics, logs, and traces in a unified format. Instrumenting once with OpenTelemetry lets you send telemetry to any backend (Prometheus, Jaeger, Loki, Datadog, New Relic) without re-instrumenting when you switch vendors. The OpenTelemetry Collector acts as a pipeline between your services and your backends, handling batching, sampling, and routing. You can send the same telemetry to multiple backends simultaneously (e.g., traces to Jaeger for the engineering team and to Datadog for the operations team) without changing application code.