0%
Applied AI Systems
Retrieval and Knowledge
Agents and Reasoning
Evaluation, Interpretability, and Safety
Production Research Engineering
Multi-Agent Systems and Case Studies
Production Observability for LLMs
The first time an LLM system goes wrong in production, you will reach for the logs. And the first thing you will discover is that your web-framework's default request log tells you almost nothing. "POST /v1/chat returned 200 in 4.2s" does not help you debug why the model just hallucinated a refund policy that your company does not offer. You need logs built for LLM-shaped systems, not logs built for CRUD apps.
The asymmetry that makes LLM logging hard is simple: too much logging is expensive (token payloads are big, and tokens cost money to re-process during analysis), too little logging makes debugging impossible (if you did not capture the exact prompt and response, you cannot reproduce the failure). This section is about finding the right middle.
The Minimum Viable Log Record
Every LLM request should produce a structured log record with at least the following fields. Think of this as the LLM equivalent of the Apache Common Log Format. If you capture less than this, debugging production failures becomes guesswork.
request_id— a UUID propagated through every downstream call. Without this, you cannot stitch traces together.user_idandsession_id— who triggered the call, inside which conversation.timestamp— ISO 8601, with timezone. Use UTC.model— the exact model identifier, including version.gpt-4o-2024-11-20, notgpt-4o. Model upgrades silently change behavior, and if you log only the family name you cannot tell whether a regression coincided with a provider rollout.prompt— the full rendered prompt string sent to the model, after template substitution. If you only log the template, you lose the actual input values and cannot reproduce.response— the full model output, including any stop reason or tool calls.input_tokens,output_tokens— counts, for cost accounting.latency_ms— end-to-end wall clock from request arrival to response completion.time_to_first_token_ms— separately from total latency. For streaming UIs this matters more than total latency.status— success, model_error, timeout, content_filter_block, rate_limited, etc. Do not collapse all errors into HTTP 500.cost_usd— computed at log time from token counts and the model's price card. Do not compute this from logs later; the price card drifts and you will lose historical accuracy.
That is the core. Everything else is optional but often valuable.
Logging Retrieval and Tool Calls
If your system uses RAG, the retrieved documents are part of the input. You cannot debug a RAG hallucination without knowing which chunks the retriever returned. Log:
retrieval_query— the query sent to the retriever (which may differ from the user's raw input if you do query rewriting).retrieved_doc_ids— stable identifiers, not the raw text. Dedupe against a document store for the actual content.retrieval_scores— the similarity scores for the top-k results. Low scores on the winning document is a useful leading indicator of bad context.num_docs_returned— if you hit zero results, that is a specific failure mode.
For tool-calling agents, log every tool invocation as its own record:
tool_name,tool_args,tool_result,tool_latency_ms,tool_error.
A useful mental model: think of tool calls and retrievals as child spans of the parent LLM request, not as opaque side effects. If the parent request fails, you want to be able to walk down into any of its children without going to a separate system.
What Not to Log
Logging the full prompt and full response is expensive in three ways.
- Storage. A typical chat response is 500-5000 tokens. Logging 1M requests a day at an average of 2000 tokens per record, that is 2 billion tokens per day, roughly 10 GB of raw text. Over a month you are paying for 300 GB of storage per service.
- Privacy. User prompts often contain PII (names, addresses, medical details) and sometimes credentials. If you ship these logs to a third-party observability backend, you have created a data-exfiltration surface. Redact before logging, not after.
- Replay cost. If you later want to re-run an old request through a new model, you will want to keep the prompt. But if you keep only a hash of the prompt, you lose that replay capability.
The pragmatic compromise many teams adopt: log the full prompt and response for a sampled fraction of traffic (say 1-10%), and log only metadata (tokens, latency, status) for the rest. This keeps debug capability for a representative sample while keeping storage and privacy risk bounded. For high-severity errors (5xx, content filter trips), always log the full payload regardless of sampling.
A common mistake is to log the prompt and response as JSON-escaped strings inside a single big log line. This works until you try to grep them. Log them as separate typed fields in a structured backend like Loki, BigQuery, or ClickHouse. You will thank yourself the first time you need to find every request where the response contained the phrase 'I don't have access to that information'.
Sampling and Privacy
Sampling is not optional at scale, but naive sampling loses exactly the requests you care about. Three patterns work well.
Head-based sampling: decide at request time whether to log this request in full, with some fixed probability. Fast and simple, but you might miss the one request per thousand that failed catastrophically.
Tail-based sampling: keep the full payload in an in-memory buffer, and at response time decide whether to persist based on outcome. Always persist errors, slow requests, and content filter trips; sample the rest. This is more expensive to implement but catches more interesting traffic.
Privacy-aware redaction: run a PII scrubber over the prompt before logging. Replace detected names, emails, phone numbers, and SSNs with typed placeholders ([PERSON], [EMAIL]). This preserves the prompt structure for debugging while removing the sensitive payloads. Presidio and similar libraries do this reasonably well out of the box. Just be aware that a scrubber is never perfect, so treat the redacted logs as still-sensitive data and protect accordingly.
Concrete Example
Here is a minimal Python logging shim that captures the core fields for a single LLM call.
This is 30 lines of code and it gets you 80% of the debuggability you need. Start here, then graduate to OpenTelemetry (next section) when you want spans, propagation, and cross-service stitching.
Pick your log schema before you ship to production. Retrofitting a new field into a year of historical logs is painful, so over-specify on day one. The ten fields listed above are the minimum I would not compromise on, even for an MVP.