The GenAI Field Guide

Observability

See what happened in each request, why it failed, and what it cost.

Observability is the ability to answer questions about a running system from the data it emits, including questions nobody thought to ask in advance. For ordinary services that means logs, metrics and traces. An LLM application needs all three plus something extra, because its worst failures do not raise errors. A request that returns HTTP 200 in 900 ms can still contain a fabricated refund policy, a citation to the wrong document, or an agent that called the same search tool eleven times. Without the prompt, the retrieved context, the tool calls and the output recorded together, nobody can explain what happened or prove that a fix worked.

The mental model is a trace per request: a tree of timed spans for each model call, retrieval, tool call and validation step, each carrying the version of the prompt and model, the token counts and the outcome. Aggregate those traces and you get the operational numbers (latency percentiles, time to first token, cost per task, error rates). Score a sample of them and you get the quality numbers (groundedness, retrieval relevance, user feedback). The same traces then feed the eval set, so production failures become regression tests.

The questions build in that order. The Basic questions define traces, spans, what to log, tokens, time to first token and latency percentiles. The Advanced questions cover tracing agents, monitoring retrieval, detecting hallucinations, closing the loop with evals, OpenTelemetry, dashboards, and three practical additions: keeping sensitive data safe inside traces, using user feedback as a signal, and alerting on failures that never throw an exception.

What is LLM observability?

LLM observability is collecting the traces, metrics, logs and quality signals needed to explain what an AI application did on any request, why it failed, how long it took and what it cost.

Classic monitoring asks whether a service is up and fast. LLM observability adds a third question: was the answer any good? That matters because language model failures are usually silent. The API returns a well-formed response, the status code is 200, and the content is wrong, unsupported by the sources, off-policy or in the wrong format. An error-rate graph will never show it.

In practice it means capturing, for each request, the full chain of work: the user input, the assembled prompt (or a reference to its template and version), the model and parameters, any retrieved documents, every tool call with arguments and results, the final output, token usage, timing, and any validation or guardrail outcome. These are stored as traces so an engineer can open one request and replay its story step by step.

On top of the per-request record sit aggregates: latency percentiles, time to first token, cost per task, tool failure rates, retrieval hit rates, and quality scores from automated graders, sampled human review and user feedback. Each of these can be sliced by prompt version, model, customer segment or feature, which is how you find that a regression only affects long documents or one tenant.

The trade-off is cost and risk. Full prompts and outputs are large, so storage grows fast, and they often contain personal or confidential data. Mature setups keep full payloads for a limited window or a sample, keep structured metadata for everything, and restrict who can read the raw text.

What is a trace?

A trace is the complete, linked record of everything one request did: every model call, retrieval, tool call and check, with timing and outcome, arranged as a tree under a single trace id.

When a user asks a question, the application may rewrite the query, search a vector store, rerank results, call the model, validate the output and maybe retry. A trace ties all of those steps together under one identifier so you can see the whole journey instead of scattered log lines. Each step inside the trace is a span, and spans nest: a retrieval span might contain an embedding span and a search span.

The trace id is created at the entry point (the API handler or the job runner) and passed to every downstream call. In distributed systems it travels in request headers; the W3C Trace Context standard defines a traceparent header for this. If any component drops the id, the trace breaks into fragments and the slow step disappears from view.

Traces answer questions that metrics cannot. A p95 latency graph tells you some requests are slow; a trace tells you that this one was slow because the reranker timed out twice and the retry policy waited four seconds. A cost graph says spend doubled; traces show that one prompt change made the agent call the search tool six times instead of two.

For LLM systems, a useful trace also carries the semantic content: inputs, outputs and attributes such as the model name, prompt version and token counts. That is what separates an LLM trace from a generic HTTP trace, and it is why teams often add an LLM-aware layer on top of their existing tracing.

What is a span?

A span is one timed unit of work inside a trace, such as a model call or a vector search, with a start time, a duration, a parent, a status and attributes describing what happened.

If a trace is the story of a request, spans are its sentences. Each span records a name (vector_search), a start timestamp, an end timestamp, a status (ok or error), a pointer to its parent span, and a set of key-value attributes. Spans can also carry events, timestamped notes inside the span such as "first token received" or "retry 2 started".

The parent links turn a flat list into a tree. The root span covers the whole request; children cover the steps; grandchildren cover the sub-steps. Reading the tree shows both sequence and concurrency: two tool calls that overlap in time ran in parallel, and the parent span's duration is set by the slowest one.

Choosing span boundaries is a design decision. Wrap every network call and every model call, because those are where time and money go. Wrap meaningful logical steps ("plan", "retrieve", "validate") so the tree reads like the algorithm. Do not wrap every function, or the trace becomes noise and the tracing overhead becomes measurable.

For LLM work, the attributes are where the value lives: model name, temperature, maximum tokens, input and output token counts, finish reason, prompt template id and version, and for retrieval the number of results and their ids. OpenTelemetry's generative AI semantic conventions define standard names for many of these so that different tools can read them.

What should you log?

Log enough to reproduce and attribute any request: ids, versions of prompt and model, parameters, timing, token usage, retrieved context ids, tool calls and results, validation outcomes and feedback, with sensitive text redacted or access-controlled.

The test for an LLM log is reproduction. Given a bad answer reported by a user, can you rebuild exactly what the model saw and why the system acted as it did? That requires the identifiers (trace id, session id, a pseudonymous user id, tenant), the versions (prompt template version, model identifier as returned by the provider, tool schema version, index version), the parameters (temperature, maximum tokens, tools offered), and the context (retrieved chunk ids and scores, conversation history or a pointer to it).

Then record what happened: the output, the finish reason (a natural stop, the length limit, a content filter, or a tool call), every tool call with its arguments, result status and duration, every retry and its cause, guardrail and validation outcomes, and token usage split into input, cached input, output and, where the provider reports them, reasoning tokens. Finally attach outcome signals when they arrive: thumbs up or down, an edit, an escalation to a human, or whether the user completed the task.

Structure matters more than volume. Write logs as structured events with consistent field names, not free text, so you can query "all requests on prompt v14 where the finish reason was length". Keep large payloads (full prompts, documents, outputs) in a separate store keyed by trace id, with shorter retention and tighter access than the metadata.

What not to log is part of the answer. Secrets, raw credentials and API keys must never appear. Personal data should be redacted or tokenized before storage where the use case allows, and retention should match your privacy commitments and any provider agreements.

What are input and output tokens?

Input tokens are the tokens the model reads (system prompt, history, retrieved context, tool definitions, user message); output tokens are the tokens it generates. Both are billed, usually at different rates, and both drive latency.

A token is the unit a model reads and writes, typically a word fragment of a few characters in English. Every call has two counts. Input tokens (also called prompt tokens) cover everything sent: the system prompt, tool and schema definitions, conversation history, retrieved documents and the user's message. Output tokens (completion tokens) cover what the model generates, including tool-call arguments.

They behave differently. Input tokens are processed in parallel during the prefill phase, so a large input mainly raises time to first token. Output tokens are generated one at a time during decoding, so output length largely sets total generation time. Most providers price output tokens several times higher than input tokens; check the provider's current pricing for the actual ratio.

Two refinements matter for monitoring. Many providers offer prompt caching, where a repeated prefix is billed at a discount and reported as cached input tokens. Reasoning models may also produce reasoning tokens that are billed as output but not shown in the visible answer; some providers report them separately and some do not expose them at all. A dashboard that only counts visible output will under-estimate cost for these models.

Always take token counts from the provider's usage field in the response rather than estimating with a local tokenizer. Tokenizers differ between model families, the provider adds formatting tokens you cannot see, and the billed count is the one that matters. In observability, record all four numbers per call and roll them up per task, because one user action often triggers several calls.

What is time to first token?

Time to first token (TTFT) is the delay between sending a request and receiving the first generated token. With streaming it is the wait a user feels before text appears, so it shapes perceived speed more than total time.

A model response has two phases. During prefill the model processes the whole input and builds its internal cache; during decode it produces tokens one by one. TTFT covers network time, any queueing at the provider, and prefill. After that, the stream rate (tokens per second, or its inverse, time per output token) determines how quickly the rest arrives. Total latency is roughly TTFT plus output tokens divided by the stream rate.

TTFT grows with input length, because prefill work scales with the number of input tokens. It also grows with provider load, cold starts on self-hosted servers, and for reasoning models that think before answering, since the hidden reasoning tokens are generated before the first visible one. Prompt caching can cut TTFT noticeably on long repeated prefixes because the cached part does not need to be recomputed.

Measure it where the user is. The provider-side number excludes your own pre-processing (retrieval, guardrails, prompt assembly), which often adds hundreds of milliseconds. For a chat interface, the meaningful metric is the time from the user pressing send to the first character on screen, measured in the client or at least at your API edge. Record both so you can tell whether a slowdown is yours or the provider's.

TTFT matters most for interactive features. For background jobs that wait for the complete output, total latency and throughput matter instead, and streaming adds nothing.

What are p50, p95, and p99 latency?

They are percentiles: p50 is the time within which half of requests finish, p95 covers 95% and p99 covers 99%. The high percentiles describe the slow experiences an average hides.

Sort a window of request durations from fastest to slowest. The p50 (the median) is the value halfway down the list; the p95 is the value 95% of the way down; the p99 is 99% of the way. A p95 of 4 seconds means one request in twenty took longer than 4 seconds. Percentiles are preferred over averages because latency distributions are skewed: a few very slow requests pull the mean up while most users see something faster, and the mean describes nobody's actual experience.

LLM latency has especially long tails. Output length varies per request, agents take a variable number of steps, retries add whole seconds, and provider queues fluctuate with load. It is common to see a p50 of 2 seconds and a p99 of 20 or more. The tail matters because heavy users hit it often: someone who sends 50 messages a day will probably see the p99 at least once a day.

Two technical points trip teams up. First, percentiles cannot be averaged: the mean of p95 across ten servers is not the fleet p95. Store latencies as histograms (counts per duration bucket) which can be merged, then compute percentiles from the merged histogram. Second, small samples make high percentiles noisy; a p99 over 200 requests rests on the two slowest requests.

Always slice percentiles by something meaningful: feature, model, prompt version, input size band. A blended p95 across a fast classification endpoint and a slow agent endpoint describes neither.

How do you trace an agent?

Give each task one root span and nest a span for every loop iteration, model call, tool call, retrieval, handoff and validation, recording the decision, arguments, results, tokens and stop reason, so you can replay the trajectory and see where it went wrong.

An agent is a loop: the model decides an action, the harness executes it, the result goes back into context, and the model decides again until it stops or hits a budget. Tracing it means making that loop legible. The root span represents the whole task and carries the goal, the agent and prompt versions, the budget limits and the final outcome. Under it, each iteration gets a span, and inside each iteration sit the model call span and the tool or retrieval spans it triggered.

The most useful attributes are about decisions, not only timing. On the model call record which tools were offered, which tool was chosen and its arguments, the finish reason and the token usage. On the tool span record the result status, a summary or reference to the result, the duration and any error. On the root record the stop reason: task complete, step limit, token or cost budget exceeded, human escalation or unrecoverable error. A trace without a stop reason cannot distinguish "finished" from "gave up".

Multi-agent systems and long-running tasks add two needs. Handoffs between agents must propagate the trace context so the sub-agent's spans attach under the parent rather than starting a new trace. Tasks that pause for human approval or resume after a restart need a durable task id stored alongside the trace id; spans emitted after the resume should be linked to the original trace (OpenTelemetry supports span links for this) so the whole history reads as one story.

With this in place, the common agent pathologies become queries rather than investigations: loops where the same tool is called with near-identical arguments, steps where a tool error was ignored, tasks where cost per step grew because context kept expanding, and trajectories that reached the right answer by an unacceptable path. Those trajectories are also the input for trajectory evaluation, which grades the path as well as the result.

How do you monitor retrieval quality?

Log what was retrieved for every query, then track cheap operational signals (empty or low-score results, source mix, index freshness) continuously and judge relevance and answer support on a sample with model graders and periodic human review.

Retrieval fails quietly. When the right chunk is not retrieved, the model usually answers anyway from whatever it got or from its own training, and the answer looks plausible. So the first requirement is that every RAG trace records the query (and any rewritten query), the retrieved chunk ids, their scores and ranks, which chunks survived reranking and which were actually cited in the answer.

From that record, several signals are cheap enough to compute on every request. The zero-hit or low-score rate: the share of queries where nothing scored above a relevance threshold, often a sign of missing content or a broken index. Score distribution drift: if median top scores fall after an embedding model or chunking change, something regressed. Source mix: which documents and collections are being returned, which catches an outdated page suddenly dominating. Index freshness: the age of the newest indexed document and ingestion error counts. Citation rate: how often the answer cites any retrieved chunk at all.

Relevance itself needs judgement. Sample a few percent of traffic, and all traffic flagged by users, and ask a model grader two questions: was each retrieved chunk relevant to the query, and is the answer supported by the retrieved chunks. The first measures retrieval precision; the second measures groundedness. Recall (did we miss the right document) cannot be measured from production alone because you do not know what was missed; for that, maintain a labelled query set with known relevant documents and run recall@K against it on a schedule and after every index change.

Combine the views by slice. Retrieval often fails for a specific category: questions about a product launched last month, queries in another language, or a tenant whose documents failed to ingest. Breaking the signals down by topic, language and tenant is what turns a slightly lower average into an actionable finding.

Can hallucinations be detected in production?

Partly. Grounding checks against retrieved sources, sampled model judges, consistency checks, user feedback and human review each catch some hallucinations; none catches all, so layer them and focus effort on high-risk answers.

A hallucination is output that is not supported by the provided sources or by fact, presented as if it were. Detecting it in production is hard because there is usually no reference answer: you cannot compare against the truth when you do not know it. What you can check is narrower and more useful.

Grounding checks work when the system has sources. Split the answer into claims and ask whether each claim is supported by the retrieved context, using a model judge or a natural language inference (entailment) model. This catches the most common RAG failure, answering beyond the evidence, but it cannot catch a claim that is faithfully copied from a wrong source. Structural checks catch cheap cases deterministically: citations that point to chunks that were never retrieved, quoted numbers or ids that do not appear in the context, URLs or product names that do not exist in your catalogue.

Consistency checks sample the same question several times or ask the model to verify its own answer; disagreement signals uncertainty. Methods in the spirit of the SelfCheckGPT paper use this idea. They cost several extra calls, so they suit high-stakes routes rather than all traffic. Token log-probabilities, where the provider exposes them, give a weak confidence signal for short factual answers but are poorly calibrated for long text.

Human signals close the gap: thumbs-down, corrections, escalations and support tickets, plus a regular expert review of a stratified sample. Report a measured unsupported-claim rate from the sample with its confidence interval rather than claiming the system does not hallucinate.

Detection is only useful if it leads somewhere. Online, a failed grounding check can block the answer, rewrite it with a caveat, or route it to a human. Offline, every confirmed hallucination becomes an eval case and is tagged by cause: retrieval miss, ignored context, outdated source or pure invention. The cause tells you whether to fix the index, the prompt or the product's promises.

How do evals connect to observability?

They form one loop: production traces reveal failures, failures become eval cases, evals gate the next change, and the same graders run online on sampled traffic so offline scores and production quality can be compared.

Offline evals answer "is this change better on the cases we know about". Observability answers "what is happening on the traffic we actually get". Each is weak alone. An eval set written once drifts away from real usage within weeks, and production monitoring without evals tells you something broke but gives you no safe way to verify a fix. Connected, they become a quality flywheel.

The loop works like this. Traces from production, especially those with negative feedback, failed checks or unusual patterns, are reviewed in error analysis sessions where someone reads them and labels the failure type. Confirmed failures are copied into the eval dataset with the expected behaviour written down, after removing personal data. The fix (a prompt change, a retrieval tweak, a new guardrail) is developed against that dataset, and the regression eval gates the release so the old failure cannot quietly return.

The reverse direction matters as much. The graders used offline (format checks, groundedness judges, rubric judges) can be run online on a sample of production traces and the scores attached to the traces. That gives a production quality metric on the same scale as the offline one. If offline scores rise but online scores do not, the eval set no longer represents traffic, and that is a signal to refresh it from recent traces.

Shared infrastructure makes this cheap: the same trace format for eval runs and production, so an eval run is browsable like production; stable case ids that link back to the source trace; and version attributes on both so a dashboard can compare prompt v14 offline against prompt v14 live. Shadow tests and canary releases sit in the middle of the loop, running a candidate on real traffic and comparing its online scores before full rollout.

What does OpenTelemetry provide?

OpenTelemetry is a vendor-neutral open standard with APIs, SDKs, a wire protocol (OTLP) and a collector for traces, metrics and logs, plus emerging semantic conventions for generative AI, so LLM spans can flow to any compatible backend.

OpenTelemetry (often shortened to OTel) is a Cloud Native Computing Foundation project that standardizes how applications produce telemetry. It has four parts that matter here. The API and SDKs in most major languages let code create spans, record metrics and emit logs. Context propagation carries the trace id across process boundaries, using the W3C Trace Context headers by default. OTLP, the OpenTelemetry Protocol, is the standard format for sending telemetry. The Collector is a separate service that receives telemetry, processes it (batching, sampling, filtering, redaction) and exports it to one or more backends.

The practical benefit is decoupling. You instrument once and choose or change the backend later: a general observability platform, an open-source tracing store, or an LLM-specific tool that accepts OTLP. Your LLM spans also land in the same trace as your HTTP handlers, database queries and queues, so a slow answer can be attributed to the model, the vector store or your own code in one view.

For LLM work, OpenTelemetry has generative AI semantic conventions: standard attribute names such as gen_ai.operation.name, gen_ai.request.model, gen_ai.response.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, plus conventions for agent and tool spans and for recording message content. As of late 2026 these conventions are still evolving and some parts are marked experimental, so attribute names can change between versions and tools vary in which version they read. Pin the convention version you emit and check what your backend expects.

What OpenTelemetry does not provide is the LLM-specific analysis: prompt management, eval scoring, trace review interfaces, or dataset curation. Those come from the backend or LLM observability tool. A common architecture is OTel instrumentation in the application, an OTel Collector that redacts and samples, and two exporters: one to the general platform used by operations, one to the LLM tool used for quality work.

Instrumentation libraries exist that patch popular model SDKs and frameworks to emit these spans automatically. They are a fast start, but check what they capture by default, since some record full prompts and outputs as span content.

What should an AI dashboard show?

Show outcomes first (task success and quality scores), then the operational health behind them (errors, latency percentiles and TTFT, cost per task, tool and retrieval health), each sliceable by feature, model and prompt version with change markers.

A good dashboard answers three questions in order: are users getting good results, is the system healthy, and what changed. Most AI dashboards invert this and lead with request counts and token totals, which are easy to compute and rarely tell you anything is wrong.

The outcome row should hold task success (the share of tasks that reached a successful end state, defined per feature: ticket resolved without escalation, extraction passed validation, user accepted the draft), online quality scores from sampled judges, user feedback rates, and escalation or human-override rates. The health row holds error rate split by cause (provider errors, timeouts, rate limits, validation failures, guardrail blocks), latency as p50, p95 and p99 plus TTFT for streaming, and cost per completed task with tokens per task. The component row holds tool success rates and latency per tool, retrieval signals such as the zero-hit rate, agent steps per task and the share of step-limit stops, and fallback or retry rates.

Two features make a dashboard useful in an incident. Slicing: every chart filters by feature, model identifier, prompt version, tenant and region, because regressions are usually local. Change annotations: vertical markers for deployments, prompt releases, model updates, index rebuilds and provider incidents, so a step change in cost or quality lines up with its cause at a glance.

Different audiences need different views. Engineers on call need the health and component rows at minute resolution. Product owners need outcomes and cost per task over days and weeks. Finance needs spend by feature and tenant. One crowded page for everyone serves nobody; three focused views built on the same metrics work better.

Drift deserves a place but needs care: shifts in input topic mix, language, input length or retrieval score distributions often precede quality drops. Show them as trends next to quality, not as alarms on their own.

How do you keep sensitive data safe in traces?

Decide per field what is captured, redact or pseudonymize personal data before it leaves the application, store full payloads separately with short retention and strict access, and make sure your observability vendor's terms match your privacy commitments.

LLM traces are unusually sensitive. They hold what users typed (often names, addresses, account details, health or financial information), the documents retrieved on their behalf, and the model's outputs about them. A tracing store that keeps all of that for a year, searchable by every engineer, can easily become the largest concentration of personal data in the company and a target for exfiltration. Privacy has to be designed into observability rather than added after an audit.

Start by classifying fields. Metadata (ids, versions, timings, token counts, statuses, chunk ids) is low risk and can be kept for every request. Content (prompts, messages, documents, outputs, tool arguments that include customer data) is high risk. Capture content only where it earns its place: on a sample, on flagged traces, or for features where debugging requires it. Many teams keep content for days or weeks and metadata for months.

Redact as early as possible, ideally in the application or in a collector you run, before data reaches any third-party backend. Pattern-based detection handles structured identifiers such as emails, phone numbers and card numbers reliably; names and free-text details need a named-entity recognition model and will still be imperfect. Where you need to correlate a user's requests, replace identifiers with stable pseudonyms (a keyed hash) rather than storing the raw value. Never log secrets: strip authorization headers and API keys from tool-call spans.

Then control the store. Restrict raw-content access to a small group with audit logging, support deletion so a user's data-erasure request also reaches traces, and set retention automatically. If you use a hosted observability product, review where data is stored, whether it is used for any other purpose, and whether that matches the commitments in your provider agreements, especially if you rely on a zero-data-retention arrangement with the model provider. Sending the same prompts to an observability vendor that retains them can quietly undo that arrangement.

How do you use user feedback as a quality signal?

Capture explicit feedback (ratings, comments) and implicit behaviour (edits, retries, copies, escalations, abandonment) on the trace, treat it as a biased but valuable signal, and use it to find failures rather than to measure accuracy directly.

Explicit feedback is what users tell you: thumbs up or down, a star rating, a reason code, a free-text comment. It is easy to collect and rare in volume; often only a small percentage of interactions get any rating, and the users who rate skew toward strong reactions, mostly negative. That makes explicit feedback good at pointing to failures and poor at estimating overall quality.

Implicit feedback comes from behaviour and is available for nearly every interaction. Useful signals include: the user regenerated or rephrased the same question (dissatisfaction), copied the answer or accepted a suggestion (likely useful), edited a generated draft heavily before sending (partial failure, and the edit distance is measurable), asked to talk to a human or opened a ticket (failure), or left mid-conversation. Each is ambiguous on its own; a copied answer might be copied to complain about it. Combined and validated against labelled samples, they become strong signals.

Attach every feedback event to the trace id of the response it refers to, not just to the session. That is what lets you open the exact prompt, context and output behind a thumbs-down, and slice feedback rates by prompt version, model or retrieval outcome. Store the reason codes and comments as structured data so they can be grouped.

Use feedback in three ways. As a triage queue: negative and escalated traces go first into error analysis. As a trend: changes in feedback rates after a release are an early warning, even if the absolute level is biased. As eval material: reviewed failures become regression cases. Do not use raw ratings as the target of automatic optimization without care, because users reward confident, agreeable and longer answers, which can make a system worse at telling them uncomfortable truths.

How do you alert on an LLM system?

Alert on symptoms users feel and on budget breaches: error and timeout rates, latency objectives, cost per task and spend rate, failed validations, step-limit stops and sharp drops in quality proxies, with thresholds tied to volume so alerts stay rare and actionable.

The usual rule for alerting applies: page on symptoms that hurt users, not on every internal metric that wiggles. For LLM systems, the list of symptoms is longer because failures do not always throw errors. A provider can return valid but empty responses, a guardrail can start blocking a third of answers, or an agent can start looping and multiply spend without a single exception.

Group alerts into four families. Availability: provider error and timeout rates, rate-limit responses, and fallback activation, which tells you the primary model is struggling even when users are still served. Latency: p95 against an objective, and TTFT for streaming paths, evaluated over windows long enough to be stable. Cost: spend rate per hour against a budget, cost per task against its normal band, and tokens per task, which catches runaway loops and prompt bloat within minutes rather than at month end. Quality proxies: schema validation failure rate, guardrail block rate, empty or truncated outputs (finish reason equal to length), retrieval zero-hit rate, step-limit stops, and sharp rises in negative feedback.

Quality proxies need care. They are noisy, so alert on large relative changes against a baseline from the same time of day or week, not on fixed thresholds, and require a minimum volume before an alert can fire. Slower signals, such as sampled judge scores, belong in a daily report or a ticket rather than a page, because by the time enough samples arrive the trend is better handled in working hours.

Every alert should link to a dashboard pre-filtered to the affected feature and version, and to example traces. An alert that says "validation failures up 4x on refund_assistant since prompt v15" with five sample traces is fixable in minutes; one that says "quality anomaly detected" is not.