Why this matters now: LangChain's survey of 1,340 practitioners found that 89 percent of organizations have implemented some form of agent observability, while only 37.3 percent run evaluations against live traffic. Gartner expects LLM observability to be part of 50 percent of GenAI deployments by 2028, up from 15 percent today. Teams that ship agents in 2026 without production quality checks are trusting the one signal that cannot tell them when the agent is wrong.

Picture the Dashboard Everyone Is Proud Of

Response time sits at 400 milliseconds. Error rate is under one percent. Token usage is flat. The on-call channel has been quiet for three weeks. Then a customer success lead forwards a screenshot of the agent telling a client that a refund window is 90 days when the policy says 30, and it turns out the agent has been saying that for a month.

Nothing on the dashboard was broken, because the dashboard was built for a different kind of system. Traditional monitoring asks whether a service is up, fast, and free of exceptions. An agent can be up, fast, and exception free while retrieving a stale policy document, misreading a tool result, or drifting after a silent model update. The failure is in the content of the answer, and HTTP status codes do not carry content.

That is the core of the observability gap. Infrastructure telemetry tells you the agent ran. It says nothing about whether the run was any good. Closing that gap takes three things layered on top of what you already have: traces that show every step the agent took, scores that judge the output on live traffic, and alerts that fire when those scores move. Most teams have the first in partial form, the second rarely, and the third almost never.

A green dashboard proves the agent is running. It has never once proved the agent is right.
89%
Of organizations have implemented some form of observability for their agents. Source: LangChain, State of Agent Engineering, 1,340 respondents, November–December 2025.
37.3%
Run online evaluations on live traffic, versus 52.4% who evaluate offline on test sets. Source: LangChain, State of Agent Engineering, 2025.
4 layers
Traces, quality scores, cost and drift signals, and alert ownership. The four things 10decoders checks first in every agent quality review.

Where Agent Observability Goes Blind

Blind spotWhat closes itSeverity
Latency and error dashboards stay green while answers are wrongQuality scoring on a sample of live production trafficCritical
Only the final model call is logged, not tool calls or retrievalEnd to end traces with a span for every step and handoffCritical
A model or prompt change ships with no before and after comparisonVersion tags on every trace and score trends split by versionHigh
Cost is reviewed monthly, after a runaway loop has already runPer request token and tool spend with a hard alert thresholdHigh
Traces store raw prompts containing customer dataRedaction at capture and short retention on sensitive fieldsModerate
Agent telemetry lives in a separate tool the platform team never seesOpenTelemetry based export into the existing observability stackLower

Not sure what your agent monitoring can actually see?

10decoders reviews your agent traces, quality checks, and alerting against what a production incident would need. You get a ranked list of blind spots and a scoped plan to close them.

Book a Free AI Assessment →

Tracing Tells You What Happened. Scoring Tells You Whether It Was Good

Teams often treat these as one project and stop after tracing. Tracing is worth doing first, because a single user request in an agent system can involve several model calls, tool invocations, vector lookups, and a handoff between agents. When something goes wrong, you need the whole chain in front of you. Among teams that already run agents in production, LangChain reports that 71.5 percent have full tracing that lets them inspect individual steps and tool calls.

A trace is a record, though, not a verdict. Someone still has to read it and decide the answer was wrong, and nobody reads thousands of traces a day. Scoring is what automates that judgment: a grader, whether a rule, a model, or a human review queue, rates a sample of production outputs for correctness, groundedness, or task completion, and the scores become a time series you can alert on. Gartner's analysts describe the split the same way. Explainability shows why a model answered as it did, while observability checks that the behavior holds up over time and can be relied on.

The practical consequence is that quality becomes a metric with a threshold, like error rate. If the share of grounded answers drops from 94 percent to 86 percent after a Thursday deployment, you want that visible on Friday morning, not in a customer escalation two weeks later. Offline test sets still matter for release gating, but they are frozen samples. Production traffic changes daily, and the only way to see how the agent handles it is to measure the traffic itself.

Stage 1
Where most teams start

Infrastructure Monitoring

Uptime, latency, and error rates are tracked. The agent is treated like any other service, so wrong answers only surface through user complaints.

Stage 2
Where most teams stall

Tracing Without Judgment

Full traces exist and engineers can debug a reported failure. Nothing scores live outputs, so problems are found only after someone reports them.

Stage 3
Where mature teams operate

Traced, Scored, and Alerting

Live traffic is sampled and graded, scores are split by model and prompt version, and a quality drop pages a named owner with a written first step.

Production Agent Observability Checklist

Run this against every agent that touches a customer, a payment, or a regulated record.

Before You Call an Agent Observable

Every LLM call, tool call, and retrieval step is a span in one traceA request that touches five model calls and two vector lookups should read as one timeline, not seven log lines.
Prompts, model version, and config are recorded on each traceWithout the exact inputs and version, a bad answer cannot be reproduced next Tuesday.
A sample of live traffic is scored for correctness on a scheduleOffline test sets tell you the agent worked in staging. Scoring live traffic tells you it works now.
Quality scores have alert thresholds, like error rates doA drop in groundedness or task success pages someone, the same way a spike in 500s would.
Token and tool spend is tracked per request and per agentA loop that retries forever shows up as a cost curve before it shows up as an invoice.
Traces use OpenTelemetry conventions where the stack allows itYour agent telemetry should land in the tools your platform team already runs.
Sensitive fields are redacted before traces are storedFull prompt logging without redaction turns your observability store into a data leak.
Every alert has a named owner and a written first stepA dashboard nobody is assigned to read is decoration.
Add the quality score to the dashboard before the next release, or wait for a customer to add it for you.

What to Do This Week

01 Pick one agent and trace a single real request end to end

Choose the agent closest to a customer or a financial decision, send one real request through it, and check whether you can see every model call, tool call, and retrieval result in order. Whatever is missing from that view is your first instrumentation ticket, and it is usually the tool calls.

02 Define one quality score you can compute automatically

Start narrow. For a retrieval agent that might be whether the answer is supported by the retrieved passages. For a workflow agent it might be whether the final task state matches the request. Write the rule down, run it on fifty recent production traces, and read every failure by hand to check the grader is not fooling you.

03 Tag every trace with the model version and prompt version

This is a small change with large payoff. Once traces carry version tags, you can compare scores before and after any deployment or vendor model update and see quickly whether a change helped or hurt. Without tags, a quality drop has no obvious cause and the investigation starts from zero.

04 Set one alert on the quality score and name the person it wakes

Pick a threshold from the fifty traces you just reviewed, route the alert to a named engineer, and write two lines describing what that person does first. An alert with no owner and no first step will be muted within a month, which returns you to the green dashboard problem.

Let 10decoders Close the Observability Gap in Your AI Agents

We review your traces, live quality scoring, cost signals, and alert ownership, then deliver a ranked gap list and a plan to reach production grade observability without rebuilding your stack.