agent-observability

v2026.09.24

Instrument LLM agents with traces, metrics, and replay. Use when an agent in production is silently failing, regressing, drifting, or burning tokens. Selects LangSmith / Langfuse / Phoenix, defines node-level spans, attaches evals to traces, and enables session replay.

GitHub
Install command
npx skhub add akillness/agent-observability
Markdown
SKILL.md

Agent Observability

Overview

Production agents fail differently than services: bad tool calls, hallucinated arguments, runaway loops, silent quality regression. This skill picks an observability backend, instruments node-level spans, attaches evals to traces, and sets up replay — so failures are observable, not just guessed at.

When to use

  • Agent works in dev, breaks in prod with no logs that explain why
  • Need to compare prompt/model changes against a baseline (eval-in-trace)
  • Token/latency cost is growing and you don't know which node is the culprit
  • A user reports a bad answer and you need to replay the exact session
  • Multiple agents/subagents — you need a single trace tree, not interleaved logs

Platform selection

BackendPick whenHosting
LangSmithLangChain/LangGraph stack, want managedSaaS
LangfuseOpen-source, self-host required, multi-frameworkSaaS or self-host
Arize PhoenixOTel-native, embed/eval drift, OSS-firstLocal or self-host

Default: Langfuse when you need self-host, LangSmith when you're already on LangGraph, Phoenix when OTel is mandated.

Instrumentation pattern (node-level spans)

# LangGraph + Langfuse
from langfuse.decorators import observe
from langfuse.openai import openai  # auto-traces tool calls

@observe(name="planner_node")
def planner(state):
    return {"plan": llm.invoke(state["task"])}

@observe(name="tool_executor")
def tool_executor(state):
    return {"observation": run_tool(state["action"])}

Required span attributes:

  • input / output (full, not truncated)
  • model, temperature, max_tokens
  • tool_name, tool_args, tool_result_status
  • tokens_in, tokens_out, cost_usd
  • session_id, user_id, trace_id

Eval-in-trace

Attach automated graders to each span so regressions surface in the same UI as latency:

from langfuse import Langfuse
langfuse = Langfuse()
langfuse.score(
    trace_id=trace_id,
    name="answer_correctness",
    value=0.92,
    comment="LLM-as-judge vs golden"
)

Common scores: correctness, tool_call_validity, groundedness, harm, latency_sla.

Replay pattern

  1. Log full state at every node entry/exit (Langfuse: metadata={"state": state})
  2. On bug report, fetch trace by trace_id
  3. Rehydrate state, re-run from any node — diff outputs

Sampling at scale

  • 100% trace error/HITL paths
  • 10% sample happy path
  • Tail-based sampling for spans > p95 latency
  • Always log: tool failures, guardrail blocks, budget caps hit

Further reading

  • LangSmith docs — datasets, evals, trace replay
  • Langfuse docs — self-host compose, OTel exporter
  • Arize Phoenix — embed drift, OSS LLM evals
  • OpenTelemetry gen_ai semantic conventions (2026)
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

Not specified

Source path

.agent-skills/agent-observability

Default branch

main

Latest commit

f579bfe

Tree SHA

34a09b3