monitoring-observability

v2026.09.24

Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (as_type, score_current_span, should_export_span, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.

GitHub
Install command
npx skhub add yonatangross/monitoring-observability
Markdown
SKILL.md

Monitoring & Observability

A wrap around Prometheus, Grafana, OpenTelemetry and Langfuse, not a re-teaching of them. This skill carries OrchestKit's delta (version floors, house decisions, scars) and points at the vendor for everything else. Start at references/ork-delta.md.

Upstream coverage (do not restate)

These topics are fully covered first-party. Read the source, do not add a local copy.

TopicFirst-party source
Prometheus metric types, RED method, cardinality, PromQLhttps://prometheus.io/docs/practices/
Alertmanager grouping, inhibition, escalation, runbookshttps://prometheus.io/docs/alerting/latest/configuration/
Grafana dashboards, Loki and LogQL, Promtailhttps://grafana.com/docs/
OpenTelemetry spans, sampling, context propagationhttps://opentelemetry.io/docs/
Langfuse Python SDK (@observe, as_type, score_current_span, should_export_span, LangfuseMedia)https://langfuse.com/docs/sdk/python
Langfuse v2 to v4 Python and v3 to v5 JS migration pathshttps://langfuse.com/docs/sdk/python/v4-migration
Langfuse self-hosting (ClickHouse, Redis, S3, Helm)https://langfuse.com/docs/deployment/self-host
Langfuse cost tracking, model pricing, Metrics API v2https://langfuse.com/docs/model-usage-and-cost
Langfuse scores, online evaluators, annotation queues, prompt managementhttps://langfuse.com/docs/scores/overview
Langfuse framework integrations (LangChain, LangGraph, CrewAI, Pydantic AI, Bedrock, LiveKit)https://langfuse.com/docs/integrations
Agent Graphs, observation types, rendered tool callshttps://langfuse.com/docs/tracing-features/agent-graphs
PSI, KS test, KL and JS divergence, Wasserstein, embedding drifthttps://www.evidentlyai.com/blog/data-drift-detection-large-datasets
EWMA control chartshttps://www.itl.nist.gov/div898/handbook/pmc/section3/pmc324.htm
structlog, Winston, correlation IDs, log samplinghttps://www.structlog.org/en/stable/

Quick Reference

CategoryRulesImpactWhen to Use
Infrastructure Monitoring1CRITICALGrafana dashboards, Golden Signals, SLO/SLI
LLM Observability1HIGHLangfuse tracing, observation types, agent graphs
Silent Failures3HIGHTool skipping, quality degradation, loop/token spike alerting

Total: 5 rules across 3 categories. Drift detection, cost tracking, eval scoring, Prometheus instrumentation and alert-rule authoring moved to the upstream sources listed above.

Quick Start

# Langfuse v4 LLM tracing: semantic as_type plus inline scoring
from langfuse import observe, get_client

@observe(as_type="generation", name="analyze_content")
async def analyze_content(content: str):
    get_client().update_current_trace(
        user_id="user_123", session_id="session_abc",
        tags=["production", "orchestkit"],
    )
    result = await llm.generate(content)
    get_client().score_current_span(name="response_quality", value=0.85)
    return result
# Prometheus RED method, wired the way this repo expects (bounded labels only)
from prometheus_client import Counter, Histogram

http_requests = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
http_duration = Histogram('http_request_duration_seconds', 'Request latency',
    buckets=[0.01, 0.05, 0.1, 0.5, 1, 2, 5])

Infrastructure Monitoring

Dashboard and health-check patterns. Metric instrumentation and alert-rule syntax are upstream.

RuleFileKey Pattern
Grafana Dashboardsrules/monitoring-grafana.mdGolden Signals, SLO/SLI, health checks

CC 2.1.161 — OTEL resource attributes as metric labels: OTEL_RESOURCE_ATTRIBUTES values are now attached as labels on metric datapoints, so usage metrics can be sliced by custom dimensions (team, repo, environment). Add label selectors to dashboards for multi-tenant / per-team cost and usage tracking.

LLM Observability

Langfuse-based tracing for LLM applications. Cost tracking, scoring and drift statistics are upstream; what stays here is how this repo wires traces.

RuleFileKey Pattern
Langfuse Tracesrules/llm-langfuse-traces.md@observe decorator, OTEL spans, agent graphs

Silent Failures

Detection and alerting for silent failures in LLM agents.

RuleFileKey Pattern
Tool Skippingrules/silent-tool-skipping.mdExpected vs actual tool calls, Langfuse traces
Quality Degradationrules/silent-degraded-quality.mdHeuristics + LLM-as-judge, z-score baselines
Silent Alertingrules/silent-alerting.mdLoop detection, token spikes, escalation workflow

CC 2.1.169 — OTEL client-cert paths require trust: untrusted project settings can no longer set OTEL client-certificate paths without a trust confirmation. If your OTEL exporter uses client certs configured in project .claude/settings.json, expect a one-time trust prompt on first use in an untrusted project — telemetry silently not flowing after 2.1.169 is usually this gate, not the collector.

Key Decisions

DecisionRecommendationRationale
Metric methodologyRED method (Rate, Errors, Duration)Industry standard, covers essential service health
Log formatStructured JSONMachine-parseable, supports log aggregation
TracingOpenTelemetryVendor-neutral, auto-instrumentation, broad ecosystem
LLM observabilityLangfuse (not LangSmith)Open-source, self-hosted, built-in prompt management
LLM tracing API@observe(as_type=...) + score_current_span()v4: semantic types, inline scoring, span filtering
Langfuse APIsObservations API v2 + Metrics API v2v4 (Mar 2026): faster querying, aggregations at scale
Hook telemetry transportJSONL under ~/.claude/analytics/, never an SDK in-processHooks are per-event processes; SDK init would be paid on every spawn (references/ork-delta.md)

Detailed Documentation

ResourceDescription
references/ork-delta.mdStart here. Floors, house decisions and scars that upstream docs do not carry
references/langfuse-js-v5.mdJS/TS SDK v5 delta from Python 4.x: package map, phantom packages, SpanProcessor vs exporter. Read before writing any JS Langfuse code
references/experiments-api.mdLangfuse experiments and dataset runs as this repo uses them
references/evaluation-scores.mdScore shapes and scoring pipeline wiring
references/session-tracking.mdSession and user grouping across multi-step workflows
references/metrics-collection.mdClaude Code OTEL metric inventory and collector-side joins
references/dashboards.mdDashboard layout conventions
references/structured-logging.mdStructured log field conventions
references/dev-agent-lens.mdLiteLLM proxy layer for API-boundary observability
examples/orchestkit-monitoring-dashboard.mdWorked monitoring dashboard example
scripts/Templates: Prometheus, OpenTelemetry, health checks, Langfuse

Related Skills

  • defense-in-depth - Layer 8 observability as part of security architecture
  • devops-deployment - Observability integration with CI/CD and Kubernetes
  • resilience-patterns - Monitoring circuit breakers and failure scenarios
  • llm-evaluation - Evaluation patterns that integrate with Langfuse scoring
  • caching - Caching strategies that reduce costs tracked by Langfuse
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

src/skills/monitoring-observability

Default branch

main

Latest commit

43c04fa

Tree SHA

29981ce