monitoring-observability

v2026.09.24

Monitoring and observability patterns for Prometheus metrics, Grafana dashboards, Langfuse v4 LLM tracing (as_type, score_current_span, should_export_span, LangfuseMedia), and drift detection. Use when adding logging, metrics, distributed tracing, LLM cost tracking, or quality drift monitoring.

GitHub
安装命令
npx skhub add yonatangross/monitoring-observability
Markdown
SKILL.md

Monitoring & Observability

A wrap around Prometheus, Grafana, OpenTelemetry and Langfuse, not a re-teaching of them. This skill carries OrchestKit's delta (version floors, house decisions, scars) and points at the vendor for everything else. Start at references/ork-delta.md.

Upstream coverage (do not restate)

These topics are fully covered first-party. Read the source, do not add a local copy.

TopicFirst-party source
Prometheus metric types, RED method, cardinality, PromQLhttps://prometheus.io/docs/practices/
Alertmanager grouping, inhibition, escalation, runbookshttps://prometheus.io/docs/alerting/latest/configuration/
Grafana dashboards, Loki and LogQL, Promtailhttps://grafana.com/docs/
OpenTelemetry spans, sampling, context propagationhttps://opentelemetry.io/docs/
Langfuse Python SDK (@observe, as_type, score_current_span, should_export_span, LangfuseMedia)https://langfuse.com/docs/sdk/python
Langfuse v2 to v4 Python and v3 to v5 JS migration pathshttps://langfuse.com/docs/sdk/python/v4-migration
Langfuse self-hosting (ClickHouse, Redis, S3, Helm)https://langfuse.com/docs/deployment/self-host
Langfuse cost tracking, model pricing, Metrics API v2https://langfuse.com/docs/model-usage-and-cost
Langfuse scores, online evaluators, annotation queues, prompt managementhttps://langfuse.com/docs/scores/overview
Langfuse framework integrations (LangChain, LangGraph, CrewAI, Pydantic AI, Bedrock, LiveKit)https://langfuse.com/docs/integrations
Agent Graphs, observation types, rendered tool callshttps://langfuse.com/docs/tracing-features/agent-graphs
PSI, KS test, KL and JS divergence, Wasserstein, embedding drifthttps://www.evidentlyai.com/blog/data-drift-detection-large-datasets
EWMA control chartshttps://www.itl.nist.gov/div898/handbook/pmc/section3/pmc324.htm
structlog, Winston, correlation IDs, log samplinghttps://www.structlog.org/en/stable/

Quick Reference

CategoryRulesImpactWhen to Use
Infrastructure Monitoring1CRITICALGrafana dashboards, Golden Signals, SLO/SLI
LLM Observability1HIGHLangfuse tracing, observation types, agent graphs
Silent Failures3HIGHTool skipping, quality degradation, loop/token spike alerting

Total: 5 rules across 3 categories. Drift detection, cost tracking, eval scoring, Prometheus instrumentation and alert-rule authoring moved to the upstream sources listed above.

Quick Start

# Langfuse v4 LLM tracing: semantic as_type plus inline scoring
from langfuse import observe, get_client

@observe(as_type="generation", name="analyze_content")
async def analyze_content(content: str):
    get_client().update_current_trace(
        user_id="user_123", session_id="session_abc",
        tags=["production", "orchestkit"],
    )
    result = await llm.generate(content)
    get_client().score_current_span(name="response_quality", value=0.85)
    return result
# Prometheus RED method, wired the way this repo expects (bounded labels only)
from prometheus_client import Counter, Histogram

http_requests = Counter('http_requests_total', 'Total requests', ['method', 'endpoint', 'status'])
http_duration = Histogram('http_request_duration_seconds', 'Request latency',
    buckets=[0.01, 0.05, 0.1, 0.5, 1, 2, 5])

Infrastructure Monitoring

Dashboard and health-check patterns. Metric instrumentation and alert-rule syntax are upstream.

RuleFileKey Pattern
Grafana Dashboardsrules/monitoring-grafana.mdGolden Signals, SLO/SLI, health checks

CC 2.1.161 — OTEL resource attributes as metric labels: OTEL_RESOURCE_ATTRIBUTES values are now attached as labels on metric datapoints, so usage metrics can be sliced by custom dimensions (team, repo, environment). Add label selectors to dashboards for multi-tenant / per-team cost and usage tracking.

LLM Observability

Langfuse-based tracing for LLM applications. Cost tracking, scoring and drift statistics are upstream; what stays here is how this repo wires traces.

RuleFileKey Pattern
Langfuse Tracesrules/llm-langfuse-traces.md@observe decorator, OTEL spans, agent graphs

Silent Failures

Detection and alerting for silent failures in LLM agents.

RuleFileKey Pattern
Tool Skippingrules/silent-tool-skipping.mdExpected vs actual tool calls, Langfuse traces
Quality Degradationrules/silent-degraded-quality.mdHeuristics + LLM-as-judge, z-score baselines
Silent Alertingrules/silent-alerting.mdLoop detection, token spikes, escalation workflow

CC 2.1.169 — OTEL client-cert paths require trust: untrusted project settings can no longer set OTEL client-certificate paths without a trust confirmation. If your OTEL exporter uses client certs configured in project .claude/settings.json, expect a one-time trust prompt on first use in an untrusted project — telemetry silently not flowing after 2.1.169 is usually this gate, not the collector.

Key Decisions

DecisionRecommendationRationale
Metric methodologyRED method (Rate, Errors, Duration)Industry standard, covers essential service health
Log formatStructured JSONMachine-parseable, supports log aggregation
TracingOpenTelemetryVendor-neutral, auto-instrumentation, broad ecosystem
LLM observabilityLangfuse (not LangSmith)Open-source, self-hosted, built-in prompt management
LLM tracing API@observe(as_type=...) + score_current_span()v4: semantic types, inline scoring, span filtering
Langfuse APIsObservations API v2 + Metrics API v2v4 (Mar 2026): faster querying, aggregations at scale
Hook telemetry transportJSONL under ~/.claude/analytics/, never an SDK in-processHooks are per-event processes; SDK init would be paid on every spawn (references/ork-delta.md)

Detailed Documentation

ResourceDescription
references/ork-delta.mdStart here. Floors, house decisions and scars that upstream docs do not carry
references/langfuse-js-v5.mdJS/TS SDK v5 delta from Python 4.x: package map, phantom packages, SpanProcessor vs exporter. Read before writing any JS Langfuse code
references/experiments-api.mdLangfuse experiments and dataset runs as this repo uses them
references/evaluation-scores.mdScore shapes and scoring pipeline wiring
references/session-tracking.mdSession and user grouping across multi-step workflows
references/metrics-collection.mdClaude Code OTEL metric inventory and collector-side joins
references/dashboards.mdDashboard layout conventions
references/structured-logging.mdStructured log field conventions
references/dev-agent-lens.mdLiteLLM proxy layer for API-boundary observability
examples/orchestkit-monitoring-dashboard.mdWorked monitoring dashboard example
scripts/Templates: Prometheus, OpenTelemetry, health checks, Langfuse

Related Skills

  • defense-in-depth - Layer 8 observability as part of security architecture
  • devops-deployment - Observability integration with CI/CD and Kubernetes
  • resilience-patterns - Monitoring circuit breakers and failure scenarios
  • llm-evaluation - Evaluation patterns that integrate with Langfuse scoring
  • caching - Caching strategies that reduce costs tracked by Langfuse
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

src/skills/monitoring-observability

默认分支

main

最新提交

43c04fa

Tree SHA

29981ce