qa-resilience

v2026.09.24

Designs and tests distributed-system resilience. Use when adding retries, deadlines, hedging, circuit breakers, overload protection, chaos experiments, or SLO reliability gates.

GitHub
Install command
npx skhub add vasilyu1983/qa-resilience
Markdown
SKILL.md

QA Resilience

Use this skill when reliability work is about failure behavior, overload protection, degraded mode, or resilience testing. The goal is not "add retries everywhere." The goal is predictable failure handling, clear ownership, and testable recovery behavior.

Quick Reference

SymptomStart With
slow or hanging dependencydeadline and timeout budget
transient dependency failurebounded retry with jitter and retry budget
sustained dependency failurecircuit breaker and fallback
rate-limited dependencyhonor Retry-After, expose degraded behavior, and test quota paths intentionally
one bad host in a healthy pooloutlier detection or endpoint ejection
queue or pool saturationbulkheads, concurrency limits, load shedding
non-critical feature outagegraceful degradation or feature flag fallback
resilience validationdeterministic fault injection before chaos

When to Use This Skill

  • retries, deadlines, hedging, breakers, bulkheads, and overload protection
  • degraded-mode UX or API behavior
  • service-mesh or gateway resilience policy
  • chaos engineering, game days, DR drills, and fault injection
  • release gates based on failure behavior, not only happy-path load tests

Route Elsewhere


Workflow

  1. Identify the critical user journeys and the dependencies that can break them.
  2. Define the contract per dependency:
    • timeout or deadline budget
    • retry ownership
    • breaker or outlier policy
    • concurrency and queue limits
    • degraded behavior if the dependency is unavailable
  3. Decide where the policy lives:
    • app code
    • client library
    • mesh or gateway
  4. Test in stages:
    • design review or checker output establishes policy coherence only
    • deterministic fault injection proves the response to a named controlled fault
    • staged chaos in non-production proves bounded behavior and recovery in that environment
    • narrow prod canary or game day proves only the observed target path and window, with guardrails
  5. Define pass or fail signals:
    • burn rate
    • p95 or p99
    • fallback rate
    • breaker transitions
    • shed volume
    • recovery time
  6. Record the injected fault, steady-state hypothesis, blast-radius limit, abort condition, observation window, recovery result, and dependencies simulated rather than observed.

Selective verification and hazard analysis

  • Use formal methods when the question is whether a recovery/failover protocol preserves an invariant under concurrent faults, or when a counterexample is needed. Return the property, initial states, transitions/environment assumptions, checked bound, result/trace, and model-to-code limits. A finite invariant check does not establish liveness or recovery time. Skip for routine timeout/retry configuration and named fault tests whose behavior is already directly observable.
  • Use safety engineering when retries, rollback, failover, operators, or other functioning controls can interact to cause a specified unacceptable loss, such as duplicate financial effects or destructive recovery. Return loss → hazard → unsafe control action and context → constraint → owner → verification evidence, including residual unknowns. Skip for ordinary availability or latency tuning without a harmful interaction question. Uptime and a hazard diagram do not prove safety.

Keep resilience policy, fault-injection execution, recovery evidence, and release recommendations with this skill; use the foundation only for the named gap.

Pattern Rules

  • retries happen at one layer only
  • deadlines come before retries
  • hedging is only for idempotent or cancellation-safe reads
  • overload handling must shed early instead of collapsing late
  • readiness and liveness must stay bounded and shallow
  • resilience behaviors belong in targeted checks; do not let rate-limit or degraded-mode coverage leak into unrelated happy-path suites
  • prod experiments require blast-radius limits, abort criteria, dashboards, and owners

Failure Modes to Validate

  • timeouts and deadline propagation
  • retry storms and duplicate side effects
  • partial dependency outages
  • slow downstreams and long-tail latency
  • queue buildup and connection-pool exhaustion
  • one-bad-host behavior inside a pool
  • degraded-mode responses and stale-data fallbacks
  • rate-limit handling, Retry-After, and client backoff expectations
  • visible state convergence after recovery or backend resets
  • failover and failback behavior
  • metastable failure: a self-sustaining feedback loop (retry amplification, cache-miss stampede, queue backlog, connection-pool churn) that does not self-resolve after the original trigger clears — see references/cascading-failure-prevention.md

Testing Ladder

Deterministic first

  • inject latency, errors, timeouts, malformed payloads, and unavailable endpoints in controlled tests
  • verify the intended control activates and the wrong controls do not

Chaos second

  • start in non-production
  • use a small blast radius and a fixed time window
  • stop immediately on error-budget or customer-impact breach

DR and game days last

  • validate RTO and RPO claims explicitly
  • rehearse recovery ownership, not only technical failover

Operational Guardrails

  • every experiment needs a stated hypothesis and steady-state metric
  • every run should capture timestamps, targets, blast radius, and dashboard links
  • telemetry fields for retries, breaker transitions, hedging, shedding, and fallback are part of the resilience contract
  • releases should be gated on reliability behavior, not only resource usage

Anti-Patterns

  • no timeouts
  • retries at every hop
  • fixed-interval or unbounded retries
  • fixed per-call retry count with no system-wide retry budget (does not tighten as error rate rises)
  • retries without idempotency
  • hedging unsafe writes
  • no bulkheads or queue bounds
  • deep readiness or liveness checks
  • silent degraded mode
  • untested failover plans
  • happy-path-only load testing

Scripts

ScriptPurpose
scripts/resilience_checker.pyScores resilience pattern coverage and reports gaps

Typical usage:

python scripts/resilience_checker.py assess --input data/sample-service-profile.json
python scripts/resilience_checker.py gaps --input data/sample-service-profile.json
python scripts/resilience_checker.py report --input data/sample-service-profile.json --output resilience-report.md

See scripts/README.md for the input format and scoring logic.

ASCII Flow

Resilience request
  -> Identify failure mode: slow, transient, sustained, overloaded, or degraded
  -> Set deadlines, retry budgets, isolation, and fallback ownership
  -> Add telemetry for saturation, errors, latency, and degraded behavior
  -> Validate with deterministic fault injection before chaos experiments
  -> Gate release on recovery evidence, SLO impact, and rollback path
  -> Document runbook actions and residual risk

Navigation

Foundation applied recipes

Core references

Operational resources

Related Skills

Fact-Checking

  • Known bugs, regressions, framework/compiler/runtime footguns, and version-specific crash or workaround guidance must be verified against current primary web sources before being treated as current fact.
  • Verify current platform features, mesh capabilities, and vendor-specific behavior before final answers when the recommendation depends on a live product.
  • Prefer primary docs for runtime or tooling specifics.
  • If web access is unavailable, keep external product guidance marked as unverified.

Learnings Loop

When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.

After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

frameworks/shared-skills/skills/qa-resilience

Default branch

main

Latest commit

8dc5de4

Tree SHA

700bf67