chaos-engineering-expert

v2026.09.24

Expert in chaos engineering principles, failure injection, resilience testing, Chaos Monkey, Gremlin, and building fault-tolerant systems. Use when the user mentions reliability, testing, SRE, resilience, failure injection, or resilience testing, or when the task involves Chaos Engineering Principles, Failure Types, Tools & Platforms, or Chaos Toolkit Experiment.

GitHub
Install command
npx skhub add personamanagmentlayer/chaos-engineering-expert
Markdown
SKILL.md

Chaos Engineering Expert

Core Concepts

Chaos Engineering Principles

  • Hypothesis-Driven - Define expected system behavior
  • Production Testing - Test in real environments
  • Minimize Blast Radius - Start small, expand gradually
  • Automation - Continuous chaos experiments
  • Learn and Improve - Build resilience iteratively
  • Observability - Monitor system behavior

Failure Types

  • Network Failures - Latency, packet loss, partitions
  • Resource Exhaustion - CPU, memory, disk
  • Service Failures - Process crashes, unavailability
  • Data Corruption - Corrupt files, bad data
  • Time Drift - Clock skew, NTP failures
  • Dependency Failures - Third-party service outages

Tools & Platforms

  • Chaos Monkey - Netflix's random termination tool
  • Gremlin - Enterprise chaos engineering platform
  • Chaos Toolkit - Open-source chaos experiments
  • Litmus - Kubernetes chaos engineering
  • Pumba - Docker chaos testing
  • Toxiproxy - Network condition simulation

Best Practices

Experiment Design

  • Start with hypothesis
  • Define steady-state metrics
  • Begin with small blast radius
  • Test in staging first
  • Automate experiments
  • Document learnings

Safety Measures

  • Implement circuit breakers
  • Set up monitoring/alerting
  • Have rollback procedures
  • Limit blast radius
  • Run during business hours initially
  • Get stakeholder buy-in

Observability

  • Monitor golden signals
  • Track error rates
  • Measure latency (p50, p95, p99)
  • Monitor resource utilization
  • Log all chaos events
  • Correlate metrics

Culture

  • Foster blameless culture
  • Share learnings openly
  • Make chaos regular practice
  • Train teams on chaos engineering
  • Start with game days
  • Celebrate failures as learning

Anti-Patterns

Common Mistakes

  • Testing in production without preparation
  • No rollback plan
  • Ignoring blast radius
  • Running attacks during incidents
  • No monitoring in place
  • Blaming teams for failures

Experiment Design Issues

  • No clear hypothesis
  • Undefined success criteria
  • Too broad scope initially
  • Missing steady-state verification
  • No automation
  • Poor documentation

Cultural Problems

  • Blame-focused culture
  • Resistance to controlled failure
  • Lack of stakeholder support
  • No learning from experiments
  • Chaos for chaos sake
  • Security concerns ignored

Reference Documentation

Detailed material lives alongside this skill and is read on demand:

  • Implementation Examples — Chaos Toolkit Experiment, Gremlin Attack Scenarios, Custom Chaos Tool (Python), Kubernetes Chaos with Litmus

Resources

Official Documentation

Learning Resources

Tools & Platforms

Community Resources

Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

Apache-2.0

Source path

stdlib/qa/chaos-engineering-expert

Default branch

main

Latest commit

79ccaa9

Tree SHA

d3a3f94