distributed-systems

v2026.09.24

Distributed systems patterns for locking, resilience, idempotency, and rate limiting. Use when implementing distributed locks, circuit breakers, retry policies, idempotency keys, token bucket rate limiters, or fault tolerance patterns.

GitHub
Install command
npx skhub add yonatangross/distributed-systems
Markdown
SKILL.md

Distributed Systems Patterns

Comprehensive patterns for building reliable distributed systems. Each category has individual rule files in rules/ loaded on-demand.

Quick Reference

CategoryRulesImpactWhen to Use
Distributed Locks1CRITICALFencing tokens, owner validation; Redis/Redlock and Postgres advisory via upstream docs
Resilience3CRITICALCircuit breakers, retry with backoff, bulkhead isolation
Idempotency1HIGHIdempotency keys; dedup and database-backed storage via upstream docs
Rate Limiting2HIGHToken bucket, sliding window; SlowAPI integration via upstream docs
Edge Computing2HIGHEdge workers, V8 isolates, CDN caching, geo-routing
Event-Driven2HIGHEvent sourcing, CQRS, transactional outbox, sagas

Total: 11 rules across 6 categories. Removed topics point at first-party sources in Upstream coverage; ork-specific scars live in references/ork-delta.md.

Quick Start

# Redis distributed lock with Lua scripts
async with RedisLock(redis_client, "payment:order-123"):
    await process_payment(order_id)

# Circuit breaker for external APIs
@circuit_breaker(failure_threshold=5, recovery_timeout=30)
@retry(max_attempts=3, base_delay=1.0)
async def call_external_api():
    ...

# Idempotent API endpoint
@router.post("/payments")
async def create_payment(
    data: PaymentCreate,
    idempotency_key: str = Header(..., alias="Idempotency-Key"),
):
    return await idempotent_execute(db, idempotency_key, "/payments", process)

# Token bucket rate limiting
limiter = TokenBucketLimiter(redis_client, capacity=100, refill_rate=10)
if await limiter.is_allowed(f"user:{user_id}"):
    await handle_request()

Distributed Locks

Coordinate exclusive access to resources across multiple service instances.

RuleFileKey Pattern
Fencing Tokensrules/locks-fencing-tokens.mdOwner validation, TTL, heartbeat extension

Redis single-node locks, Redlock quorum, and PostgreSQL advisory locks are first-party documented; see Upstream coverage.

Resilience

Production-grade fault tolerance for distributed systems.

RuleFileKey Pattern
Circuit Breakerrules/resilience-circuit-breaker.mdCLOSED/OPEN/HALF_OPEN states, sliding window
Retry & Backoffrules/resilience-retry-backoff.mdExponential backoff, jitter, error classification
Bulkhead Isolationrules/resilience-bulkhead.mdSemaphore tiers, rejection policies, queue depth

Idempotency

Ensure operations can be safely retried without unintended side effects.

RuleFileKey Pattern
Idempotency Keysrules/idempotency-keys.mdDeterministic hashing, Stripe-style headers

Event-consumer dedup and database-backed idempotency storage follow the Stripe pattern; see Upstream coverage.

Rate Limiting

Protect APIs with distributed rate limiting using Redis.

RuleFileKey Pattern
Token Bucketrules/ratelimit-token-bucket.mdRedis Lua scripts, burst capacity, refill rate
Sliding Windowrules/ratelimit-sliding-window.mdSorted sets, precise counting, no boundary spikes

SlowAPI + Redis wiring and tiered limits are first-party documented; see Upstream coverage.

Edge Computing

Edge runtime patterns for Cloudflare Workers, Vercel Edge, and Deno Deploy.

RuleFileKey Pattern
Edge Workersrules/edge-workers.mdV8 isolate constraints, Web APIs, geo-routing, auth at edge
Edge Cachingrules/edge-caching.mdCache-aside at edge, CDN headers, KV storage, stale-while-revalidate

Event-Driven

Event sourcing, CQRS, saga orchestration, and reliable messaging patterns.

RuleFileKey Pattern
Event Sourcingrules/event-sourcing.mdEvent-sourced aggregates, CQRS read models, optimistic concurrency
Event Messagingrules/event-messaging.mdTransactional outbox, saga compensation, idempotent consumers

Upstream coverage (do not restate)

These topics were removed from this skill on 2026-07-31 (wrap-plus-delta thinning) because a first-party source maintains them. Consult the source; do not re-add tutorials here. Ork-specific scars for these topics live in references/ork-delta.md.

TopicFirst-party source
Redis single-node locks, Redlock algorithm and quorumhttps://redis.io/docs/latest/develop/use/patterns/distributed-locks/
PostgreSQL advisory locks (session and transaction level)https://www.postgresql.org/docs/current/explicit-locking.html#ADVISORY-LOCKS
Circuit breaker pattern, thresholds, setup and rollout guideshttps://learn.microsoft.com/azure/architecture/patterns/circuit-breaker
Bulkhead pattern deep dive (thread pool, semaphore, tiers)https://learn.microsoft.com/azure/architecture/patterns/bulkhead
Retry strategies, exponential backoff, jitter, retry budgetshttps://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/ and https://tenacity.readthedocs.io
HTTP and LLM provider error classification (retryable vs not)https://docs.claude.com/en/api/errors and https://platform.openai.com/docs/guides/error-codes
Idempotency keys, request dedup, database-backed idempotencyhttps://docs.stripe.com/api/idempotent_requests
Token bucket algorithm and Redis rate-limiting patternshttps://redis.io/glossary/rate-limiting/
FastAPI distributed rate limiting (SlowAPI middleware, tiers)https://slowapi.readthedocs.io/
LLM fallback chains, provider failover, cost trackinghttps://vercel.com/docs/ai-gateway; model ids and pricing at https://docs.claude.com/en/docs/about-claude/pricing

Key Decisions

DecisionRecommendation
Lock backendRedis for speed, PostgreSQL if already using it, Redlock for HA
Lock TTL2-3x expected operation time
Circuit breaker recoveryHalf-open probe with sliding window
Retry algorithmExponential backoff + full jitter
Bulkhead isolationSemaphore-based tiers (Critical/Standard/Optional)
Idempotency storageRedis (speed) + DB (durability), 24-72h TTL
Rate limit algorithmToken bucket for most APIs, sliding window for strict quotas
Rate limit storageRedis (distributed, atomic Lua scripts)

When NOT to Use

No separate event-sourcing/saga/CQRS skills exist; they are rules within distributed-systems. But most projects never need them.

PatternInterviewHackathonMVPGrowthEnterpriseSimpler Alternative
Event sourcingOVERKILLOVERKILLOVERKILLOVERKILLWHEN JUSTIFIEDAppend-only table with status column
Saga orchestrationOVERKILLOVERKILLOVERKILLSELECTIVEAPPROPRIATESequential service calls with manual rollback
Circuit breakerOVERKILLOVERKILLBORDERLINEAPPROPRIATEREQUIREDTry/except with timeout
Distributed locksOVERKILLOVERKILLBORDERLINEAPPROPRIATEREQUIREDDatabase row-level lock (SELECT FOR UPDATE)
CQRSOVERKILLOVERKILLOVERKILLOVERKILLWHEN JUSTIFIEDSingle model for read/write
Transactional outboxOVERKILLOVERKILLOVERKILLSELECTIVEAPPROPRIATEDirect publish after commit
Rate limitingOVERKILLOVERKILLSIMPLE ONLYAPPROPRIATEREQUIREDNginx rate limit or cloud WAF

Rule of thumb: If you have a single server process, you do not need distributed systems patterns. Use in-process alternatives. Add distribution only when you actually have multiple instances.

Anti-Patterns (FORBIDDEN)

# LOCKS: Never forget TTL (causes deadlocks)
await redis.set(f"lock:{name}", "1")  # WRONG - no expiry!

# LOCKS: Never release without owner check
await redis.delete(f"lock:{name}")  # WRONG - might release others' lock

# RESILIENCE: Never retry non-retryable errors
@retry(max_attempts=5, retryable_exceptions={Exception})  # Retries 401!

# RESILIENCE: Never put retry outside circuit breaker
@retry  # Would retry when circuit is open!
@circuit_breaker
async def call(): ...

# IDEMPOTENCY: Never use non-deterministic keys
key = str(uuid.uuid4())  # Different every time!

# IDEMPOTENCY: Never cache error responses
if response.status_code >= 400:
    await cache_response(key, response)  # Errors should retry!

# RATE LIMITING: Never use in-memory counters in distributed systems
request_counts = {}  # Lost on restart, not shared across instances

Detailed Documentation

ResourceDescription
scripts/Templates: lock implementations, circuit breaker, rate limiter
references/ork-delta.mdOrk-specific scars and house decisions kept after the wrap-plus-delta thinning

Related Skills

  • caching - Redis caching patterns, cache as fallback
  • background-jobs - Job deduplication, async processing with retry
  • observability-monitoring - Metrics and alerting for circuit breaker state changes
  • error-handling-rfc9457 - Structured error responses for resilience failures
  • auth-patterns - API key management, authentication integration
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

src/skills/distributed-systems

Default branch

main

Latest commit

43c04fa

Tree SHA

29981ce