juicebox-incident-runbook

v2026.09.24

Juicebox incident response. Trigger: "juicebox incident", "juicebox outage".

GitHub
安装命令
npx skhub add jeremylongshore/juicebox-incident-runbook
Markdown
SKILL.md

Juicebox Incident Runbook

Overview

Incident response procedures for Juicebox AI analysis platform integration failures. Covers analysis timeouts, dataset corruption, quota exhaustion, and export failures. Juicebox powers AI-driven people search and candidate analysis, so incidents disrupt recruiting pipelines, talent intelligence workflows, and automated sourcing. Classify severity immediately using the matrix below and follow the corresponding playbook.

Severity Levels

LevelDefinitionResponse TimeExample
P1 - CriticalFull API outage or dataset corruption15 minHealth endpoint returns 5xx, analysis results missing
P2 - HighAnalysis timeouts or export failures30 minSearch queries hang beyond 30s, CSV exports fail
P3 - MediumQuota exhaustion or rate limiting2 hours429 responses, account quota at 100% usage
P4 - LowPartial data or degraded result quality8 hoursSearch returns fewer results than expected

Diagnostic Steps

# Check API health
curl -s -o /dev/null -w "HTTP %{http_code}\n" \
  -H "Authorization: Bearer $JUICEBOX_API_KEY" \
  https://api.juicebox.ai/v1/health

# Check account quota usage
curl -s -H "Authorization: Bearer $JUICEBOX_API_KEY" \
  https://api.juicebox.ai/v1/account/quota | jq '.used, .limit, .remaining'

# Test a minimal search request
curl -s -w "\nHTTP %{http_code}\n" \
  -H "Authorization: Bearer $JUICEBOX_API_KEY" \
  -H "Content-Type: application/json" \
  -X POST https://api.juicebox.ai/v1/search \
  -d '{"query": "software engineer", "limit": 1}'

Incident Playbooks

API Outage

  1. Confirm via health endpoint and status.juicebox.ai
  2. Activate fallback mode — serve cached search results to active users
  3. Pause any automated sourcing pipelines to avoid wasting quota on retries
  4. Notify recruiting team that live search is temporarily unavailable
  5. Monitor status page and resume operations once health check passes

Authentication Failure

  1. Verify API key is set: echo $JUICEBOX_API_KEY | wc -c
  2. Test with health endpoint (see diagnostics above)
  3. If 401: API key may be revoked — regenerate in Juicebox dashboard
  4. If 403: check account tier permissions for the requested endpoint
  5. Deploy new key and verify search requests succeed

Data Sync Failure

  1. Check if recent analysis results are returning stale or incomplete data
  2. Verify export endpoints are responding — test a small CSV export
  3. If exports fail: check if the analysis job completed successfully first
  4. For dataset corruption: re-trigger the analysis with fresh parameters
  5. Contact Juicebox support with job IDs and error responses

Communication Template

**Incident**: Juicebox Integration [Outage/Degradation]
**Status**: [Investigating/Identified/Mitigating/Resolved]
**Started**: YYYY-MM-DD HH:MM UTC
**Impact**: [Search unavailable / exports failing / quota exhausted / N recruiting workflows paused]
**Current action**: [Cached results active / quota upgrade requested / re-analysis running]
**Next update**: HH:MM UTC

Post-Incident

  • Document timeline from detection to resolution
  • Identify root cause (Juicebox outage / quota burn / export bug / auth expiry)
  • Calculate impact: missed candidates, paused pipelines, stale data duration
  • Add quota usage alerting at 80% threshold to prevent exhaustion
  • Implement request caching to reduce redundant API calls
  • Review automated pipeline frequency to avoid quota spikes

Error Handling

Incident TypeDetectionResolution
Analysis timeoutRequests exceeding 30s SLAReduce query complexity, retry with smaller scope
Dataset corruptionMissing or inconsistent analysis resultsRe-trigger analysis job, verify input parameters
Quota exhaustion429 responses, quota endpoint shows 0 remainingPause automation, request quota increase, optimize usage
Export failureCSV/JSON export returns error or empty payloadVerify analysis job completed, retry export with job ID

Prerequisites

  • An incident owner, synthetic lead fixture, approved support path, source-authority record, and a redaction rule for contact details and enrichment data.

Instructions

  1. Open a timestamped incident, classify impact, and reproduce once with the synthetic fixture.
  2. Isolate identity, source authority, enrichment, suppression, destination, quota, or export failure using status and opaque correlation IDs only.
  3. Freeze nonessential jobs, apply one reversible remediation, and stop any export when source, scope, or destination cannot be verified.
  4. Verify recovery with least-privilege access and a suppression check; delete staged artifacts and document rollback before closure.

Output

Produce an incident receipt with severity, UTC window, affected integration, redacted error class, source/destination/suppression state, remediation/rollback decision, and next action. Exclude names, contact data, enrichment values, and credentials.

Examples

P2; source=synthetic-prospects; error=429; action=bounded-backoff; destination=held; suppression=pass; contacts_exported=0; rollback=not-needed is a safe incident summary.

Resources

  • Juicebox Status
  • Juicebox API Docs

Next Steps

See juicebox-observability for monitoring setup and quota tracking dashboards.

发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

2026年9月24日

分类

未分类

许可证

MIT

源路径

skills/.curated/juicebox-incident-runbook

默认分支

main

最新提交

e5a6c3b

Tree SHA

c2dc8e8