sc-log-fix

v2026.09.25

Analyze application logs, identify errors, perform root cause analysis, and interactively fix issues with a learned knowledge base. Use when debugging from logs, triaging errors, or fixing recurring issues.

GitHub
Install command
npx skhub add tony363/sc-log-fix
Markdown
SKILL.md

Log-Fix: Iterative Bug Fixing from Log Analysis

Analyze logs from multiple sources, identify errors, perform root cause analysis, and interactively fix issues with pattern learning.

Quick Start

# Analyze recent errors from all log sources
/sc:log-fix

# Filter by time and severity
/sc:log-fix --time 1h --level ERROR

# Target specific log source
/sc:log-fix --source backend --component app.services

# Trace a specific request
/sc:log-fix --request-id abc-123-def

# Dry run - analyze without fixing
/sc:log-fix --dry-run

Behavioral Flow

  1. Discover - Detect available log sources (files, docker, systemd)
  2. Load Knowledge - Check for learned patterns from previous fixes
  3. Collect - Parse and normalize logs from detected sources
  4. Triage - Cluster errors, rank by severity, match known patterns
  5. Analyze - Root cause analysis with stack traces and code inspection
  6. Fix - Interactive fix loop with user approval
  7. Learn - Save successful fix patterns to knowledge base
  8. Report - Summary of session with remaining issues

Flags

FlagTypeDefaultDescription
--timestring1hTime range: 30m, 1h, 1d, 7d
--sourcestringallLog source: backend, frontend, docker, all
--levelstringERRORMinimum severity: ERROR, WARNING, all
--componentstring-Filter by logger/module name
--request-idstring-Trace specific request across logs
--dry-runboolfalseAnalyze and report without applying fixes
--fixbooltrueEnable interactive fix loop

Phase 0: Log Source Discovery

Auto-detect all available log sources before analysis.

Detection Strategy

SourceDetection MethodCommon Locations
File logsGlob for logs/**/*.log, *.loglogs/, var/log/
DockerCheck docker compose psdocker-compose.yml
SystemdCheck journalctl availabilitySystem services
PM2Check pm2 listNode.js apps

Log Format Detection

Auto-detect log format:

  • JSON (structured) - Parse with jq
  • Text (unstructured) - Parse with regex patterns
  • Combined (nginx/apache style) - Parse with format-specific regex
# Detect format by reading first line
head -1 logs/app.log | python3 -c "import sys,json; json.loads(sys.stdin.read()); print('json')" 2>/dev/null || echo "text"

Phase 1: Knowledge Loading

Load historical fix knowledge for enhanced analysis.

Knowledge File Structure

{
  "version": "1.0",
  "error_patterns": [
    {
      "id": "pattern-001",
      "signature": {
        "message_pattern": "Connection refused.*port \\d+",
        "component": "app.services.database",
        "exception_type": "ConnectionError"
      },
      "fixes": [{
        "description": "Check database is running, verify connection string",
        "file_changed": "src/config.py",
        "validation_command": "pytest tests/test_db.py -v",
        "success_count": 3,
        "confidence": 0.9
      }]
    }
  ],
  "component_knowledge": {},
  "metadata": {
    "total_fixes": 0,
    "updated_at": null
  }
}

Location: .claude/knowledge/log-fix-knowledge.json

When knowledge is loaded:

  • Triage: Show "Known issue" badge for previously seen errors
  • Root Cause: Suggest fixes with historical success rates
  • Fix Loop: Offer one-click application of proven fixes

Phase 2: Triage and Clustering

Incident Clustering

Group by:

  1. Same error type (identical/similar messages)
  2. Same component (logger/service name)
  3. Same request (request_id correlation)
  4. Same time window (errors within 5s)

Severity Ranking

PriorityCriteria
P0Stack trace, service crash, database errors
P1ERROR level, API failures, auth issues
P2WARNING level, deprecation, performance
P3Informational, cleanup suggestions

Triage Summary Output

## Error Triage Summary

**Time Range**: [start] to [end]
**Total Errors**: N | **Unique Types**: M | **Known Patterns**: K

### Top Issues by Frequency

| # | Count | Component | Error Pattern | Priority | Historical |
|---|-------|-----------|---------------|----------|------------|
| 1 | 15 | app.services.api | Connection timeout | P1 | Known: 3 fixes, 90% |
| 2 | 8 | app.models | Validation failed | P1 | New pattern |

Phase 3: Root Cause Analysis

For each incident cluster:

  1. Stack Trace Analysis - Extract file paths, identify failing line, map to components
  2. Log Correlation - Follow request_id across services
  3. Pattern Matching - Check knowledge base, then use code-level analysis
  4. Code Inspection - Read relevant source files, identify the bug

Root Cause Output

## Root Cause: [Error Type]

**Occurrences**: N | **Component**: [logger]

### Error Details
[Full error and stack trace]

### Relevant Code
**File**: `src/services/api.py:145`
[Code context with highlighted line]

### Probable Cause
[Analysis explaining why the error occurs]

### Suggested Fix
[Specific code change with diff preview]

Phase 4: Interactive Fix Loop

Fix Workflow

  1. Present - Show error, root cause, affected files
  2. Show diff - Display exact code changes proposed
  3. Ask permission - User approves, skips, or modifies
  4. Apply if approved - Use Edit tool
  5. Validate - Run relevant tests
  6. Learn - Offer to save pattern to knowledge base

User Options

ActionDescription
ApplyApply the proposed fix
SkipSkip this issue, move to next
EditModify the proposed fix before applying
Test firstWrite a test before fixing (invoke /sc:tdd)
DetailsShow more context about the error
QuitExit fix loop

Validation After Fix

Fix TypeValidation
Python codepytest -k "test_function" -v
JavaScriptnpm test -- --testPathPattern=file
ConfigRestart and verify health
DatabaseRun migrations

Phase 5: Knowledge Learning

After a successful fix, offer to save the pattern:

  1. Extract error signature - Convert specific values to regex patterns
  2. Record fix details - What changed, where, validation command
  3. Update knowledge file - Append or update existing pattern
  4. Report confidence - Initial confidence 1.0, adjusted over time

Learning Prompt

Fix applied and validated!

Save this pattern for future similar errors?
  [Y] Yes, save pattern and fix
  [n] No, this was a one-off fix
  [e] Yes, but let me edit the description first

Phase 6: Session Report

## Log-Fix Session Summary

**Logs Analyzed**: [sources] | **Time Range**: [range]
**Errors Found**: N | **Unique Issues**: M

### Issues Addressed

| # | Component | Issue | Status | Action |
|---|-----------|-------|--------|--------|
| 1 | app.services | Connection timeout | FIXED | Added retry config |
| 2 | app.models | Validation error | SKIPPED | User deferred |

### Files Modified
- src/services/api.py
- src/config.py

### Patterns Learned
- 1 new pattern saved to knowledge base

### Next Steps
- Run `/sc:pr-check` to validate all changes
- Consider `/sc:tdd` for issues without test coverage

MCP Integration

PAL MCP (Debugging & Analysis)

ToolWhen to UsePurpose
mcp__pal__debugComplex bugsMulti-stage root cause analysis
mcp__pal__thinkdeepUnclear patternsDeep investigation of recurring issues
mcp__pal__codereviewFix validationReview proposed fix quality
mcp__pal__apilookupDependency errorsGet current docs for version issues

PAL Usage Patterns

# Debug complex error pattern
mcp__pal__debug(
    step="Investigating recurring connection timeout in API layer",
    hypothesis="Connection pool exhaustion under load",
    confidence="medium",
    relevant_files=["src/services/api.py", "src/config.py"]
)

# Deep analysis of unclear pattern
mcp__pal__thinkdeep(
    step="Why do these errors only occur during peak hours?",
    hypothesis="Race condition in connection pooling",
    confidence="low"
)

Rube MCP (Notifications)

ToolWhen to UsePurpose
mcp__rube__RUBE_SEARCH_TOOLSExternal loggingFind logging service tools
mcp__rube__RUBE_MULTI_EXECUTE_TOOLNotificationsPost fix summary to Slack/Jira

Tool Coordination

  • Bash - Log parsing (jq, grep), test execution, docker commands
  • Glob - Log file discovery
  • Grep - Error pattern search, request ID tracing
  • Read - Source code inspection, log file reading
  • Edit - Apply code fixes
  • Write - Knowledge base updates, session reports

Related Skills

  • /sc:tdd - Write tests before fixing (TDD approach)
  • /sc:pr-check - Validate all changes before PR
  • /sc:analyze - Deeper code analysis when root cause is unclear
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.25

Published

Sep 25, 2026

Category

Uncategorized

License

MIT

Source path

.claude/skills/sc-log-fix

Default branch

main

Latest commit

6634f8e

Tree SHA

993acfd