macos-disk-usage

v2026.09.24

macOS disk-usage forensics and space recovery on APFS. Use when df reports a disk near-full, hunting what's eating space, reclaiming OrbStack/Docker, or pruning caches and local snapshots.

GitHub
Install command
npx skhub add laurigates/macos-disk-usage
Markdown
SKILL.md

macOS Disk Usage & Space Recovery

When to Use This Skill

Use this skill when...Use something else when...
df reports a volume near-full and you need the honest figureThe root disk on Linux is full — use the Linux-focused disk-full-recovery rule
Hunting "what's eating my disk" on macOSGeneral process forensics — use ps/top directly
Reclaiming space from OrbStack/Docker, caches, or local snapshotsA GUI hang/freeze occurred — see macos-incident-postmortem
du totals and free space disagree (purgeable snapshots, or CoW clones)launchservicesd DB bloat specifically — see launchservices-health
Deleting a big tree freed far less than du promised—

Platform Guard

This skill is macOS-only. APFS containers, diskutil apfs, tmutil local snapshots, and BSD du block-counting are Darwin-specific. Refuse to act if uname -s is not Darwin.

test "$(uname -s)" = "Darwin" || { echo "macos-plugin: not Darwin, refusing"; exit 1; }

This skill cross-references, does not duplicate, the Linux-focused disk-full-recovery rule (root partition, ~/.cache, apt/journalctl). The APFS / Time Machine / OrbStack angle below is the macOS-specific complement.

Core Expertise

The single most-important macOS fact: APFS shares free space across every volume in a container. A per-volume Capacity % is therefore misleading — the sealed system snapshot volume routinely reads 99% while the container has tens of GB free. Trust the Avail column (or diskutil apfs list), never the Capacity %, or the investigation goes down the wrong path before it starts.

The second fact: du lies about reclaimable space. tmutil local snapshots hold purgeable space that du never counts, so du totals and reported free space can disagree by tens of GB. Check snapshots first when the numbers don't reconcile.

The third fact: du also over-counts in the other direction — copy-on-write clones. APFS reports a cloned file at full st_blocks even though its blocks are shared and cost nothing, so a du total can far exceed what deleting the tree would actually free. This is not fixable by switching tools: dust, gdu, and ncdu all read st_blocks and double-count identically. See CoW clones below.

Step 1: Measure honestly (built-in CLI)

# Trust Avail, NOT Capacity % — APFS shares free space container-wide
df -h /

# The container's real free space, per-volume roles, and snapshot count
diskutil apfs list

# Top-level consumers, one level deep, staying on one filesystem (-x)
du -hx -d1 / 2>/dev/null | sort -h
du -hx -d1 ~ 2>/dev/null | sort -h

BSD du counts allocated blocks, so it reports the real on-disk size of sparse files (e.g. OrbStack's data.img.raw reads its true 33 GB, not its apparent size). Trust the number for sparse files — but not as a general rule: allocated blocks are also what makes du over-count clones (below), and they never include purgeable snapshot space.

CoW clones: du over-counts what deleting would free

A copy-on-write clone (cp -c, clonefile(2)) shares blocks with its source, so it costs ~nothing — but every tool reading st_blocks bills it at full size. Measured on a 200 MB file plus one clone (ground truth: 200 MB):

ToolReports
BSD du, GNU du, dust, gdu400M
dust -s (apparent)600M — also counts hardlinks
df delta0 MB for the clone ✅

All of them dedupe hardlinks by inode; a clone has a distinct inode, so that logic never fires.

This is not academic — it is the default on this platform for two common package managers:

ToolGlobal cache2nd project's copy costsdu claims
uv → .venv~/.cache/uv0 bytes (clonefile)full size
bun → node_modules~/.bun/install/cache0 bytes (clonefile)full size
cargo → target/~/.cargo/registry (sources only)full size — each project compiles its own copyaccurate

So du is honest for target/ and inflated for node_modules/.venv. Deleting a clone frees real space only where it holds the last reference to those blocks.

When a size gates a decision ("how much will this free?"), measure a df delta — it is the only clone-aware measurement available:

before=$(df -k / | awk 'NR==2{print $4}'); rm -rf <target>; sync
echo "freed $(( ($(df -k / | awk 'NR==2{print $4}') - before) / 1048576 )) GB"

Real cost of skipping this: a reclaim sweep summing du predicted 68.7 GB where the df delta was 17.7 GB — a 4× overstatement (2026-08). Clones are detectable via fcntl(F_LOG2PHYS_EXT) — a clone shares physical device offsets while having a distinct inode, and stat/MetadataExt carries no clone signal at all — but no general-purpose tool implements it yet (bootandy/dust#590).

Purgeable snapshot space (the thing du hides)

When du totals and df free space disagree, local Time Machine snapshots are the usual cause:

# List local snapshots holding purgeable space
tmutil listlocalsnapshots /

# Thin them (reclaims purgeable space invisible to du); needs sudo
sudo tmutil thinlocalsnapshots / 999999999999 4

Step 1.5: Fast-path — the usual-suspects scan

Step 1 gives the honest total; this gives where it went without re-deriving the hunt by hand each time. scripts/scan-suspects.sh reads a public catalog of known reclaim targets (scripts/suspects.tsv — dev caches, build-artifact dirs, VM images), dus the ones that exist on this machine, and emits a ranked rollup with the cleanup command and safety tier already attached:

# Whole home (thorough; slower — du walks every tree)
bash "${CLAUDE_SKILL_DIR}/scripts/scan-suspects.sh" --home-dir "$HOME" --root "$HOME"

# Bound the build-artifact search to where projects live, raise the noise floor
bash "${CLAUDE_SKILL_DIR}/scripts/scan-suspects.sh" --root ~/repos --min-mb 1000

Output follows the structured SUSPECT id=… tier=… size_kb=… cmd="…" convention, plus a ranked human table and per-tier totals (TIER_SAFE_KB=…, RECLAIMABLE_SAFE=…). Work the tiers exactly as Step 4 — exhaust safe before decision, never blind-delete userdata.

⚠️ Those totals are du sums, so they are an upper bound wherever CoW clones apply — notably the node_modules / .venv suspects, which uv and bun clone from a global cache. Treat RECLAIMABLE_* as "at most this much"; confirm with a df delta after acting. See CoW clones.

This augments Step 1, it does not replace it. The scan is pure measurement; the judgment stays with the agent — reading the APFS Avail-not-Capacity % picture, checking purgeable snapshots when du and df disagree, and deciding which decision/userdata items to actually remove.

The UNKNOWN lines are the feedback loop. Any big directory not in the catalog is surfaced as UNKNOWN size_kb=… path=…. When one recurs, add it to suspects.tsv as a pattern (never a machine-specific path — this repo is public) and open a PR; the catalog grows from real findings. Catalog columns and placeholders ({dir}, {parent}) are documented in the file's header.

Two operational notes:

  • Runtime scales with disk size — du walks every tree, so a full disk can take a minute or two. Bound with --root <project-dir> and --min-mb; swap in dust (Step 3) if you want the walk parallelised.
  • The catalog is the shared artifact; the scan output is machine-specific — keep it in scratch, never commit it.

Step 2: The OrbStack / Docker recovery play

OrbStack/Docker is frequently the dominant consumer (100+ GB seen). The VM image (data.img.raw) grows but never shrinks on its own. The recovery is: prune inside the VM, then OrbStack auto-shrinks the host image (unlike Docker Desktop, which needs manual reclaim).

# See what's reclaimable before pruning
docker system df

# Prune unused images + build cache (NOT volumes — see warning below)
docker system prune -a -f

A real session reclaimed 85 GB this way (159 images / 57 GB reclaimable, 42 GB build cache), after which OrbStack auto-shrank the host image 103 GB → 33 GB.

Never blind-prune named volumes

docker volume ls -f dangling=true lists named volumes (olmap_postgres_data, docker_neo4j_data, …) that are "unused" only because no container is currently running — they hold live project databases. A blind docker volume prune would wipe them.

# Inspect first — named volumes are project data, anonymous-hash ones are scratch
docker volume ls -f dangling=true

Prune only anonymous-hash volumes (long hex names), never the human-named ones, and only after confirming with the user.

Step 3: Third-party tooling

Install via the mise aqua: backend (checksum-verified standalone binaries):

mise use -g aqua:bootandy/dust   # `dust` — fast tree-map, the one to lead with
mise use -g aqua:Byron/dua-cli   # `dua`  — interactive TUI via `dua i`
ToolInstallStrength
dustaqua:bootandy/dustFast visual tree; lead with this. dust -r reverse, -d N depth, -X <glob> exclude, -s apparent size, -j JSON to stdout
duaaqua:Byron/dua-cliInteractive deletion TUI (dua i)
gdu / ncduaqua:dundee/gdu, ncduTUI disk usage analyzers
diskonautaqua:imsnif/diskonautSpatial treemap navigator

GUI options (mention, don't install): DaisyDisk, GrandPerspective, OmniDiskSweeper.

Step 4: Tiered cleanup

Work top-down — exhaust the safe tier before touching anything that needs a decision.

TierWhatHowNotes
Safe / regenerableHomebrew downloadsbrew cleanup -sRebuilds on next install
Go module cachego clean -modcacheRe-downloads on next build
pip cachepip cache purge
uv cacheuv cache cleanFails with a lock error if another uv process / active session is running
Playwright / Yarn / node-gyp cachesrm -rf ~/Library/Caches/{ms-playwright,Yarn,node-gyp}Regenerable
Needs a decisionDocker images + build cachedocker system prune -a -fOrbStack auto-shrinks host image after
Local Time Machine snapshotssudo tmutil thinlocalsnapshots /Loses local restore points
User dataNamed Docker volumesconfirm eachLive project databases — never blind-prune
Anything in ~/Documents, ~/Downloadsuser confirmsIrreplaceable

Agentic Optimizations

ContextCommand
Ranked usual-suspects rollupbash "${CLAUDE_SKILL_DIR}/scripts/scan-suspects.sh" --root ~/repos
Honest free spacedf -h / | awk 'NR==2{print $4" avail"}'
Container free spacediskutil apfs list | grep -i 'Capacity In Use|Free'
Top home consumersdu -hx -d1 ~ 2>/dev/null | sort -h | tail -15
Snapshot counttmutil listlocalsnapshots / | grep -c com.apple
Docker reclaimabledocker system df
Fast tree (dust)dust -d2 -r ~
Machine-readable sizes (dust)dust -j -o b -d1 <dir> | jq -r '.children[] | [(.size | rtrimstr("B") | tonumber), .name] | @tsv' | sort -rn — -j emits a {size, name, children} tree to stdout; -o b makes sizes bytes ("12345B") instead of human strings

Quick Reference

NeedCommand
Fast-path suspect scanbash "${CLAUDE_SKILL_DIR}/scripts/scan-suspects.sh"
Honest free spacedf -h / (read Avail, not Capacity %)
APFS container truthdiskutil apfs list
Disk hog scandu -hx -d1 <dir> or dust -d2 -r <dir>
Purgeable snapshotstmutil listlocalsnapshots /
Thin snapshotssudo tmutil thinlocalsnapshots / 999999999999 4
Docker reclaimdocker system df then docker system prune -a -f

Error Handling

SymptomCauseFix
df shows 99% but space "missing"APFS per-volume Capacity % on the sealed snapshot volumeRead Avail, or diskutil apfs list
du total ≪ used spacePurgeable local snapshotstmutil listlocalsnapshots /; thin them
Deleted a big tree, df barely movedCoW clones — blocks were shared, du billed them per copyExpected for .venv/node_modules; measure a df delta, and prune the global cache to free the shared blocks
uv cache clean lock errorAnother uv process / active session runningClose other uv work and retry
Docker image still huge after pruneReclaim ran inside VM; host image not yet shrunkOrbStack auto-shrinks shortly; Docker Desktop needs manual reclaim
volume prune wiped a databaseNamed volume pruned while no container ranRestore from backup; only prune anonymous-hash volumes

Related Skills

  • launchservices-health — when launchservicesd DB bloat (not disk usage) is the concern
  • macos-incident-postmortem — when a freeze/panic, not a full disk, is the symptom
  • disk-full-recovery (user-global rule) — the Linux-focused complement to this APFS-specific skill
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

macos-plugin/skills/macos-disk-usage

Default branch

main

Latest commit

1668324

Tree SHA

b2d4cc3