io-bound-data-processing

v2026.09.24

Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs random, mmap, async), data formats (Parquet vs CSV vs JSON, predicate pushdown), chunking & batching, spill-to-disk (external merge sort, DuckDB/Polars), pipelining (bounded queues, backpressure, checkpointing), codec selection (zstd/lz4/gzip), concurrency for I/O-bound workloads (asyncio, threads, prefetch), and observability (iowait vs CPU%, rows/sec, py-spy/strace). Trigger on "process a large file", "stream this", "out-of-core", "OOM kill", "this is slow", or code with `pd.read_csv` of multi-GB files, `requests.get(...).content` on big bodies, `BytesIO` on unbounded inputs, per-row INSERTs, sequential `requests.get` loops, falling `tqdm` rates — even if I/O or memory isn't mentioned. Complement to computer-science-algorithms.

GitHub
Install command
npx skhub add pproenca/io-bound-data-processing
Markdown
SKILL.md

Community I/O-bound data processing on constrained resources Best Practices

A reference for engineers processing datasets larger than RAM on a single low-compute box. Organized by execution-lifecycle impact: rules near the top of the table govern whether the job runs at all; rules near the bottom shave the last 10 %. Optimize from the top of the waterfall.

Scope: the patterns that show up in real ETL / data-engineering / batch work on a laptop, a 2-vCPU container, or a Raspberry Pi-class node — streaming, formats, chunking, spill, backpressure, codecs, and the concurrency model that actually matches an I/O-bound bottleneck. Out of scope (covered elsewhere): the algorithmic primitives themselves (see computer-science-algorithms), distributed compute beyond a single box (use Spark/Dask), and database-engine internals (see official docs).

Distilled from Apache Arrow / Parquet docs, Polars User Guide, DuckDB docs, pandas — Scaling to large datasets, Linux man pages (mmap(2), sendfile(2), posix_fadvise(2)), Brendan Gregg's USE method and Systems Performance, Kleppmann's Designing Data-Intensive Applications, and the zstd / lz4 reference benchmarks.

When to Apply

Reach for these rules when:

  • A job OOM-kills, swaps, or runs much slower than expected on a small box
  • Input is larger than RAM and you need to scan, filter, aggregate, sort, or join it
  • A pipeline has unbounded buffers between stages, or memory grows linearly during a "streaming" job
  • You see one-row-per-RTT writes (INSERT per row, requests.get per URL, f.read(32) per record)
  • You're picking a format/codec/serializer and the choice matters at scale
  • A top shows low CPU and high iowait, or you don't know which it is
  • "It's slow but I don't know why" — start at the obs- category

Rule Categories By Priority

#CategoryPrefixImpactWhy it cascades
1Memory Disciplinemem-CRITICALSlurping into RAM defeats every downstream technique on a constrained box
2I/O Access Patternsio-CRITICALDisk/net are 10³–10⁶× slower than RAM; access pattern dominates wall-clock
3Data Format & Encodingfmt-HIGHFormat fixes the lower bound on I/O volume + decode cost before any logic runs
4Chunking & Batchingbatch-HIGHGranularity controls peak memory and amortizes per-item overhead
5Spill-to-Disk & External Memoryspill-HIGHWhen data > RAM, the choice is "spill cleanly" or "OOM"
6Pipelining & Backpressurepipe-MEDIUM-HIGHUnbounded buffers between fast producers and slow sinks = OOM
7Compression & Serializationcodec-MEDIUMTrades CPU for I/O; right codec saves orders of magnitude
8Concurrency for I/O-Bound Workloadsconc-MEDIUMAsync / threads / processes are different tools; wrong model wastes CPU
9Observability & Throughput Tuningobs-LOW-MEDIUMCan't tune what you don't measure; iowait ≠ CPU-bound

Quick Reference

1. Memory Discipline (CRITICAL)

2. I/O Access Patterns (CRITICAL)

3. Data Format & Encoding (HIGH)

4. Chunking & Batching (HIGH)

5. Spill-to-Disk & External Memory (HIGH)

6. Pipelining & Backpressure (MEDIUM-HIGH)

7. Compression & Serialization (MEDIUM)

8. Concurrency for I/O-Bound Workloads (MEDIUM)

9. Observability & Throughput Tuning (LOW-MEDIUM)

How to Use

Start with the question that matches the problem:

Code examples are in Python (most readable across audiences). The reasoning generalizes — equivalent libraries in other ecosystems (Arrow C++/Rust/Java, Polars Rust, DuckDB everywhere, libuv-style async) follow the same patterns.

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for adding new rules
metadata.jsonVersion and reference information
AGENTS.mdAuto-built TOC navigation

Related Skills

  • computer-science-algorithms — Algorithmic primitives this skill builds on (external merge sort, sketches, hash partitioning, sampling)
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/.experimental/io-bound-data-processing

Default branch

master

Latest commit

cf93c57

Tree SHA

afbb575