io-bound-data-processing

v2026.09.24

Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs random, mmap, async), data formats (Parquet vs CSV vs JSON, predicate pushdown), chunking & batching, spill-to-disk (external merge sort, DuckDB/Polars), pipelining (bounded queues, backpressure, checkpointing), codec selection (zstd/lz4/gzip), concurrency for I/O-bound workloads (asyncio, threads, prefetch), and observability (iowait vs CPU%, rows/sec, py-spy/strace). Trigger on "process a large file", "stream this", "out-of-core", "OOM kill", "this is slow", or code with `pd.read_csv` of multi-GB files, `requests.get(...).content` on big bodies, `BytesIO` on unbounded inputs, per-row INSERTs, sequential `requests.get` loops, falling `tqdm` rates — even if I/O or memory isn't mentioned. Complement to computer-science-algorithms.

GitHub
安装命令
npx skhub add pproenca/io-bound-data-processing
Markdown
SKILL.md

Community I/O-bound data processing on constrained resources Best Practices

A reference for engineers processing datasets larger than RAM on a single low-compute box. Organized by execution-lifecycle impact: rules near the top of the table govern whether the job runs at all; rules near the bottom shave the last 10 %. Optimize from the top of the waterfall.

Scope: the patterns that show up in real ETL / data-engineering / batch work on a laptop, a 2-vCPU container, or a Raspberry Pi-class node — streaming, formats, chunking, spill, backpressure, codecs, and the concurrency model that actually matches an I/O-bound bottleneck. Out of scope (covered elsewhere): the algorithmic primitives themselves (see computer-science-algorithms), distributed compute beyond a single box (use Spark/Dask), and database-engine internals (see official docs).

Distilled from Apache Arrow / Parquet docs, Polars User Guide, DuckDB docs, pandas — Scaling to large datasets, Linux man pages (mmap(2), sendfile(2), posix_fadvise(2)), Brendan Gregg's USE method and Systems Performance, Kleppmann's Designing Data-Intensive Applications, and the zstd / lz4 reference benchmarks.

When to Apply

Reach for these rules when:

  • A job OOM-kills, swaps, or runs much slower than expected on a small box
  • Input is larger than RAM and you need to scan, filter, aggregate, sort, or join it
  • A pipeline has unbounded buffers between stages, or memory grows linearly during a "streaming" job
  • You see one-row-per-RTT writes (INSERT per row, requests.get per URL, f.read(32) per record)
  • You're picking a format/codec/serializer and the choice matters at scale
  • A top shows low CPU and high iowait, or you don't know which it is
  • "It's slow but I don't know why" — start at the obs- category

Rule Categories By Priority

#CategoryPrefixImpactWhy it cascades
1Memory Disciplinemem-CRITICALSlurping into RAM defeats every downstream technique on a constrained box
2I/O Access Patternsio-CRITICALDisk/net are 10³–10⁶× slower than RAM; access pattern dominates wall-clock
3Data Format & Encodingfmt-HIGHFormat fixes the lower bound on I/O volume + decode cost before any logic runs
4Chunking & Batchingbatch-HIGHGranularity controls peak memory and amortizes per-item overhead
5Spill-to-Disk & External Memoryspill-HIGHWhen data > RAM, the choice is "spill cleanly" or "OOM"
6Pipelining & Backpressurepipe-MEDIUM-HIGHUnbounded buffers between fast producers and slow sinks = OOM
7Compression & Serializationcodec-MEDIUMTrades CPU for I/O; right codec saves orders of magnitude
8Concurrency for I/O-Bound Workloadsconc-MEDIUMAsync / threads / processes are different tools; wrong model wastes CPU
9Observability & Throughput Tuningobs-LOW-MEDIUMCan't tune what you don't measure; iowait ≠ CPU-bound

Quick Reference

1. Memory Discipline (CRITICAL)

2. I/O Access Patterns (CRITICAL)

3. Data Format & Encoding (HIGH)

4. Chunking & Batching (HIGH)

5. Spill-to-Disk & External Memory (HIGH)

6. Pipelining & Backpressure (MEDIUM-HIGH)

7. Compression & Serialization (MEDIUM)

8. Concurrency for I/O-Bound Workloads (MEDIUM)

9. Observability & Throughput Tuning (LOW-MEDIUM)

How to Use

Start with the question that matches the problem:

Code examples are in Python (most readable across audiences). The reasoning generalizes — equivalent libraries in other ecosystems (Arrow C++/Rust/Java, Polars Rust, DuckDB everywhere, libuv-style async) follow the same patterns.

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for adding new rules
metadata.jsonVersion and reference information
AGENTS.mdAuto-built TOC navigation

Related Skills

  • computer-science-algorithms — Algorithmic primitives this skill builds on (external merge sort, sketches, hash partitioning, sampling)
发现
标签

此技能尚未发布标签。

版本
最新版本元数据

版本

v2026.09.24

发布时间

Sep 24, 2026

分类

未分类

许可证

MIT

源路径

skills/.experimental/io-bound-data-processing

默认分支

master

最新提交

cf93c57

Tree SHA

afbb575