Community I/O-bound data processing on constrained resources Best Practices
A reference for engineers processing datasets larger than RAM on a single low-compute box. Organized by execution-lifecycle impact: rules near the top of the table govern whether the job runs at all; rules near the bottom shave the last 10 %. Optimize from the top of the waterfall.
Scope: the patterns that show up in real ETL / data-engineering / batch work on a laptop, a 2-vCPU container, or a Raspberry Pi-class node — streaming, formats, chunking, spill, backpressure, codecs, and the concurrency model that actually matches an I/O-bound bottleneck. Out of scope (covered elsewhere): the algorithmic primitives themselves (see computer-science-algorithms), distributed compute beyond a single box (use Spark/Dask), and database-engine internals (see official docs).
Distilled from Apache Arrow / Parquet docs, Polars User Guide, DuckDB docs, pandas — Scaling to large datasets, Linux man pages (mmap(2), sendfile(2), posix_fadvise(2)), Brendan Gregg's USE method and Systems Performance, Kleppmann's Designing Data-Intensive Applications, and the zstd / lz4 reference benchmarks.
When to Apply
Reach for these rules when:
- A job OOM-kills, swaps, or runs much slower than expected on a small box
- Input is larger than RAM and you need to scan, filter, aggregate, sort, or join it
- A pipeline has unbounded buffers between stages, or memory grows linearly during a "streaming" job
- You see one-row-per-RTT writes (
INSERTper row,requests.getper URL,f.read(32)per record) - You're picking a format/codec/serializer and the choice matters at scale
- A
topshows low CPU and high iowait, or you don't know which it is - "It's slow but I don't know why" — start at the obs- category
Rule Categories By Priority
| # | Category | Prefix | Impact | Why it cascades |
|---|---|---|---|---|
| 1 | Memory Discipline | mem- | CRITICAL | Slurping into RAM defeats every downstream technique on a constrained box |
| 2 | I/O Access Patterns | io- | CRITICAL | Disk/net are 10³–10⁶× slower than RAM; access pattern dominates wall-clock |
| 3 | Data Format & Encoding | fmt- | HIGH | Format fixes the lower bound on I/O volume + decode cost before any logic runs |
| 4 | Chunking & Batching | batch- | HIGH | Granularity controls peak memory and amortizes per-item overhead |
| 5 | Spill-to-Disk & External Memory | spill- | HIGH | When data > RAM, the choice is "spill cleanly" or "OOM" |
| 6 | Pipelining & Backpressure | pipe- | MEDIUM-HIGH | Unbounded buffers between fast producers and slow sinks = OOM |
| 7 | Compression & Serialization | codec- | MEDIUM | Trades CPU for I/O; right codec saves orders of magnitude |
| 8 | Concurrency for I/O-Bound Workloads | conc- | MEDIUM | Async / threads / processes are different tools; wrong model wastes CPU |
| 9 | Observability & Throughput Tuning | obs- | LOW-MEDIUM | Can't tune what you don't measure; iowait ≠ CPU-bound |
Quick Reference
1. Memory Discipline (CRITICAL)
mem-stream-dont-slurp— Iterate sources chunk-by-chunk; peak RAM = chunk size, not file sizemem-prefer-generators-over-lists-for-pipelines— Generators flow; lists materializemem-shrink-dtypes-before-loading— Narrow ints, categoricals; 2-8× memory reduction at load timemem-use-views-not-copies— Slicing without copying; NumPy/Arrow zero-copy semanticsmem-bound-the-working-set— Chunk size = budget ÷ row-size × amplification, not a round numbermem-release-references-explicitly— Drop intermediates so peak ≠ N × chunk
2. I/O Access Patterns (CRITICAL)
io-prefer-sequential-over-random— Sort offsets, advise the kernel, let readahead helpio-buffer-explicitly-for-small-records—BufferedReadercollapses 1000× syscallsio-stream-http-bodies-with-iter-content—stream=True+iter_contentinstead of.contentio-mmap-for-random-or-shared-large-files— Zero-copy + on-demand paging for random accessio-async-for-many-concurrent-streams— One thread, thousands of awaitsio-batch-and-pipeline-network-roundtrips—COPY, pipelines, multi-key endpoints, HTTP/2io-zero-copy-when-moving-bytes-as-is—sendfile,copy_file_range,shutil.copyfile
3. Data Format & Encoding (HIGH)
fmt-columnar-for-analytical-scans— Parquet/Arrow for filter+project workloadsfmt-line-delimited-for-streaming-row-ingest— NDJSON over JSON-array for streamingfmt-push-predicates-into-the-reader— Row-group statistics skip whole chunksfmt-prefer-schema-on-write-when-possible— Typed columns beat schema-on-read every timefmt-avoid-deeply-nested-json-for-hot-paths— Flat schema, or binary on hot pipes
4. Chunking & Batching (HIGH)
batch-pick-chunk-size-by-memory-budget— Compute from budget, not a constantbatch-use-vectorized-apis-not-row-loops— NumPy / Polars / Arrow kernels, notiterrowsbatch-process-with-stable-iterators—iter_batches,chunksize=,collect(streaming=True)batch-coalesce-writes-with-buffered-output—COPY/executemany, sized write buffersbatch-keyset-pagination-over-offset—WHERE id > $last_id, neverOFFSET Non deep cursors
5. Spill-to-Disk & External Memory (HIGH)
spill-external-merge-sort-when-data-exceeds-ram— Out-of-core sort; delegate tosort/ DuckDB when possiblespill-partition-by-hash-for-out-of-core-groupby-join— Hash-partition both sides; process per partitionspill-use-temp-files-not-process-memory—SpooledTemporaryFile, never unboundedBytesIOspill-use-engines-that-spill-automatically— DuckDB / Polars / Dask manage spill for you
6. Pipelining & Backpressure (MEDIUM-HIGH)
pipe-use-bounded-queues-for-producer-consumer— BoundedQueueis the backpressure mechanismpipe-apply-backpressure-from-slow-stages— Slow sink throttles the fast sourcepipe-prefer-pull-iteration-over-push-callbacks— Pull backpressures naturally; push needs policypipe-checkpoint-progress-for-resumability— Atomic checkpoint after each batch; redo bounded
7. Compression & Serialization (MEDIUM)
codec-zstd-or-lz4-as-defaults-not-gzip— Pick codec by access pattern; gzip is legacycodec-dictionary-encoding-for-repetitive-strings— 5-50× on low-cardinality columnscodec-prefer-binary-protocols-over-json-for-rpc— Protobuf / Arrow IPC / MsgPack on hot wirescodec-train-a-zstd-dictionary-for-many-small-payloads—zstd --trainfor sub-1 KB messages
8. Concurrency for I/O-Bound Workloads (MEDIUM)
conc-asyncio-for-many-network-streams-not-for-cpu— Asyncio multiplexes waits; useless for computeconc-thread-pools-for-blocking-io-libraries— GIL releases during blocking I/O; threads workconc-overlap-compute-with-prefetch— One batch ahead via background thread / coroutineconc-tune-parallelism-to-the-bottleneck— Match worker count to the binding resource
9. Observability & Throughput Tuning (LOW-MEDIUM)
obs-measure-iowait-not-just-cpu—iostat -x,vmstat, psutil — find the real bottleneckobs-instrument-throughput-rows-per-second— Rates normalize across runs; catch regressionsobs-profile-with-py-spy-or-strace-for-syscall-storms— Attach, don't guess
How to Use
Start with the question that matches the problem:
- "It OOMs / RAM grows linearly during a streaming job" →
mem-andpipe-(likely an unbounded buffer or per-iteration accumulation) - "Disk is 100 % busy but CPU is idle" →
io-(likely random access, unbuffered I/O, or wrong format) - "Why is reading this 10 GB file so slow?" → start with
fmt-columnar-for-analytical-scansandio-buffer-explicitly-for-small-records - "How big should the chunk be?" →
batch-pick-chunk-size-by-memory-budget - "Data doesn't fit in RAM" →
spill-(and prefer an engine that does it for you:spill-use-engines-that-spill-automatically) - "Lots of small network calls" →
io-batch-and-pipeline-network-roundtrips,io-async-for-many-concurrent-streams - "Adding workers didn't help" →
conc-tune-parallelism-to-the-bottleneck,obs-measure-iowait-not-just-cpu - "It's slow but I don't know why" →
obs-first; profile before changing anything
Code examples are in Python (most readable across audiences). The reasoning generalizes — equivalent libraries in other ecosystems (Arrow C++/Rust/Java, Polars Rust, DuckDB everywhere, libuv-style async) follow the same patterns.
Reference Files
| File | Description |
|---|---|
| references/_sections.md | Category definitions and ordering |
| assets/templates/_template.md | Template for adding new rules |
| metadata.json | Version and reference information |
| AGENTS.md | Auto-built TOC navigation |
Related Skills
computer-science-algorithms— Algorithmic primitives this skill builds on (external merge sort, sketches, hash partitioning, sampling)