compiler-optimizations-deep

v2026.09.24

Deep compiler optimizations skill for RA, ISel, and PGO. Use when explaining register allocation, instruction selection, LICM, vectorization limits, or profile-guided optimization beyond -O3. Activates on queries about register allocation, instruction selection, LICM, auto-vectorization failure, PGO, or BOLT.

GitHub
Install command
npx skhub add mohitmishra786/compiler-optimizations-deep
Markdown
SKILL.md

Compiler Optimizations (Deep)

Purpose

Explain optimization phases beyond flags: mid-level IR opts, register allocation, instruction selection/scheduling, vectorization boundaries, PGO, and post-link BOLT — bridging skills/compilers/pgo and LLVM/GCC internals.

When to Use

  • -O3 did not vectorize a hot loop
  • Teaching why register pressure causes spills
  • Planning PGO or BOLT deployment
  • Understanding pass interaction (e.g., LICM before vectorize)

Workflow

1. Compiler pipeline map

Frontend → LLVM IR / GCC GIMPLE
├── Mid-level: DCE, GVN, LICM, inlining
├── Loop opts: unroll, vectorize
├── Codegen prep: legalize types
├── Instruction selection (DAG → machine ops)
├── Register allocation (greedy, linear scan)
└── Peephole / scheduling

2. Vectorization failure triage

clang -O3 -Rpass=loop-vectorize -Rpass-missed=loop-vectorize foo.c
Miss reasonTypical fix
Unknown trip countpeel loop; assert count
Dependencereorder / separate accumulators
Function call in loopinline or outline
Alignment unknown__builtin_assume_aligned

3. Register allocation intuition

When live ranges exceed physical registers, the allocator spills to stack slots — costly loads/stores. Reducing live ranges (splitting variables, rematerialization) helps.

GCC/LLVM both use graph coloring variants (LLVM "greedy regalloc").

4. PGO workflow (Clang)

clang -fprofile-instr-generate -O2 -o app foo.c
./app   # training workload
llvm-profdata merge default.profraw -o default.profdata
clang -fprofile-instr-use=default.profdata -O2 -o app_pgo foo.c

Improves branch layout, inlining, and vectorization thresholds.

See skills/compilers/pgo for GCC and BOLT.

5. BOLT (post-link)

llvm-bolt -instrument app -o app.inst
./app.inst
llvm-bolt -data=perf.fdata -reorder-blocks=+ -o app.bolt app

Optimizes layout after linker — needs relocations (-Wl,--emit-relocs).

6. LICM example

Loop-invariant code motion hoists x * scale out of inner loop when legal — reduces work per iteration.

7. Agent usage

/compiler-optimizations-deep Why did LLVM fail to vectorize this reduction loop?

Common Problems

SymptomCauseFix
PGO no gainUnrepresentative trainingMatch production input
BOLT crashStripped binaryKeep symbols + relocs
Spills in asmRegister pressureSimplify live ranges
-O3 slowerCode bloat / cacheTry -O2 or PGO
Different GCC/ClangPass ordering differsCompare IR + asm

Related Skills

  • skills/compilers/pgo — PGO and BOLT detail
  • skills/compiler-internals/llvm-ir-and-passes — IR-level opts
  • skills/compiler-internals/code-generation-and-backends — ISel and backends
  • skills/computer-architecture/cpu-pipelines-and-hazards — scheduling context
  • skills/low-level-programming/simd-intrinsics — manual vectorization
Discovery
Tags

No tags published for this skill.

Version
Latest version metadata

Version

v2026.09.24

Published

Sep 24, 2026

Category

Uncategorized

License

MIT

Source path

skills/compiler-internals/compiler-optimizations-deep

Default branch

main

Latest commit

bdc5847

Tree SHA

1178323