PivCo-Huffman “merge” operations
The post surveys existing approaches to parallel Huffman decoding—multi-stream, interleaved (e.g. GDeflate), and speculative brute force—and explains why each carries costs like gather-heavy access, magic interleave constants baked into wire formats, or discarding 80%+ of work. PivCo-Huffman sidesteps these by recasting decoding as list merges reducible to prefix sums, scaling to any vector width. The bulk derives merge kernels for AVX-512 VBMI2, SSE4.2/AVX2, and NEON, trading lookup-table size against instruction count; results are mainly on high-end hardware.