Skip to content

FSE: fuse symbol spreading with table building on the fast path - #4836

Open
deadcaf3 wants to merge 1 commit into
facebook:devfrom
deadcaf3:fse-fused-table-build
Open

deadcaf3 wants to merge 1 commit into
facebook:devfrom
deadcaf3:fse-fused-table-build

Conversation

@deadcaf3

Copy link
Copy Markdown

What and why

FSE_buildCTable_wksp() and ZSTD_buildFSETable() build their tables in two passes: spread the symbols into the table, then walk it in position order and read each symbol back. This PR fuses the two passes on the fast path (no low-probability symbols).

FSE_TABLESTEP() is odd for every valid tableLog, so the spread is an invertible permutation: position u holds the symbol at sorted index (u * inv) & (tableSize-1), with inv the inverse of the step modulo tableSize. Walking positions in order and reading symbols from the sorted layout removes the scattered writes and the read-back. FSE_invTableStep() (new, in fse.h) computes inv and asserts it.

Tables are bit-identical, so the format and the compressed output are unchanged. The low-probability path is left as is. No API change.

Performance

Default make flags. Apple M1 (clang 21), Intel Xeon 8481C and AMD EPYC 9B14 (gcc 13.3, Ubuntu 24.04).

Table build, ns per table, 256 random distributions with tableLog 5 to 9, min of 7 rounds, base and new interleaved:

CPU FSE_buildCTable_wksp ZSTD_buildFSETable
Apple M1 449 → 365 (−19%) 509 → 389 (−24%)
Intel 8481C 677 → 512 (−24%) 600 → 524 (−13%)
AMD 9B14 547 → 460 (−16%) 628 → 511 (−19%)

Decompression speed, zstd -b1 -B<chunk> -i3 on 3.4 MB of C source, two interleaved runs:

chunk Intel AMD M1
4 KB +0.9% +2.1% +1.6%
16 KB +0.7% +2.1% +1.9%
32 KB +0.1% +1.1% +0.8%
64 KB −0.2% +0.3% 0%

Compression speed within ±1%, compressed sizes identical. Not measured: other data sets, levels above 1, MSVC, 32-bit.

Verification

  • New unit test in tests/fuzzer.c compares both builders against a plain spread-then-read-back reference for every tableLog, with and without low-probability symbols. It fails on three mutants, including a wrong but shared inverse that keeps encoder and decoder consistent.
  • fuzzer, zstreamtest and decodecorpus -t pass with asserts enabled. make staticAnalyze reports the same findings on dev and on this branch.
  • Applying the same fusion to FSE_buildDTable_internal() regressed 5 to 23% on Sapphire Rapids with gcc, so it is left out.

Checklist

  • I searched open pull requests and issues: this problem is not already being addressed.
  • Bug fix: it comes with a reproducer or a test that fails without the fix. (not a bug fix)
  • Performance change: numbers from repeated runs are included, with CPU, compiler, flags and data.
  • I stated below whether AI or other automated tools were used.

FSE_buildCTable_wksp() and ZSTD_buildFSETable_body() spread symbols
into the table in one pass, then walk the table in position order to
build the state transitions, reading each symbol back.

The spread step, tableSize/2 + tableSize/8 + 3, is odd for every valid
tableLog, so the spread is an invertible permutation: position u holds
the symbol at sorted index (u * inv) & (tableSize-1), where inv is the
inverse of step modulo tableSize. Walking positions in order and
reading symbols straight from the sorted layout fuses the two passes,
removing the scattered writes into the table and the read-back.

Only the fast path (no low-probability symbols) changes. The
low-probability spread skips positions and is left as is. Tables are
bit-identical, so compressed output is unchanged.

Fast-path table build time, ns/table, mixed tableLog 5-9:

                        FSE_buildCTable_wksp   ZSTD_buildFSETable
  Apple M1, clang 21        449 -> 365            509 -> 389
  Xeon 8481C, gcc 13        677 -> 512            600 -> 524
  EPYC 9B14, gcc 13         547 -> 460            628 -> 511

End-to-end decompression of independent 4 KB to 32 KB chunks improves
by 1-2% on all three machines; 64 KB and larger blocks are unchanged.
@meta-cla meta-cla Bot added the CLA Signed label Oct 10, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant