perf(parquet): reuse page-index bytes through the reader cache - #285
Open
wangyong9999 wants to merge 3 commits into
Open
perf(parquet): reuse page-index bytes through the reader cache#285wangyong9999 wants to merge 3 commits into
wangyong9999 wants to merge 3 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Repeated Parquet point reads reuse serialized footers through
ReaderBuilder::WithCache, but reload column and offset indexes when each reader is rebuilt. On remote storage this adds index range requests to every lookup.Reuse immutable page-index bytes through the caller-provided cache. The footer determines eligible ranges; URI, offset and length identify each entry. Page-index entries share the existing
DATA_FILE_FOOTERcache budget. Data-page reads and snapshot discovery are unchanged.Readers, streams and decryptors remain query-local. No process-wide cache, reader pool or new capacity setting is introduced. Missing cache/URI uses the existing uncached path; read and cache errors propagate. Cached footer and page-index allocations retain their memory pool until eviction, since the cache may outlive a reader's custom pool.
No separate issue or design document.
Tests
Added to
paimon-parquet-format-test:TestPageIndexBytesSurviveReaderClose: multiple row groups, second reader avoids index storage reads, data reads bypass cache, invalidation reloads indexes, and a temporary pool survives until eviction.TestPointReadReusesFooterAndPageIndexes: two independent point readers return the same row without reloading metadata.TestCachedFooterKeepsAllocatorAliveUntilEviction: footer allocation lifetime with a temporary pool.Local C++17 syntax validation passed before the allocator-lifetime follow-up. Full compilation and test execution are in progress; CI results will be followed up. No latency improvement is claimed before end-to-end measurement.
API and Format
No public API or storage-format change. Reuses the existing Cache interface and metadata budget.
Documentation
Inline comments explain range eligibility and lifetime. No new user option.
Generative AI tooling
Generated-by: OpenAI Codex