Repository navigation
API: Add the file type - #18229
Closed
RussellSpitzer wants to merge 1 commit into
Closed
API: Add the file type#18229RussellSpitzer wants to merge 1 commit into
RussellSpitzer wants to merge 1 commit into
Conversation
Add a `file` type for referencing a range of bytes stored inline or in an external file, as defined by the Parquet FILE logical type. The type is its own NestedType with its own TypeID rather than a struct subclass, so visitors cannot silently treat it as a struct where that would be wrong, while the default visitor behavior still falls back to struct handling for the majority of visitors that need no special casing. A file's six nested fields are derived from the ID of the field that holds the type rather than stored in the schema, which keeps the structure immutable. ID assignment reserves the derived block so that a later column cannot be given an ID that the file already owns. The type is gated to format version 4. Generated-by: Cursor
This was referenced Sep 24, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
First PR in a series implementing the Iceberg
filetype (#17919, spec in #17918). This one adds the type itself and nothing else, so the remaining changes can be reviewed in small pieces.Types.FileTypeis its ownNestedTypewith its ownTypeID.FILE, rather than a subclass ofStructType. Visitors that do not care about the distinction fall back to struct behavior:That fallback is what keeps this series small: most visitors, and all the engine integrations, keep compiling and working without a change.
The six nested fields (
uri,offset,size,content_type,checksum,inline) are derived from the enclosing field ID rather than stored, so a file column at ID n owns IDs n+1 through n+6. This PR only teaches ID assignment to reserve that block so a file type survives a round trip. Validating that no other column claims an ID from the block is PR 2.A file column requires format version 4, gated through
Schema.MIN_FORMAT_VERSIONS.Series outline
Each of these is a separate follow-up PR; only this one is open right now.
Engine support (Spark, Flink) and the Parquet
FILElogical-type annotation (which needs parquet-java#3608) are deferred to separate issues.One open question for reviewers
TypeUtil.SchemaVisitor.file()falls back tostruct(), butTypeUtil.CustomOrderSchemaVisitor.file()throwsUnsupportedOperationException. The inconsistency is deliberate for now — the custom-order visitors that need file handling override it explicitly in later PRs — but it may be worth making both consistent.AI Disclosure
filetype into a series of small, independently reviewable PRs, carve each from the integration tree, and verify each with the affected modules' test suites.