Skip to content

API: Add the file type - #18229

Closed
RussellSpitzer wants to merge 1 commit into
apache:mainfrom
RussellSpitzer:file-type-1-api
Closed

RussellSpitzer wants to merge 1 commit into
apache:mainfrom
RussellSpitzer:file-type-1-api

Conversation

@RussellSpitzer

Copy link
Copy Markdown
Member

First PR in a series implementing the Iceberg file type (#17919, spec in #17918). This one adds the type itself and nothing else, so the remaining changes can be reviewed in small pieces.

Types.FileType is its own NestedType with its own TypeID.FILE, rather than a subclass of StructType. Visitors that do not care about the distinction fall back to struct behavior:

public T file(Types.FileType file, List<T> fieldResults) {
  return struct(file.asStruct(), fieldResults);
}

That fallback is what keeps this series small: most visitors, and all the engine integrations, keep compiling and working without a change.

The six nested fields (uri, offset, size, content_type, checksum, inline) are derived from the enclosing field ID rather than stored, so a file column at ID n owns IDs n+1 through n+6. This PR only teaches ID assignment to reserve that block so a file type survives a round trip. Validating that no other column claims an ID from the block is PR 2.

A file column requires format version 4, gated through Schema.MIN_FORMAT_VERSIONS.

Series outline

Each of these is a separate follow-up PR; only this one is open right now.

  1. API: Add the file type (this PR)
  2. API, Core: Reserve and validate the derived field ID block
  3. Core: Serialize the file type in schema JSON
  4. API, Core, Data: Handle the file type in type switches and struct views
  5. Core: Read and write file columns in Avro
  6. Parquet: Map the file type to a group
  7. ORC: Support the file type in schema conversion and visitors
  8. Hive, Kafka Connect, AWS, Arrow: Render a file column as a struct

Engine support (Spark, Flink) and the Parquet FILE logical-type annotation (which needs parquet-java#3608) are deferred to separate issues.

One open question for reviewers

TypeUtil.SchemaVisitor.file() falls back to struct(), but TypeUtil.CustomOrderSchemaVisitor.file() throws UnsupportedOperationException. The inconsistency is deliberate for now — the custom-order visitors that need file handling override it explicitly in later PRs — but it may be worth making both consistent.


AI Disclosure

  • Model: Claude Opus 4.5
  • Platform/Tool: Cursor
  • Human Oversight: partially reviewed
  • Prompt Summary: Split an existing single-commit implementation of the Iceberg file type into a series of small, independently reviewable PRs, carve each from the integration tree, and verify each with the affected modules' test suites.

Add a `file` type for referencing a range of bytes stored inline or in an
external file, as defined by the Parquet FILE logical type. The type is its own
NestedType with its own TypeID rather than a struct subclass, so visitors cannot
silently treat it as a struct where that would be wrong, while the default
visitor behavior still falls back to struct handling for the majority of
visitors that need no special casing.

A file's six nested fields are derived from the ID of the field that holds the
type rather than stored in the schema, which keeps the structure immutable. ID
assignment reserves the derived block so that a later column cannot be given an
ID that the file already owns.

The type is gated to format version 4.

Generated-by: Cursor
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant