Skip to content

[core] Complete nested updates and partition-aware temporal upserts - #973

Merged
JingsongLi merged 2 commits into
apache:mainfrom
JingsongLi:codex/native-nested-updates
Sep 27, 2026
Merged

JingsongLi merged 2 commits into
apache:mainfrom
JingsongLi:codex/native-nested-updates

Conversation

@JingsongLi

@JingsongLi JingsongLi commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Complete the Rust core capabilities needed for PyPaimon Native nested updates and temporal-key upserts, and restrict upsert data reads to the input partitions.

Dependent PyPaimon integration and end-to-end tests: apache/paimon#10242. Its Native CI continues to build apache/paimon-rust@main, so this core PR should merge first.

The existing implementation rejects equivalent Arrow layouts such as list<item: int> versus Paimon's list<element: int>, lacks several row-ID constructor conversions, rejects temporal upsert keys, and reads unrelated partitions during key matching.

Brief change log

  • Normalize complete upsert batches with the ordinary write normalizer before key matching or staging files. Normalize MAP children by key/value position, including custom Arrow child names.
  • Implement row-ID constructor conversions for top-level lists of pairs to MAP, ROW to MAP, and MAP to ROW in Rust core. Preserve slices, empty collections, NULL parents and nested NULL values.
  • Retain the distinction between safe predicate assignments and whole-column row-ID constructor fallback. Validate ROW field order and nullability, reject ambiguous top-level dict NULL values, and support dictionary logical NULLs and list/view layouts.
  • Accept TIME, TIMESTAMP and TIMESTAMP_LTZ upsert keys through the existing Arrow key encoder, preserving storage-unit precision, NULL equality, last-source-wins and duplicate-target matching.
  • Filter upsert data splits using exact source partition tuples after planning, preserving global row IDs and snapshot selection. Metadata planning is unchanged.
  • Normalize manifest partition values before comparison: Java can reserve storage for NULL high-precision timestamps and decimals, so raw BinaryRow bytes are not a reliable logical identity.
  • Add public core tests for nested and temporal upserts, duplicate keys, appends, untouched fields, invalid-input cleanup and partition read scope. Removing unrelated data files proves that upsert matching does not read them. A Java-layout regression covers NULL variable storage and the shifted offset of a following non-NULL field.

All conversion and upsert execution remains in Rust core. Invalid UTF-8 MAP field names are rejected rather than reproducing PyArrow's silent NULL result for malformed names.

Tests

  • Core unit tests: table::update_input::tests (28), table::write_batch_normalize::tests (4), table::upsert_key_matcher::tests (3), table::table_upsert::tests (1).
  • Integration tests: table_update_test, table_update_nested_test, table_update_paths_test (32).
  • Total: 68 passed.
  • cargo clippy --locked --all-targets --workspace --features fulltext,vortex -- -D warnings and cargo fmt --all --check passed.
  • Built the Python extension and ran the dependent PyPaimon regression suite with PyArrow 18.1 and all five Native CI flags: 973 passed, 2 skipped, 16 subtests passed. The two skips are Python-only MAP view-key cases without an Arrow take kernel; their Native counterparts pass.

API and Format

No public API or persisted format changes. Existing update and upsert APIs accept additional compatible inputs.

Documentation

No new configuration or public methods.

@JingsongLi JingsongLi changed the title [core] Support nested values in table updates and upserts [core] Complete nested updates and partition-aware temporal upserts Sep 27, 2026

@leaves12138 leaves12138 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed fab8376 together with the PyPaimon integration in apache/paimon#10242 at 32fc7188c744e7ce2efb7f07051971d22477277b. No blocking regression found in the reviewed changes.

Checked the distinction between safe predicate casts and whole-column row-ID constructor fallback, nested NULL/slice handling, temporal key precision and duplicate matching, and exact partition filtering without changing global row IDs or the pinned snapshot. Also checked the Java-compatible partition normalization and sequence-ordering integration.

Local validation:

  • 36 focused core unit tests and 32 public integration tests passed (68 total).
  • Built the Python extension from this exact revision. With PyArrow 18.1 and all five Native CI flags, the paired Python suites produced 973 passed, 2 expected Python-only skips, and 16 passed subtests.
  • 12 additional constructor/partition boundary cases passed. Rustfmt and changed-file Python lint passed.

All 14 current Rust CI checks are successful. Merge this core change before the dependent Python PR and rerun its Native CI against main.

@QuakeWang QuakeWang left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@JingsongLi
JingsongLi merged commit 952f6b8 into apache:main Sep 27, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants