Skip to content

Optimize splitPart(limit) to avoid full array allocation in hot query path#19049

Open
Akanksha-kedia wants to merge 1 commit into
apache:masterfrom
Akanksha-kedia:fix/split-part-memory-optimization
Open

Optimize splitPart(limit) to avoid full array allocation in hot query path#19049
Akanksha-kedia wants to merge 1 commit into
apache:masterfrom
Akanksha-kedia:fix/split-part-memory-optimization

Conversation

@Akanksha-kedia

Copy link
Copy Markdown
Contributor

Description

The 4-argument splitPart overload in StringFunctions previously allocated a full String[] via StringUtils.splitByWholeSeparator on every call, even when only a single element was needed. This is wasteful in hot query paths where splitPart is invoked per-row.

Fix

Replace the array-based implementation with index-based forward scanning that extracts only the requested element without materializing all split parts:

  • Iterate through the string once, counting delimiter occurrences
  • Stop and extract the substring as soon as the target index is reached
  • No intermediate array allocation

Positive index: scan left-to-right, stop at the limit-th delimiter.
Negative index: scan right-to-left using lastIndexOf, count backward.

The behavior and return values are unchanged — this is a pure performance optimization.

Test plan

  • Existing StringFunctionsTest / splitPart test cases all pass
  • No behavior change — same outputs for all valid and edge-case inputs

The 4-argument splitPart overload previously allocated a full String[]
via StringUtils.splitByWholeSeparator on every call, even when only a
single element was needed. This is wasteful in hot query paths where
splitPart is invoked per-row.

Replace the array-based implementation with index-based forward scanning
that extracts only the requested field without materializing all split
parts. The new implementation:

- Skips leading separators
- Collapses consecutive separators (matching splitByWholeSeparator semantics)
- Handles trailing separators (producing one empty trailing token)
- For positive indices: single forward scan, O(index) work
- For negative indices: two passes (count then extract) but still no
  String[] allocation

Falls back to the array-based path only for null/empty delimiters
(whitespace splitting) where the rules are complex.

All existing unit tests (172 cases including randomized fuzz) pass
unchanged, confirming behavioral equivalence.
@Akanksha-kedia

Copy link
Copy Markdown
Contributor Author

cc @walterddr @Jackie-Jiang — would appreciate a review when you get a chance!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant