Skip to content

Hash aggregate with a few large groups fails after spilling: spill batches are sized by rows, not bytes #25851

Description

@andygrove

Describe the bug

A grouped aggregate whose groups carry large state, such as array_agg over a low-cardinality key, fails with ResourcesExhausted once it spills, even though the pool is free when it fails. With the same data spread over many small groups, the same query at the same memory limit spills and succeeds.

The spill writes the aggregate's state in batches of at most batch_size rows, whatever their size in bytes (stream.rs#L355-L359, called from spill.rs#L222-L230). With fewer groups than batch_size, the whole table goes into one batch, and everything after the spill has to hold that batch at once:

To Reproduce

On main at c1786f7, save this as repro.sql:

SET datafusion.execution.target_partitions = 2;
SELECT count(*) AS groups, sum(n) AS values_collected FROM (
  SELECT k, cardinality(array_agg(v)) AS n
  FROM (
    SELECT value % 16 AS k, concat(CAST(value AS VARCHAR), repeat('x', 100)) AS v
    FROM generate_series(1, 3000000)
  )
  GROUP BY k
);

datafusion-cli -m 128M -f repro.sql fails on every run:

Resources exhausted: Additional allocation failed for FinalHashAggregateStream[0] with top memory consumers (across reservations) as:
  AggregateStream[1]#9(can spill: false) consumed 48.0 B, peak 48.0 B,
  AggregateStream[0]#5(can spill: false) consumed 48.0 B, peak 48.0 B,
  AggregateStream[0]#2(can spill: false) consumed 48.0 B, peak 48.0 B.
Error: Failed to allocate additional 214.5 MB for FinalHashAggregateStream[0] with 0.0 B already allocated for this reservation - 128.0 MB remain available for the total memory pool: greedy(used: 144.0 B, pool_size: 128.0 MB)

Controls with the same binary and the same data volume:

  • Without -m, the query returns 16 groups and 3,000,000 values.
  • With value % 500000 (500,000 groups) and -m 128M, it returns 500,000 groups and 3,000,000 values. EXPLAIN ANALYZE shows the FinalPartitioned aggregate spilled 32 times (569.0 MB).

To see where the memory goes, I added temporary eprintln! probes to sort_and_spill, the merge's admission, the final aggregate's out-of-memory branch and the replay loop. They are not needed for the repro. For each final partition, the output is (abbreviated):

final-oom: reserved=0 wanted=224915992 groups=8 input_rows=80
agg-spill: max_record_batch_memory=183944547
merge: spill_files=1 in_memory_streams=0
replay: input_rows=8 table_groups=8 resize_to=224914200 spillable=false

Here the partial aggregates hand each final partition an 80-row batch that already holds all of its state. So the final aggregate's first reservation fails, and it spills the whole table as one 8-row, 184 MB batch. Reading that back needs 225 MB in a 128 MB pool, although each group is only about 25 MB.

Expected behavior

The query succeeds under the memory limit, as it does when the same data is spread over many groups. A spilled run of a few large groups should be written in batches small enough to read back within the budget, for example by capping each spill batch by bytes as well as by rows. That can't help a single group that is larger than the budget, but here each group is about a fifth of the pool.

Additional context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions