Is your feature request related to a problem or challenge?
DataFusion represents row-count aggregates such as COUNT(*) as COUNT(1). During execution, the scalar 1 is expanded into a full Int64Array for every input batch, although the accumulator only needs the number of rows. Similarly, COUNT(non_nullable_column) unnecessarily keeps the column in the scan.
This issue comes from profiling in #25536 (comment).
Describe the solution you'd like
Simplify non-DISTINCT COUNT calls whose arguments are safe to elide and provably non-null (for example, non-null literals and direct non-nullable columns) to a nullary COUNT(), while preserving the original output name.
Teach aggregate execution to pass the input row count explicitly so nullary COUNT works without materializing an argument array in grouped and ungrouped aggregation.
Nullable arguments, DISTINCT, and arbitrary expressions that may error or be volatile should not be rewritten.
Is your feature request related to a problem or challenge?
DataFusion represents row-count aggregates such as
COUNT(*)asCOUNT(1). During execution, the scalar1is expanded into a fullInt64Arrayfor every input batch, although the accumulator only needs the number of rows. Similarly,COUNT(non_nullable_column)unnecessarily keeps the column in the scan.This issue comes from profiling in #25536 (comment).
Describe the solution you'd like
Simplify non-
DISTINCTCOUNTcalls whose arguments are safe to elide and provably non-null (for example, non-null literals and direct non-nullable columns) to a nullaryCOUNT(), while preserving the original output name.Teach aggregate execution to pass the input row count explicitly so nullary
COUNTworks without materializing an argument array in grouped and ungrouped aggregation.Nullable arguments,
DISTINCT, and arbitrary expressions that may error or be volatile should not be rewritten.