Skip to content

Validate cached file statistics and Parquet metadata with e_tag, not just size and mtime #25841

Description

@andygrove

Is your feature request related to a problem or challenge?

The file statistics cache and the file metadata (Parquet footer) cache treat a cached entry as valid if the file's size and last modified time still match (plus the schema fingerprint, for statistics):

ObjectMeta also has e_tag and version, but neither is compared. S3 reports Last-Modified with one-second precision, so a file overwritten in place at the same size within the same second passes the check, and the entry for the old contents is used. Stale statistics give wrong answers for COUNT(*), MIN and MAX, which can be answered from statistics alone. A stale footer would make the reader decode the new contents with the old file's row group and page offsets.

The keys don't include the object store either. Statistics are keyed by TableScopedPath, a table reference plus a store-relative path, and footers by the store-relative Path alone. So s3://bucket-a/data/part-0.parquet and s3://bucket-b/data/part-0.parquet share a footer cache entry. They share a statistics entry too when they're read through the same table name, or anonymously through read_parquet, which uses the shared cache with no table reference. Only size and modification time tell them apart.

Both cases need a size and a modification time to match exactly, so they're rare. But the exposure grows with the lifetime of the RuntimeEnv. A stronger check only helps where the file listing itself is fresh, for example with list_files_cache_ttl set or the list files cache turned off.

This came up in Ballista, where apache/datafusion-ballista#2498 shares one file statistics cache across all the sessions a scheduler creates, while each job still lists its own files. @comphead pointed out the gap in review.

Describe the solution you'd like

  1. In both is_valid_for methods, also compare e_tag and version when the cached and the current ObjectMeta both have them, and fall back to size and modification time otherwise. S3 listings include the ETag, and LocalFileSystem derives one from the inode, modification time and size. Two different objects with the same content can have the same ETag, but then their statistics and footers are the same too.
  2. Include the object store URL in the cache keys, or store it in each entry and check it. TableScopedPath is public, so this is an API change. With (1) in place it mainly matters for stores that don't report ETags.

Describe alternatives you've considered

Documenting the limitation instead. That's cheaper, but it's hard for users to make sure a rewrite always changes the size or the second-granularity modification time.

Additional context

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions