Skip to content

Bloom filters collide for distinct column paths with the same dot string #3826

Description

@costas-db

Problem

Parquet column paths are component-based, but ParquetFileWriter stores Bloom filters in a map keyed by String.join(".", descriptor.getPath()).

These distinct paths therefore collide:

Top-level field named `a.b`: ["a.b"]
Nested field `b` in `a`:     ["a", "b"]

When both columns have Bloom filters, the nested filter overwrites the top-level filter under the shared "a.b" key. During footer serialization, both columns can receive the nested column filter.

Reproduction

A test writes disjoint values (top-* and nested-*) to the two string columns and checks each footer Bloom filter for its own values.

Current result:

Top-level field path ["a.b"] contains its own Bloom values: false
Nested field path ["a", "b"] contains its own Bloom values: true

Expected: both results are true.

Affected code

  • ParquetFileWriter.writeColumnChunk
  • ParquetFileWriter.serializeBloomFilters

Similar dot-string path handling also exists in dictionary lookup and row-group copying, but this issue includes a minimal Bloom-filter reproduction.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions