Problem
Parquet column paths are component-based, but ParquetFileWriter stores Bloom filters in a map keyed by String.join(".", descriptor.getPath()).
These distinct paths therefore collide:
Top-level field named `a.b`: ["a.b"]
Nested field `b` in `a`: ["a", "b"]
When both columns have Bloom filters, the nested filter overwrites the top-level filter under the shared "a.b" key. During footer serialization, both columns can receive the nested column filter.
Reproduction
A test writes disjoint values (top-* and nested-*) to the two string columns and checks each footer Bloom filter for its own values.
Current result:
Top-level field path ["a.b"] contains its own Bloom values: false
Nested field path ["a", "b"] contains its own Bloom values: true
Expected: both results are true.
Affected code
ParquetFileWriter.writeColumnChunk
ParquetFileWriter.serializeBloomFilters
Similar dot-string path handling also exists in dictionary lookup and row-group copying, but this issue includes a minimal Bloom-filter reproduction.
Problem
Parquet column paths are component-based, but
ParquetFileWriterstores Bloom filters in a map keyed byString.join(".", descriptor.getPath()).These distinct paths therefore collide:
When both columns have Bloom filters, the nested filter overwrites the top-level filter under the shared
"a.b"key. During footer serialization, both columns can receive the nested column filter.Reproduction
A test writes disjoint values (
top-*andnested-*) to the two string columns and checks each footer Bloom filter for its own values.Current result:
Expected: both results are
true.Affected code
ParquetFileWriter.writeColumnChunkParquetFileWriter.serializeBloomFiltersSimilar dot-string path handling also exists in dictionary lookup and row-group copying, but this issue includes a minimal Bloom-filter reproduction.