Skip to content

Per-column Hadoop configuration cannot distinguish dotted and nested column paths #3832

Description

@costas-db

Problem

Per-column Hadoop configuration currently identifies a column with a flattened dot string such as a.b. This cannot distinguish two valid, different Parquet paths:

Top-level field named `a.b`: ["a.b"]
Nested field `b` in `a`:     ["a", "b"]

ColumnConfigParser passes the suffix to ParquetProperties as a string, and ColumnProperty interprets it with ColumnPath.fromDotString. Consequently, a setting intended for the top-level dotted field cannot be represented and may instead apply to the nested field. This is datatype-independent and affects per-column settings such as Bloom filters and statistics.

Reproduction

Use two ordinary string columns with the paths above, disable Bloom filters and statistics globally, and enable them only for the top-level ["a.b"] path. On current master, the top-level column still has no Bloom filter or value statistics because the structured override is not recognized.

Proposed direction

Add a backward-compatible structured-path key form that encodes each path component independently, plus path-aware parser and builder overloads. Keep existing dot-string keys working for compatibility, while allowing callers to target dotted field names unambiguously.

This changes only configuration lookup identity; it does not alter the Parquet file format.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions