Skip to content

fix(serde): reject invalid schema and partition IDs in JSON - #940

Open
kamcheungting-db wants to merge 11 commits into
apache:mainfrom
kamcheungting-db:strict-schema-spec-json-ids
Open

kamcheungting-db wants to merge 11 commits into
apache:mainfrom
kamcheungting-db:strict-schema-spec-json-ids

Conversation

@kamcheungting-db

@kamcheungting-db kamcheungting-db commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Require schema and partition IDs to be actual, in-range JSON integers.

Generic JSON conversion could previously coerce invalid values. For an ID field, this could turn:

1.5

into:

1

This PR accepts values such as:

"schema-id": 10

and rejects:

"schema-id": 1.5
"schema-id": true
"schema-id": 2147483648

The validation covers schema IDs, field IDs, identifier field IDs, nested list/map IDs, partition source/field IDs, partition spec IDs, and the table-metadata current-schema-id / default-spec-id references.

Missing partition field IDs remain supported for legacy v1 metadata; an explicitly present invalid value, including null, is rejected.

Stack

This is 4 of 4 and depends on #939:

  1. Move DataFile serde into core (refactor(serde): move DataFile JSON handling from REST to core #974)
  2. Java-compatible DataFile JSON (fix(serde): make DataFile JSON compatible with Java Iceberg #975)
  3. Core FileScanTask JSON serde (feat(serde): support standalone FileScanTask JSON in core #939)
  4. This PR: strict schema and partition ID parsing

Each PR targets main; this PR contains all preceding commits and should be reviewed last.

Testing

  • Added schema and partition ID validation tests covering fractional, boolean, and out-of-range values, plus legacy v1 null and fractional table-metadata references.
  • Full stack: 145 JSON serde tests passed (1 disabled).
  • Full stack: 332 REST tests passed; 1 existing test skipped.

@kamcheungting-db
kamcheungting-db force-pushed the strict-schema-spec-json-ids branch from 0c99aa2 to 54f4541 Compare September 30, 2026 08:09
@kamcheungting-db kamcheungting-db changed the title fix: validate schema and partition IDs as JSON integers fix(serde): validate schema and partition IDs as JSON integers Sep 30, 2026
@kamcheungting-db

Copy link
Copy Markdown
Contributor Author

This PR is now the final layer in the serde stack:

  1. refactor(serde): move DataFile JSON handling from REST to core #974 — move DataFile serde into core
  2. fix(serde): make DataFile JSON compatible with Java Iceberg #975 — Java-compatible DataFile JSON
  3. feat(serde): support standalone FileScanTask JSON in core #939 — core FileScanTask serde
  4. fix(serde): reject invalid schema and partition IDs in JSON #940 — strict schema and partition ID parsing

For the incremental diff that belongs only to this PR, use:
kamcheungting-db/iceberg-cpp@filescan-task-serde...strict-schema-spec-json-ids

In short, an ID like 10 remains valid, while 1.5, true, and values outside int32_t are rejected. Missing v1 partition field IDs remain supported.

This PR contains the earlier stack commits because all Apache PRs target main; it should be reviewed last.

@kamcheungting-db
kamcheungting-db force-pushed the strict-schema-spec-json-ids branch from 54f4541 to a0ad0f9 Compare September 30, 2026 10:26
@kamcheungting-db
kamcheungting-db force-pushed the strict-schema-spec-json-ids branch from a0ad0f9 to 62d582c Compare September 30, 2026 10:55
@kamcheungting-db kamcheungting-db changed the title fix(serde): validate schema and partition IDs as JSON integers fix(serde): reject invalid schema and partition IDs in JSON Sep 30, 2026
@kamcheungting-db
kamcheungting-db force-pushed the strict-schema-spec-json-ids branch from 62d582c to 28851ed Compare September 30, 2026 20:33

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant