Conversation
Parse normalized GMT offsets, including seconds, without requiring TZDB files. Add boundary checks and regression coverage using Java-written ORC files. Fixes apache#2715
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes were proposed in this pull request?
Add support for Java fixed-offset timezone IDs (GMT±HH:MM and GMT±HH:MM:SS) in the C++ timezone resolver, using cached fixed-offset timezones without requiring TZDB files.
Why are the changes needed?
The Java writer can record IDs such as GMT-00:00 or GMT+05:30 in the stripe’s writerTimezone. The C++ reader treats these IDs as timezone database filenames and fails when reading
timestamp columns because the corresponding files do not exist.
This change allows the C++ reader to resolve these offsets and read the timestamps correctly.
Fixes ORC-2222 and #2715.
How was this patch tested?
Added tests covering:
• Positive and negative offsets, signed zero, minute and second precision, and invalid formats.
• Epoch calculation, time conversion, caching, and operation without TZDB files.
• Reading Java-written ORC files with four fixed-offset timezones, including pre-1970 timestamps and nanosecond precision.
Rebuilt C++ and ran make -C build test-out: all 763 core tests and 87 tool tests passed. Also verified that the original file attached to #2715 reads successfully.
Was this patch authored or co-authored using generative AI tooling?
Generated-by: OpenAI Codex (GPT-6)