feat: add Parquet content-defined chunking writer support - #3889
Open
kszucs wants to merge 1 commit into
Open
Conversation
Contributor
There was a problem hiding this comment.
Warning
Copilot couldn't run its full agentic review because it didn't start before the timeout. Make sure your repository has a runner available, or add a copilot-code-review.yml file specifying one with the runs-on attribute. See the docs for more details.
Pull request overview
Adds support for PyArrow Parquet content-defined chunking (CDC) in the PyIceberg write path, exposing it via new table properties and guarding usage by the minimum supported PyArrow version.
Changes:
- Add new table properties for enabling/configuring Parquet CDC (enabled/min/max/norm-level).
- Wire CDC properties through
_get_parquet_writer_kwargs, including a shared_require_pyarrow_versionguard. - Add unit + integration-style tests and document the new properties.
Reviewed changes
Copilot reviewed 4 out of 4 changed files in this pull request and generated 4 comments.
| File | Description |
|---|---|
| tests/io/test_pyarrow.py | Adds tests for CDC kwargs generation, version gating, and end-to-end wiring into pq.ParquetWriter. |
| pyiceberg/table/init.py | Introduces new CDC-related TableProperties constants and defaults. |
| pyiceberg/io/pyarrow.py | Adds _require_pyarrow_version helper, reuses it for Azure FS guard, and forwards CDC options to PyArrow writer kwargs. |
| mkdocs/docs/configuration.md | Documents new CDC table properties and their defaults/requirements. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
PyArrow's ParquetWriter has supported content-defined chunking natively since 21.0.0, producing stable page boundaries across appends for content-addressable storage. Wire this through as write.parquet.content-defined-chunking.* table properties, mirroring the property names and defaults already used by iceberg-rust. PyArrow validates the chunk sizes itself and raises a clear error, so pyiceberg doesn't duplicate those checks. Requesting CDC on an older PyArrow raises an ImportError from a shared _require_pyarrow_version helper, which also replaces the existing Azure filesystem guard.
kszucs
force-pushed
the
feat/parquet-cdc-writer-support
branch
from
September 1, 2026 19:53
2a93c6d to
d7f61aa
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reopens #3608, which was closed by the stale bot. Rebased on current
main; GitHub refused to reopen the original PR after the branch was updated.Rationale for this change
PyArrow's
ParquetWriterhas natively supported content-defined chunking (CDC) since 21.0.0, producing stable page boundaries across appends (useful for content-addressable storage / dedup). PyIceberg's write path already funnels through a single kwargs builder (_get_parquet_writer_kwargs), so this wires CDC through aswrite.parquet.content-defined-chunking.*table properties, mirroring the property names and defaults iceberg-rust already uses for cross-engine consistency. Apyarrow>=21.0.0version guard raises a clearImportErrorif CDC is requested on an older PyArrow (extracted into a shared_require_pyarrow_versionhelper, reused by the existing Azure-filesystem version guard).Are these changes tested?
Yes: unit tests for
_get_parquet_writer_kwargs(disabled by default, enabled with defaults, enabled with custom values, unsupported PyArrow version) and an integration-style test that writes a table with CDC enabled end-to-end and reads it back.Are there any user-facing changes?
Yes: four new table properties (
write.parquet.content-defined-chunking.enabled,.min-chunk-size,.max-chunk-size,.norm-level), documented inconfiguration.md.