Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
252 commits
Select commit Hold shift + click to select a range
7e93f6b
[Flink] Report commit metrics via the Delta-Tables API (#7246)
timothyw553 Aug 5, 2026
e82d3bc
[Flink] Document why catalog browsing stays on the UC OpenAPI client …
timothyw553 Aug 5, 2026
3d52075
[PROTOCOL] Interval Types RFC Update (#7308)
rajayush143 Aug 5, 2026
937dfc7
[Spark] Promote a single AMT manifest to the root instead of writing …
ctring Aug 5, 2026
8aa15c8
Populate LastManifestCommit in CommitInfo and VersionChecksum upon co…
yumingxuanguo-db Aug 6, 2026
ed5cec6
[Kernel] Promote Snapshot/Scan/CommitRange accessors used by Delta Sp…
chiinlquah Aug 6, 2026
2c2be67
Enabling partitioned column mapping on batch write path in DSv2 (#7343)
sotikoug83 Aug 6, 2026
107a285
Add usage logging for @-syntax path-based time travel (#7375)
sotikoug83 Aug 6, 2026
67f0100
Use DeltaV2Snapshot data skipping for v2 connector file selection (#7…
huan233usc Aug 6, 2026
f907436
[Flink] Fix Unity Catalog SQL metadata handling (#7349)
timothyw553 Aug 6, 2026
72a1cb2
[Spark] Add a log-commit-only action reader for SingleCommit and simp…
geraintcjy Aug 7, 2026
7dc4181
[Spark] Fix DSv2 read of null partition values (#7381)
sotikoug83 Aug 7, 2026
182e0f4
[Spark] Various DSv2 Suite refactors (#7380)
BrooksWalls Aug 7, 2026
5f78ddd
Introduce AMTPassthrough to AddFile action (#7390)
geraintcjy Aug 7, 2026
bf0416e
[Infra] Use MuleSoft as a Maven fallback (#7355)
timothyw553 Aug 7, 2026
52ed6e8
Add kernel factory for constructing Default Engine (#7384)
OussamaSaoudi Aug 7, 2026
28b8550
[Flink] Support narrow integer types in Unity Catalog tables (#7389)
timothyw553 Aug 7, 2026
43133b5
[Flink] Resolve VOID columns in catalog schemas (#7392)
timothyw553 Aug 7, 2026
a2846e5
[Spark] Use immutable partition values in V2 scan descriptors (#7405)
hariuserx Aug 8, 2026
760b343
Support startingTimestamp for Delta Sharing streaming when a version …
littlegrasscao Aug 8, 2026
5dba2c3
[Build] Make Spark 4.2 the default version (#7363)
timothyw553 Aug 10, 2026
39e95eb
[Kernel][Defaults] Reuse the file's Hadoop Configuration in the Parqu…
charlenelyu-db Aug 10, 2026
f044e7e
[Flink] Propagate writer completion failures (#7398)
timothyw553 Aug 10, 2026
a9a1469
[Spark] Rename DeltaSnapshotManager to DeltaV2SnapshotManager (#7391)
seewishnew Aug 10, 2026
4df35d4
[Protocol RFC] Nanosecond timestamps feature (#6082)
itamarst Aug 11, 2026
b9b923f
[Kernel] Retry when winning commits are unavailable (#7402)
timothyw553 Aug 11, 2026
cb97bdb
[Flink] Preserve final checkpoint lineage (#7401)
timothyw553 Aug 11, 2026
85bd136
Extend CDC testing: route DeltaCDCScala/SQLSuite through the V2 chang…
SanJSp Aug 11, 2026
b1bb7e0
[Spark] Add retries to post-commit log listing (#7393)
IAmAleksei Aug 11, 2026
d22546f
[Spark] Add AMTIncrementalWriteSuite and deduplicate leaf MDV positio…
prakharjain09 Aug 11, 2026
2d3625a
[AMT] Persist All-Files-in-CRC only for root-only AMT trees (no backR…
PhilPlato Aug 11, 2026
8d1d9ed
[Spark] Increase concurrent commit coordinator test wait (#7430)
murali-db Aug 11, 2026
9853f78
Enrich _last_checkpoint for adaptive metadata checkpoints (#7418)
yumingxuanguo-db Aug 11, 2026
eabbf4a
[Spark] Allow decimal partition columns in IcebergCompatV2 partition …
michaelzenz Aug 11, 2026
04ef251
[Spark] Rebuild adaptive metadata deletion vectors as absolute-path d…
geraintcjy Aug 11, 2026
a726ac8
Use error subclasses for DELTA_UNIVERSAL_FORMAT_VIOLATION (#7428)
ala Aug 12, 2026
6b876e7
Move NullType-columns-dropped text out of DELTA_NON_PARTITION_COLUMN_…
ala Aug 12, 2026
fd631b7
Fix DataManifestEntry tracking.status across AMT write paths (#7443)
prakharjain09 Aug 12, 2026
e1351f5
[Spark] Support scan-reported statistics in Delta connector scans (#7…
yyanyy Aug 12, 2026
00038f0
[Spark] Persist typed Iceberg partition values in adaptive metadata m…
ctring Aug 13, 2026
db56392
[Spark] Move DELTA_INVALID_IDEMPOTENT_WRITES_OPTIONS explanation into…
ala Aug 13, 2026
78e3352
Add withCheckpointDir hook to CatalogManagedStreamingSuite for config…
EricTsengTy Aug 15, 2026
f226bbb
[Kernel][Defaults] Reuse the already-read Parquet footer in the reade…
charlenelyu-db Aug 15, 2026
5cd424c
Install the correct AMT checkpoint provider upon cold snapshot init (…
yumingxuanguo-db Aug 15, 2026
ef808a4
[AMT] Clear backReference on superseding AddFile in removeRows (#7469)
PhilPlato Aug 15, 2026
ce4d836
[Spark] Introduce the "r" relative-path deletion vector storage type …
geraintcjy Aug 17, 2026
152b734
Add DeltaTableProvider test trait to parameterize table format in tes…
tmaire111 Aug 17, 2026
e92dbc9
Revert "[Kernel] Promote Snapshot/Scan/CommitRange accessors used by …
chiinlquah Aug 17, 2026
13de71c
[Flink] Add internal schema update support for Kernel-backed tables (…
KaiqiJinWow Aug 17, 2026
f27ced4
[Spark] Rename UUID-relative deletion vector descriptor helpers (#7473)
geraintcjy Aug 18, 2026
be2c96a
[Flink] Copy rows before buffering sink writes (#7474)
timothyw553 Aug 18, 2026
56cf0a0
[Spark] DSv2 Connector Metadata-only Delete Support for Partitioning …
ZiyaZa Aug 18, 2026
9be5f28
[SPARK] test refactor: Resolve delta._ imports in columnmapping tests…
fwc Aug 18, 2026
cccb39b
[RFC] Clarify backreference ownership for DV updates (#7471)
anoopj Aug 18, 2026
907d12f
Fix DataEntry tracking and CDF on the incremental AMT path (#7481)
prakharjain09 Aug 18, 2026
9442152
[Flink] Publish versioned artifacts for Flink 2.0 through 2.3 (#7421)
timothyw553 Aug 18, 2026
4d63bae
[Iceberg V4 AMT RFC] Require int64 TIMESTAMP(MICROS) timestamps (#7485)
anoopj Aug 19, 2026
8856b3d
[Spark] Fix flaky identity column streaming test (#7488)
murali-db Aug 19, 2026
3b04f11
[Spark] Fix flaky AMT partition values test (#7487)
murali-db Aug 19, 2026
67a6389
Add string-based path helpers to AMTUtils (#7486)
geraintcjy Aug 19, 2026
36c0c59
Install the correct AMT checkpoint provider on cold init when _last_c…
yumingxuanguo-db Aug 19, 2026
1d3fac7
Populate the five manifest_info row counts on a written AMT leaf (#7490)
prakharjain09 Aug 20, 2026
bae9765
[Build] Allow local overwrites during release publishing (#7495)
timothyw553 Aug 20, 2026
f4b0e0f
Introduce object-identity unique id for deletion vectors (#7489)
geraintcjy Aug 20, 2026
6737eba
[Protocol] Clarify tags are not required by extendedFileMetadata (#7496)
dengsh12 Aug 20, 2026
3f15ff4
[Storage] Fix S3 fast listing to exclude nested keys (#7451)
foss-contributor Aug 21, 2026
992b52f
Add dataChange as a top-level CommitInfo field (#7500)
PhilPlato Aug 21, 2026
f5902a4
Install the correct AMT checkpoint provider upon warm snapshot update…
yumingxuanguo-db Aug 22, 2026
3ef726e
[Spark] Add scoped profiling for Delta V2 planning (#7505)
murali-db Aug 22, 2026
bb1057f
[Spark] Preserve AddFile tags in Kernel-backed snapshots (#7508)
hariuserx Aug 22, 2026
0bb9738
[Spark] Support Kernel-backed size in bytes, file-size histogram, and…
HaoSun1993 Aug 24, 2026
14af3b7
[Spark] Make Row Tracking backfill errors actionable (#7514)
c27kwan Aug 24, 2026
cc83061
[Spark] Extract hasCapacityFor for DeltaSource admit function (#7516)
zifeif2 Aug 24, 2026
cb8e56a
[Spark] Add streaming partitioned-write tests for column-mapped table…
PorridgeSwim Aug 24, 2026
f81bd13
Split AMTIncrementalWriteSuite into a shared base and topical suites …
prakharjain09 Aug 25, 2026
250aa90
[Spark] DeltaV2OptimisticTransaction: Kernel-backed blind-append comm…
HaoSun1993 Aug 25, 2026
dc064f6
[Spark] Reconcile deletion vectors by object identity in log replay (…
geraintcjy Aug 25, 2026
0d2cf23
Record CommitInfo.dataChange on the commitLarge path (#7524)
PhilPlato Aug 26, 2026
b65766c
Add required version field to ContentRoot (#7530)
PhilPlato Aug 26, 2026
3ceda94
[Kernel-Spark] Unify DeltaV2SnapshotManager on the V1 Snapshot facade…
seewishnew Aug 26, 2026
d9b9029
[Spark] Add sort_order_id, key_metadata, split_offsets to AMTPassthro…
geraintcjy Aug 26, 2026
f2ab013
Fix clusteringColumns example in PROTOCOL.md to use path arrays (#7531)
anoopj Aug 26, 2026
4cc0f9d
Use DeltaTableProvider indirection for format literals in DeltaInsert…
tmaire111 Aug 27, 2026
5bdcc23
[Spark] Block unsupported maintenance on catalog-managed tables (#7525)
timothyw553 Aug 27, 2026
ad30d1e
[Spark] Extract DeltaFileSystemOptions from DeltaLog (#7504)
seewishnew Aug 27, 2026
d78feae
[Spark] Add V1 and V2 static file selection profiling (#7509)
murali-db Aug 27, 2026
53b8481
[Spark] Support conflict resolution for AMT log-only commits (#7521)
prakharjain09 Aug 27, 2026
1cbcb85
Give metadata-domain conflicts their own error class (#7539)
ala Aug 27, 2026
61edf2b
Split DELTA_REPLACE_WHERE_MISMATCH into subclasses to stop embedding …
ala Aug 27, 2026
f210c07
Revert "[Spark] DeltaV2OptimisticTransaction: Kernel-backed blind-app…
HaoSun1993 Aug 27, 2026
049a591
[Spark] Add scoped profiling frames to Delta V2 scan planning (#7542)
murali-db Aug 27, 2026
923adcf
[Spark] Write AMT content stats as typed Iceberg V4 values (#7535)
ctring Aug 27, 2026
c6f47bf
Discover the AMT checkpoint on cold/warm snapshot init via parallel C…
yumingxuanguo-db Aug 28, 2026
c498837
[Spark] Centralize the AMT-enabled check in AMTUtils.amtEnabled (#7545)
prakharjain09 Aug 28, 2026
b556699
[Spark] Retain the snapshot used for Delta V2 scan planning (#7548)
hariuserx Aug 28, 2026
abc2991
[Spark] Narrow two internal TahoeFileIndex listing helpers to private…
gengliangwang Aug 28, 2026
f14f31a
[Spark] AMT: field-id manifest reads survive column rename; write tim…
ctring Aug 28, 2026
aa5496c
[Spark] Make Delta DSV2 initialOffset idempotent (#7547)
PorridgeSwim Aug 29, 2026
24d4b21
[Spark] DeltaV2OptimisticTransaction: Kernel-backed blind-append comm…
HaoSun1993 Aug 29, 2026
d899dca
Preserve AddFile tags in the V4 AMT checkpoint (#7554)
PhilPlato Aug 29, 2026
e07f174
[Spark] AMT tests: reuse shared manifest-read helpers (#7555)
ctring Aug 30, 2026
10abc48
[Spark] Support conflict resolution when a log commit wins over an in…
prakharjain09 Aug 31, 2026
ccb85f7
[Spark] Add KernelActionUtils to convert Kernel actions to V1 Delta a…
HaoSun1993 Aug 31, 2026
4cb7d8e
[Storage] Synchronize local HDFS log writes across instances (#7567)
timothyw553 Aug 31, 2026
d7b8a44
[Spark] Convert DeltaV2Batch from Java to Scala (#7566)
murali-db Aug 31, 2026
b12db4d
[Kernel] Add PartitionKeyType to Transaction.getWriteContext for phys…
HaoSun1993 Sep 1, 2026
b224ab6
[Spark] Preserve catalog table in DeltaV2Snapshot (#7569)
seewishnew Sep 1, 2026
852f37b
Resolve the inherited file sequence number in the AMT reader (#7575)
PhilPlato Sep 1, 2026
987ad07
Move hard-coded messages of DELTA_UNIVERSAL_FORMAT_CONVERSION_FAILED …
ala Sep 1, 2026
6644310
[Spark] Abstract the AMT manifest deletion-vector bitmap behind a Man…
prakharjain09 Sep 1, 2026
b2fe753
[Spark] Support AMT deletion vectors in shallow clone (#7552)
geraintcjy Sep 1, 2026
62b1165
[Spark] Compute SnapshotState directly from CRC when possible (#6929)
felipepessoto Sep 1, 2026
d34cbe4
[Spark] Support decoding SetTransaction (TXN) actions from Delta Kern…
HaoSun1993 Sep 1, 2026
953746d
[Spark] Ignore staged commits in filesystem listings (#7550)
timothyw553 Sep 2, 2026
4f39d63
[Spark] Make the metrics schema output not nullable (#7582)
leonwind Sep 2, 2026
9360da7
Move closing parenthesis to its own line in OptimisticTransaction doC…
ChengJiX Sep 2, 2026
67ef998
[Spark] Refactor V2 scan E2E tests in Scala (#7589)
hariuserx Sep 3, 2026
be6cd8c
[Spark] Release cached DataFrame on commit in DeltaV2MicroBatchStream…
PorridgeSwim Sep 3, 2026
9b1ec79
Refine the ManifestBitmap API for manifest deletion vectors (#7591)
prakharjain09 Sep 3, 2026
ad6bafc
[Spark] Convert KernelEngineFactory from Java to Scala (#7590)
murali-db Sep 3, 2026
0a30ee8
[Spark] Update DeltaSourceSuite cases to work with DSv2 (#7587)
BrooksWalls Sep 3, 2026
c3afce0
Allow raising a bare Delta error class when subclasses exist (#7596)
ala Sep 3, 2026
4a16101
[Spark] Add standalone DSv2 table-manager cache skeleton (#7511)
seewishnew Sep 4, 2026
1d04c93
[UniForm] Remove Delta version to Iceberg snapshot ID mapping (#7606)
ChengJiX Sep 4, 2026
bde4d06
Abstract IncrementalAMTWriter old-AMT inputs behind an actions provid…
prakharjain09 Sep 4, 2026
0bb49e5
[Spark] Use relative path in DataEntry (#7608)
geraintcjy Sep 4, 2026
1d8adbc
AMT conflict resolution: winning AMT commit vs losing log or inline-t…
prakharjain09 Sep 4, 2026
ac1c99e
[Other] Add dev container configuration (#7581)
felipepessoto Sep 4, 2026
b7c43ec
[Spark] Add KernelContext for deferred Hadoop configuration (#7568)
seewishnew Sep 7, 2026
0d9fe2b
[Spark] Revert staged commit filesystem filtering (#7616)
timothyw553 Sep 7, 2026
2706eea
[SPARK] Disable inferring partition predicates from generated columns…
Omar-Saleh Sep 8, 2026
d29fe04
[Kernel] Wire JUnit 5 into kernel-api and convert two leaf suites (#7…
slachiewicz Sep 8, 2026
c4a8b4f
[Spark] Add DSv2 table API source tests (#7605)
BrooksWalls Sep 8, 2026
8d19b37
[Spark] Test refactor: less SharedSparkSession, protect beforeAll, mi…
fwc Sep 8, 2026
770664c
Clarify non-file action logging requirement for manifest commits (#7543)
anoopj Sep 8, 2026
6e1cdad
[Protocol RFC] Iceberg v4: use Int for backreference pos (#7586)
anoopj Sep 9, 2026
e277ff3
[Kernel] Fix kernel corrupt checkpoint bug (#7221)
austinshores Sep 9, 2026
0ae17d9
[SPARK] Add new expressions to AllowedExpressionAllowlist (#7634)
Omar-Saleh Sep 9, 2026
a2f9ac4
[Spark] Build default engines from KernelContext (#7611)
seewishnew Sep 9, 2026
3190dd5
Add Kernel-backed concurrent-writer conflict checking for DeltaV2Opti…
HaoSun1993 Sep 9, 2026
e0d317d
[Spark] Assert no V1 fallback in DeltaV2SourceRowTrackingSuite (#7636)
PorridgeSwim Sep 9, 2026
f053084
Retain table and column comments in CREATE OR REPLACE TABLE (#7035)
geraintcjy Sep 9, 2026
4c1dbe9
[Spark] Make the AdaptiveMetadata table feature removable (#7637)
geraintcjy Sep 10, 2026
66aea55
Pull commit-preparation logic out of OptimisticTransaction.doCommit (…
prakharjain09 Sep 10, 2026
e69e1e3
[Spark] Port catalog managed streaming suite for DSv2 (#7597)
BrooksWalls Sep 10, 2026
47bcf98
[Spark] Canonicalize deletion vector paths for AMT checksum validatio…
KaiqiJinWow Sep 10, 2026
06acae1
[PROTOCOL RFC] User-Defined Types (UDT) (#7560)
sanujbasu Sep 11, 2026
fc75759
[Spark] Ignore staged commits in filesystem listings (#7643)
timothyw553 Sep 11, 2026
dbc6b2e
Round-trip stats column names after DROP COLUMN (#7649)
zsxwing Sep 11, 2026
94883b8
Move column mapping advice for DROP/RENAME COLUMN errors into error s…
ala Sep 14, 2026
ffab8ee
[Spark] Implement DeltaV2TableManager with uncached snapshots (#7651)
seewishnew Sep 14, 2026
45bb624
[Spark] Rebalance file listing across tasks in CONVERT TO DELTA schem…
cravani Sep 14, 2026
efb67b2
[Spark] Extend AMTBackReferenceSuite to cover TRUNCATE and CONVERT ba…
PhilPlato Sep 15, 2026
8338b2b
[Spark] Add withConf helper to DeltaSQLCommandTest, use in INSERT tes…
fwc Sep 15, 2026
b330420
[Spark] Route DeltaV2Table through table manager cache (#7653)
seewishnew Sep 15, 2026
b672ebe
Structure the ALTER TABLE REPLACE COLUMNS unsupported-reason error (#…
ala Sep 15, 2026
0c81f33
[Spark] Emit an AMT checkpoint from deltaLog.checkpoint() and commitL…
prakharjain09 Sep 15, 2026
04698ba
[Spark] Re-verify AMT back references after conflict resolution (#7657)
PhilPlato Sep 15, 2026
18ba6b4
[Spark] Read the commit-level dataChange where a commit summary suffi…
PhilPlato Sep 15, 2026
79f8a92
Pad the log segment with the correct non-compacted deltas when a comp…
yumingxuanguo-db Sep 16, 2026
02084b3
[Spark] Extract first-batch start version resolution into DeltaStream…
PorridgeSwim Sep 16, 2026
c1ef568
Let DeltaErrors docs-link helpers accept a read-only SparkConf (#7669)
ala Sep 16, 2026
c6193cf
DeltaInsertIntoTest: don't set conf that's enabled by default (#7672)
fwc Sep 16, 2026
d07092a
[Spark] DeltaInsertIntoSchemaEvolutionSuite: withSQLConf -> withConf …
fwc Sep 16, 2026
9b74e80
[Spark] Update streaming schema mismatch error (#7664)
BrooksWalls Sep 16, 2026
debf4b2
Reuse or regenerate a losing full AMT checkpoint that conflicts with …
prakharjain09 Sep 17, 2026
8795732
[Spark] Centralize executeDml test helper in DeltaSQLCommandTest (#7675)
PorridgeSwim Sep 17, 2026
1d70393
[Protocol RFC] File data type (#7148)
dejankrak-db Sep 17, 2026
052429d
Update SQLSTATE for DELTA_FAIL_RELATIVIZE_PATH (#7674)
anniedde Sep 17, 2026
1242d60
[Spark] Add cached snapshot manager (#7565)
seewishnew Sep 17, 2026
283a5c5
[Spark] Log additional streaming source options (#7679)
zsxwing Sep 17, 2026
58fd061
[Delta] Stabilize metadata evolution schema assertions (#7684)
chiinlquah Sep 17, 2026
0097b87
[Spark] Add effective file sequence number to AddFile (#7681)
PhilPlato Sep 17, 2026
9e3d197
[Spark] Honor catalog maintenance permissions (#7564)
timothyw553 Sep 18, 2026
3b1b99c
Skip the Kernel-owned row-tracking domain metadata when committing vi…
HaoSun1993 Sep 18, 2026
bae5aa3
[Spark] Add AMT variants of the MERGE INTO test suites (#7688)
ctring Sep 20, 2026
e86931d
Add AMT conflict-resolution round metrics (#7694)
prakharjain09 Sep 21, 2026
aea3dcf
[Spark] Support snapshot-backed incremental AMT bootstrap (#7712)
KaiqiJinWow Sep 22, 2026
f1e17ad
[Delta] Log adaptive metadata tree invariant violations (#7693)
prakharjain09 Sep 22, 2026
5c9fc2d
Add query context for Delta V2 snapshot operations (#7702)
seewishnew Sep 22, 2026
ca18978
Add read-only Mumbling bitmap reader and PFOR codec (#7715)
prakharjain09 Sep 22, 2026
59ee981
[Spark] Extract DeletionVectorDescriptor object-identity helpers into…
geraintcjy Sep 22, 2026
ac61e9b
[Spark] Populate AMT data manifest tracking (#7691)
PhilPlato Sep 22, 2026
1e13fb0
Skip or regenerate a losing AMT OPTIMIZE checkpoint that conflicts wi…
prakharjain09 Sep 22, 2026
41bacdb
[Spark] Add DataSkipping V1 test coverage for AMT manifest tables (#7…
ctring Sep 22, 2026
fd744bc
[Spark] Expand DeltaLoggingProvider usage in DeltaLogging fns (#7714)
fwc Sep 23, 2026
e749699
Move emitAMTCheckpoint into the AMTWriterManager object (#7723)
prakharjain09 Sep 23, 2026
4d34254
Keep BEGIN/END index markers for AddCDCFile commits in CDF streaming …
zhu-tom Sep 23, 2026
3cca169
[Spark] Use object-identity-aware deletion vector ids in the duplicat…
geraintcjy Sep 23, 2026
47906ce
[Spark] Resolve inherited first_row_id via row-index filter providers…
PhilPlato Sep 24, 2026
680d045
[Spark] Remove unused identity conflict SQL hook (#7729)
shlokj Sep 24, 2026
ef3e91b
Move DELTA_PROTOCOL_CHANGED additional info into error subclasses (#7…
ala Sep 24, 2026
692fc65
[Spark] Handle missing checkpoint RDD block ID during MERGE (#7697)
tharun026 Sep 24, 2026
09f6b27
Move code related to table history in VacuumCommand.scala to DeltaHis…
rajeshparangi Sep 24, 2026
5e29e97
[Spark] Support AMT deletion vectors in RESTORE (#7733)
geraintcjy Sep 24, 2026
14174f0
Add query-context Delta V2 snapshot manager APIs (#7703)
seewishnew Sep 24, 2026
944e622
[Spark] Truncate LogSegment string representation to avoid driver OOM…
dhruvarya-db Sep 24, 2026
f02bab9
[Spark] Add fileType-preview table feature (#7678)
dejankrak-db Sep 25, 2026
9856e96
[Spark] Extract Delta insert test helpers for suite variants (#7736)
raulandrei00 Sep 25, 2026
f60368e
[Spark] Clean up test exclusions of the AMT data skipping test base (…
ctring Sep 26, 2026
ca2e195
Make ManifestBitmap immutable with an explicit mutable phase (#7745)
prakharjain09 Sep 26, 2026
178124f
[Spark] Strengthen AMT content-stats round-trip tests (#7749)
ctring Sep 28, 2026
f47ad0f
Include conflicting commit in DELTA_PROTOCOL_CHANGED.WRITE_TO_EMPTY_D…
ala Sep 28, 2026
8dc0c56
[RFC] Iceberg V4: commitInfo.dataChange requirement and contentRoot t…
PhilPlato Sep 28, 2026
b15ffe4
Replace free-form explanation in DELTA_ILLEGAL_OPTION with structured…
ala Sep 28, 2026
d9fd17e
Assert error conditions via checkError/getCondition instead of messag…
ala Sep 28, 2026
1bea2e4
[DELTA] Require concrete query context in snapshot APIs (#7744)
seewishnew Sep 28, 2026
5aeee34
[Spark] Handle unbackfilled AMT manifest commits and backfill precedi…
yumingxuanguo-db Sep 28, 2026
1d45ef3
[Spark] Validate committed dataChange against operation expectations …
jiahao-db Sep 29, 2026
b83b1d8
[Spark] Respect delta.enableVariantShredding in the DSv2 write path (…
viirya Sep 29, 2026
cf1d6a6
[Spark] Serialize Delta file statistics with the built-in to_json exp…
OscarZhou0107 Sep 29, 2026
d7337f7
[Flink] Fix flink to delta conversion for boolean and short values (#…
sebbegg Sep 29, 2026
599fdd6
[Spark] Use relative path in DataEntry deletion vector location (#7761)
geraintcjy Sep 29, 2026
94e7a05
[Spark] Add overridable hook for winning transaction SQL in IdentityC…
Soufiane-Barrada Sep 29, 2026
b474125
Add a mutable bitmap and writer to the Mumbling deletion vector forma…
prakharjain09 Sep 29, 2026
fe0fa36
[Spark] Add a TestBarrier utility and AMT conflict-resolution concurr…
prakharjain09 Sep 29, 2026
95a44f6
Pass query context through Delta V2 table snapshots (#7704)
seewishnew Sep 30, 2026
516dc91
Revert "[Spark] Add a TestBarrier utility and AMT conflict-resolution…
prakharjain09 Sep 30, 2026
6e8152c
[Spark] Reapply TestBarrier utility and AMT conflict-resolution concu…
prakharjain09 Sep 30, 2026
83ff213
[RFC] Align manifest min_sequence_number with Iceberg V4 (#7763)
PhilPlato Sep 30, 2026
6a62fd8
[Protocol] Accept Materialize Partition Columns RFC (#7650)
emkornfield Sep 30, 2026
01bbfed
Support column mapping in Delta V2 optimistic transactions (#7764)
HaoSun1993 Sep 30, 2026
da45be1
Revert "[Spark] Serialize Delta file statistics with the built-in to_…
OscarZhou0107 Sep 30, 2026
e53c7bc
[Spark] Replace cached snapshot after latest commit deletion (#7755)
seewishnew Sep 30, 2026
68d5659
Route cached Delta V2 snapshots with query context (#7708)
seewishnew Sep 30, 2026
d0f7ab2
Revert "[Spark] Respect delta.enableVariantShredding in the DSv2 writ…
viirya Oct 1, 2026
f27a6ad
[Spark] Preserve AMT back references during minor compaction (#7780)
yumingxuanguo-db Oct 1, 2026
4bc20f7
Recreate a losing incremental OPTIMIZE AMT checkpoint in-transaction …
prakharjain09 Oct 1, 2026
1149d81
[Spark] Populate manifest minimum sequence number (#7781)
PhilPlato Oct 1, 2026
5a211bb
[Delta] Pass Base Snapshot through table write path to replace Kernel…
HaoSun1993 Oct 1, 2026
5056f01
[Spark] Fix REORG TABLE APPLY (PURGE) failing on executors without a …
s-maling-celonis Oct 1, 2026
2d60b99
[PROTOCOL RFC] Add Concurrent Identity Columns table feature (#7272)
mkroll-db Oct 1, 2026
f4f7720
Use configured timeout for CDC AvailableNow test (#7787)
chiinlquah Oct 1, 2026
8915384
[Spark] Respect delta.enableVariantShredding in the DSv2 write path (…
viirya Oct 2, 2026
85a0b04
[Spark] Pass the Snapshot facade through the Delta V2 read path (#7793)
murali-db Oct 2, 2026
c0a99cd
[Build] Bump version to 4.5.0-SNAPSHOT (#7770)
seewishnew Oct 2, 2026
3d1969e
[Spark] Write adaptive metadata trees before Iceberg metadata generat…
KaiqiJinWow Oct 2, 2026
f03f45b
[Spark] Use mapPartitions for AMT leaf writes and pass catalog table …
geraintcjy Oct 3, 2026
52fc23e
[Spark] Support portable manifest row-index filter providers (#7746)
PhilPlato Oct 3, 2026
4eea023
Bump Iceberg from 1.11.0 to 1.12.0 (#7772)
nastra Oct 5, 2026
1a86e64
[Spark] Defer physical input-file construction in the Delta V2 scan p…
huan233usc Oct 5, 2026
b32d291
[SPARK] DeltaInsertIntoImplicitCastSuite: remove superfluous SET CONF…
fwc Oct 5, 2026
d99585d
Preserve row tracking watermark in Delta V2 commits (#7801)
HaoSun1993 Oct 5, 2026
2465323
[Kernel] Type future-version failures for UC reads (#7771)
seewishnew Oct 5, 2026
71e9845
[Spark] Reuse DeltaSourceSnapshot for V2 streaming initial snapshots …
PorridgeSwim Oct 5, 2026
15e0d40
Consolidate Snapshot state accessors onto a single checksum source (#…
dhruvarya-db Oct 5, 2026
b23bdb9
[Spark] Preserve literal string predicates in Delta overwrite filters…
longvu-db Oct 6, 2026
9a9e518
[Spark] Extract partition filtering helpers into DeltaLogUtils (#7809)
PorridgeSwim Oct 6, 2026
cb53c50
[Spark] Use AMTParquetFileFormat for AMT manifest writes (#7810)
ctring Oct 6, 2026
6ccac38
[Spark] Support ALTER TABLE ... REPLACE PARTITIONED BY WITH CLUSTER BY
sezruby Oct 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
20 changes: 20 additions & 0 deletions .devcontainer/devcontainer.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
{
"name": "Delta Lake Dev",
"image": "mcr.microsoft.com/devcontainers/java:17-bookworm",
"features": {
"ghcr.io/devcontainers/features/python:1": {
"version": "3.10"
}
},
"customizations": {
"vscode": {
"extensions": [
"scalameta.metals"
],
"settings": {
"metals.sbtScript": "${containerWorkspaceFolder}/build/sbt"
}
}
},
"postCreateCommand": "echo 'Dev container ready. Run: build/sbt compile'"
}
8 changes: 4 additions & 4 deletions .github/workflows/build.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -39,12 +39,12 @@ jobs:
key: delta-sbt-cache-cross-spark-${{ hashFiles('project/scripts/setup_unitycatalog_main.sh') }}

# publishM2 compiles every aggregated project, including storage, which has
# unitycatalog-client as a compile-scope dependency. test_cross_spark_publish.py also
# unitycatalog-client as a compile-scope dependency. test_published_artifacts.py also
# iterates over released Spark versions (sbt -DsparkVersion=<X.Y>), so we need UC's
# spark connector published for each variant Delta will resolve. Discover the version list
# at runtime from project/spark-versions.json (same source the matrix workflows use) and
# invoke the setup script per variant; cache hits (the steady state) make the script a
# no-op. Snapshot versions are skipped by the cross-Spark test, so only released ones are
# no-op. Snapshot versions are skipped by the published-artifact test, so only released ones are
# published here.
- name: Set up pinned Unity Catalog for each released Spark version
run: |
Expand All @@ -54,8 +54,8 @@ jobs:
echo "::endgroup::"
done

- name: Run cross-Spark build test
run: python project/tests/test_cross_spark_publish.py
- name: Test published artifacts
run: python project/tests/test_published_artifacts.py

- name: Save SBT cache
if: github.event_name == 'push' && github.ref == 'refs/heads/master'
Expand Down
39 changes: 35 additions & 4 deletions .github/workflows/flink_test.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6,13 +6,19 @@ on:
paths:
- 'flink/**'
- 'kernel/**'
- 'build.sbt'
- 'project/CrossFlinkVersions.scala'
- '.github/workflows/flink_test.yaml'
- '!**/*.md'
- '!**/*.txt'
pull_request:
branches: [master, branch-*]
paths:
- 'flink/**'
- 'kernel/**'
- 'build.sbt'
- 'project/CrossFlinkVersions.scala'
- '.github/workflows/flink_test.yaml'
- '!**/*.md'
- '!**/*.txt'

Expand All @@ -26,9 +32,34 @@ env:
SBT_OPTS: "-Dsbt.coursier.home-dir=/home/runner/.cache/coursier -Dsbt.ivy.home=/home/runner/.ivy2"

jobs:
generate-matrix:
name: "Generate Flink versions matrix"
runs-on: ubuntu-24.04
outputs:
flink-versions: ${{ steps.versions.outputs.flink-versions }}
steps:
- name: Checkout code
uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4.3.1
- name: Install Java
uses: actions/setup-java@c1e323688fd81a25caa38c78aa6df2d33d3e20d9 # v4.8.0
with:
distribution: "zulu"
java-version: "17"
- name: Generate matrix
id: versions
run: |
build/sbt exportFlinkVersionsJson
versions="$(jq -c '[.[].fullVersion]' target/flink-versions.json)"
echo "flink-versions=$versions" >> "$GITHUB_OUTPUT"

test:
name: "DF"
name: "DF: Flink ${{ matrix.flink }}"
needs: generate-matrix
runs-on: ubuntu-24.04
strategy:
fail-fast: false
matrix:
flink: ${{ fromJson(needs.generate-matrix.outputs.flink-versions) }}
steps:
- name: Show runner specs
run: |
Expand Down Expand Up @@ -56,7 +87,7 @@ jobs:
~/.sbt
~/.ivy2
~/.cache/coursier
key: sbt-flink
key: sbt-flink-${{ matrix.flink }}
- name: Check cache status
run: |
if [ "${{ steps.cache-sbt.outputs.cache-hit }}" == "true" ]; then
Expand All @@ -70,7 +101,7 @@ jobs:
uses: ./.github/actions/setup-unitycatalog
- name: Run unit tests
run: |
build/sbt flinkGroup/test
build/sbt -DflinkVersion=${{ matrix.flink }} flinkGroup/test
- name: Save SBT cache
if: github.event_name == 'push' && github.ref == 'refs/heads/master'
uses: actions/cache/save@0057852bfaa89a56745cba8c7296529d2fc39830 # v4.3.0
Expand All @@ -79,4 +110,4 @@ jobs:
~/.sbt
~/.ivy2
~/.cache/coursier
key: sbt-flink
key: sbt-flink-${{ matrix.flink }}
40 changes: 36 additions & 4 deletions PROTOCOL.md
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,7 @@
- [Generated Columns](#generated-columns)
- [Default Columns](#default-columns)
- [Identity Columns](#identity-columns)
- [Materialize Partition Columns](#materialize-partition-columns)
- [Writer Version Requirements](#writer-version-requirements)
- [Requirements for Readers](#requirements-for-readers)
- [Reader Version Requirements](#reader-version-requirements)
Expand Down Expand Up @@ -621,7 +622,7 @@ Field Name | Data Type | Description | optional/required
path| String | A relative path to a file from the root of the table or an absolute path to a file that should be removed from the table. The path is a URI as specified by [RFC 2396 URI Generic Syntax](https://www.ietf.org/rfc/rfc2396.txt), which needs to be decoded to get the data file path. | required
deletionTimestamp | Option[Long] | The time the deletion occurred, represented as milliseconds since the epoch | optional
dataChange | Boolean | When `false` the records in the removed file must be contained in one or more `add` file actions in the same version | required
extendedFileMetadata | Boolean | When `true` the fields `partitionValues`, `size`, and `tags` are present | optional
extendedFileMetadata | Boolean | When `true` the fields `partitionValues` and `size` are present | optional
partitionValues| Map[String, String] | A map from partition column to value for this file. See also [Partition Value Serialization](#Partition-Value-Serialization) | optional
size| Long | The size of this data file in bytes | optional
stats | [Statistics Struct](#Per-file-Statistics) | Contains statistics (e.g., count, min/max values for columns) about the data in this logical file | optional
Expand Down Expand Up @@ -1836,7 +1837,7 @@ The following is an example for the `domainMetadata` action definition of a tabl
{
"domainMetadata": {
"domain": "delta.clustering",
"configuration": "{\"clusteringColumns\":[\"col-daadafd7-7c20-4697-98f8-bff70199b1f9\", \"col-5abe0e80-cf57-47ac-9ffc-a861a3d1077e\"]}",
"configuration": "{\"clusteringColumns\":[[\"col-daadafd7-7c20-4697-98f8-bff70199b1f9\"], [\"col-5abe0e80-cf57-47ac-9ffc-a861a3d1077e\"]]}",
"removed": false
}
}
Expand All @@ -1845,11 +1846,12 @@ The example above converts `configuration` field into JSON format, including esc
```json
{
"clusteringColumns": [
"col-daadafd7-7c20-4697-98f8-bff70199b1f9",
"col-5abe0e80-cf57-47ac-9ffc-a861a3d1077e"
["col-daadafd7-7c20-4697-98f8-bff70199b1f9"],
["col-5abe0e80-cf57-47ac-9ffc-a861a3d1077e"]
]
}
```
Each entry in `clusteringColumns` is the name path of a clustering column: a single segment for a top-level column, and multiple segments for a nested column (for example, `["user", "address", "city"]`). If [Column Mapping](#column-mapping) is enabled, physical names are used for each segment.


# Variant Data Type
Expand Down Expand Up @@ -2490,6 +2492,34 @@ When `delta.identity.allowExplicitInsert` is false, writers should meet the foll
- Overflow when calculating generated Identity values should be detected and such writes should not be allowed.
- `delta.identity.highWaterMark` should be updated to the new highest value when the write operation commits.

## Materialize Partition Columns

When this feature is supported, partition columns are physically written to Parquet files alongside the data columns. To support this feature:
- The table must be on Writer Version 7, and a feature name `materializePartitionColumns` must exist in the table `protocol`'s `writerFeatures`.

Unlike most writer features, `materializePartitionColumns` has no associated `delta.enable*` table property and defines no additional metadata requirements. It is therefore [active](#active-features) whenever it is [supported](#supported-features): its presence in the `protocol`'s `writerFeatures` alone forces the writer requirements below. Hence, for this feature, the terms *supported*, *enabled*, and *active* (including in the decision matrix below) all refer to the same state.

When supported:
- When the writer feature `materializePartitionColumns` is set in the protocol, writers must materialize partition columns into every newly added Parquet data file referenced by an `AddFile` action. Each materialized partition column value must equal the corresponding logical partition value recorded in that file's `AddFile` `partitionValues`. This mimics the same partition column materialization requirement from [IcebergCompatV1](#iceberg-compatibility-v1) and [IcebergCompatV2](#iceberg-compatibility-v2). As such, the `materializePartitionColumns` feature can be seen as a subset of the requirements imposed by those features, providing the partition column materialization guarantee independently without requiring full Iceberg compatibility.
- When the writer feature `materializePartitionColumns` is not set in the table protocol, writers are not required to write partition columns to data files. Note that other features might still require materialization of partition values, such as [IcebergCompatV1](#iceberg-compatibility-v1).

When [Column Mapping](#column-mapping) is enabled, materialized partition columns are written to the Parquet data file using their assigned physical column names and field IDs, the same as data columns.

This feature does not impose any requirements on readers. All Delta readers must be able to read the table regardless of whether partition columns are materialized in the data files. If partition values are present in both parquet and AddFile metadata, Delta readers should continue to read partition values from AddFile metadata. Also, [file-level statistics](#per-file-statistics) should not be written for the partition column as it would repeat information already present in an AddFile's `partitionValues`.

Note that this table feature, as well as [IcebergCompatV1](#iceberg-compatibility-v1) (and related table features that require partition column materialization), if enabled, take priority over the `delta.writePartitionColumnsToParquet` table property. In other words, if a table feature is enabled that requires materialization of partition columns, and table metadata contains a `false` value for `delta.writePartitionColumnsToParquet`, partition columns must be materialized.

| Table feature enablement | Value of `delta.writePartitionColumnsToParquet` table property | Writer requirement |
| ------------------------ | -------------------------------------------------------------- | ------------------ |
| A writer feature requiring partition column materialization (eg. `materializePartitionColumns`) is *enabled* | `false` | Partition columns *must* be materialized in parquet data files |
| A writer feature requiring partition column materialization (eg. `materializePartitionColumns`) is *enabled* | `true` | Partition columns *must* be materialized in parquet data files |
| A writer feature requiring partition column materialization (eg. `materializePartitionColumns`) is *enabled* | unset | Partition columns *must* be materialized in parquet data files |
| No writer feature requiring partition column materialization is enabled | `false` | Partition columns *should not* be materialized in parquet data files |
| No writer feature requiring partition column materialization is enabled | `true` | Partition columns *should* be materialized in parquet data files |
| No writer feature requiring partition column materialization is enabled | unset | No requirement on partition column materialization |

The value of having both the table feature `materializePartitionColumns` and the table property `delta.writePartitionColumnsToParquet` supported is that not every table is going to need the heightened requirement of only allowing writes from writers that understand `materializePartitionColumns`. In other words, `materializePartitionColumns` imposes a writer compatibility edge that `delta.writePartitionColumnsToParquet` does not.

## Writer Version Requirements

The requirements of the writers according to the protocol versions are summarized in the table below. Each row inherits the requirements from the preceding row.
Expand Down Expand Up @@ -2525,6 +2555,7 @@ Property | Description | Details
`delta.parquet.compression.codec` | Compression codec writers SHOULD use for new Parquet data and checkpoint files. Changing this property does not affect existing files; a table may contain files written with different codecs, which is a normal and expected state. | Widely supported values (matched case-insensitively): `uncompressed`/`none` (no compression), `snappy`, `gzip`, `lz4` (deprecated, Hadoop framing), `lz4_raw` ([LZ4 block format](https://parquet.apache.org/docs/file-format/data-pages/compression/#lz4_raw)), `zstd`.<br><br>When absent, writers SHOULD default to `zstd`. If a writer does not support or recognize the specified codec, it SHOULD abort with an appropriate error or fall back to a default codec.<br><br>Readers SHOULD support all codecs listed above regardless of the current property value. Parquet files written with other [parquet-supported codecs](https://parquet.apache.org/docs/file-format/data-pages/compression/) may also exist; readers MAY support reading these files.
`delta.parquet.format.version` | Parquet data page format writers SHOULD use for new data and checkpoint files. This property is a directive to writers only; readers do not need to consult it, as Parquet pages are self-describing via the `PageType` field in each page header. Changing this property does not affect existing files; a table MAY contain files written with different data page versions, which is a normal and expected state. | Valid values: `1.0.0` (DataPageV1) and `2.x.x` (DataPageV2, where `x.x` is any minor.patch version). Recommended values are `1.0.0` and `2.12.0`.<br><br>When absent, writers SHOULD default to `1.0.0`. Writers SHOULD validate this property and abort if the value does not match `1.0.0` or `2.MINOR.PATCH`.<br><br>Readers SHOULD support both DataPageV1 and DataPageV2 pages regardless of this property's value. Tables intended for access by engines beyond the Delta Lake connectors SHOULD use `1.0.0`, as DataPageV2 support varies across the broader Parquet ecosystem.
`delta.enableVariantShredding` | When `true`, writers could write variant data to parquet files in [shredded](#variant-shredding) format. | Valid values: `true` (shredding allowed) and `false` (shredding not allowed).<br><br>When enabled, writers must ensure that the `variantShredding` table feature is present in the table `protocol`'s `readerFeatures` and `writerFeatures`.
`delta.writePartitionColumnsToParquet` | Controls whether writers SHOULD write partition columns in newly written data parquet files, in the absence of any writer features that necessitate writing of partition columns (eg. `IcebergCompatV1`). In other words, if no writer feature is enabled that requires materialization of partition columns, writers should read this property to decide whether to materialize partition columns in data parquet files or not. Writer features requirements take precedence over this property's value. Readers should continue to read partition values off of AddFile actions, regardless of the presence of partition values in data files. File-level statistics should not be present for partition columns in partitioned tables in any case. This setting does not apply to writers of files of any other file format. | Boolean field, with valid values `false` and `true`.

# Appendix

Expand All @@ -2551,6 +2582,7 @@ Feature | Name | Readers or Writers?
[Clustered Table](#clustered-table) | `clustering` | Writers only
[VACUUM Protocol Check](#vacuum-protocol-check) | `vacuumProtocolCheck` | Readers and Writers
[In-Commit Timestamps](#in-commit-timestamps) | `inCommitTimestamp` | Writers only
[Materialize Partition Columns](#materialize-partition-columns) | `materializePartitionColumns` | Writers only

## Deletion Vector Format

Expand Down
Loading
Loading