Repository navigation
Row-level concurrency for OSS Delta: reconcile the writes that provably cannot conflict (umbrella) #3
Description
Activity
- added a commit that references this issue
on Jul 8, 2026 - changed the title
[-][Umbrella] Reduce unnecessary write conflicts (row-level concurrency)[/-][+]Row-level concurrency for OSS Delta: reconcile the writes that provably cannot conflict (umbrella)[/+]on Aug 3, 2026 Live e2e validation — Case 4 (OPTIMIZE ⟂ DML), both directions
Ran a concurrent end-to-end test: two engines committing to the same Delta table at the same time — one running this fork's build on Apache Spark 4.0, the other Databricks Runtime 16.4 — racing on the shared
_delta_log(coordinated only by sentinel marker files, so these are real atomic-commit version races, not replayed conflicts).Confirmed DBR-compatible: both Case 4 directions reconciled live instead of aborting, and each engine independently verified the table stayed consistent — matching row counts, no duplicate keys, every committed row-level change visible.
Case 4a — forward remap (
sezruby/delta#9): a losing OPTIMIZE absorbs the winning DML's in-place deletion vector into its compacted output. Run once per winning operation:DELETEwinner: 14 forward remaps (winningOp=DELETE). OPTIMIZE side: 48 commits / 2 aborts; DELETE side: 40 / 0 abort / 0 straggler; final rows agree (999,960), no dup — both PASS.- merge-on-read
UPDATEwinner: 11 forward remaps (winningOp=UPDATE), each a single-winner remap. OPTIMIZE side: 18 commits / 0 abort; UPDATE side: 20 / 0 abort / every update visible; final rows agree (200,000), no dup — both PASS. - merge-on-read
MERGE(matched-DELETE) winner: 1 forward remap (winningOp=MERGE). OPTIMIZE side: 52 commits / 0 abort; MERGE side: 20 / 0 abort / 0 straggler; final rows agree (199,980), no dup — both PASS.
(A merge-on-read
UPDATEwrites a new small file each iteration, so OPTIMIZE keeps finding work to compact and remaps repeatedly; a matched-DELETEMERGE, like aDELETE, is deletion-vector-only, so OPTIMIZE converges after the first compaction and the later windows are no-ops — hence the single remap, not a limitation.)Case 4b — reverse remap (
sezruby/delta#11): a losing UPDATE rebases its deletion vector onto the winner OPTIMIZE's compacted output. 3 reverse remaps (winningOp=OPTIMIZE). UPDATE side: 33 commits / 0 abort / every update visible; OPTIMIZE side: 30 commits (8 real compactions); final rows agree (1,000,000), no dup — both PASS.Reconcile events are the existing
delta.optimize.conflictReconciliation.remapped/.reverseRemappedrecords.Scope note: this reconcile is for deletion-vector-enabled tables, where merge-on-read is the default (
UPDATE_USE_PERSISTENT_DELETION_VECTORSdefaults on), so the winner masks its changed rows with a deletion vector in place —RemoveFile(P)+AddFile(P, dv)at the same path. Forward remap (4a) reconciles that regardless of the operation that wrote the DV —DELETE, merge-on-readUPDATE, and merge-on-readMERGEalike (the three runs above); any coexisting new file a merge-on-readUPDATE/MERGEalso writes is left untouched.OptimizeConflictReconciliationSuitecovers these in-unit. The genuinely-deferred cases remain the umbrella's backlog rows 5/6 — reclustering /ZORDER⟂ DML (rows permuted across files, so no derivable offset) andMERGEnet-new inserts / not-matched-by-source — which permute or reclassify rows and so need row tracking.
Background
Delta's documented row-level concurrency states that several concurrent-write pairs on a table with
deletion vectors cannot conflict — two DML statements editing disjoint rows, a compaction
OPTIMIZEvs a concurrentDELETE/UPDATE, a blindINSERTagainst anything. OSS Delta todaystill aborts the loser in most of these:
ConflictCheckerconservatively treats a file that aconcurrent commit removed, or added a DV to, as a conflict — even when the two commits are logically
independent.
This umbrella tracks closing that gap without any protocol change and without row tracking, as a
series of small, independently reviewable, opt-in (default-off) additions to
ConflictChecker.Organizing principle
Rather than enumerating special cases, the work is driven by one property:
Everything reconstructable from that information is in scope (Phase 1); anything needing a
whole-table read, a row permutation, or row tracking is backlog. This lands exactly at parity with
the documented contract, because the deferred cases are the ones the reference implementation does
not reconcile either.
Phase 1 — opt-in, no row tracking
OPTIMIZE(dataChange=false) ⟂ non-blind appenddataChange=false(relocation) adds from the added-files conflict checkOPTIMIZEloses to a concurrent DMLOPTIMIZEOPTIMIZEpersisted and remap the losing DML's deletions onto the compacted outputTracked issues / PRs
Each case carries its own design in its own issue.
OPTIMIZEvs non-blind append — [RLC] Case 3 — exclude relocation (dataChange=false) files from the append-conflict check #6 · upstreamed as [File-level concurrency] Exclude no-data-change (OPTIMIZE) files from append-only conflict checks delta-io/delta#7331 (fork PR [RLC] Case 3 — exclude relocation (dataChange=false) files from the append-conflict check #10 closed)OPTIMIZEloses to DML (forward remap) — [RLC] Case 4 — OPTIMIZE ⟂ DML reconciliation (compaction offset remap) #7 · PR [RLC] Case 4a — OPTIMIZE loses to DML (forward compaction offset remap) #9OPTIMIZE(reverse remap) — [RLC] Case 4 — OPTIMIZE ⟂ DML reconciliation (compaction offset remap) #7 · PR [RLC] Case 4b — DML loses to OPTIMIZE (reverse compaction offset remap) #11Dependencies. Case 4 reuses Case 2's deletion-vector helpers and splits into a shared write side
(PR #12) that captures & persists the composition, plus two reconcile readers that consume it —
forward (4a) and reverse (4b) — giving the stacking chain Case 2 → 4 (capture) → 4a → 4b. Cases
1 and 3 are independent. Suggested review order: this issue → 1 → 2 → 4 (capture) → 4a → 4b.
How the compaction cases stay cheap (Case 4)
Compaction concatenates each source file's live rows into one contiguous run in the output without
permuting within a file, so the whole problem reduces to where each source's run starts in the
output. That offset isn't derivable from the read plan (scan splits are sorted before packing, a
shuffle reorders rows, a non-Spark engine's order is opaque), so it is observed at write time: a
write-stage operator reads the per-row source-file identity that
input_file_name()already exposesand records, on each removed source's
RemoveFiletombstone, the run it contributed —compactedInto(the output path) andcompactionInfo({rowOffsetInTarget, sourceNumPhysicalRecords}).Rows pass through unchanged: no
_metadata.row_index, no helper column, no row copy.sourceNumPhysicalRecordsis a physical count, kept self-consistent with the tombstone's own DV;the live run length is
sourceNumPhysicalRecords − |sourceDV|, so a source that carried a read-timeDV has its non-contiguous live rows reconstructed lazily at conflict time. The tags live on the
short-lived tombstone rather than the output
AddFile(which snapshot reconstruction replays onevery read), so the read hot path is untouched; tombstone retention (~7 days default) comfortably
outlives the conflict window (a concurrent transaction's runtime), so the tag is still present when a
losing DML reads it.
Safety model
today's abort, so a bug can never produce wrong data — only a spurious (but correct) abort.
affected row maps within its source's run; otherwise the loser aborts exactly as it does today.
aborts.
Backlog — needs row tracking, not part of this proposal
ZORDER⟂ DMLNOT MATCHED BY SOURCEThe reference contract does not reconcile these either, so Phase 1 already reaches full parity;
closing them would require load-bearing row tracking and a scan inside the otherwise-cheap,
synchronous conflict path.
Related work
Case 2; this umbrella's Case 2 proposes a deletion-vector-only path that does not require row
tracking, and the remaining cases (compaction, append, reader-side skipping) are additive. Happy
to coordinate.
Notes
Clean-room: this proposal cites only Delta's public row-level-concurrency contract and our own
black-box observations of its published behavior; it copies no proprietary source. Case 4's
composition carrier uses the
compactedInto/compactionInfotombstone-tag format so anOPTIMIZEwritten by another engine can be reconciled against on a shared table; that interop is modeled on the
published format and is best-effort, not a verified guarantee (an unrecognized tag only forgoes the
reconcile — never wrong data).