From 1a891a8b194e706a5027029ec86bbf1c437914c7 Mon Sep 17 00:00:00 2001 From: bobbyiliev Date: Mon, 6 Jul 2026 22:05:27 +0300 Subject: [PATCH 01/19] design: Add design doc for Subscribe SDKs and external sink framework --- .../design/20260706_subscribe_sdk.md | 805 ++++++++++++++++++ 1 file changed, 805 insertions(+) create mode 100644 doc/developer/design/20260706_subscribe_sdk.md diff --git a/doc/developer/design/20260706_subscribe_sdk.md b/doc/developer/design/20260706_subscribe_sdk.md new file mode 100644 index 0000000000000..2b7a1a5871886 --- /dev/null +++ b/doc/developer/design/20260706_subscribe_sdk.md @@ -0,0 +1,805 @@ +# Subscribe SDKs: canonical clients for `SUBSCRIBE` and a framework for external sinks + +- Associated: + - (TBD: link the Subscribe SDK program epic once filed) + - [MaterializeIncLabs/mz-redis-sync](https://github.com/MaterializeIncLabs/mz-redis-sync) (prior art) + - [MaterializeIncLabs/novu-materialize-integration](https://github.com/MaterializeIncLabs/novu-materialize-integration) (prior art) + - [Durable subscriptions pattern](https://materialize.com/docs/transform-data/patterns/durable-subscriptions/) (documented protocol this work packages) + - [database-issues#5182](https://github.com/MaterializeInc/database-issues/issues/5182) (errors poison subscribe dataflows, relevant server limitation) + +## The Problem + +`SUBSCRIBE` is the primitive that turns Materialize from a database you query +into a database that drives your systems. It is how users push incrementally +maintained results into caches, notification services, feature stores, +websockets, and anything else that is not reachable by `CREATE SINK`. Today we +give users the primitive and a pattern document, and then every user rebuilds +the same client-side machinery from scratch. + +That machinery is genuinely hard to get right. The correct protocol for +consuming `SUBSCRIBE` durably is: + +1. Subscribe `WITH (PROGRESS, SNAPSHOT true)`. +2. Buffer updates until a progress message proves a timestamp is closed. +3. Apply the closed batch and persist its frontier atomically. +4. On restart, resume `WITH (PROGRESS, SNAPSHOT false) AS OF frontier - 1`, + because `SNAPSHOT false` emits updates strictly *after* the `AS OF` + (`src/compute/src/sink/subscribe.rs`, the `should_emit` closure), so + subtracting one is what makes the resume gap-free. +5. Keep `RETAIN HISTORY` on the subscribed object wider than your worst-case + downtime, and handle the compaction-horizon error when it is not. + +Each step has a failure mode that is invisible in testing and destructive in +production. We know this because every existing implementation we could find, +including our own demos, gets a different step wrong: + +| Codebase | Commit discipline | Resume | Idempotency | Retractions | +| --- | --- | --- | --- | --- | +| mz-redis-sync (Labs) | Correct: commits data and frontier atomically at progress rows | Omits `SNAPSHOT` on resume, which defaults to `true`, so every restart replays the full snapshot (contradicting its own commit message) | Inherent (keyed writes) | Delegated to `ENVELOPE UPSERT`, but `key_violation` crashes the process and poisons restarts | +| novu-materialize-integration (Labs) | Persists the timestamp of the *first row* of a batch. A crash mid-batch permanently loses the rest of that timestamp | Wall-clock guess of whether the checkpoint is still inside the retention window | Content-hash keys (false-dedups identical payloads across time) | Maps `mz_diff = -1` to a notification revoke, a great idea, implemented with two bugs | +| novu-materialize-poc (Labs) | Correct idea: buffer until the timestamp advances | No `PROGRESS`, so batch boundaries are a heuristic | None, duplicates on replay | Ignored, deletes trigger spurious notifications | + +(One fairness note on the first row: because that sink writes keyed upserts +transactionally, the snapshot replay is wasteful rather than incorrect, +every restart rewrites the full view into Redis. It still contradicts the +commit message that claims to avoid it, which is its own kind of evidence.) + +None of the three reconnects on failure. None has a single test. One declared +a retry library as a dependency and never imported it. These are demos, +written quickly, some AI-assisted, and that is precisely the point: this is +what subscribe consumers look like when they are written the way real +integrations get written, by whoever shows up with a deadline. The +correctness burden currently sits on the application side of an API this +subtle, so this is the result. Throughout this document, the demos serve +only as evidence of the failure modes. Every claim about correct behavior is +anchored in the Materialize codebase and documentation, which are the ground +truth this design is built on. + +Beyond correctness, there are stability hazards that clients must be built +around and currently are not: + +- The coordinator buffers subscribe output in an unbounded channel with no + backpressure (`src/adapter/src/active_compute_sink.rs`, acknowledged + TODO). A slow consumer translates into `environmentd` memory growth. + Separately, subscribe output is subject to `max_result_size` in the + compute-client layer (`PendingSubscribe::stash`), so an initial snapshot + larger than that limit terminates the subscription with a "total result + exceeds max size" error, a sizing hazard no demo handles. +- Once a subscribe dataflow produces an error it repeats that error forever + (database-issues#5182). Recovery is always reconnect-shaped, and clients + that do not know this retry into the same error. +- The compaction-horizon error is distinguishable only by matching the error + message text, under a generic SQLSTATE (`DATA_EXCEPTION`, shared with + other failures). Worse, the text is not stable: the timestamp-selection + rework (#34712) replaced "Timestamp (X) is not valid for all inputs" with + "could not find a valid timestamp for the query". mz-redis-sync regexes + the old text, so its one bespoke error handler is already broken against + current Materialize, and its non-matching branch silently swallows the + error and re-enters `FETCH` inside an aborted transaction, an infinite + loop. This is what message-text dispatch does over time. + +Finally, our docs teach the naive version. Every per-language client page +(`doc/user/content/integrations/client-libraries/*.md`) shows a bare +`DECLARE`/`FETCH` loop with no progress handling, no resume, and no +reconnection. The durable-subscriptions pattern doc is correct and complete, +but it is prose describing a protocol, and prose does not compose into +applications. + +The problem, stated in one line: **we make users hand-implement a consistency +protocol that we already know how to implement, and they predictably get it +wrong.** + +## Success Criteria + +A solution is successful if: + +1. **Correct by construction.** A user following the happy path of the SDK + cannot commit an unclosed timestamp, cannot get the `AS OF` boundary + arithmetic wrong, and cannot silently lose or duplicate updates across a + restart. The resume token is opaque. The batch is the unit of consumption. +2. **Minutes to first stream.** `install`, paste a connection string, write + five lines, see typed updates. The SDK works against Materialize Cloud, + self-managed, and the emulator without configuration differences. +3. **Stable under failure.** Network blips, cluster restarts, and process + crashes are recovered automatically or surfaced as one of a small set of + typed, documented errors with remediation guidance. Memory use is bounded + and documented, including during snapshots. +4. **General over destinations.** The sink framework cleanly supports the two + families of destinations, keyed-state stores (Redis, DynamoDB, Postgres + upsert tables, search indexes, caches) and event receivers (webhooks, + queues, notification services). Redis ships first as the reference + implementation, not as a special case. +5. **Consistent across languages.** TypeScript, Python, and Rust expose the + same model with idiomatic surfaces, verified by a shared conformance suite + rather than by hope. Behavior is specified in a written protocol document. +6. **Honest about guarantees.** The SDK names its delivery guarantees + precisely, exactly-once state application for transactional sinks and + at-least-once with idempotency keys for side-effecting sinks, and makes + the user pick one explicitly. +7. **Testable by users.** Sink authors get a testing story from us, recorded + stream fixtures and an emulator harness, so their sinks can be tested + without a live environment. + +Quantitative proxies: replace mz-redis-sync with an SDK-based equivalent at +functional parity plus the known bugs fixed. A chaos suite that kills the +process at arbitrary points and verifies destination convergence passes for +every language. The per-language docs pages can be rewritten on top of the SDK +with fewer lines than the naive loops they replace. + +## Out of Scope + +- **Rewriting subscribe server internals.** The SDK is buildable on today's + semantics, and a scoped set of server-side companion changes is planned as + part of this program (see "Materialize-side workstream"). Anything beyond + that scoped list, such as a general subscribe protocol redesign, is out of + scope. +- **Exactly-once side effects.** No client library can make an HTTP call + exactly once. We provide at-least-once delivery plus idempotency-key + helpers, and we say so plainly. +- **A browser/edge SDK.** The WebSocket endpoint (`/api/experimental/sql`) + does support `SUBSCRIBE`, but it is experimental. A browser client is + future work gated on stabilizing that surface. +- **Go, Java, and other languages in v1.** The spec and conformance suite are + designed to make additional languages cheap, but the v1 program is Rust, + Python, and TypeScript, in that order. GA can precede the TypeScript + release (see the Delivery plan). +- **Private-preview subscribe features.** `ENVELOPE DEBEZIUM` and + `WITHIN TIMESTAMP ORDER BY` are behind feature flags. The SDK's model + leaves room for them but v1 does not depend on them. +- **Scale-out consumption.** Sharding one subscription's updates across + multiple parallel workers has no server-side support and is not attempted. + High availability itself is in scope (fencing in v1, active-passive + failover shortly after, see Layer 2), the exclusion here is only parallel + fan-out of a single subscription for throughput. +- **General driver features.** The SDK consumes `SUBSCRIBE` (and the small + amount of SQL needed around it). It is not a query builder, ORM, or general + Materialize client. + +## Solution Proposal + +Ship a family of client libraries, one written behavior spec, and one shared +conformance suite, under a working title of the **Materialize Subscribe SDK**: + +``` + +--------------------------------------+ + | Layer 3: sink kit + built-in sinks | + | (Redis reference sink, webhook sink)| + +--------------------------------------+ + | Layer 2: durable subscribe | + | (checkpoints, guarantees, policies) | + +--------------------------------------+ + | Layer 1: subscribe client | + | (typed stream of closed batches) | + +--------------------------------------+ + | native pg driver per language | + | (node-postgres / psycopg / tokio-pg)| + +--------------------------------------+ +``` + +The layers are strictly separated so each is independently useful. A user who +only wants live updates in a websocket handler uses Layer 1 and never sees a +checkpoint. A user syncing Redis uses Layer 3 and never sees a `FETCH`. + +### The model: five opinions + +The SDK's UX comes from five opinions applied uniformly across languages: + +1. **The unit of consumption is the closed timestamp, never the bare row.** + The stream yields consistent batches. Each batch contains every update for + an interval of timestamps that a progress message has proven complete, + plus the frontier and a resume token. Users physically cannot observe a + half-delivered timestamp. (An advanced `raw()` mode exposes per-row + delivery for power users, clearly documented as forfeiting the batch + guarantees.) +2. **Resume tokens are opaque.** A token encapsulates the frontier, the + `SNAPSHOT false AS OF frontier - 1` arithmetic, the query fingerprint, and + a fencing epoch. Users store bytes and hand them back. There is no + timestamp arithmetic in user code. +3. **Guarantees are named and chosen, not implied.** Durable consumption + requires the user to pick `transactional` (checkpoint commits atomically + with effects, exactly-once state application) or `at_least_once` (effects + first, checkpoint after, idempotency keys provided). There is no default + that quietly picks one. +4. **Reconnection is the SDK's job.** Transient failures are retried with + jittered backoff and automatic resume. Everything else surfaces as one of + a small typed error set, each carrying remediation guidance. +5. **One spec, one conformance suite, three implementations.** Language + surfaces are idiomatic, behavior is identical, and identical is checked by + machines. + +### Layer 1: the subscribe client + +Layer 1 wraps the ecosystem-native PostgreSQL driver and runs the +`DECLARE`/`FETCH` loop, presenting a typed stream. + +Illustrative TypeScript (all sketches in this doc are illustrative, not final +API commitments): + +```typescript +import { SubscribeClient } from "@materializeinc/subscribe"; + +const client = await SubscribeClient.connect(process.env.MZ_URL!); + +const stream = client.subscribe({ + query: "SELECT id, amount FROM winning_bids", + envelope: { upsert: { key: ["id"] } }, +}); + +for await (const batch of stream) { + // batch.frontier: bigint -- every update with timestamp < this is present + // batch.resumeToken: ResumeToken -- opaque, serializable + // batch.isSnapshot: boolean + // batch.updates: Upsert[] | Delete[] | KeyViolation[] + render(batch.updates); +} +``` + +Illustrative Python: + +```python +from materialize_subscribe import connect, UpsertEnvelope + +client = connect(dsn) +for batch in client.subscribe( + "SELECT id, amount FROM winning_bids", + envelope=UpsertEnvelope(key=["id"]), +): + for update in batch.updates: + ... +``` + +Illustrative Rust: + +```rust +let client = mz_subscribe::Client::connect(&dsn).await?; +let mut stream = client + .subscribe(Subscribe::query("SELECT id, amount FROM winning_bids") + .envelope_upsert(["id"])) + .await?; +while let Some(batch) = stream.try_next().await? { + apply(batch.updates()); +} +``` + +Responsibilities and design points: + +- **Batching by progress.** The client requests `PROGRESS` always. It buffers + updates and emits a batch when the frontier advances. Progress-only + advancement (idle periods) is surfaced as an empty batch carrying a fresh + resume token, so downstream checkpoints keep advancing during quiet hours. + This directly fixes the failure mode where an idle subscription's + checkpoint ages out of the retention window. +- **Envelope decoding.** Three modes map to typed events: + - default diff envelope: `Insert{row, diff}` / `Retract{row, diff}` with + multiplicities preserved (a `diff` of -3 is three retractions and the + type says so), + - `ENVELOPE UPSERT`: `Upsert{key, value}` / `Delete{key}` / + `KeyViolation{key}`. Key violations are a first-class event, not an + exception, because a crash-restart loop cannot fix them. +- **Snapshot streaming with bounded memory.** The initial snapshot is one + giant timestamp and can exceed client memory if buffered whole (the + mz-redis-sync failure mode). The client streams it as chunked batches + flagged `partial: true`, with the closing chunk carrying the resume token. + Consumers that need atomic snapshot visibility get help from Layer 3. + Server-side, a snapshot exceeding `max_result_size` terminates the + subscription with a "total result exceeds max size" error regardless of + client behavior. The SDK surfaces that as a typed error with sizing + remediation rather than as an opaque stream failure. +- **Fetch pacing.** `FETCH c WITH (timeout ...)` with adaptive sizing. + The client always drains the server promptly and applies backpressure to + the application from its own bounded buffer, because unread results buffer + without limit in `environmentd`. If the application cannot keep up, the + documented options are a larger bound, spill-to-disk (opt-in), or letting + the buffer block the FETCH loop and accepting server-side growth. The SDK + makes the trade-off visible instead of implicit. +- **Typed errors.** All failures map to a small taxonomy: + + | Error | Meaning | Default behavior | + | --- | --- | --- | + | `Transient` | network blip, cluster restart, replica loss | auto-reconnect and resume | + | `CompactionHorizon` | resume point older than retained history | surface with policy hook (see Layer 2) | + | `DependencyDropped` | subscribed object or its cluster dropped | surface, policy hook can `refollow` recreated objects (see Layer 2) | + | `StreamPoisoned` | dataflow error (e.g. division by zero in the view), repeats forever per database-issues#5182 | surface with explanation, reconnect will not fix until the data or view changes | + | `SchemaMismatch` | resumed subscription's columns differ from checkpoint fingerprint | surface with policy hook | + | `Fatal` | auth, TLS, SQL errors in the user's query | surface immediately | + + The mapping rules key on structured error codes, which is the first item + in the Materialize-side workstream, sequenced to land before the SDK's + stable release. Because the SDK must also work against server versions + that predate those codes, a stable release may carry a message-text + fallback in exactly one internal function, gated on server version, + covered by conformance tests pinned to each legacy message (the text has + already changed once across releases, see the Problem section), and + removed when pre-code versions age out of support. No other + message-string behavior ships anywhere. +- **One direct connection per subscription.** A subscription owns a dedicated + connection for its lifetime. The SDK documents, and detects where it can, + that transaction-mode poolers (PgBouncer and friends) break the + `DECLARE`/`FETCH` loop. Applications running many subscriptions get an + explicit connection budget instead of a surprise. +- **Cluster targeting and resume cost.** The subscribe options include the + target cluster (`SET cluster`), and the docs shipped with the SDK + recommend a dedicated serving cluster for subscriptions. The SDK surfaces + the cost asymmetry on resume: `SUBSCRIBE ` on a materialized view + or table resumes cheaply from storage, while `SUBSCRIBE (SELECT ...)` + rebuilds a dataflow over retained history. Guidance: materialize the view + you sink. +- **Ordering is per timestamp, not within it.** Updates within a closed batch + are unordered (server-side `WITHIN TIMESTAMP ORDER BY` exists but is + private preview). Keyed-state sinks are insensitive to this. Event sinks + that need deterministic order within a timestamp can sort client-side via + a batch option, and the docs say when that matters. +- **Bounded subscriptions.** `UP TO` is exposed as a first-class option, which + makes deterministic reads of a timestamp window possible. This is useful in + its own right and is how much of the SDK's own test suite drives itself. +- **Session hygiene.** Sets `application_name`, suppresses the welcome notice + where drivers require it, exposes the connection's cluster and role in a + startup log line, and pre-validates the query with + `SELECT * FROM () WHERE FALSE LIMIT 0` for fail-fast schema and + permission errors (an mz-redis-sync trick worth canonizing). +- **Type mapping contract.** Each language documents a total mapping from + Materialize types to SDK values, including `numeric` precision, temporal + types, arrays, `jsonb`, and NULLs. The conformance vectors include every + type. + +### Layer 2: durable subscribe + +Layer 2 adds checkpointing and delivery guarantees on top of Layer 1. + +```typescript +import { durableSubscribe, redisCheckpoints } from "@materializeinc/subscribe"; + +await durableSubscribe(client, { + name: "bids-to-webhook", // checkpoint identity + query: "SELECT id, amount FROM winning_bids", + checkpoints: redisCheckpoints(redis), // pluggable store + guarantee: "at_least_once", + onCompactionHorizon: "resnapshot", // or "fail", or callback + handler: async (batch, ctx) => { + for (const event of batch.updates) { + await sendWebhook(event, { idempotencyKey: ctx.idempotencyKey(event) }); + } + }, +}); +``` + +Design points: + +- **Checkpoint stores are pluggable and tiny.** The interface is + `load(name) -> ResumeToken | null` and `store(name, token)`. Ships with: + the destination itself (the strongly recommended default, see Layer 3), + a Materialize table (zero extra infrastructure, the Novu demos' good idea), + Postgres, and a local file (development). +- **Two named guarantees.** + - `transactional`: the handler receives the batch and the token and must + commit both atomically (Layer 3 sinks do this for you). Result: + exactly-once state application. The destination always equals the source + at some real Materialize timestamp. + - `at_least_once`: effects run first, the checkpoint commits after the + handler returns. Result: replays possible after a crash, and + `ctx.idempotencyKey(event)` provides a stable key derived from + subscription name, `mz_timestamp`, key columns, and an ordinal within + the batch. This fixes both observed idempotency failures: content-only + hashes that false-dedup distinct events, and no key at all. +- **Retraction policy for event sinks.** Side-effecting consumers declare + what a retraction means: `ignore`, or `compensate(fn)` (the Novu demo's + retraction-to-revoke mapping, promoted to a supported concept). + Compensation needs a correlation identity that is stable between an + insert and its later retraction, which the delivery idempotency key + deliberately is not (it includes the timestamp and ordinal). The SDK + therefore provides `ctx.correlationKey(event)`, derived from the row's + key columns only, for exactly this purpose. The Novu demo collapsed both + identities into one content hash, which made revokes work but false-dedups + distinct events. Separating the two identities is the fix. +- **Compaction-horizon policy.** When the checkpoint is older than retained + history, policy decides: `fail` (default, explicit), `restart_from_now` + (accept a gap, for notification-style consumers), or `resnapshot` (rebuild + destination state, meaningful mainly for Layer 3 keyed sinks which know how + to reconcile). The SDK detects the condition from the typed error, not from + wall-clock guessing against `RETAIN HISTORY` config. +- **Query fingerprinting.** The token embeds a fingerprint of the query text + and output schema. Resuming with a changed query surfaces `SchemaMismatch` + and the policy hook chooses between failing and re-snapshotting. No demo + handled this. Real deployments hit it on their first schema change. Tokens + themselves carry a format version so stored checkpoints survive SDK + upgrades. +- **Object swaps and blue/green deployments.** Recreating or swapping the + subscribed object (the recommended zero-downtime deploy pattern, and what + tooling like mz-deploy automates) terminates the subscription with + `DependencyDropped`. The policy hook offers `refollow`: re-resolve the + object by name, resume from the checkpoint if the new object's schema + fingerprint and retained history allow it, otherwise fall through to the + re-snapshot policy. Without this, every view deploy is a paging incident + for whoever runs the sink. +- **Fencing and high availability.** The token carries an epoch. Checkpoint + stores implement compare-and-set on `(epoch, frontier)`, refusing writes + from a stale epoch and refusing frontier regression. In v1 this makes the + two-instances mistake loud instead of silently corrupting. The production + HA story builds on the same primitive: active-passive failover where + standby instances contend for a lease in the checkpoint store, the lease + holder bumps the epoch on acquisition, and the fencing rule guarantees at + most one writer even across partitions. This is the Kafka Connect and + Debezium recovery model, no consensus service required beyond the + checkpoint store's compare-and-set. If real deployments show the + client-side lease is not enough, the Materialize-side workstream leaves + room for a server-assisted primitive (for example, named subscriptions + with server-enforced single ownership), designed properly rather than + worked around. +- **Liveness watchdog.** Optional max-staleness on progress. If the frontier + stops advancing (hung connection, wedged cluster) the SDK reconnects or + surfaces, with correct units, unlike the demo whose watchdog compared + seconds to milliseconds and could never fire. +- **Graceful shutdown.** On signal: finish the in-flight batch, checkpoint, + close the cursor, disconnect. + +### Layer 3: the sink kit and built-in sinks + +Layer 3 answers "I want this view synced into X" with a small interface per +destination family and the hard problems solved once. + +Positioning: the SDK core (Layers 1 through 3's interfaces) is the fully +supported product surface with a day-one stability commitment. Built-in +sinks are supported reference implementations of those interfaces, Redis +first. The framework is the product, destinations are instances of it, and +users building their own sinks are as much the audience as users running +ours. + +**Keyed-state sinks** (Redis, DynamoDB, Postgres tables, search indexes). +Consume `ENVELOPE UPSERT`. The contract: + +```python +class KeyedStateSink(Protocol): + def load_token(self) -> ResumeToken | None: ... + def apply(self, batch: UpsertBatch, token: ResumeToken) -> None: + """Apply updates and persist token atomically. Must be idempotent.""" + def resync_begin(self, generation: int) -> None: ... + def resync_end(self, generation: int) -> None: + """Make generation current and sweep keys from older generations.""" +``` + +The kit drives the state machine: initial snapshot, incremental batches, +resume, and re-snapshot after compaction-horizon or state loss. The +`resync_*` hooks solve the two problems every demo either hit or documented +away: + +- *Atomic snapshot visibility.* Large snapshots arrive as partial batches. A + sink can stage them under a generation marker and flip visibility at + `resync_end`, or accept eventually-visible snapshots. The kit supports + both, the sink declares which. +- *Orphan reconciliation.* Re-snapshotting after state loss only upserts + currently-live keys. Keys deleted while offline linger forever unless swept + (mz-redis-sync's README admits exactly this gap). Generation-tagged sweep + at `resync_end` closes it. + +**Event sinks** (webhooks, queues, notification services). Consume the diff +envelope through Layer 2's `at_least_once` mode with idempotency keys and +retraction policy. The kit ships a generic webhook sink as the reference for +this family. + +**The Redis reference sink.** First concrete sink, the successor to +mz-redis-sync: + +- Data models: string (`SET key value`), hash (row columns as fields), and + JSON value encoding. Multi-column keys via a documented key template. + NULLs handled per a documented rule instead of crashing the driver. +- Writes: one `MULTI`/`EXEC` per closed batch containing the data commands + plus the checkpoint `SET`. This is the atomic + data-plus-progress-in-one-transaction rule that the durable-subscriptions + pattern doc prescribes, applied with the destination as the transaction + boundary (mz-redis-sync instantiates the same idea). Large batches chunk + under a generation marker with a sweep, trading atomic visibility for + bounded transactions, per the sink's declared mode. +- Checkpoint key namespaced under the sink's prefix (the demo's frontier key + bypassed its own prefixing and could collide with data). +- Fencing via a Lua compare-and-set on `(epoch, frontier)`. +- `key_violation` events surface through a policy hook (log-and-skip or + fail) instead of crashing into a poison-restart loop. + +**Standalone runner as a reference example.** The repo ships a small +config-file-driven runner built on the Rust implementation (working name +`mz-sink`), positioned as an example application, not a core deliverable: +it lives in `examples/`, demonstrates end-to-end sink deployment for +non-Rust users, replaces mz-redis-sync as the thing we point demos at, and +serves as the long-running soak target for the chaos suite. The SDK is the +product people build with however they please, the runner shows one good +way. + +### Language strategy: native implementations, one spec + +Three native implementations, not a shared Rust core with bindings: + +- TypeScript on `pg` (node-postgres), Python on `psycopg` (v3, async-capable, + sync facade), Rust on `tokio-postgres`. +- The hard, subtle part of this project is a small protocol state machine. + It is precisely the kind of logic a written spec plus shared test vectors + can pin down across implementations. +- The expensive part of a shared-core approach is everything else: TLS, + auth, connection lifecycle, event-loop integration in Node, asyncio + integration in Python, packaging native modules for every platform. The + ecosystem drivers already solve all of it, natively and idiomatically, and + users already trust them. + +The spec (working name **Subscribe Consumption Protocol**, versioned, +`spec/` in the SDK repo) defines: batching and progress semantics, resume +token contents and boundary arithmetic, guarantee modes and their crash +matrices, error taxonomy and mapping rules, envelope decoding, and the type +mapping. Each language's README links its conformance report. + +This is the proven model for exactly this kind of SDK. MongoDB drivers are +the canonical example: the major drivers are per-language native +implementations, unified by the public `mongodb/specifications` repo of +prose specs plus JSON test files, and MongoDB change streams use an opaque +resume token with the same role as ours. Kafka demonstrates both paths at +once: the Java client is the spec-bearing reference, native implementations +displaced the C-binding clients where binding pain was highest (kafka-go +over cgo bindings, KafkaJS over node-rdkafka), while librdkafka bindings +persist where they work well enough. The lesson is that bindings are a +per-ecosystem cost gamble, native is uniformly safe. Debezium shows the +cost of skipping multi-language entirely, its embedded engine is JVM-only, +which is a large part of why non-JVM teams never adopted it directly. + +Delivery order: **Rust and Python first.** Rust is the reference +implementation, written against the spec as the spec is written, and it is +the language of the team that must vouch for the semantics. Python is the +fastest path to validating the UX with real integration builders and +carries the prototype. **TypeScript follows** once the spec has survived two +implementations, **Go later**. The spec makes each additional language a +mechanical, conformance-checked project. + +Suggested packaging (open question below for final naming): + +| Language | Core | Redis sink | +| --- | --- | --- | +| TypeScript | `@materializeinc/subscribe` | `@materializeinc/sink-redis` | +| Python | `materialize-subscribe` | extra: `materialize-subscribe[redis]` | +| Rust | `mz-subscribe` (crates.io) | feature: `mz-subscribe/redis`, binary `mz-sink` | + +One monorepo (`MaterializeInc/subscribe-sdk` or Labs equivalent) holding +`spec/`, `conformance/`, and the three implementations, so a spec change and +its cross-language fallout land in one PR. + +### Testing + +The demos shipped zero tests between them. The SDK inverts this, and the test +infrastructure is a deliverable users get too: + +1. **Conformance vectors.** Language-agnostic recorded subscribe streams + (JSON) covering: progress interleavings, multiplicities beyond one, + upsert/delete/key_violation, partial snapshots, idle progress, + every mapped type, and error frames. Every implementation must replay + them to identical decoded events and tokens. +2. **Emulator integration suite.** Docker-based (testcontainers) scenarios + against the Materialize emulator: snapshot then stream, resume across + client restart, resume across emulator restart, `RETAIN HISTORY` expiry + producing `CompactionHorizon`, dropped view producing `DependencyDropped`, + poisoned dataflow producing `StreamPoisoned`. +3. **Chaos suite (the flagship).** Kill the sink process with SIGKILL at + randomized points (mid-batch, between apply and checkpoint, mid-snapshot), + restart, repeat, then assert the destination equals + `SELECT ... AS OF ` at the checkpointed frontier. Run for both + guarantee modes, asserting convergence for `transactional` and + convergence-with-duplicates-absorbed for `at_least_once`. +4. **Fencing test.** Two instances against one checkpoint, assert the stale + epoch is refused and the destination stays consistent. +5. **Property tests.** Random diff streams through envelope decoding and + batching, asserting consolidation and frontier invariants. +6. **User-facing test kit.** The fixture player from (1) is exported so sink + authors can unit-test their `apply` implementations offline. + +### Observability + +Built-in, consistent across languages: frontier lag (wall clock minus +frontier, the one metric every operator wants and no demo had), batches and +updates applied, reconnect count, checkpoint age, buffer occupancy. Exposed +as callbacks/hooks in the libraries and as Prometheus metrics plus health +endpoint in `mz-sink`. OpenTelemetry spans for connect, snapshot, batch +apply, and checkpoint, so a sink shows up in the same traces as the +application it feeds. + +### Documentation integration + +The per-language client pages in `doc/user/content/integrations/` currently +teach the naive loop. Once the SDK exists, each page leads with the SDK and +keeps the raw `DECLARE`/`FETCH` version as an appendix for driver-only +environments. The durable-subscriptions pattern doc becomes the conceptual +explanation behind the SDK's design, linking to it as the implementation. + +### Materialize-side workstream (planned with the SDK, done properly) + +The SDK must not paper over server gaps with client-side workarounds. Where +the correct behavior needs the database's help, the database change is part +of this program's plan, sequenced so the SDK's stable release builds on real +primitives. Each item below gets its own issue and, where non-trivial, its +own design doc. In sequence: + +1. **Structured error codes** (dedicated SQLSTATEs) for the + compaction-horizon error, dependency-dropped, and dataflow errors. Today + the compaction-horizon failure surfaces as generic `DATA_EXCEPTION` + shared with other errors, and its message text has already changed once + (#34712), which silently broke the one existing client that dispatched + on it. Small, high leverage, and a hard prerequisite for SDK GA: the + typed error taxonomy must key on codes, never on message text. +2. **Docs**: cross-link the durable-subscriptions pattern from every + client-library page. Can land immediately, independent of everything + else. +3. **Subscribe output backpressure** in the coordinator (the existing + unbounded-channel TODO in `active_compute_sink.rs`). Protects + `environmentd` from slow consumers regardless of which client they use. + The SDK's prompt-drain design reduces exposure but only the server can + bound it. +4. **Progress cadence control**, so consumers can trade update granularity + for faster checkpoint advancement and cheaper idle streams. +5. **Non-poisoning subscribe errors** (database-issues#5182), so a transient + dataflow error does not permanently wedge a subscription. +6. **WebSocket `SUBSCRIBE` stabilization**, the gate for the browser SDK. +7. **Server-assisted subscription ownership** (named subscriptions with + single-writer enforcement), contingent on evidence from HA deployments + that the client-side lease is insufficient. + +Items 1 and 2 are cheap and land before or with SDK GA. Items 3 through 5 +improve every subscribe consumer and proceed on their own track with the SDK +as the motivating consumer. Items 6 and 7 are demand-gated. + +## Minimal Viable Prototype + +Prototype = **Python Layer 1 + Layer 2 + the Redis sink, chaos-tested against +the emulator**, living in the SDK monorepo from day one: + +1. Layer 1 client on psycopg with batching, envelopes, typed errors, + reconnect-and-resume. +2. Checkpoints in Redis, `transactional` mode only. +3. The Redis sink at functional parity with mz-redis-sync plus the known + fixes: correct resume statement, `key_violation` handling, orphan sweep on + resync, namespaced checkpoint key, NULL handling. +4. The kill-at-random-points chaos test asserting convergence against + `SELECT ... AS OF`. + +This validates the three riskiest claims early: that the batch/token model is +pleasant to use, that exactly-once state application holds under crash +testing, and that the emulator is a sufficient CI target. The Rust reference +implementation starts from the same spec as soon as the prototype stabilizes +the batch and token model, and the first conformance vectors are extracted +from the prototype's test suite. The Novu integration is then rebuilt on +Layer 2's `at_least_once` mode as the second validation, exercising +idempotency keys and retraction policy with a real side-effecting +destination. + +## Delivery plan + +Phased, each phase with an exit criterion, and no phase starts before the +previous one's criterion passes. + +**Phase 0, foundations.** The SDK monorepo skeleton (`spec/`, +`conformance/`, `python/`, `rust/`, `examples/`), CI wiring, and the +emulator test harness. Spec v0 drafted from this design. In parallel, the +two cheap Materialize-side items land: the structured error code change and +the docs cross-links. + +**Phase 1, the Python MVP.** As described under Minimal Viable Prototype, +Layers 1 and 2 plus the Redis sink. Exit: the chaos suite passes (the +destination equals `SELECT ... AS OF ` under arbitrary kill +schedules), the fencing test passes, and every typed error is reproducible +in integration tests. + +**Phase 2, the spec becomes real.** Conformance vectors extracted from +Phase 1's tests, then the Rust reference implementation built against spec +and vectors. Exit: both implementations pass identical vectors and emulator +scenarios, and spec v1 is published. The Novu-style at-least-once rebuild +validates the event-sink surface in this phase. + +**Phase 3, breadth.** TypeScript, the generic webhook sink (pending open +question 3), the `mz-sink` example runner, and the docs-site integration +that replaces the naive per-language loops. + +**GA gate.** A stable release requires: Phase 2 complete, the +structured-error-code change released in Materialize (with the +version-gated fallback for older servers), and the public docs rewritten on +the SDK. HA failover (the lease design, open question 6) ships in the first +post-GA minor release. + +## Future work + +Deliberately excluded from v1, recorded so reviewers can see the growth path: + +- **Fan-out helper.** One subscription demultiplexed to many in-process + consumers by key (the websocket-server pattern, thousands of clients fed + from one `SUBSCRIBE`). Layer 1's batch model supports it, a first-class + helper makes it a five-line feature. +- **Per-key coalescing and debounce.** Event sinks often want "at most one + webhook per key per interval" to absorb flapping. A Layer 2 option with a + max-delay bound, at the cost of intermediate updates, which the diff model + makes safe to drop. +- **Browser SDK** on the WebSocket endpoint once it stabilizes, implementing + the same spec. +- **Additional languages** (Go first, given existing docs coverage) and + additional built-in sinks chosen by demand, with adoption data feeding the + case for native server-side sinks for the top destinations. +- **Serverless guidance.** Long-lived subscriptions do not fit + function-per-request platforms. The pattern doc for bridging (a small + always-on consumer feeding a queue) is docs work once the SDK exists. + +## Alternatives + +**A shared Rust core with native bindings (napi-rs, PyO3).** Single +implementation of the state machine, mechanical bindings. Rejected because +the state machine is the small part. The core would own connection +management, TLS, and auth in three runtimes, integrate with two foreign +async runtimes, and complicate packaging and debugging for every user. The +per-language ecosystem drivers are more battle-tested than anything we would +ship. A spec plus conformance vectors pins down cross-language behavior at a +fraction of the cost, which is how the database-driver ecosystem itself +works. + +**Native server-side sinks (`CREATE SINK ... INTO REDIS`).** The +strongest alternative. It is the best eventual UX for the specific, +high-volume destinations, and this SDK does not preclude it. Rejected as the +*first* move because: each destination becomes a server feature with a +release train, storage/compute team ownership, and years of long-tail +destination requests (the Kafka Connect catalog is hundreds of connectors +deep). The SDK serves the long tail by construction, ships without touching +`environmentd`, and its adoption data tells us which destinations deserve +native sinks. The consistency model (frontier-atomic commits) is the same +one a native sink would implement, so the concepts transfer. + +**Sink to Kafka, use the Kafka Connect ecosystem.** Works today for users +who run Kafka. Rejected as the answer because it taxes every user with a +Kafka deployment plus Connect operational burden to reach a cache, adds a +hop of latency, and loses the direct mapping between destination state and a +Materialize timestamp unless the connector is consistency-aware. It remains +the right answer for organizations already deep in Connect, and the docs +should keep saying so. + +**Docs only, no code.** The durable-subscriptions pattern doc already +describes the protocol. The evidence section above shows what documentation +alone achieves, three implementations, three different correctness bugs, by +authors closer to Materialize than any customer will be. + +**Harden mz-redis-sync as a one-off product.** Fixes one destination in one +language, leaves every other consumer where they are, and keeps the protocol +logic welded to Redis specifics. The layered SDK subsumes it, and `mz-sink` +delivers the same operational artifact. + +## Decisions taken so far + +Settled during initial review, recorded so the remaining questions are +crisp: + +- **Official from day one.** This ships under `MaterializeInc` with a + stability commitment, not as a Labs experiment with a graduation path. +- **All three layers in scope.** The SDK core is the fully supported main + focus. Built-in sinks (Redis first) are supported reference + implementations of the framework. +- **Language order: Rust and Python first, TypeScript next, Go later.** +- **HA is planned, not deferred.** Fencing in v1, lease-based active-passive + failover shortly after. Production use is the assumption. +- **No shortcut workarounds for server gaps.** Where correct behavior needs + the database's help, the server change is planned as part of this program + (see the Materialize-side workstream), sequenced ahead of the SDK + behavior that depends on it. +- **The standalone runner is a reference example**, not a core deliverable. + +## Open questions + +1. **Final package naming.** `subscribe` as the noun + (`@materializeinc/subscribe`, `materialize-subscribe`, `mz-subscribe`) + versus a broader name that leaves room for the SDK to grow into general + client duties later. Day-one stability makes renaming expensive, so this + needs deciding before the first publish. +2. **Guarantee vocabulary.** Are we comfortable publicly branding + `transactional` mode as "exactly-once state application"? It is accurate + for destinations where checkpoint and data commit atomically, but the + term invites misreading as exactly-once side effects. Candidate framing: + "consistent sinks" versus "delivery sinks". +3. **v1 sink surface.** Redis plus the generic webhook sink, or Redis only? + Webhook is cheap on Layer 2 and exercises the second destination family + early, which argues for including it. +4. **Checkpoint store default for event sinks.** Destination-embedded + checkpoints are the opinionated default for keyed-state sinks, but no + destination-embedded store exists for a webhook. Materialize-table store + as the event-sink default? +5. **Conformance gate.** Block releases of any language on the full emulator + chaos suite, or vectors only for patch releases? +6. **HA lease parameters.** Lease duration, heartbeat cadence, and whether + the lease lives only in the checkpoint store or also surfaces in + `mz_internal.mz_subscriptions` for observability. Needs a short design of + its own before the failover release. From 0137bb17d071f1c8b9ca54ad4c15589b81250f9b Mon Sep 17 00:00:00 2001 From: bobbyiliev Date: Wed, 8 Jul 2026 13:11:33 +0300 Subject: [PATCH 02/19] design: Re-center Subscribe SDK on single-view core with cohorts as a first-class generalization --- .../design/20260706_subscribe_sdk.md | 132 +++++++++++++++--- 1 file changed, 113 insertions(+), 19 deletions(-) diff --git a/doc/developer/design/20260706_subscribe_sdk.md b/doc/developer/design/20260706_subscribe_sdk.md index 2b7a1a5871886..225967cdf37f5 100644 --- a/doc/developer/design/20260706_subscribe_sdk.md +++ b/doc/developer/design/20260706_subscribe_sdk.md @@ -188,9 +188,12 @@ The SDK's UX comes from five opinions applied uniformly across languages: The stream yields consistent batches. Each batch contains every update for an interval of timestamps that a progress message has proven complete, plus the frontier and a resume token. Users physically cannot observe a - half-delivered timestamp. (An advanced `raw()` mode exposes per-row - delivery for power users, clearly documented as forfeiting the batch - guarantees.) + half-delivered timestamp. A batch is a consistent moment of one view, and + the same engine generalizes to a *cohort* of views advancing to a joint + consistent moment (see "Consistency layering and cohorts"). The advanced + `raw()` mode that exposes the decoded per-row stream (forfeiting the batch + guarantees) is that same underlying layer, and is the substrate cohorts + build on. 2. **Resume tokens are opaque.** A token encapsulates the frontier, the `SNAPSHOT false AS OF frontier - 1` arithmetic, the query fingerprint, and a fencing epoch. Users store bytes and hand them back. There is no @@ -207,6 +210,43 @@ The SDK's UX comes from five opinions applied uniformly across languages: surfaces are idiomatic, behavior is identical, and identical is checked by machines. +### Consistency layering and cohorts + +The SDK is layered so that a single consistency engine serves both the +single-view and the multi-view case: + +1. **Decoded stream.** The transport's `FETCH` output decoded into timestamped + changes and progress markers (`Data{timestamp, change}` / + `Progress{frontier}`). This is the `raw()` layer: composable and public, and + on its own forfeiting the higher guarantees. +2. **Consistency engine.** Buffers changes and releases everything strictly + below a release frontier, consolidating per timestamp. It takes the release + frontier as an *input*: a single-view subscription passes its own frontier; + a cohort passes `min(frontier)` across its members. The single-view batcher + is literally the one-member case of the cohort engine, so there is exactly + one buffer-and-release implementation to get right and to test. +3. **Batch / moment.** The closed `ConsistentBatch` (one view) or cohort + moment (N views) handed to the consumer, with its resume token. + +**Cohorts** turn N independent subscriptions into one stream of jointly +consistent moments. Because every object in a Materialize environment shares +one logical timeline, progress timestamps from independent subscriptions are +directly comparable, so `min(frontier)` across members is a valid global +consistent cut. The SDK withholds a moment until every member has closed it, +so the consumer only ever sees jointly consistent state and never has to reason +about the cut itself. This is a genuine Materialize differentiator (Frank +McSherry's `mz-bridge-recipe` is the prior art) and it composes on the layering +above rather than being a separate code path. + +The cohort is a shipped, first-class helper, not the core. The common +single-view case pays none of its concept weight: `consume([oneView])` is +identical in feel to a single subscription, and a single view is just the +degenerate cohort of one. A cohort of N views is N subscriptions and therefore +N connections (a long-running `SUBSCRIBE` cursor holds its connection in a +`FETCH` loop, so several cannot be multiplexed onto one). That cost is linear +in views and worth stating. Dynamic cohort membership (add/drop/merge/split of +a live cohort) is deferred to future work. + ### Layer 1: the subscribe client Layer 1 wraps the ecosystem-native PostgreSQL driver and runs the @@ -269,10 +309,14 @@ Responsibilities and design points: resume token, so downstream checkpoints keep advancing during quiet hours. This directly fixes the failure mode where an idle subscription's checkpoint ages out of the retention window. -- **Envelope decoding.** Three modes map to typed events: +- **Envelope decoding.** Two modes map to typed events: - default diff envelope: `Insert{row, diff}` / `Retract{row, diff}` with - multiplicities preserved (a `diff` of -3 is three retractions and the - type says so), + multiplicities preserved (a `diff` of -3 is three retractions and the type + says so). Within a closed batch, diffs are consolidated per row (net + multiplicity, net-zero rows dropped) so a batch is a clean net delta at the + frontier, not a replay of intra-window churn. Consolidation and the sink's + upsert target key on the same content-derived row identity, so they agree + by construction (the `mz-bridge-recipe` technique). - `ENVELOPE UPSERT`: `Upsert{key, value}` / `Delete{key}` / `KeyViolation{key}`. Key violations are a first-class event, not an exception, because a crash-restart loop cannot fix them. @@ -285,13 +329,17 @@ Responsibilities and design points: subscription with a "total result exceeds max size" error regardless of client behavior. The SDK surfaces that as a typed error with sizing remediation rather than as an opaque stream failure. -- **Fetch pacing.** `FETCH c WITH (timeout ...)` with adaptive sizing. - The client always drains the server promptly and applies backpressure to - the application from its own bounded buffer, because unread results buffer - without limit in `environmentd`. If the application cannot keep up, the - documented options are a larger bound, spill-to-disk (opt-in), or letting - the buffer block the FETCH loop and accepting server-side growth. The SDK - makes the trade-off visible instead of implicit. +- **Never backpressure Materialize.** Unread `SUBSCRIBE` output buffers without + limit in `environmentd` (an acknowledged server TODO), so an unconsumed + subscription makes *Materialize* grow, not the client. The SDK therefore + decouples a continuous drain loop (`FETCH c WITH (timeout ...)`, adaptive + sizing, always draining the server promptly) from delivery to the consumer, + with a buffer in between. That buffer is bounded and fails loud when full: + the client is the thing that must fall over, never the database. If the + consumer cannot keep up the documented options are a larger bound or opt-in + spill-to-disk. The SDK never silently lets the server grow, and it never + couples draining to the consumer's pace (the mistake in the initial + scaffolding, where `next()` only fetched on demand). - **Typed errors.** All failures map to a small taxonomy: | Error | Meaning | Default behavior | @@ -378,15 +426,22 @@ Design points: - `at_least_once`: effects run first, the checkpoint commits after the handler returns. Result: replays possible after a crash, and `ctx.idempotencyKey(event)` provides a stable key derived from - subscription name, `mz_timestamp`, key columns, and an ordinal within - the batch. This fixes both observed idempotency failures: content-only - hashes that false-dedup distinct events, and no key at all. + subscription name, `mz_timestamp`, key columns, and an ordinal within the + *timestamp*. It must not depend on batch grouping: batch boundaries are + not stable across a resume (the same change can land in a + differently-grouped moment when frontiers tick at different wall-clock + instants, and only net state converges), whereas the set of updates at a + given `mz_timestamp` is deterministic. Keying on a within-batch ordinal + would silently stop deduplicating replays. This fixes both observed + idempotency failures: content-only hashes that false-dedup distinct + events, and no key at all. - **Retraction policy for event sinks.** Side-effecting consumers declare what a retraction means: `ignore`, or `compensate(fn)` (the Novu demo's retraction-to-revoke mapping, promoted to a supported concept). Compensation needs a correlation identity that is stable between an insert and its later retraction, which the delivery idempotency key - deliberately is not (it includes the timestamp and ordinal). The SDK + deliberately is not (it includes the timestamp and a within-timestamp + ordinal). The SDK therefore provides `ctx.correlationKey(event)`, derived from the row's key columns only, for exactly this purpose. The Novu demo collapsed both identities into one content hash, which made revokes work but false-dedups @@ -582,7 +637,9 @@ infrastructure is a deliverable users get too: 4. **Fencing test.** Two instances against one checkpoint, assert the stale epoch is refused and the destination stays consistent. 5. **Property tests.** Random diff streams through envelope decoding and - batching, asserting consolidation and frontier invariants. + batching, asserting consolidation and frontier invariants, including the + cohort case: `min(frontier)` release, and a laggard member holding the joint + moment until it catches up. 6. **User-facing test kit.** The fixture player from (1) is exported so sink authors can unit-test their `apply` implementations offline. @@ -604,6 +661,22 @@ keeps the raw `DECLARE`/`FETCH` version as an appendix for driver-only environments. The durable-subscriptions pattern doc becomes the conceptual explanation behind the SDK's design, linking to it as the implementation. +### Demonstrators + +Two example apps ship in `examples/`, because the strongest case for the SDK +is shown, not told: + +- **Always-fresh cache** (single view to Redis), modeled on Justin Bradley's + `mz-sink` demo: one expensive query, three read backends (recompute, + Materialize index, sink-backed cache) switchable live so the win is felt. + The messaging leads with correctness (a cache kept fresh with no invalidation + logic because it is downstream of the change stream), with latency as the + hook. +- **Consistent cohort dashboard** (N views to one consistent view): a dashboard + over several views (e.g. orders, inventory, pricing) where a naive per-view + cache visibly tears and the cohort never does. This is the demo that sells + the cross-view consistency a single-view cache demo structurally cannot show. + ### Materialize-side workstream (planned with the SDK, done properly) The SDK must not paper over server gaps with client-side workarounds. Where @@ -704,7 +777,14 @@ Deliberately excluded from v1, recorded so reviewers can see the growth path: - **Fan-out helper.** One subscription demultiplexed to many in-process consumers by key (the websocket-server pattern, thousands of clients fed from one `SUBSCRIBE`). Layer 1's batch model supports it, a first-class - helper makes it a five-line feature. + helper makes it a five-line feature. This is the inverse of a cohort + (many-into-one). Cohorts themselves ship in v1 (see "Consistency layering + and cohorts"). +- **Dynamic cohort membership.** Adding, dropping, merging, and splitting the + members of a *live* cohort without re-subscribing, plus the durable + bookkeeping of the resume position a merged or split cohort resumes from. + The fixed-cohort case ships in v1. Live topology changes are the hard, + unvalidated part of `mz-bridge-recipe` and are deferred. - **Per-key coalescing and debounce.** Event sinks often want "at most one webhook per key per interval" to absorb flapping. A Layer 2 option with a max-delay bound, at the cost of intermediate updates, which the diff model @@ -777,6 +857,20 @@ crisp: (see the Materialize-side workstream), sequenced ahead of the SDK behavior that depends on it. - **The standalone runner is a reference example**, not a core deliverable. +- **Single-view core, cohort as a first-class generalization.** The consistent + batch of one view is the core primitive. Multi-view consistency ships as a + cohort helper on the same engine (`min(frontier)`), not as the mandatory + mental model. Considered and rejected: re-centering the whole SDK on cohorts + (the demand across all prior art, the server's native unit, and Materialize's + own per-object product surface are all single-view, and the cohort stays + composable on the layering so it is not precluded). +- **One consistency engine, and never backpressure Materialize.** A single + buffer-and-release engine (single-view = cohort-of-one), fed by a continuous + drain into a bounded buffer that fails loud, so the client falls over, never + the server. +- **Idempotency keys never depend on batch grouping.** Batch boundaries are not + stable across a resume, so the delivery key is derived from `mz_timestamp` + plus within-timestamp identity, not a within-batch ordinal. ## Open questions From 2101ea15a33e9ff21fa9586d2ec76a2b61053853 Mon Sep 17 00:00:00 2001 From: bobbyiliev Date: Wed, 8 Jul 2026 16:22:37 +0300 Subject: [PATCH 03/19] design: Record cohort consistency/availability boundaries and deferred hardening --- .../design/20260706_subscribe_sdk.md | 38 +++++++++++++++++++ 1 file changed, 38 insertions(+) diff --git a/doc/developer/design/20260706_subscribe_sdk.md b/doc/developer/design/20260706_subscribe_sdk.md index 225967cdf37f5..dfeba7207ab3d 100644 --- a/doc/developer/design/20260706_subscribe_sdk.md +++ b/doc/developer/design/20260706_subscribe_sdk.md @@ -247,6 +247,29 @@ N connections (a long-running `SUBSCRIBE` cursor holds its connection in a in views and worth stating. Dynamic cohort membership (add/drop/merge/split of a live cohort) is deferred to future work. +The cohort is the SDK's strongest-consistency construct, and by the same token +its least available. It makes progress only as fast as its slowest member, and +a jointly consistent moment exists only once every member has closed it. The +single-view path and the raw stream are more available precisely because they +answer for one view alone. Three consequences are worth stating for v1, and the +unresolved policy for each is an open question below. + +- **One timeline.** `min(frontier)` is a valid cut only when every member reads + the same logical timeline, so their timestamps are comparable. That holds for + objects on the default timeline. A member on a user-defined timeline, or a + source carrying external timestamps, would make the comparison meaningless. + v1 documents this as a constraint rather than checking it. +- **A laggard is bounded, not silent.** Holding a leading member's changes until + the slowest catches up is inherent to the guarantee, so a stalled member would + otherwise buffer its peers without bound. v1 caps the total buffered across + the cohort with a lag budget and fails loud when it is exceeded, rather than + growing memory quietly. +- **Failure is all-or-nothing in v1.** Any one member erroring tears the whole + cohort down, and recovery re-subscribes every member at the joint frontier. + This is the conservative, consistent choice. Reconnecting only the failed + member without disturbing the others is an availability improvement tracked + with the HA workstream. + ### Layer 1: the subscribe client Layer 1 wraps the ecosystem-native PostgreSQL driver and runs the @@ -871,6 +894,12 @@ crisp: - **Idempotency keys never depend on batch grouping.** Batch boundaries are not stable across a resume, so the delivery key is derived from `mz_timestamp` plus within-timestamp identity, not a within-batch ordinal. +- **Cohort ships demonstrated, not hardened.** v1 proves the one engine + generalizes to multi-view consistency and bounds the laggard case with a + fail-loud lag budget. Availability hardening (per-member reconnect, timeline + validation, a configurable budget) is deliberately deferred until real + workloads show which of it matters, rather than designed against a cohort no + one runs yet. The core single-view path is where v1 invests. ## Open questions @@ -897,3 +926,12 @@ crisp: the lease lives only in the checkpoint store or also surfaces in `mz_internal.mz_subscriptions` for observability. Needs a short design of its own before the failover release. +7. **Cohort failure and lag policy.** On a member error, tear the whole cohort + down and resume every member at the joint frontier (v1 behavior), or + reconnect only the failed member? And when the lag budget is hit, fail loud + (v1), drop the laggard from the cut, or block the fast members? Both interact + with the HA workstream and want deciding before cohorts are branded stable. +8. **Cohort timeline validation.** Should the SDK detect and reject a cohort + whose members do not share a comparable timeline, or is documenting the + constraint enough for v1? Rejecting needs a way to read an object's timeline, + so it is a small server-side dependency if we want it. From b0d52d1cc24232414d4fa4c97661fdf1469caf72 Mon Sep 17 00:00:00 2001 From: bobbyiliev Date: Wed, 7 Oct 2026 14:06:22 +0300 Subject: [PATCH 04/19] design: Materialize SDK for SUBSCRIBE consumption and custom sinks --- .../design/20260706_materialize_sdk.md | 750 ++++++++++++++ .../design/20260706_subscribe_sdk.md | 937 ------------------ 2 files changed, 750 insertions(+), 937 deletions(-) create mode 100644 doc/developer/design/20260706_materialize_sdk.md delete mode 100644 doc/developer/design/20260706_subscribe_sdk.md diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md new file mode 100644 index 0000000000000..99708700b76ad --- /dev/null +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -0,0 +1,750 @@ +# Materialize SDK: correct `SUBSCRIBE` consumption and a framework for custom sinks + +- Associated: + - Product requirements: [Sink SDK PRD](https://app.notion.com/p/materialize/Sink-SDK-3dd13f48d37b8024985ce70b1d96e9cd) (PRD-85). Overlapping proposal: PRD-87 (`mzq`). + - Tracking: Linear project "Subscribe SDK" (DEX). + - [#38468](https://github.com/MaterializeInc/materialize/pull/38468): durable subscribe design, the server primitive this SDK moves onto later. + - [#37483](https://github.com/MaterializeInc/materialize/pull/37483), [#37517](https://github.com/MaterializeInc/materialize/pull/37517): prototype Rust and Python packages. + - [#37905](https://github.com/MaterializeInc/materialize/pull/37905): server-side bound on subscribe output buffering. + - SQL-528: `SUBSCRIBE ... WITH (PROGRESS) ... UP TO` emits no final progress row. + - [database-issues#5182](https://github.com/MaterializeInc/database-issues/issues/5182): dataflow errors poison a subscribe. + - [Durable subscriptions pattern](https://materialize.com/docs/transform-data/patterns/durable-subscriptions/): the manual protocol this SDK packages. + - Prior art: [mz-redis-sync](https://github.com/MaterializeIncLabs/mz-redis-sync), [novu-materialize-integration](https://github.com/MaterializeIncLabs/novu-materialize-integration), [mz-turbopuffer-sink](https://github.com/MaterializeInc/mz-turbopuffer-sink). + +## The Problem + +Customers want Materialize results in systems we do not sink to natively: +caches, search indexes, their own Postgres, notification services, live UIs, +and agents. `SUBSCRIBE` is the primitive for all of them, and every team that +uses it re-derives the same protocol by hand. The use cases behind the product +requirements are representative: + +- A cache that keeps a Valkey or Redis keyspace equal to a materialized view. + It needs the initial snapshot and a fast, gap-free restart. +- A search index that keeps two turbopuffer namespaces equal to two views, at + one consistent timestamp, exactly once. Today this runs through Kafka. +- A relational target that writes several views into an application's + Postgres, transactionally across tables. + +The protocol for consuming `SUBSCRIBE` durably is: + +1. Subscribe `WITH (PROGRESS, SNAPSHOT = true)`. +2. Buffer updates until a progress message proves a timestamp is closed. +3. Apply the closed batch and persist its frontier atomically. +4. On restart, resume `WITH (PROGRESS, SNAPSHOT = false) AS OF frontier - 1`. +5. Keep history readable at the stored frontier, and fail loudly when it is not. + +Each step has a failure mode that is invisible in testing. Every existing +implementation we could find gets a different one wrong: + +| Codebase | Commit discipline | Resume | Idempotency | Retractions | +| --- | --- | --- | --- | --- | +| mz-redis-sync | Correct: data and frontier commit together at progress rows | Omits `SNAPSHOT` on resume, so every restart replays the full snapshot | Inherent (keyed writes) | Delegated to `ENVELOPE UPSERT`, but `key_violation` crashes the process and poisons restarts | +| novu-materialize-integration | Persists the timestamp of the first row of a batch, so a crash mid-batch loses the rest of that timestamp | Wall-clock guess of whether the checkpoint is still retained | Content-hash keys, which false-dedup identical payloads across time | Maps a retraction to a revoke, with two bugs | +| A work-queue subscriber in an internal project | No `PROGRESS`, so batch boundaries are a heuristic | Re-snapshots on every restart | Relies on the target being idempotent | Ignored | + +None of them reconnects on failure, and none has tests. Our per-language client +docs teach a bare `DECLARE`/`FETCH` loop with no progress handling and no +resume. The durable-subscriptions pattern page is correct, but it is prose that +every user turns into code again. + +The rest of this section describes the server behavior a correct consumer has +to account for. Each item was checked against the code. + +### Resume arithmetic + +With `SNAPSHOT = false`, an `AS OF t` emits times strictly greater than `t` +(`src/compute/src/sink/subscribe.rs:115-122`). A consumer that holds everything +below frontier `F` resumes with `AS OF F - 1`. + +### Retention window + +`RETAIN HISTORY` sets `since = upper - window`. After an environment outage or +upgrade the upper catches up and the window slides with it, so a one-hour +window can expire during a one-hour outage. A `REFRESH EVERY` view's upper +jumps to the next refresh, so a sink that is down across a refresh can lose +history however short the downtime. Retention belongs to the view's owner, and +any `ALTER` changes every sink's margin without telling it. + +A connected subscribe holds its input readable up to what it has emitted. Only +a disconnect puts history at risk, and the consumer then resumes from its +committed frontier, which can trail what it received. + +### Indexed targets + +If the subscribed object has an index on the cluster the subscribe runs on, the +dataflow imports the index instead of the persist shard +(`src/adapter/src/optimize/dataflows.rs:326-343`, +`src/adapter/src/coord/indexes.rs:88-93`), and `AS OF` is checked against the +index's `since` (`src/adapter/src/coord/timestamp_selection.rs:277-297`). An +index's default window is one second (`src/adapter-types/src/compaction.rs:20`), +and the view's `RETAIN HISTORY` does not carry over to its indexes. So +`AS OF F - 1` fails on an indexed view, however long the view's retention. An +index on a different cluster does not count. Index-level `RETAIN HISTORY` +requires `enable_index_options`, which is off by default. + +### Snapshot elision + +`SUBSCRIBE ` with `SNAPSHOT = false` never reads the snapshot +(`src/transform/src/dataflow.rs:464-522`, +`src/persist-client/src/operators/shard_source.rs:438-448`). Any query form, +including `SELECT * FROM mv` or a projection and filter, reads the whole +snapshot at the `AS OF` before emitting anything. A non-materialized view is +inlined, so its inputs' snapshots are read too. + +### Stream shape + +Diff output is ordered by (time, row bytes) and consolidated +(`src/compute/src/sink/subscribe.rs:212-217`). Upsert and Debezium envelopes +order by (time, key, row). Retractions do not come first. +`WITHIN TIMESTAMP ORDER BY` changes the order but requires +`enable_within_timestamp_order_by_in_subscribe`, off by default. + +Everything between two progress messages arrives together, so one batch spans +many timestamps. "At most one row per key" holds per timestamp, not per batch. + +`UP TO` ends without a final progress row (SQL-528), so data after the last +progress row is never closed by the server. + +### Buffering limits + +When a subscribe's backlog in `environmentd` exceeds +`subscribe_max_buffered_bytes` (default 128 MiB), the subscribe is retired with +`SubscribeFellBehind`, SQLSTATE 53200 (#37905). A single result larger than +`max_result_size` also ends the subscription. + +### Errors + +| Condition | SQLSTATE | Message | +| --- | --- | --- | +| `AS OF` below the readable frontier | `22000` (`DATA_EXCEPTION`, shared) | "could not find a valid timestamp for the query" | +| Subscribed object or cluster dropped | `42704` | "relation 'x' was dropped" and similar | +| Client fell behind | `53200` | `SubscribeFellBehind` | +| Dataflow error | varies | the evaluation error, repeated on every retry | + +The history-loss message has changed once already (#34712), which broke the +one client that matched its text. A dataflow error repeats until the data or +the view changes (database-issues#5182), so retrying cannot fix it. + +## Success Criteria + +"PRD" marks a requirement from the product requirements. "Design" marks one +this document adds. + +| # | Requirement | Source | +| --- | --- | --- | +| R1 | Exactly-once state in the target, across restarts and crashes | PRD | +| R2 | Back off and retry while the target is unavailable | PRD | +| R3 | Dead-letter a batch after too many retries, and reject single rows the target cannot take | PRD | +| R4 | One consistent timestamp across several views | PRD | +| R5 | Transactional writes when the target supports them | PRD | +| R6 | Every change carries its `mz_timestamp` | PRD | +| R7 | Live streaming with no target (UIs, caches, agents) | Design | +| R8 | Bounded client memory, and no server-side buffering on behalf of a slow client | Design | +| R9 | A small set of typed errors, with history loss stopping the consumer by default | Design | +| R10 | Several languages with identical behavior, checked by machines | Design | +| R11 | Sink authors can test without a live environment | Design | + +A user following the happy path cannot commit an unclosed timestamp, cannot +get the `AS OF` arithmetic wrong, and cannot silently lose or duplicate updates +across a restart. The check is a sink that converges to `SELECT ... AS OF` its +committed frontier after being killed at random points. + +## Out of Scope + +- Rewriting subscribe internals. The server changes this program depends on are + listed under "Materialize-side workstream". +- Exactly-once side effects. No client can make an HTTP call exactly once, so + event targets get at-least-once delivery with idempotency keys. +- A general Materialize client. Drivers already run queries well. See "Naming". +- Scale-out of one subscription across parallel workers, which has no server + support. +- Stateless workers (functions, lambdas). They need the server to own the + resume point, which durable subscriptions (#38468) provide. +- Private-preview subscribe features (`ENVELOPE DEBEZIUM`, + `WITHIN TIMESTAMP ORDER BY`), which are behind flags. + +## Solution Proposal + +The SDK will hold the subscribe protocol in one Rust library, the protocol core, +which performs no I/O and is compiled into a thin package per language. Each +package will use that language's own database driver for the connection and +will expose two modules: `subscribe` for live and durable consumption, and +`sink` for writing to targets. Every change will keep its `mz_timestamp`. +Durable consumption will store its checkpoint in the target, fenced by an epoch, +and will read one object plus an optional projection, filter, and envelope. The +first sink will be turbopuffer. The spec, conformance vectors, and an end-to-end +suite will live in this repository and run in the nightlies. + +### Naming + +This document proposes **Materialize SDK**, with `subscribe` and `sink` as its +first modules. This is the decision we most want reviewers to push on. + +"Sink SDK" names the most visible use and excludes live UIs, cache warmers, and +agents that watch a view with no sink at all (R7). "Subscribe SDK" names the +primitive accurately, but the package will need things that are not +`SUBSCRIBE`: checking a view's retention margin, choosing a cluster, attaching +to a durable subscription, and later the WebSocket transport. "Materialize SDK" +gives those a home without renaming the package after the first publish, which +is expensive once users pin it. The risk is that the name reads as a general +client. The scope rules that out: drivers stay the way to run queries, and the +SDK only adds modules for things drivers get wrong. + +Package names will be `materialize-sdk` on PyPI, `@materializeinc/sdk` on npm, +and `materialize-sdk` on crates.io. + +### Architecture + +``` + +---------------------------------------------------------------+ + | sink: targets (turbopuffer first), retry, dead-letter | + +---------------------------------------------------------------+ + | subscribe: durable consumption (checkpoints, fencing, | + | history-loss policy) and live streams | + +---------------------------------------------------------------+ + | transport, per language: the native driver | + | (psycopg, node-postgres, tokio-postgres), owns connections | + +---------------------------------------------------------------+ + | protocol core, Rust, no I/O: decode, release engine, | + | multi-view cut, tokens, statements, error classification | + +---------------------------------------------------------------+ +``` + +The protocol core is a library, not a service. It runs inside the user's +process, compiled into whichever package they install, and never opens a +connection. The package's driver fetches rows from Materialize and passes them +to the protocol core in a function call. The protocol core returns decoded +changes, closed batches, resume tokens, SQL text, and typed errors. "Language +and generation strategy" explains the choice. + +### Batches + +A batch will hold every change below its frontier that was not in an earlier +batch, and nothing at or above it. Apart from the partial snapshot chunks +described under "Live streams", users will not be able to observe a +half-delivered timestamp. + +Every change will keep its `mz_timestamp` (R6). Changes in a batch will be +ordered by timestamp and consolidated per timestamp, never across timestamps. +Consolidating across a batch drops the timestamps that targets need to stamp +rows, to order a delete before a re-insert of the same key, and to derive +idempotency keys that survive a resume. A helper will give the net change per +key across the batch for keyed targets that only want the final state. + +Progress will be an exclusive frontier: everything below `F` is committed. The +`- 1` will exist in exactly one function in the protocol core. + +Committing and fetching will be separate calls. `commit(frontier)` will run +inside the target's transaction and `next()` outside it. A combined call would +keep the target transaction open while the SDK waits for Materialize, which can +take seconds, or hours for a `REFRESH` view. + +```mermaid +sequenceDiagram + participant M as Materialize + participant D as Driver + participant C as Protocol core + participant S as Sink code + participant T as Target + D->>M: FETCH n c + M-->>D: data rows, progress rows + D->>C: rows, one call per FETCH + Note over C: holds rows until a progress row closes them + C-->>S: batch below frontier F, each change with its mz_timestamp + rect rgba(127, 127, 127, 0.12) + Note over S,T: one target transaction + S->>T: write changes + S->>T: commit(F), conditional on the epoch + end + S->>C: next() + Note over M,T: after a crash or restart + S->>T: load checkpoint + T-->>S: F and epoch + S->>C: resume(F) + C-->>D: SUBSCRIBE ... AS OF F - 1 + D->>M: DECLARE c CURSOR +``` + +### Subscription scope + +A durable subscription, one that checkpoints and resumes, will read one object +plus an optional projection, filter, and envelope. Arbitrary SQL will be +accepted only for live streams that never resume. + +Resuming a general query rehydrates its whole dataflow and needs history on +every input of the query. Durable subscriptions (#38468) accept exactly this +narrower surface, so starting narrow makes the later move a transport change, +not an API break. Widening later is compatible, and narrowing is not. +Structured input also lets the SDK build the statement and place `ENVELOPE` +before `WITH` without parsing user SQL. + +Projections and filters still read the snapshot on resume (see "Snapshot +elision"). The docs will recommend the plain object form for large views, and +the SDK will log the resume cost at startup. + +The SDK will refuse an indexed object at startup. It will check whether the +subscribed object has an index on the subscribing cluster and fail with an error +that names the remedy: a dedicated subscribe cluster with no index on the object. + +### Live streams + +A background loop will fetch with `FETCH c WITH (timeout = ...)` into a +bounded buffer, independent of the consumer's pace. If the consumer falls +behind, the client will fail with a typed error. Pausing the fetch loop instead +would push the backlog into `environmentd` until the server retires the +subscribe (R8). + +The initial snapshot is one timestamp and can exceed client memory. It will +arrive as chunks marked partial, with the token on the closing chunk. The sink +module's generations (see "Sink module") give atomic snapshot visibility to +targets that need it. + +A frontier advance with no data will yield an empty batch with a fresh token, +so checkpoints keep moving through quiet hours. + +`UP TO` will be supported. Until SQL-528 is fixed, the SDK will release +everything below `UP TO` when a bounded stream ends without error, so no closed +data is lost. + +### Checkpoints and fencing + +A checkpoint store will have `load(name)` and `commit(name, epoch, frontier)`. +The target is the recommended store, because writing data and checkpoint in +one transaction gives exactly-once state (R1, R5). A Materialize-table store and +a local file store will ship for targets with no transaction. + +Each worker start will take the next epoch. A commit will be conditional on the +stored epoch not being newer, and a stale worker will get a typed `Fenced` +error. Without the fence, a second instance of one sink would overwrite the +first one's progress. + +Every durable subscription will require an explicit name, which is its +checkpoint identity. A default derived from a class name would give two +deployments one checkpoint, and they would fence each other. Rows from several +views will be tagged by view name, because tagging by list position remaps rows +when the list is reordered. + +The checkpoint will also record a fingerprint of the subscription: the object, +projection, filter, envelope, and output column types. A resume whose +fingerprint differs fails with `SchemaMismatch` instead of mixing rows of two +shapes in one target. + +### Delivery guarantees + +The user will choose one of two guarantees. Under `transactional`, the target +commits the batch and the checkpoint together, so state is exactly once. Under +`at_least_once`, effects run first and the checkpoint commits after. The SDK +will provide an idempotency key built from the subscription name, +`mz_timestamp`, the key columns, and an ordinal within the timestamp. An ordinal +within the batch would break deduplication, because batch boundaries move +across a resume. + +When delivery fails, a policy will return `retry` (back off and deliver the +same batch again), `skip` (the user has dead-lettered it, advance past it), or +`stall` (exit, and an operator resumes from the checkpoint). The default will be +exponential backoff from 1s to 1m for ten attempts, then `stall` (R2, R3). +`skip` moves the checkpoint past real data, so it will never be a default and +will always be logged. A target that cannot represent a single row will call +`reject(row, error)`. The row will be logged with its key and timestamp, and the +rest of the batch will commit. + +`commit_interval` will be the minimum time between commits, for targets that +prefer fewer, larger writes. + +### History loss and retention margin + +When the checkpoint is older than the readable frontier, a policy will decide: +`stall` (the default), `resnapshot` (rebuild the target, for keyed targets that +can reconcile), or `skip_to_now` (accept a gap, for notification-style +consumers). Detection will use the error type (R9). + +At startup and periodically, the SDK will compare the checkpoint to the +object's readable frontier and report the margin as a metric, with a warning +below a configured threshold. Retention has to outlast an operator's response +to a stall, not only the retry budget. + +Blue/green deploys recreate the subscribed object, which ends the stream with a +dropped-object error. A `refollow` policy will re-resolve the name and resume if +the new object's fingerprint and retained history allow it, and otherwise fall +back to the history-loss policy. + +On a signal, the SDK will finish the in-flight batch, commit, close the cursor, +and disconnect. + +### Multi-view consistency + +A multi-view subscription (R4) will release changes only up to the minimum +frontier across its members, so every released moment is a consistent cut +across all views. A Materialize timestamp is a global commit order, so one +stored frontier names that cut. Log-position systems cannot do this from +positions alone, which is worth saying in the product story. + +The SDK will pick one `AS OF` and pass it to every member, and startup will fail +if any member cannot be read there. Members that pick their own `AS OF` values +hand out one view's snapshot before another's exists. + +The lag budget will count only changes the joint frontier holds back. A member's +own snapshot waiting for its own progress row will not count. + +Each view's `RETAIN HISTORY` counts from its own upper, so a lagging view can +leave the stored cut readable on one view and compacted on another. The +retention-margin check will run per member. + +A multi-view subscription will hold one connection per member. Any member error +will end all members, and recovery will resume every member at the stored cut. + +### Typed errors + +| Error | Detected by | Default | +| --- | --- | --- | +| `Transient` | no SQLSTATE (connection-level) | reconnect and resume | +| `FellBehind` | `53200` | resume from the last commit, report a metric | +| `HistoryLost` | `22000` plus the timestamp-selection message | history-loss policy | +| `ObjectDropped` | `42704` | `refollow` policy, else stop | +| `StreamPoisoned` | dataflow error | stop | +| `SchemaMismatch` | checkpoint fingerprint | history-loss policy | +| `Fenced` | checkpoint commit | stop | +| `IndexedTarget` | startup check | stop with remedy | +| `Fatal` | anything else (auth, TLS, SQL) | stop | + +The `HistoryLost` text match will be the only message matching in the SDK, kept +in one function and gated on server version until a dedicated SQLSTATE exists. + +### Connections + +Each subscription will hold one session for its whole life. Transaction-mode +poolers break `DECLARE`/`FETCH`, and a session-mode pool gains nothing because +the session is never returned. The SDK will take a connection factory as well as +a URL, so rotating credentials (OIDC tokens, app passwords) can plug in. A +multi-view subscription of N views will use N connections. + +The WebSocket SQL API supports `SUBSCRIBE` but is experimental. It is the +natural transport for browsers and functions. The protocol core is +transport-agnostic, so adding it later will not change the model. + +### Sink module + +A sink author will implement three things: open the target, write a batch and +call `commit(frontier)` inside the target's transaction, and optionally +`reject` rows the target cannot take. The SDK will drive everything else: +snapshot, resume, retry, fencing, and the history-loss policy. + +Keyed-state targets (caches, search indexes, tables) will read the upsert +envelope. When they re-snapshot, the SDK will write the snapshot under a new +generation and, at the end, call a sweep that removes keys from older +generations. That removes keys deleted while the sink was offline, which a plain +re-snapshot leaves behind. A target that needs the snapshot to appear at once +will make the new generation visible only at the sweep. + +Event targets (webhooks, queues, notifications) will read the diff envelope +under `at_least_once`, with the idempotency key above. They will declare what a +retraction means: ignore it, or call a compensating function. Compensation will +use a correlation key built from the row's key columns only, because the +idempotency key includes the timestamp and so differs between an insert and its +later retraction. + +### turbopuffer sink + +The first sink will keep turbopuffer namespaces equal to views. It answers the +search-index use case and will run against our internal context graph, which +has no Kafka, so the existing Kafka-based sink does not fit there. + +turbopuffer's documentation states that one write request to one namespace is +applied atomically and is durable on return, and that there are no transactions +across namespaces. Conditional writes compare the stored document with the +incoming one (`$ref_new`) and silently skip rows whose condition fails. An +upsert to a document that does not exist is applied unconditionally, and a +delete's condition sees `null` for every `$ref_new` attribute. + +Conditional writes alone therefore do not make a replay safe. If a later batch +deleted a key, replaying an earlier upsert of that key finds no document and +recreates it. The sink will instead write a checkpoint document into each +namespace in the same request as that namespace's data, so a namespace's data +and its frontier commit atomically. On restart the sink will resume from the +lowest checkpoint across its namespaces, and for each namespace it will drop +every change below that namespace's own checkpoint. Batch boundaries move across +a resume, so the comparison is per change, using its `mz_timestamp`. That gives +each namespace exactly-once state. + +Upserts will also be conditional on the stored `mz_timestamp` being older, and +deletes on it being below the batch frontier, passed as a literal because a +delete's `$ref_new` values are `null`. A stale worker that slips past fencing +writes the same rows at the same timestamps as the live worker, because +Materialize output is deterministic per timestamp. The condition only has to +stop its older rows from replacing newer ones. + +Writes across namespaces are not atomic. Between the writes of one cut, a reader +can see one namespace at the new cut and another at the old one. Each namespace +converges to every cut, and a reader that needs a joint view filters on +`mz_timestamp`. The sink's docs will state this. + +Embedding cost will follow the existing sink's transform model: a transform +declares the columns it reads, and runs only for rows where those columns +changed. + +### Durable subscriptions + +Durable subscriptions (#38468) give the server a per-consumer hold advanced by +`ACKNOWLEDGE`, with a wall-clock deadline in place of a window measured from the +upper. When they land, the SDK will attach with +`SUBSCRIBE USING DURABLE SUBSCRIPTION` in place of `AS OF F - 1`, and will +acknowledge only after the target commit. The target-side checkpoint will stay +mandatory, because the server alone is at-least-once. `START AT` will migrate an +existing sink at its stored frontier without a gap. A consumer that resumes by +filtering on its own position will check the opening progress message, because +recreating the subscription can move the start past that position without an +error. Stateless workers become possible at that point, because the server owns +the resume point. + +Durable subscriptions will remove the retention sizing problem, the +`AS OF F - 1` arithmetic on the default path, and most of the retention-margin +check, because the server holds history for each consumer. The rest of the +protocol core stays. Rows still need decoding and release at progress, the +server can resend data the target already committed, a multi-view cut still +spans several subscriptions, and fencing, retry, and dead-lettering do not +depend on where history lives. + +### Security + +The SDK will hold Materialize credentials and target credentials in the user's +process. The docs will recommend a dedicated role with `SELECT` on the object and +`USAGE` on the subscribe cluster, and nothing else. The protocol core performs no +network I/O, so it adds no network surface of its own. + +A resume token is not secret but is also not authenticated. Anyone who can +write the checkpoint can move the frontier forward and make the sink skip data, +so the checkpoint needs the same write protection as the target data it sits +next to. A browser transport would put credentials in the browser, which is one +reason it waits for a stable WebSocket API with scoped credentials. + +## Language and generation strategy + +The options, from least to most shared: + +- A. Hand-written per language, one spec, shared test vectors. +- B. Types, errors, tokens, and statement templates generated from one schema, + with the engine hand-written per language. +- C. One protocol core in Rust with no I/O, compiled into each language package. +- D. SDKs generated by an agent from a precise spec, accepted when they pass the + conformance suite. + +The SDK will use C. The prototype Rust package is already most of this protocol +core: its decoder, release engine, multi-view engine, tokens, statements, and +classification are pure, and only the transport does I/O. A gives every +language its own copy of the state machine, and the prototypes already show the +drift that follows (text values in Rust, native values in Python). B removes +drift in the generated parts but leaves the engine, where the bugs found so far +live, duplicated. D puts all the weight on the spec and suite, which C needs +anyway. + +The usual objection to a shared core is that it owns TLS, authentication, and +async integration in every runtime. A protocol core with no I/O owns none of +that. Each language keeps its native driver for the connection, which users +already trust, and feeds rows into the protocol core. + +Bindings will use PyO3 and maturin (abi3 wheels) for Python, napi-rs for Node, +and WebAssembly for browsers and edge runtimes. The WebAssembly build is how the +browser and function cases, and the Console's own subscribe client, can reuse +the protocol core later. Rows will cross the boundary in one call per `FETCH`, +because a call per row pays a binding crossing per row. The protocol core will +not panic across the boundary, and its errors will become each language's +native exceptions. It will not depend on Materialize workspace crates, so it +builds and versions on its own. + +The costs are a wheel and npm build matrix per platform, Rust stack traces in +Python and Node bug reports, and a source build that needs a Rust toolchain on +unsupported platforms. Go binds through cgo, which is painful, so Go will be +hand-written against the conformance vectors. If native packaging proves too +costly, the fallback is B plus A. + +### Generated and hand-written parts + +The binding glue and each package's type definitions will be generated from the +protocol core. napi-rs emits TypeScript declarations, and the Python package +will ship type stubs generated from the PyO3 module. Type mapping lives in the +protocol core: it decodes each Materialize type into one documented value +model, and each binding converts that model to the language's native types +(for example `numeric` to `Decimal` in Python), with the conversions covered by +the vectors. Each package will hand-write only the transport over its driver, +the idiomatic API surface (iterators in Python, async iterators in Node), and +the sink modules. + +### Repository layout and releases + +``` +misc/materialize-sdk/ + core/ protocol core (Rust library) + spec/ behavior spec, versioned + conformance/ vectors, shared by every package + python/ package: PyO3 binding, psycopg transport, sinks + node/ package: napi-rs binding, node-postgres transport, sinks + test/ mzcompose end-to-end suite, run in the nightlies +``` + +The SDK will start in its own Cargo workspace under `misc/`, outside the +Materialize workspace, so its dependencies and lockfile stay independent. A +spec change and its fallout in every package then land in one PR, and server +changes run against every package in the nightlies. + +Each release will publish every package at the same version, built from the +same protocol core, so a version number means the same behavior in every +language. Versions follow semantic versioning. Resume tokens and checkpoints +carry a format version, so a checkpoint written by an older release resumes +after an upgrade or fails with a clear error. The nightlies define which +Materialize versions a release supports. + +## Testing + +The conformance suite is the contract, whatever the generation strategy. + +1. Conformance vectors: language-agnostic JSON files, each with inputs and + expected outputs for one function (decode, batch release, multi-view release, + token encoding, statement SQL, and error classification from SQLSTATE and + message). The protocol core will run them natively, and each package will run + them through its binding. +2. End-to-end suite: the spec, vectors, and end-to-end suite will live in this + repository and run in the nightly pipeline against `main` through mzcompose. + One runner will drive a thin adapter per language over stdin and stdout. + Scenarios: snapshot then stream, resume after a kill, resume after a server + restart, history loss, dropped object, poisoned dataflow, fell-behind, + indexed object refused, bounded stream. A server change that breaks the SDK, + such as a changed error message, then breaks a nightly before it reaches a + customer. +3. Convergence test: kill a sink at random points (mid-batch, between write and + commit, mid-snapshot), restart, repeat, then compare the target with + `SELECT ... AS OF` the committed frontier, for both guarantees. +4. Fencing test: two workers share one checkpoint, the stale one is refused, and + the target stays consistent. +5. User test kit: the recorded-stream player from the vectors will ship to users, + so sink authors can test their targets offline (R11). + +Keeping the vectors and the end-to-end suite next to the server is what lets a +server change fail a test before it fails a customer. See "Repository layout and +releases". + +## Observability + +The SDK will report frontier lag (wall clock minus frontier), retention margin, +batches and changes applied, retries, dead-letters, reconnects, checkpoint age, +and buffer occupancy, the same in every language. They will be exposed as +callbacks in the libraries and as OpenTelemetry metrics and spans for connect, +snapshot, apply, and commit. + +## Materialize-side workstream + +Where correct behavior needs the database, the database change is part of this +program. + +1. Dedicated SQLSTATEs for history loss and for dataflow errors. +2. SQL-528: a final progress row clamped to `UP TO`. +3. Durable subscriptions (#38468). +4. Snapshot elision for a projection and filter without a temporal predicate, + which #38468 also needs. +5. Non-poisoning subscribe errors (database-issues#5182). +6. Docs that cross-link the durable-subscriptions pattern from every client page + now, and lead with the SDK once it ships. +7. A stable WebSocket `SUBSCRIBE`, the gate for browser and function transports. + +The server-side buffering bound (#37905) has landed and needs no further work. + +## Minimal Viable Prototype + +The prototype will be the protocol core extracted from the Rust package, one +language package, and the turbopuffer sink running against the internal context +graph. It tests the three riskiest claims: + +- that a protocol core with no I/O binds into a language package without + packaging or performance problems, +- that per-namespace checkpoint documents give each turbopuffer namespace + exactly-once state under random kills, +- that the nightly end-to-end suite catches server changes that break the SDK. + +## Delivery plan + +October: the prototype above, with the end-to-end suite in the nightlies. Exit +criteria are a passing convergence test, a passing fencing test, and every typed +error reproduced in the end-to-end suite. + +November and December: the second language, durable subscriptions as they land, +and a Redis or Postgres reference sink chosen by demand. + +GA requires dedicated SQLSTATEs released, the client docs rewritten on the SDK, +and both languages passing the same vectors and end-to-end suite. + +## Alternatives + +### The product strawman + +The product requirements include a strawman API with a `Cursor` and a +`SinkConnector`. Its boundary between tracking progress and delivering data is +right, and this design keeps it. The parts that change: + +| Strawman | This design | Why | +| --- | --- | --- | +| `update()` commits and fetches | `commit()` inside the transaction, `next()` outside | A target transaction must not wait on Materialize | +| `queries` as SQL strings, split at `ENVELOPE` | Object, projection, filter, envelope as structured input | Resume cost, durable subscription compatibility, no string splicing | +| One cursor per query | One `AS OF` for all members, release at the minimum frontier | Independent cursors produce torn cuts | +| Rows stamped with the batch frontier | Every change keeps its `mz_timestamp` | Batches span timestamps | +| Retractions come first | Order by timestamp, net per key within a timestamp | Not true by default | +| `name` defaults to the class name, rows keyed by list index | Required names, rows tagged by view name | Shared checkpoints fence each other, reordering remaps rows | +| `commit_interval` in progress messages | Minimum time between commits | Progress cadence is not a user-facing unit | +| Snapshot handling left to the user | Partial snapshot chunks, generation sweep, history-loss policy | Large snapshots and orphan keys are the common failure | +| Conditional turbopuffer writes for replay safety | A checkpoint document per namespace, written with the data | A replayed upsert recreates a key a later batch deleted | + +### Build only on durable subscriptions + +The semantics would be simplest, but nothing ships until the server does, and +the target-side checkpoint is needed either way. The narrow subscription scope +lets the SDK ship now and move over without an API break. + +### Native server-side sinks per target + +This is the best eventual experience for a few high-volume targets. Each target +becomes a server feature with its own release train, and the long tail of +targets never ends. The SDK serves the long tail and shows which targets deserve +native sinks. + +### Kafka and the connector ecosystem + +This is right for organizations already running Kafka Connect. It makes everyone +else run Kafka to reach a cache, adds a hop, and loses the mapping between target +state and a Materialize timestamp unless the connector is consistency-aware. + +### A native message queue + +`mzq` (PRD-87) overlaps with this design. If it lands, it is another transport +under the same protocol core and sink API, not a replacement for them. This +needs settling with the product owner. + +### Docs only + +The pattern doc already exists. The table in "The Problem" shows what docs alone +produce. + +## Open questions + +1. Name: Materialize SDK, Subscribe SDK, or Sink SDK? The reasoning is under + "Naming". This is cheap to change now and expensive after the first publish. +2. October language: Python matches the existing turbopuffer sink, the + turbopuffer client, and the field team's tooling. Rust is where the protocol + core already lives. Recommendation: the Rust protocol core plus the Python + package. +3. Generation strategy: is a protocol core with bindings acceptable for + packaging and support, or do we start with B plus A? +4. turbopuffer checkpoint document: the per-namespace checkpoint document is what + makes replay safe, but every query must exclude it. The alternative is + tombstones in place of deletes, swept later, with the checkpoint in a + Materialize table. Which cost is easier for users? +5. Retention margin check: can a sink role read the catalog state it needs (the + object's readable frontier and the cluster's indexes) without extra grants? +6. `mzq`: a transport under this SDK, or a separate path? +7. Object swaps: how often do blue/green deploys swap a subscribed view, and is + `refollow` the behavior deploy tooling wants? +8. Egress cost: data leaves through `environmentd`. What does a sink with a large + snapshot cost a customer, and does the docs story need a sizing page? +9. Stateless workers before durable subscriptions: is there demand that cannot + wait? +10. Repository home: should the packages stay under `misc/` in this repository + for good, or move to their own repository once the spec settles, keeping + the protocol core, vectors, and end-to-end suite here? diff --git a/doc/developer/design/20260706_subscribe_sdk.md b/doc/developer/design/20260706_subscribe_sdk.md deleted file mode 100644 index dfeba7207ab3d..0000000000000 --- a/doc/developer/design/20260706_subscribe_sdk.md +++ /dev/null @@ -1,937 +0,0 @@ -# Subscribe SDKs: canonical clients for `SUBSCRIBE` and a framework for external sinks - -- Associated: - - (TBD: link the Subscribe SDK program epic once filed) - - [MaterializeIncLabs/mz-redis-sync](https://github.com/MaterializeIncLabs/mz-redis-sync) (prior art) - - [MaterializeIncLabs/novu-materialize-integration](https://github.com/MaterializeIncLabs/novu-materialize-integration) (prior art) - - [Durable subscriptions pattern](https://materialize.com/docs/transform-data/patterns/durable-subscriptions/) (documented protocol this work packages) - - [database-issues#5182](https://github.com/MaterializeInc/database-issues/issues/5182) (errors poison subscribe dataflows, relevant server limitation) - -## The Problem - -`SUBSCRIBE` is the primitive that turns Materialize from a database you query -into a database that drives your systems. It is how users push incrementally -maintained results into caches, notification services, feature stores, -websockets, and anything else that is not reachable by `CREATE SINK`. Today we -give users the primitive and a pattern document, and then every user rebuilds -the same client-side machinery from scratch. - -That machinery is genuinely hard to get right. The correct protocol for -consuming `SUBSCRIBE` durably is: - -1. Subscribe `WITH (PROGRESS, SNAPSHOT true)`. -2. Buffer updates until a progress message proves a timestamp is closed. -3. Apply the closed batch and persist its frontier atomically. -4. On restart, resume `WITH (PROGRESS, SNAPSHOT false) AS OF frontier - 1`, - because `SNAPSHOT false` emits updates strictly *after* the `AS OF` - (`src/compute/src/sink/subscribe.rs`, the `should_emit` closure), so - subtracting one is what makes the resume gap-free. -5. Keep `RETAIN HISTORY` on the subscribed object wider than your worst-case - downtime, and handle the compaction-horizon error when it is not. - -Each step has a failure mode that is invisible in testing and destructive in -production. We know this because every existing implementation we could find, -including our own demos, gets a different step wrong: - -| Codebase | Commit discipline | Resume | Idempotency | Retractions | -| --- | --- | --- | --- | --- | -| mz-redis-sync (Labs) | Correct: commits data and frontier atomically at progress rows | Omits `SNAPSHOT` on resume, which defaults to `true`, so every restart replays the full snapshot (contradicting its own commit message) | Inherent (keyed writes) | Delegated to `ENVELOPE UPSERT`, but `key_violation` crashes the process and poisons restarts | -| novu-materialize-integration (Labs) | Persists the timestamp of the *first row* of a batch. A crash mid-batch permanently loses the rest of that timestamp | Wall-clock guess of whether the checkpoint is still inside the retention window | Content-hash keys (false-dedups identical payloads across time) | Maps `mz_diff = -1` to a notification revoke, a great idea, implemented with two bugs | -| novu-materialize-poc (Labs) | Correct idea: buffer until the timestamp advances | No `PROGRESS`, so batch boundaries are a heuristic | None, duplicates on replay | Ignored, deletes trigger spurious notifications | - -(One fairness note on the first row: because that sink writes keyed upserts -transactionally, the snapshot replay is wasteful rather than incorrect, -every restart rewrites the full view into Redis. It still contradicts the -commit message that claims to avoid it, which is its own kind of evidence.) - -None of the three reconnects on failure. None has a single test. One declared -a retry library as a dependency and never imported it. These are demos, -written quickly, some AI-assisted, and that is precisely the point: this is -what subscribe consumers look like when they are written the way real -integrations get written, by whoever shows up with a deadline. The -correctness burden currently sits on the application side of an API this -subtle, so this is the result. Throughout this document, the demos serve -only as evidence of the failure modes. Every claim about correct behavior is -anchored in the Materialize codebase and documentation, which are the ground -truth this design is built on. - -Beyond correctness, there are stability hazards that clients must be built -around and currently are not: - -- The coordinator buffers subscribe output in an unbounded channel with no - backpressure (`src/adapter/src/active_compute_sink.rs`, acknowledged - TODO). A slow consumer translates into `environmentd` memory growth. - Separately, subscribe output is subject to `max_result_size` in the - compute-client layer (`PendingSubscribe::stash`), so an initial snapshot - larger than that limit terminates the subscription with a "total result - exceeds max size" error, a sizing hazard no demo handles. -- Once a subscribe dataflow produces an error it repeats that error forever - (database-issues#5182). Recovery is always reconnect-shaped, and clients - that do not know this retry into the same error. -- The compaction-horizon error is distinguishable only by matching the error - message text, under a generic SQLSTATE (`DATA_EXCEPTION`, shared with - other failures). Worse, the text is not stable: the timestamp-selection - rework (#34712) replaced "Timestamp (X) is not valid for all inputs" with - "could not find a valid timestamp for the query". mz-redis-sync regexes - the old text, so its one bespoke error handler is already broken against - current Materialize, and its non-matching branch silently swallows the - error and re-enters `FETCH` inside an aborted transaction, an infinite - loop. This is what message-text dispatch does over time. - -Finally, our docs teach the naive version. Every per-language client page -(`doc/user/content/integrations/client-libraries/*.md`) shows a bare -`DECLARE`/`FETCH` loop with no progress handling, no resume, and no -reconnection. The durable-subscriptions pattern doc is correct and complete, -but it is prose describing a protocol, and prose does not compose into -applications. - -The problem, stated in one line: **we make users hand-implement a consistency -protocol that we already know how to implement, and they predictably get it -wrong.** - -## Success Criteria - -A solution is successful if: - -1. **Correct by construction.** A user following the happy path of the SDK - cannot commit an unclosed timestamp, cannot get the `AS OF` boundary - arithmetic wrong, and cannot silently lose or duplicate updates across a - restart. The resume token is opaque. The batch is the unit of consumption. -2. **Minutes to first stream.** `install`, paste a connection string, write - five lines, see typed updates. The SDK works against Materialize Cloud, - self-managed, and the emulator without configuration differences. -3. **Stable under failure.** Network blips, cluster restarts, and process - crashes are recovered automatically or surfaced as one of a small set of - typed, documented errors with remediation guidance. Memory use is bounded - and documented, including during snapshots. -4. **General over destinations.** The sink framework cleanly supports the two - families of destinations, keyed-state stores (Redis, DynamoDB, Postgres - upsert tables, search indexes, caches) and event receivers (webhooks, - queues, notification services). Redis ships first as the reference - implementation, not as a special case. -5. **Consistent across languages.** TypeScript, Python, and Rust expose the - same model with idiomatic surfaces, verified by a shared conformance suite - rather than by hope. Behavior is specified in a written protocol document. -6. **Honest about guarantees.** The SDK names its delivery guarantees - precisely, exactly-once state application for transactional sinks and - at-least-once with idempotency keys for side-effecting sinks, and makes - the user pick one explicitly. -7. **Testable by users.** Sink authors get a testing story from us, recorded - stream fixtures and an emulator harness, so their sinks can be tested - without a live environment. - -Quantitative proxies: replace mz-redis-sync with an SDK-based equivalent at -functional parity plus the known bugs fixed. A chaos suite that kills the -process at arbitrary points and verifies destination convergence passes for -every language. The per-language docs pages can be rewritten on top of the SDK -with fewer lines than the naive loops they replace. - -## Out of Scope - -- **Rewriting subscribe server internals.** The SDK is buildable on today's - semantics, and a scoped set of server-side companion changes is planned as - part of this program (see "Materialize-side workstream"). Anything beyond - that scoped list, such as a general subscribe protocol redesign, is out of - scope. -- **Exactly-once side effects.** No client library can make an HTTP call - exactly once. We provide at-least-once delivery plus idempotency-key - helpers, and we say so plainly. -- **A browser/edge SDK.** The WebSocket endpoint (`/api/experimental/sql`) - does support `SUBSCRIBE`, but it is experimental. A browser client is - future work gated on stabilizing that surface. -- **Go, Java, and other languages in v1.** The spec and conformance suite are - designed to make additional languages cheap, but the v1 program is Rust, - Python, and TypeScript, in that order. GA can precede the TypeScript - release (see the Delivery plan). -- **Private-preview subscribe features.** `ENVELOPE DEBEZIUM` and - `WITHIN TIMESTAMP ORDER BY` are behind feature flags. The SDK's model - leaves room for them but v1 does not depend on them. -- **Scale-out consumption.** Sharding one subscription's updates across - multiple parallel workers has no server-side support and is not attempted. - High availability itself is in scope (fencing in v1, active-passive - failover shortly after, see Layer 2), the exclusion here is only parallel - fan-out of a single subscription for throughput. -- **General driver features.** The SDK consumes `SUBSCRIBE` (and the small - amount of SQL needed around it). It is not a query builder, ORM, or general - Materialize client. - -## Solution Proposal - -Ship a family of client libraries, one written behavior spec, and one shared -conformance suite, under a working title of the **Materialize Subscribe SDK**: - -``` - +--------------------------------------+ - | Layer 3: sink kit + built-in sinks | - | (Redis reference sink, webhook sink)| - +--------------------------------------+ - | Layer 2: durable subscribe | - | (checkpoints, guarantees, policies) | - +--------------------------------------+ - | Layer 1: subscribe client | - | (typed stream of closed batches) | - +--------------------------------------+ - | native pg driver per language | - | (node-postgres / psycopg / tokio-pg)| - +--------------------------------------+ -``` - -The layers are strictly separated so each is independently useful. A user who -only wants live updates in a websocket handler uses Layer 1 and never sees a -checkpoint. A user syncing Redis uses Layer 3 and never sees a `FETCH`. - -### The model: five opinions - -The SDK's UX comes from five opinions applied uniformly across languages: - -1. **The unit of consumption is the closed timestamp, never the bare row.** - The stream yields consistent batches. Each batch contains every update for - an interval of timestamps that a progress message has proven complete, - plus the frontier and a resume token. Users physically cannot observe a - half-delivered timestamp. A batch is a consistent moment of one view, and - the same engine generalizes to a *cohort* of views advancing to a joint - consistent moment (see "Consistency layering and cohorts"). The advanced - `raw()` mode that exposes the decoded per-row stream (forfeiting the batch - guarantees) is that same underlying layer, and is the substrate cohorts - build on. -2. **Resume tokens are opaque.** A token encapsulates the frontier, the - `SNAPSHOT false AS OF frontier - 1` arithmetic, the query fingerprint, and - a fencing epoch. Users store bytes and hand them back. There is no - timestamp arithmetic in user code. -3. **Guarantees are named and chosen, not implied.** Durable consumption - requires the user to pick `transactional` (checkpoint commits atomically - with effects, exactly-once state application) or `at_least_once` (effects - first, checkpoint after, idempotency keys provided). There is no default - that quietly picks one. -4. **Reconnection is the SDK's job.** Transient failures are retried with - jittered backoff and automatic resume. Everything else surfaces as one of - a small typed error set, each carrying remediation guidance. -5. **One spec, one conformance suite, three implementations.** Language - surfaces are idiomatic, behavior is identical, and identical is checked by - machines. - -### Consistency layering and cohorts - -The SDK is layered so that a single consistency engine serves both the -single-view and the multi-view case: - -1. **Decoded stream.** The transport's `FETCH` output decoded into timestamped - changes and progress markers (`Data{timestamp, change}` / - `Progress{frontier}`). This is the `raw()` layer: composable and public, and - on its own forfeiting the higher guarantees. -2. **Consistency engine.** Buffers changes and releases everything strictly - below a release frontier, consolidating per timestamp. It takes the release - frontier as an *input*: a single-view subscription passes its own frontier; - a cohort passes `min(frontier)` across its members. The single-view batcher - is literally the one-member case of the cohort engine, so there is exactly - one buffer-and-release implementation to get right and to test. -3. **Batch / moment.** The closed `ConsistentBatch` (one view) or cohort - moment (N views) handed to the consumer, with its resume token. - -**Cohorts** turn N independent subscriptions into one stream of jointly -consistent moments. Because every object in a Materialize environment shares -one logical timeline, progress timestamps from independent subscriptions are -directly comparable, so `min(frontier)` across members is a valid global -consistent cut. The SDK withholds a moment until every member has closed it, -so the consumer only ever sees jointly consistent state and never has to reason -about the cut itself. This is a genuine Materialize differentiator (Frank -McSherry's `mz-bridge-recipe` is the prior art) and it composes on the layering -above rather than being a separate code path. - -The cohort is a shipped, first-class helper, not the core. The common -single-view case pays none of its concept weight: `consume([oneView])` is -identical in feel to a single subscription, and a single view is just the -degenerate cohort of one. A cohort of N views is N subscriptions and therefore -N connections (a long-running `SUBSCRIBE` cursor holds its connection in a -`FETCH` loop, so several cannot be multiplexed onto one). That cost is linear -in views and worth stating. Dynamic cohort membership (add/drop/merge/split of -a live cohort) is deferred to future work. - -The cohort is the SDK's strongest-consistency construct, and by the same token -its least available. It makes progress only as fast as its slowest member, and -a jointly consistent moment exists only once every member has closed it. The -single-view path and the raw stream are more available precisely because they -answer for one view alone. Three consequences are worth stating for v1, and the -unresolved policy for each is an open question below. - -- **One timeline.** `min(frontier)` is a valid cut only when every member reads - the same logical timeline, so their timestamps are comparable. That holds for - objects on the default timeline. A member on a user-defined timeline, or a - source carrying external timestamps, would make the comparison meaningless. - v1 documents this as a constraint rather than checking it. -- **A laggard is bounded, not silent.** Holding a leading member's changes until - the slowest catches up is inherent to the guarantee, so a stalled member would - otherwise buffer its peers without bound. v1 caps the total buffered across - the cohort with a lag budget and fails loud when it is exceeded, rather than - growing memory quietly. -- **Failure is all-or-nothing in v1.** Any one member erroring tears the whole - cohort down, and recovery re-subscribes every member at the joint frontier. - This is the conservative, consistent choice. Reconnecting only the failed - member without disturbing the others is an availability improvement tracked - with the HA workstream. - -### Layer 1: the subscribe client - -Layer 1 wraps the ecosystem-native PostgreSQL driver and runs the -`DECLARE`/`FETCH` loop, presenting a typed stream. - -Illustrative TypeScript (all sketches in this doc are illustrative, not final -API commitments): - -```typescript -import { SubscribeClient } from "@materializeinc/subscribe"; - -const client = await SubscribeClient.connect(process.env.MZ_URL!); - -const stream = client.subscribe({ - query: "SELECT id, amount FROM winning_bids", - envelope: { upsert: { key: ["id"] } }, -}); - -for await (const batch of stream) { - // batch.frontier: bigint -- every update with timestamp < this is present - // batch.resumeToken: ResumeToken -- opaque, serializable - // batch.isSnapshot: boolean - // batch.updates: Upsert[] | Delete[] | KeyViolation[] - render(batch.updates); -} -``` - -Illustrative Python: - -```python -from materialize_subscribe import connect, UpsertEnvelope - -client = connect(dsn) -for batch in client.subscribe( - "SELECT id, amount FROM winning_bids", - envelope=UpsertEnvelope(key=["id"]), -): - for update in batch.updates: - ... -``` - -Illustrative Rust: - -```rust -let client = mz_subscribe::Client::connect(&dsn).await?; -let mut stream = client - .subscribe(Subscribe::query("SELECT id, amount FROM winning_bids") - .envelope_upsert(["id"])) - .await?; -while let Some(batch) = stream.try_next().await? { - apply(batch.updates()); -} -``` - -Responsibilities and design points: - -- **Batching by progress.** The client requests `PROGRESS` always. It buffers - updates and emits a batch when the frontier advances. Progress-only - advancement (idle periods) is surfaced as an empty batch carrying a fresh - resume token, so downstream checkpoints keep advancing during quiet hours. - This directly fixes the failure mode where an idle subscription's - checkpoint ages out of the retention window. -- **Envelope decoding.** Two modes map to typed events: - - default diff envelope: `Insert{row, diff}` / `Retract{row, diff}` with - multiplicities preserved (a `diff` of -3 is three retractions and the type - says so). Within a closed batch, diffs are consolidated per row (net - multiplicity, net-zero rows dropped) so a batch is a clean net delta at the - frontier, not a replay of intra-window churn. Consolidation and the sink's - upsert target key on the same content-derived row identity, so they agree - by construction (the `mz-bridge-recipe` technique). - - `ENVELOPE UPSERT`: `Upsert{key, value}` / `Delete{key}` / - `KeyViolation{key}`. Key violations are a first-class event, not an - exception, because a crash-restart loop cannot fix them. -- **Snapshot streaming with bounded memory.** The initial snapshot is one - giant timestamp and can exceed client memory if buffered whole (the - mz-redis-sync failure mode). The client streams it as chunked batches - flagged `partial: true`, with the closing chunk carrying the resume token. - Consumers that need atomic snapshot visibility get help from Layer 3. - Server-side, a snapshot exceeding `max_result_size` terminates the - subscription with a "total result exceeds max size" error regardless of - client behavior. The SDK surfaces that as a typed error with sizing - remediation rather than as an opaque stream failure. -- **Never backpressure Materialize.** Unread `SUBSCRIBE` output buffers without - limit in `environmentd` (an acknowledged server TODO), so an unconsumed - subscription makes *Materialize* grow, not the client. The SDK therefore - decouples a continuous drain loop (`FETCH c WITH (timeout ...)`, adaptive - sizing, always draining the server promptly) from delivery to the consumer, - with a buffer in between. That buffer is bounded and fails loud when full: - the client is the thing that must fall over, never the database. If the - consumer cannot keep up the documented options are a larger bound or opt-in - spill-to-disk. The SDK never silently lets the server grow, and it never - couples draining to the consumer's pace (the mistake in the initial - scaffolding, where `next()` only fetched on demand). -- **Typed errors.** All failures map to a small taxonomy: - - | Error | Meaning | Default behavior | - | --- | --- | --- | - | `Transient` | network blip, cluster restart, replica loss | auto-reconnect and resume | - | `CompactionHorizon` | resume point older than retained history | surface with policy hook (see Layer 2) | - | `DependencyDropped` | subscribed object or its cluster dropped | surface, policy hook can `refollow` recreated objects (see Layer 2) | - | `StreamPoisoned` | dataflow error (e.g. division by zero in the view), repeats forever per database-issues#5182 | surface with explanation, reconnect will not fix until the data or view changes | - | `SchemaMismatch` | resumed subscription's columns differ from checkpoint fingerprint | surface with policy hook | - | `Fatal` | auth, TLS, SQL errors in the user's query | surface immediately | - - The mapping rules key on structured error codes, which is the first item - in the Materialize-side workstream, sequenced to land before the SDK's - stable release. Because the SDK must also work against server versions - that predate those codes, a stable release may carry a message-text - fallback in exactly one internal function, gated on server version, - covered by conformance tests pinned to each legacy message (the text has - already changed once across releases, see the Problem section), and - removed when pre-code versions age out of support. No other - message-string behavior ships anywhere. -- **One direct connection per subscription.** A subscription owns a dedicated - connection for its lifetime. The SDK documents, and detects where it can, - that transaction-mode poolers (PgBouncer and friends) break the - `DECLARE`/`FETCH` loop. Applications running many subscriptions get an - explicit connection budget instead of a surprise. -- **Cluster targeting and resume cost.** The subscribe options include the - target cluster (`SET cluster`), and the docs shipped with the SDK - recommend a dedicated serving cluster for subscriptions. The SDK surfaces - the cost asymmetry on resume: `SUBSCRIBE ` on a materialized view - or table resumes cheaply from storage, while `SUBSCRIBE (SELECT ...)` - rebuilds a dataflow over retained history. Guidance: materialize the view - you sink. -- **Ordering is per timestamp, not within it.** Updates within a closed batch - are unordered (server-side `WITHIN TIMESTAMP ORDER BY` exists but is - private preview). Keyed-state sinks are insensitive to this. Event sinks - that need deterministic order within a timestamp can sort client-side via - a batch option, and the docs say when that matters. -- **Bounded subscriptions.** `UP TO` is exposed as a first-class option, which - makes deterministic reads of a timestamp window possible. This is useful in - its own right and is how much of the SDK's own test suite drives itself. -- **Session hygiene.** Sets `application_name`, suppresses the welcome notice - where drivers require it, exposes the connection's cluster and role in a - startup log line, and pre-validates the query with - `SELECT * FROM () WHERE FALSE LIMIT 0` for fail-fast schema and - permission errors (an mz-redis-sync trick worth canonizing). -- **Type mapping contract.** Each language documents a total mapping from - Materialize types to SDK values, including `numeric` precision, temporal - types, arrays, `jsonb`, and NULLs. The conformance vectors include every - type. - -### Layer 2: durable subscribe - -Layer 2 adds checkpointing and delivery guarantees on top of Layer 1. - -```typescript -import { durableSubscribe, redisCheckpoints } from "@materializeinc/subscribe"; - -await durableSubscribe(client, { - name: "bids-to-webhook", // checkpoint identity - query: "SELECT id, amount FROM winning_bids", - checkpoints: redisCheckpoints(redis), // pluggable store - guarantee: "at_least_once", - onCompactionHorizon: "resnapshot", // or "fail", or callback - handler: async (batch, ctx) => { - for (const event of batch.updates) { - await sendWebhook(event, { idempotencyKey: ctx.idempotencyKey(event) }); - } - }, -}); -``` - -Design points: - -- **Checkpoint stores are pluggable and tiny.** The interface is - `load(name) -> ResumeToken | null` and `store(name, token)`. Ships with: - the destination itself (the strongly recommended default, see Layer 3), - a Materialize table (zero extra infrastructure, the Novu demos' good idea), - Postgres, and a local file (development). -- **Two named guarantees.** - - `transactional`: the handler receives the batch and the token and must - commit both atomically (Layer 3 sinks do this for you). Result: - exactly-once state application. The destination always equals the source - at some real Materialize timestamp. - - `at_least_once`: effects run first, the checkpoint commits after the - handler returns. Result: replays possible after a crash, and - `ctx.idempotencyKey(event)` provides a stable key derived from - subscription name, `mz_timestamp`, key columns, and an ordinal within the - *timestamp*. It must not depend on batch grouping: batch boundaries are - not stable across a resume (the same change can land in a - differently-grouped moment when frontiers tick at different wall-clock - instants, and only net state converges), whereas the set of updates at a - given `mz_timestamp` is deterministic. Keying on a within-batch ordinal - would silently stop deduplicating replays. This fixes both observed - idempotency failures: content-only hashes that false-dedup distinct - events, and no key at all. -- **Retraction policy for event sinks.** Side-effecting consumers declare - what a retraction means: `ignore`, or `compensate(fn)` (the Novu demo's - retraction-to-revoke mapping, promoted to a supported concept). - Compensation needs a correlation identity that is stable between an - insert and its later retraction, which the delivery idempotency key - deliberately is not (it includes the timestamp and a within-timestamp - ordinal). The SDK - therefore provides `ctx.correlationKey(event)`, derived from the row's - key columns only, for exactly this purpose. The Novu demo collapsed both - identities into one content hash, which made revokes work but false-dedups - distinct events. Separating the two identities is the fix. -- **Compaction-horizon policy.** When the checkpoint is older than retained - history, policy decides: `fail` (default, explicit), `restart_from_now` - (accept a gap, for notification-style consumers), or `resnapshot` (rebuild - destination state, meaningful mainly for Layer 3 keyed sinks which know how - to reconcile). The SDK detects the condition from the typed error, not from - wall-clock guessing against `RETAIN HISTORY` config. -- **Query fingerprinting.** The token embeds a fingerprint of the query text - and output schema. Resuming with a changed query surfaces `SchemaMismatch` - and the policy hook chooses between failing and re-snapshotting. No demo - handled this. Real deployments hit it on their first schema change. Tokens - themselves carry a format version so stored checkpoints survive SDK - upgrades. -- **Object swaps and blue/green deployments.** Recreating or swapping the - subscribed object (the recommended zero-downtime deploy pattern, and what - tooling like mz-deploy automates) terminates the subscription with - `DependencyDropped`. The policy hook offers `refollow`: re-resolve the - object by name, resume from the checkpoint if the new object's schema - fingerprint and retained history allow it, otherwise fall through to the - re-snapshot policy. Without this, every view deploy is a paging incident - for whoever runs the sink. -- **Fencing and high availability.** The token carries an epoch. Checkpoint - stores implement compare-and-set on `(epoch, frontier)`, refusing writes - from a stale epoch and refusing frontier regression. In v1 this makes the - two-instances mistake loud instead of silently corrupting. The production - HA story builds on the same primitive: active-passive failover where - standby instances contend for a lease in the checkpoint store, the lease - holder bumps the epoch on acquisition, and the fencing rule guarantees at - most one writer even across partitions. This is the Kafka Connect and - Debezium recovery model, no consensus service required beyond the - checkpoint store's compare-and-set. If real deployments show the - client-side lease is not enough, the Materialize-side workstream leaves - room for a server-assisted primitive (for example, named subscriptions - with server-enforced single ownership), designed properly rather than - worked around. -- **Liveness watchdog.** Optional max-staleness on progress. If the frontier - stops advancing (hung connection, wedged cluster) the SDK reconnects or - surfaces, with correct units, unlike the demo whose watchdog compared - seconds to milliseconds and could never fire. -- **Graceful shutdown.** On signal: finish the in-flight batch, checkpoint, - close the cursor, disconnect. - -### Layer 3: the sink kit and built-in sinks - -Layer 3 answers "I want this view synced into X" with a small interface per -destination family and the hard problems solved once. - -Positioning: the SDK core (Layers 1 through 3's interfaces) is the fully -supported product surface with a day-one stability commitment. Built-in -sinks are supported reference implementations of those interfaces, Redis -first. The framework is the product, destinations are instances of it, and -users building their own sinks are as much the audience as users running -ours. - -**Keyed-state sinks** (Redis, DynamoDB, Postgres tables, search indexes). -Consume `ENVELOPE UPSERT`. The contract: - -```python -class KeyedStateSink(Protocol): - def load_token(self) -> ResumeToken | None: ... - def apply(self, batch: UpsertBatch, token: ResumeToken) -> None: - """Apply updates and persist token atomically. Must be idempotent.""" - def resync_begin(self, generation: int) -> None: ... - def resync_end(self, generation: int) -> None: - """Make generation current and sweep keys from older generations.""" -``` - -The kit drives the state machine: initial snapshot, incremental batches, -resume, and re-snapshot after compaction-horizon or state loss. The -`resync_*` hooks solve the two problems every demo either hit or documented -away: - -- *Atomic snapshot visibility.* Large snapshots arrive as partial batches. A - sink can stage them under a generation marker and flip visibility at - `resync_end`, or accept eventually-visible snapshots. The kit supports - both, the sink declares which. -- *Orphan reconciliation.* Re-snapshotting after state loss only upserts - currently-live keys. Keys deleted while offline linger forever unless swept - (mz-redis-sync's README admits exactly this gap). Generation-tagged sweep - at `resync_end` closes it. - -**Event sinks** (webhooks, queues, notification services). Consume the diff -envelope through Layer 2's `at_least_once` mode with idempotency keys and -retraction policy. The kit ships a generic webhook sink as the reference for -this family. - -**The Redis reference sink.** First concrete sink, the successor to -mz-redis-sync: - -- Data models: string (`SET key value`), hash (row columns as fields), and - JSON value encoding. Multi-column keys via a documented key template. - NULLs handled per a documented rule instead of crashing the driver. -- Writes: one `MULTI`/`EXEC` per closed batch containing the data commands - plus the checkpoint `SET`. This is the atomic - data-plus-progress-in-one-transaction rule that the durable-subscriptions - pattern doc prescribes, applied with the destination as the transaction - boundary (mz-redis-sync instantiates the same idea). Large batches chunk - under a generation marker with a sweep, trading atomic visibility for - bounded transactions, per the sink's declared mode. -- Checkpoint key namespaced under the sink's prefix (the demo's frontier key - bypassed its own prefixing and could collide with data). -- Fencing via a Lua compare-and-set on `(epoch, frontier)`. -- `key_violation` events surface through a policy hook (log-and-skip or - fail) instead of crashing into a poison-restart loop. - -**Standalone runner as a reference example.** The repo ships a small -config-file-driven runner built on the Rust implementation (working name -`mz-sink`), positioned as an example application, not a core deliverable: -it lives in `examples/`, demonstrates end-to-end sink deployment for -non-Rust users, replaces mz-redis-sync as the thing we point demos at, and -serves as the long-running soak target for the chaos suite. The SDK is the -product people build with however they please, the runner shows one good -way. - -### Language strategy: native implementations, one spec - -Three native implementations, not a shared Rust core with bindings: - -- TypeScript on `pg` (node-postgres), Python on `psycopg` (v3, async-capable, - sync facade), Rust on `tokio-postgres`. -- The hard, subtle part of this project is a small protocol state machine. - It is precisely the kind of logic a written spec plus shared test vectors - can pin down across implementations. -- The expensive part of a shared-core approach is everything else: TLS, - auth, connection lifecycle, event-loop integration in Node, asyncio - integration in Python, packaging native modules for every platform. The - ecosystem drivers already solve all of it, natively and idiomatically, and - users already trust them. - -The spec (working name **Subscribe Consumption Protocol**, versioned, -`spec/` in the SDK repo) defines: batching and progress semantics, resume -token contents and boundary arithmetic, guarantee modes and their crash -matrices, error taxonomy and mapping rules, envelope decoding, and the type -mapping. Each language's README links its conformance report. - -This is the proven model for exactly this kind of SDK. MongoDB drivers are -the canonical example: the major drivers are per-language native -implementations, unified by the public `mongodb/specifications` repo of -prose specs plus JSON test files, and MongoDB change streams use an opaque -resume token with the same role as ours. Kafka demonstrates both paths at -once: the Java client is the spec-bearing reference, native implementations -displaced the C-binding clients where binding pain was highest (kafka-go -over cgo bindings, KafkaJS over node-rdkafka), while librdkafka bindings -persist where they work well enough. The lesson is that bindings are a -per-ecosystem cost gamble, native is uniformly safe. Debezium shows the -cost of skipping multi-language entirely, its embedded engine is JVM-only, -which is a large part of why non-JVM teams never adopted it directly. - -Delivery order: **Rust and Python first.** Rust is the reference -implementation, written against the spec as the spec is written, and it is -the language of the team that must vouch for the semantics. Python is the -fastest path to validating the UX with real integration builders and -carries the prototype. **TypeScript follows** once the spec has survived two -implementations, **Go later**. The spec makes each additional language a -mechanical, conformance-checked project. - -Suggested packaging (open question below for final naming): - -| Language | Core | Redis sink | -| --- | --- | --- | -| TypeScript | `@materializeinc/subscribe` | `@materializeinc/sink-redis` | -| Python | `materialize-subscribe` | extra: `materialize-subscribe[redis]` | -| Rust | `mz-subscribe` (crates.io) | feature: `mz-subscribe/redis`, binary `mz-sink` | - -One monorepo (`MaterializeInc/subscribe-sdk` or Labs equivalent) holding -`spec/`, `conformance/`, and the three implementations, so a spec change and -its cross-language fallout land in one PR. - -### Testing - -The demos shipped zero tests between them. The SDK inverts this, and the test -infrastructure is a deliverable users get too: - -1. **Conformance vectors.** Language-agnostic recorded subscribe streams - (JSON) covering: progress interleavings, multiplicities beyond one, - upsert/delete/key_violation, partial snapshots, idle progress, - every mapped type, and error frames. Every implementation must replay - them to identical decoded events and tokens. -2. **Emulator integration suite.** Docker-based (testcontainers) scenarios - against the Materialize emulator: snapshot then stream, resume across - client restart, resume across emulator restart, `RETAIN HISTORY` expiry - producing `CompactionHorizon`, dropped view producing `DependencyDropped`, - poisoned dataflow producing `StreamPoisoned`. -3. **Chaos suite (the flagship).** Kill the sink process with SIGKILL at - randomized points (mid-batch, between apply and checkpoint, mid-snapshot), - restart, repeat, then assert the destination equals - `SELECT ... AS OF ` at the checkpointed frontier. Run for both - guarantee modes, asserting convergence for `transactional` and - convergence-with-duplicates-absorbed for `at_least_once`. -4. **Fencing test.** Two instances against one checkpoint, assert the stale - epoch is refused and the destination stays consistent. -5. **Property tests.** Random diff streams through envelope decoding and - batching, asserting consolidation and frontier invariants, including the - cohort case: `min(frontier)` release, and a laggard member holding the joint - moment until it catches up. -6. **User-facing test kit.** The fixture player from (1) is exported so sink - authors can unit-test their `apply` implementations offline. - -### Observability - -Built-in, consistent across languages: frontier lag (wall clock minus -frontier, the one metric every operator wants and no demo had), batches and -updates applied, reconnect count, checkpoint age, buffer occupancy. Exposed -as callbacks/hooks in the libraries and as Prometheus metrics plus health -endpoint in `mz-sink`. OpenTelemetry spans for connect, snapshot, batch -apply, and checkpoint, so a sink shows up in the same traces as the -application it feeds. - -### Documentation integration - -The per-language client pages in `doc/user/content/integrations/` currently -teach the naive loop. Once the SDK exists, each page leads with the SDK and -keeps the raw `DECLARE`/`FETCH` version as an appendix for driver-only -environments. The durable-subscriptions pattern doc becomes the conceptual -explanation behind the SDK's design, linking to it as the implementation. - -### Demonstrators - -Two example apps ship in `examples/`, because the strongest case for the SDK -is shown, not told: - -- **Always-fresh cache** (single view to Redis), modeled on Justin Bradley's - `mz-sink` demo: one expensive query, three read backends (recompute, - Materialize index, sink-backed cache) switchable live so the win is felt. - The messaging leads with correctness (a cache kept fresh with no invalidation - logic because it is downstream of the change stream), with latency as the - hook. -- **Consistent cohort dashboard** (N views to one consistent view): a dashboard - over several views (e.g. orders, inventory, pricing) where a naive per-view - cache visibly tears and the cohort never does. This is the demo that sells - the cross-view consistency a single-view cache demo structurally cannot show. - -### Materialize-side workstream (planned with the SDK, done properly) - -The SDK must not paper over server gaps with client-side workarounds. Where -the correct behavior needs the database's help, the database change is part -of this program's plan, sequenced so the SDK's stable release builds on real -primitives. Each item below gets its own issue and, where non-trivial, its -own design doc. In sequence: - -1. **Structured error codes** (dedicated SQLSTATEs) for the - compaction-horizon error, dependency-dropped, and dataflow errors. Today - the compaction-horizon failure surfaces as generic `DATA_EXCEPTION` - shared with other errors, and its message text has already changed once - (#34712), which silently broke the one existing client that dispatched - on it. Small, high leverage, and a hard prerequisite for SDK GA: the - typed error taxonomy must key on codes, never on message text. -2. **Docs**: cross-link the durable-subscriptions pattern from every - client-library page. Can land immediately, independent of everything - else. -3. **Subscribe output backpressure** in the coordinator (the existing - unbounded-channel TODO in `active_compute_sink.rs`). Protects - `environmentd` from slow consumers regardless of which client they use. - The SDK's prompt-drain design reduces exposure but only the server can - bound it. -4. **Progress cadence control**, so consumers can trade update granularity - for faster checkpoint advancement and cheaper idle streams. -5. **Non-poisoning subscribe errors** (database-issues#5182), so a transient - dataflow error does not permanently wedge a subscription. -6. **WebSocket `SUBSCRIBE` stabilization**, the gate for the browser SDK. -7. **Server-assisted subscription ownership** (named subscriptions with - single-writer enforcement), contingent on evidence from HA deployments - that the client-side lease is insufficient. - -Items 1 and 2 are cheap and land before or with SDK GA. Items 3 through 5 -improve every subscribe consumer and proceed on their own track with the SDK -as the motivating consumer. Items 6 and 7 are demand-gated. - -## Minimal Viable Prototype - -Prototype = **Python Layer 1 + Layer 2 + the Redis sink, chaos-tested against -the emulator**, living in the SDK monorepo from day one: - -1. Layer 1 client on psycopg with batching, envelopes, typed errors, - reconnect-and-resume. -2. Checkpoints in Redis, `transactional` mode only. -3. The Redis sink at functional parity with mz-redis-sync plus the known - fixes: correct resume statement, `key_violation` handling, orphan sweep on - resync, namespaced checkpoint key, NULL handling. -4. The kill-at-random-points chaos test asserting convergence against - `SELECT ... AS OF`. - -This validates the three riskiest claims early: that the batch/token model is -pleasant to use, that exactly-once state application holds under crash -testing, and that the emulator is a sufficient CI target. The Rust reference -implementation starts from the same spec as soon as the prototype stabilizes -the batch and token model, and the first conformance vectors are extracted -from the prototype's test suite. The Novu integration is then rebuilt on -Layer 2's `at_least_once` mode as the second validation, exercising -idempotency keys and retraction policy with a real side-effecting -destination. - -## Delivery plan - -Phased, each phase with an exit criterion, and no phase starts before the -previous one's criterion passes. - -**Phase 0, foundations.** The SDK monorepo skeleton (`spec/`, -`conformance/`, `python/`, `rust/`, `examples/`), CI wiring, and the -emulator test harness. Spec v0 drafted from this design. In parallel, the -two cheap Materialize-side items land: the structured error code change and -the docs cross-links. - -**Phase 1, the Python MVP.** As described under Minimal Viable Prototype, -Layers 1 and 2 plus the Redis sink. Exit: the chaos suite passes (the -destination equals `SELECT ... AS OF ` under arbitrary kill -schedules), the fencing test passes, and every typed error is reproducible -in integration tests. - -**Phase 2, the spec becomes real.** Conformance vectors extracted from -Phase 1's tests, then the Rust reference implementation built against spec -and vectors. Exit: both implementations pass identical vectors and emulator -scenarios, and spec v1 is published. The Novu-style at-least-once rebuild -validates the event-sink surface in this phase. - -**Phase 3, breadth.** TypeScript, the generic webhook sink (pending open -question 3), the `mz-sink` example runner, and the docs-site integration -that replaces the naive per-language loops. - -**GA gate.** A stable release requires: Phase 2 complete, the -structured-error-code change released in Materialize (with the -version-gated fallback for older servers), and the public docs rewritten on -the SDK. HA failover (the lease design, open question 6) ships in the first -post-GA minor release. - -## Future work - -Deliberately excluded from v1, recorded so reviewers can see the growth path: - -- **Fan-out helper.** One subscription demultiplexed to many in-process - consumers by key (the websocket-server pattern, thousands of clients fed - from one `SUBSCRIBE`). Layer 1's batch model supports it, a first-class - helper makes it a five-line feature. This is the inverse of a cohort - (many-into-one). Cohorts themselves ship in v1 (see "Consistency layering - and cohorts"). -- **Dynamic cohort membership.** Adding, dropping, merging, and splitting the - members of a *live* cohort without re-subscribing, plus the durable - bookkeeping of the resume position a merged or split cohort resumes from. - The fixed-cohort case ships in v1. Live topology changes are the hard, - unvalidated part of `mz-bridge-recipe` and are deferred. -- **Per-key coalescing and debounce.** Event sinks often want "at most one - webhook per key per interval" to absorb flapping. A Layer 2 option with a - max-delay bound, at the cost of intermediate updates, which the diff model - makes safe to drop. -- **Browser SDK** on the WebSocket endpoint once it stabilizes, implementing - the same spec. -- **Additional languages** (Go first, given existing docs coverage) and - additional built-in sinks chosen by demand, with adoption data feeding the - case for native server-side sinks for the top destinations. -- **Serverless guidance.** Long-lived subscriptions do not fit - function-per-request platforms. The pattern doc for bridging (a small - always-on consumer feeding a queue) is docs work once the SDK exists. - -## Alternatives - -**A shared Rust core with native bindings (napi-rs, PyO3).** Single -implementation of the state machine, mechanical bindings. Rejected because -the state machine is the small part. The core would own connection -management, TLS, and auth in three runtimes, integrate with two foreign -async runtimes, and complicate packaging and debugging for every user. The -per-language ecosystem drivers are more battle-tested than anything we would -ship. A spec plus conformance vectors pins down cross-language behavior at a -fraction of the cost, which is how the database-driver ecosystem itself -works. - -**Native server-side sinks (`CREATE SINK ... INTO REDIS`).** The -strongest alternative. It is the best eventual UX for the specific, -high-volume destinations, and this SDK does not preclude it. Rejected as the -*first* move because: each destination becomes a server feature with a -release train, storage/compute team ownership, and years of long-tail -destination requests (the Kafka Connect catalog is hundreds of connectors -deep). The SDK serves the long tail by construction, ships without touching -`environmentd`, and its adoption data tells us which destinations deserve -native sinks. The consistency model (frontier-atomic commits) is the same -one a native sink would implement, so the concepts transfer. - -**Sink to Kafka, use the Kafka Connect ecosystem.** Works today for users -who run Kafka. Rejected as the answer because it taxes every user with a -Kafka deployment plus Connect operational burden to reach a cache, adds a -hop of latency, and loses the direct mapping between destination state and a -Materialize timestamp unless the connector is consistency-aware. It remains -the right answer for organizations already deep in Connect, and the docs -should keep saying so. - -**Docs only, no code.** The durable-subscriptions pattern doc already -describes the protocol. The evidence section above shows what documentation -alone achieves, three implementations, three different correctness bugs, by -authors closer to Materialize than any customer will be. - -**Harden mz-redis-sync as a one-off product.** Fixes one destination in one -language, leaves every other consumer where they are, and keeps the protocol -logic welded to Redis specifics. The layered SDK subsumes it, and `mz-sink` -delivers the same operational artifact. - -## Decisions taken so far - -Settled during initial review, recorded so the remaining questions are -crisp: - -- **Official from day one.** This ships under `MaterializeInc` with a - stability commitment, not as a Labs experiment with a graduation path. -- **All three layers in scope.** The SDK core is the fully supported main - focus. Built-in sinks (Redis first) are supported reference - implementations of the framework. -- **Language order: Rust and Python first, TypeScript next, Go later.** -- **HA is planned, not deferred.** Fencing in v1, lease-based active-passive - failover shortly after. Production use is the assumption. -- **No shortcut workarounds for server gaps.** Where correct behavior needs - the database's help, the server change is planned as part of this program - (see the Materialize-side workstream), sequenced ahead of the SDK - behavior that depends on it. -- **The standalone runner is a reference example**, not a core deliverable. -- **Single-view core, cohort as a first-class generalization.** The consistent - batch of one view is the core primitive. Multi-view consistency ships as a - cohort helper on the same engine (`min(frontier)`), not as the mandatory - mental model. Considered and rejected: re-centering the whole SDK on cohorts - (the demand across all prior art, the server's native unit, and Materialize's - own per-object product surface are all single-view, and the cohort stays - composable on the layering so it is not precluded). -- **One consistency engine, and never backpressure Materialize.** A single - buffer-and-release engine (single-view = cohort-of-one), fed by a continuous - drain into a bounded buffer that fails loud, so the client falls over, never - the server. -- **Idempotency keys never depend on batch grouping.** Batch boundaries are not - stable across a resume, so the delivery key is derived from `mz_timestamp` - plus within-timestamp identity, not a within-batch ordinal. -- **Cohort ships demonstrated, not hardened.** v1 proves the one engine - generalizes to multi-view consistency and bounds the laggard case with a - fail-loud lag budget. Availability hardening (per-member reconnect, timeline - validation, a configurable budget) is deliberately deferred until real - workloads show which of it matters, rather than designed against a cohort no - one runs yet. The core single-view path is where v1 invests. - -## Open questions - -1. **Final package naming.** `subscribe` as the noun - (`@materializeinc/subscribe`, `materialize-subscribe`, `mz-subscribe`) - versus a broader name that leaves room for the SDK to grow into general - client duties later. Day-one stability makes renaming expensive, so this - needs deciding before the first publish. -2. **Guarantee vocabulary.** Are we comfortable publicly branding - `transactional` mode as "exactly-once state application"? It is accurate - for destinations where checkpoint and data commit atomically, but the - term invites misreading as exactly-once side effects. Candidate framing: - "consistent sinks" versus "delivery sinks". -3. **v1 sink surface.** Redis plus the generic webhook sink, or Redis only? - Webhook is cheap on Layer 2 and exercises the second destination family - early, which argues for including it. -4. **Checkpoint store default for event sinks.** Destination-embedded - checkpoints are the opinionated default for keyed-state sinks, but no - destination-embedded store exists for a webhook. Materialize-table store - as the event-sink default? -5. **Conformance gate.** Block releases of any language on the full emulator - chaos suite, or vectors only for patch releases? -6. **HA lease parameters.** Lease duration, heartbeat cadence, and whether - the lease lives only in the checkpoint store or also surfaces in - `mz_internal.mz_subscriptions` for observability. Needs a short design of - its own before the failover release. -7. **Cohort failure and lag policy.** On a member error, tear the whole cohort - down and resume every member at the joint frontier (v1 behavior), or - reconnect only the failed member? And when the lag budget is hit, fail loud - (v1), drop the laggard from the cut, or block the fast members? Both interact - with the HA workstream and want deciding before cohorts are branded stable. -8. **Cohort timeline validation.** Should the SDK detect and reject a cohort - whose members do not share a comparable timeline, or is documenting the - constraint enough for v1? Rejecting needs a way to read an object's timeline, - so it is a small server-side dependency if we want it. From e33028097f41b182dd75577ab19e3c3a4d9a67c0 Mon Sep 17 00:00:00 2001 From: bobbyiliev Date: Wed, 7 Oct 2026 14:16:01 +0300 Subject: [PATCH 05/19] design: Add future work on driver quirks to the Materialize SDK doc --- .../design/20260706_materialize_sdk.md | 26 +++++++++++++++++++ 1 file changed, 26 insertions(+) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 99708700b76ad..0ff251c4841a0 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -672,6 +672,30 @@ and a Redis or Postgres reference sink chosen by demand. GA requires dedicated SQLSTATEs released, the client docs rewritten on the SDK, and both languages passing the same vectors and end-to-end suite. +## Future work + +The name leaves room for modules beyond `subscribe` and `sink`. The candidates +are places where an ordinary driver behaves unexpectedly against Materialize: + +- After its first query, a read transaction is confined to one time domain, and + reading objects outside it fails with SQLSTATE 25000 + (`RelationOutsideTimeDomain` in `src/adapter/src/error.rs`). Drivers and ORMs + that open a transaction for every query hit this. A module could default reads + to autocommit and offer one call that reads several views at one timestamp. +- A transaction becomes write-only after its first write, and some DDL refuses to + run inside a transaction. Typed errors could name the fix for each. +- A cursor cannot outlive its transaction (`WITH HOLD` is rejected in + `src/sql-parser/src/parser.rs`), and only `FETCH ... WITH (timeout = ...)` lets a + cursor loop idle. `subscribe` already handles both. +- Drivers return types they do not know, such as `mz_timestamp`, as text. The + protocol core's value model could serve plain queries too. + +Each module adds maintenance: it tracks server behavior, needs vectors and +end-to-end scenarios, and multiplies across languages. A module will be added +only when support or field evidence shows users hitting the problem, and only +with nightly end-to-end coverage. Plain query execution, DDL management (owned +by mz-deploy and the Terraform provider), and ORM integration stay out of scope. + ## Alternatives ### The product strawman @@ -748,3 +772,5 @@ produce. 10. Repository home: should the packages stay under `misc/` in this repository for good, or move to their own repository once the spec settles, keeping the protocol core, vectors, and end-to-end suite here? +11. Future-work modules: are any of them worth their maintenance cost, and what + evidence should trigger one? From c096fbbea4145c3fc7541aac7fff7313a1228f97 Mon Sep 17 00:00:00 2001 From: bobbyiliev Date: Wed, 7 Oct 2026 14:33:38 +0300 Subject: [PATCH 06/19] design: Release SDK packages like dbt-materialize and build Node from WebAssembly --- .../design/20260706_materialize_sdk.md | 50 +++++++++++++------ 1 file changed, 34 insertions(+), 16 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 0ff251c4841a0..64ba1ab3f7a18 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -543,25 +543,31 @@ async integration in every runtime. A protocol core with no I/O owns none of that. Each language keeps its native driver for the connection, which users already trust, and feeds rows into the protocol core. -Bindings will use PyO3 and maturin (abi3 wheels) for Python, napi-rs for Node, -and WebAssembly for browsers and edge runtimes. The WebAssembly build is how the -browser and function cases, and the Console's own subscribe client, can reuse -the protocol core later. Rows will cross the boundary in one call per `FETCH`, -because a call per row pays a binding crossing per row. The protocol core will -not panic across the boundary, and its errors will become each language's -native exceptions. It will not depend on Materialize workspace crates, so it -builds and versions on its own. - -The costs are a wheel and npm build matrix per platform, Rust stack traces in -Python and Node bug reports, and a source build that needs a Rust toolchain on -unsupported platforms. Go binds through cgo, which is painful, so Go will be -hand-written against the conformance vectors. If native packaging proves too -costly, the fallback is B plus A. +The protocol core will ship in two builds. Python will bind the native build +through PyO3 and maturin (abi3 wheels), and the Rust package will use the +library directly. Node will use the WebAssembly build, the same one browsers and +edge runtimes use, so the Node package is one artifact for every platform. This +repository already publishes Rust compiled to WebAssembly to npm +(`ci/deploy/npm.py`). WebAssembly runs slower than native code, which matters +little here because rows will cross into the protocol core in one call per +`FETCH`, not one call per row. The protocol core will not panic across the +boundary, and its errors will become each language's native exceptions. It will +not depend on Materialize workspace crates, so it builds and versions on its +own. + +The WebAssembly build also covers later languages without native builds. Go can +run it through wazero and JVM languages through Chicory, both written in their +own language, which avoids cgo and JNI. If that proves too slow, the language is +hand-written against the conformance vectors. + +The costs are a Python wheel per platform, Rust stack traces in Python bug +reports, and a source build that needs a Rust toolchain wherever no wheel +exists. If native packaging proves too costly, the fallback is B plus A. ### Generated and hand-written parts The binding glue and each package's type definitions will be generated from the -protocol core. napi-rs emits TypeScript declarations, and the Python package +protocol core. wasm-bindgen emits TypeScript declarations, and the Python package will ship type stubs generated from the PyO3 module. Type mapping lives in the protocol core: it decodes each Materialize type into one documented value model, and each binding converts that model to the language's native types @@ -578,7 +584,8 @@ misc/materialize-sdk/ spec/ behavior spec, versioned conformance/ vectors, shared by every package python/ package: PyO3 binding, psycopg transport, sinks - node/ package: napi-rs binding, node-postgres transport, sinks + node/ package: WebAssembly build, node-postgres transport, sinks + rust/ package: protocol core directly, tokio-postgres transport, sinks test/ mzcompose end-to-end suite, run in the nightlies ``` @@ -587,6 +594,17 @@ Materialize workspace, so its dependencies and lockfile stay independent. A spec change and its fallout in every package then land in one PR, and server changes run against every package in the nightlies. +Releases will work the way `dbt-materialize` releases do. A pull request bumps +the SDK version and merges to `main`. The deploy pipeline, which runs on every +`main` build (`ci/deploy/pipeline.template.yml`), publishes each package whose +version is not yet on its registry, as `ci/deploy/pypi.py` does today. Releases +need no tags. Python wheels will be built on the existing Linux x86-64, Linux +ARM, and macOS ARM Buildkite queues, with macOS x86-64 cross-compiled on the ARM +agent. Windows has no queue, so it gets the source package, which needs a Rust +toolchain to install. Publishing to crates.io needs a new deploy step. Go is the +exception to the pattern: Go modules are versioned by git tags, so a Go package +needs tags in this repository or a repository of its own. + Each release will publish every package at the same version, built from the same protocol core, so a version number means the same behavior in every language. Versions follow semantic versioning. Resume tokens and checkpoints From 4e147d558dfe5070ffd058b08925383febf83fe8 Mon Sep 17 00:00:00 2001 From: bobbyiliev Date: Wed, 7 Oct 2026 14:40:36 +0300 Subject: [PATCH 07/19] design: Sharpen the object swap open question --- doc/developer/design/20260706_materialize_sdk.md | 7 +++++-- 1 file changed, 5 insertions(+), 2 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 64ba1ab3f7a18..95e7bf2424e11 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -781,8 +781,11 @@ produce. 5. Retention margin check: can a sink role read the catalog state it needs (the object's readable frontier and the cluster's indexes) without extra grants? 6. `mzq`: a transport under this SDK, or a separate path? -7. Object swaps: how often do blue/green deploys swap a subscribed view, and is - `refollow` the behavior deploy tooling wants? +7. Object swaps: when a blue/green deploy swaps a view a sink reads, should the + sink follow the new view (`refollow`), re-snapshot the target, or stop for an + operator? The new view's history starts at its creation, so `refollow` works + only if that history covers the sink's checkpoint. Can deploy tooling + guarantee that, for example with `RETAIN HISTORY` on staged views? 8. Egress cost: data leaves through `environmentd`. What does a sink with a large snapshot cost a customer, and does the docs story need a sizing page? 9. Stateless workers before durable subscriptions: is there demand that cannot From 9b8fde3a0a5ed32c6f6622391d3f8fbbb48f278f Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Thu, 8 Oct 2026 17:37:35 +0300 Subject: [PATCH 08/19] design: Address review on limits, errors, turbopuffer tombstones, and durable attach --- .../design/20260706_materialize_sdk.md | 467 +++++++++++++----- 1 file changed, 344 insertions(+), 123 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 95e7bf2424e11..ddc180dc0f81e 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -66,9 +66,10 @@ jumps to the next refresh, so a sink that is down across a refresh can lose history however short the downtime. Retention belongs to the view's owner, and any `ALTER` changes every sink's margin without telling it. -A connected subscribe holds its input readable up to what it has emitted. Only -a disconnect puts history at risk, and the consumer then resumes from its -committed frontier, which can trail what it received. +A connected subscribe holds its input readable up to what it has emitted, not up +to what the consumer has committed. The data between those two points needs +retention even while the consumer is connected, because after a disconnect the +consumer resumes from its committed frontier. ### Indexed targets @@ -110,8 +111,17 @@ progress row is never closed by the server. When a subscribe's backlog in `environmentd` exceeds `subscribe_max_buffered_bytes` (default 128 MiB), the subscribe is retired with -`SubscribeFellBehind`, SQLSTATE 53200 (#37905). A single result larger than -`max_result_size` also ends the subscription. +`SubscribeFellBehind` (#37905). + +`environmentd` also collects every update between two progress messages before +passing them on, and fails the subscribe with "total result exceeds max size" +once that collection exceeds `max_result_size`, a system setting with a default +of 1 GB (`PendingSubscribe::stash` in `src/compute-client/src/service.rs:537-556`, +which also runs in each `clusterd` process). +An initial snapshot is one timestamp, so the whole snapshot must fit. A resume +after long downtime can hit the same limit when the frontier advances in one +step. The failure is deterministic, so a retry repeats it, and nothing the client +does with chunks can avoid it. ### Errors @@ -119,12 +129,21 @@ When a subscribe's backlog in `environmentd` exceeds | --- | --- | --- | | `AS OF` below the readable frontier | `22000` (`DATA_EXCEPTION`, shared) | "could not find a valid timestamp for the query" | | Subscribed object or cluster dropped | `42704` | "relation 'x' was dropped" and similar | -| Client fell behind | `53200` | `SubscribeFellBehind` | -| Dataflow error | varies | the evaluation error, repeated on every retry | - -The history-loss message has changed once already (#34712), which broke the -one client that matched its text. A dataflow error repeats until the data or -the view changes (database-issues#5182), so retrying cannot fix it. +| Client fell behind | `53200` (shared with the adapter's result-size error) | `SubscribeFellBehind` | +| Result over `max_result_size` | `XX000` (`INTERNAL_ERROR`, shared) | "total result exceeds max size of ..." | +| Dataflow error | `XX000` (`INTERNAL_ERROR`, shared) | the evaluation error, repeated on every retry | + +An error raised inside a running subscribe reaches the client as an +unstructured adapter error, which maps to `XX000` +(`src/adapter/src/active_compute_sink.rs:283`, `src/adapter/src/error.rs:1066`). +So a dataflow error, a result-size failure, and a genuine internal error share +one code. `53200` covers both `SubscribeFellBehind` and the adapter's +result-size error (`error.rs:1017-1018`), but a subscribe's size error takes the +`XX000` path, so during a subscribe `53200` means the client fell behind. The +history-loss message has +changed once already (#34712), which broke the one client that matched its +text. A dataflow error repeats until the data or the view changes +(database-issues#5182), so retrying cannot fix it. ## Success Criteria @@ -145,10 +164,16 @@ this document adds. | R10 | Several languages with identical behavior, checked by machines | Design | | R11 | Sink authors can test without a live environment | Design | +R4 is met in full by targets that can write the whole cut atomically, such as +one Postgres transaction. A target split into parts with no transaction across +them, such as several turbopuffer namespaces, gets R4 per part only (see +"turbopuffer sink"). + A user following the happy path cannot commit an unclosed timestamp, cannot get the `AS OF` arithmetic wrong, and cannot silently lose or duplicate updates -across a restart. The check is a sink that converges to `SELECT ... AS OF` its -committed frontier after being killed at random points. +across a restart. The check is a sink that, after being killed at random points, +holds exactly `SELECT ... AS OF F - 1`, where `F` is its committed frontier. +The frontier is exclusive, so `AS OF F` would also include changes at `F`. ## Out of Scope @@ -159,10 +184,13 @@ committed frontier after being killed at random points. - A general Materialize client. Drivers already run queries well. See "Naming". - Scale-out of one subscription across parallel workers, which has no server support. -- Stateless workers (functions, lambdas). They need the server to own the - resume point, which durable subscriptions (#38468) provide. +- Stateless workers (functions, lambdas). They need history kept across idle + gaps without a hand-sized `RETAIN HISTORY` window, which durable subscriptions + (#38468) provide. - Private-preview subscribe features (`ENVELOPE DEBEZIUM`, - `WITHIN TIMESTAMP ORDER BY`), which are behind flags. + `WITHIN TIMESTAMP ORDER BY`), which are behind flags. Durable subscriptions + support both envelopes across a resume, so `ENVELOPE DEBEZIUM` can move in once + its flag lifts. ## Solution Proposal @@ -172,8 +200,9 @@ package will use that language's own database driver for the connection and will expose two modules: `subscribe` for live and durable consumption, and `sink` for writing to targets. Every change will keep its `mz_timestamp`. Durable consumption will store its checkpoint in the target, fenced by an epoch, -and will read one object plus an optional projection, filter, and envelope. The -first sink will be turbopuffer. The spec, conformance vectors, and an end-to-end +and will read one storage collection (a table, materialized view, or source) +plus an optional projection, filter, and envelope. The first sink will be +turbopuffer. The spec, conformance vectors, and an end-to-end suite will live in this repository and run in the nightlies. ### Naming @@ -268,24 +297,30 @@ sequenceDiagram ### Subscription scope -A durable subscription, one that checkpoints and resumes, will read one object -plus an optional projection, filter, and envelope. Arbitrary SQL will be -accepted only for live streams that never resume. +A durable subscription, one that checkpoints and resumes, will read one storage +collection (a table, materialized view, or source) plus an optional projection, +filter, and envelope. The filter must not call `mz_now()`. Plain views, indexes +as targets, and temporal filters will be rejected. Arbitrary SQL will be accepted +only for live streams that never resume. Resuming a general query rehydrates its whole dataflow and needs history on -every input of the query. Durable subscriptions (#38468) accept exactly this -narrower surface, so starting narrow makes the later move a transport change, -not an API break. Widening later is compatible, and narrowing is not. -Structured input also lets the SDK build the statement and place `ENVELOPE` -before `WITH` without parsing user SQL. +every input of the query. A plain view is inlined, so resuming it costs the +same, and a temporal filter needs the snapshot on every resume. Durable +subscriptions (#38468) accept exactly this narrower surface, so starting with it +makes the later move a transport change, not an API break. Widening later is +compatible, and narrowing is not. Structured input also lets the SDK build the +statement and place `ENVELOPE` before `WITH` without parsing user SQL. Projections and filters still read the snapshot on resume (see "Snapshot -elision"). The docs will recommend the plain object form for large views, and -the SDK will log the resume cost at startup. +elision"). The docs will recommend the plain collection form for large +collections, and the SDK will log the resume cost at startup. The SDK will refuse an indexed object at startup. It will check whether the subscribed object has an index on the subscribing cluster and fail with an error that names the remedy: a dedicated subscribe cluster with no index on the object. +An index created later makes the next resume fail with the timestamp-selection +error, so the SDK will check again before it treats that error as history loss. +Durable subscriptions attach to storage, so this check applies only before them. ### Live streams @@ -295,24 +330,38 @@ behind, the client will fail with a typed error. Pausing the fetch loop instead would push the backlog into `environmentd` until the server retires the subscribe (R8). +Durable consumption will use the same fetch loop, so a slow target write fills +the client buffer, not `environmentd`. When the client buffer fills, or the +server retires the subscribe with `FellBehind`, the SDK will resume from the last +commit with exponential backoff and report a metric. Each cycle restarts a +dataflow, so a target that stays slower than the stream needs a larger buffer +or a longer `commit_interval`. + The initial snapshot is one timestamp and can exceed client memory. It will -arrive as chunks marked partial, with the token on the closing chunk. The sink +arrive as chunks marked partial, with the token on the closing chunk. Chunks +bound client memory only. A snapshot or catch-up larger than `max_result_size` +fails on the server before the first chunk arrives (see "Buffering limits"), and +the SDK will report it as `ResultTooLarge` with the remedies: subscribe to a +narrower projection, or have an administrator raise `max_result_size`. Server-side +chunking would remove the limit (see "Materialize-side workstream"). The sink module's generations (see "Sink module") give atomic snapshot visibility to targets that need it. A frontier advance with no data will yield an empty batch with a fresh token, so checkpoints keep moving through quiet hours. -`UP TO` will be supported. Until SQL-528 is fixed, the SDK will release -everything below `UP TO` when a bounded stream ends without error, so no closed -data is lost. +`UP TO` will be supported. On server versions without the SQL-528 fix, the SDK +will release everything below `UP TO` when a bounded stream ends without error, +so no closed data is lost. Gating on the server version switches the workaround +off once the fix ships. ### Checkpoints and fencing A checkpoint store will have `load(name)` and `commit(name, epoch, frontier)`. The target is the recommended store, because writing data and checkpoint in one transaction gives exactly-once state (R1, R5). A Materialize-table store and -a local file store will ship for targets with no transaction. +a local file store will ship for targets with no transaction. Neither commits +atomically with the target, so both give `at_least_once` only. Each worker start will take the next epoch. A commit will be conditional on the stored epoch not being newer, and a stale worker will get a typed `Fenced` @@ -325,10 +374,15 @@ deployments one checkpoint, and they would fence each other. Rows from several views will be tagged by view name, because tagging by list position remaps rows when the list is reordered. -The checkpoint will also record a fingerprint of the subscription: the object, -projection, filter, envelope, and output column types. A resume whose -fingerprint differs fails with `SchemaMismatch` instead of mixing rows of two -shapes in one target. +The checkpoint will also record a fingerprint of the subscription: the object's +name, catalog id, and storage shard (from `mz_internal.mz_storage_shards`, joined +through `mz_internal.mz_object_global_ids`), and +the projection, filter, envelope, and output column types. A resume compares it +with the object the name resolves to now, and "Object identity" decides what +each difference means. The shard matters because blue/green deploys use +`ALTER SCHEMA ... SWAP`, which moves names and leaves ids unchanged. After a swap +the same name points at a different object with the same columns, and without +the shard a reconnect would switch objects silently. ### Delivery guarantees @@ -361,13 +415,47 @@ consumers). Detection will use the error type (R9). At startup and periodically, the SDK will compare the checkpoint to the object's readable frontier and report the margin as a metric, with a warning -below a configured threshold. Retention has to outlast an operator's response -to a stall, not only the retry budget. - -Blue/green deploys recreate the subscribed object, which ends the stream with a -dropped-object error. A `refollow` policy will re-resolve the name and resume if -the new object's fingerprint and retained history allow it, and otherwise fall -back to the history-loss policy. +below a configured threshold. A connected subscribe holds history only up to +what it has emitted, not up to what the sink has committed (see "Retention +window"). The margin therefore has to cover the commit lag (the client buffer, +`commit_interval`, and the target's write time) as well as the retry budget and +an operator's response to a stall. An object with the default one-second window +fails every resume, so the SDK will refuse to start durable consumption when the +object's retention is below the commit lag. + +Under durable subscriptions the margin becomes the time since the last +acknowledgement against the subscription's `ACKNOWLEDGE WITHIN` deadline. +Acknowledgements advance only when the frontier does. A `REFRESH EVERY` view +between refreshes, a paused source, or a cluster with no replicas sends no +progress, and #38468 treats an acknowledgement at the current position as a +no-op, so the deadline keeps running. The deadline must therefore exceed the +retry budget, an operator's response to a stall, and the longest time the +object's frontier can stand still. The SDK will warn at startup when it does +not. + +### Object identity + +The checkpoint fingerprint lets the SDK tell three kinds of change apart: + +| Change | Example | Default | With `refollow` | +| --- | --- | --- | --- | +| Different output columns | the view's definition changed | stop with `SchemaMismatch` | stop | +| Same columns, same storage shard, new catalog id | `ALTER MATERIALIZED VIEW ... APPLY REPLACEMENT` | resume | resume | +| Same columns, new storage shard | `ALTER SCHEMA ... SWAP` in a blue/green deploy | stop | resume if the new object's history covers the checkpoint, else the history-loss policy | +| The name resolves to nothing | the object was dropped while the sink was offline | stop with `ObjectDropped` | stop | + +The SDK will look the name up in the catalog before each `SUBSCRIBE`. Without +that lookup, a name that resolves to nothing fails planning with `XX000`, which +would be misread as `StreamPoisoned` (`src/adapter/src/error.rs:1003`). `42704` +arrives only for a drop during a running stream. + +A replacement keeps the storage shard and gives the view a new catalog id +(`src/adapter/src/catalog/transact.rs:1334`), so its history is continuous. A +name swap moves the name to a different object. The running stream keeps +reading the old object until the deploy drops it, which ends the stream with +`ObjectDropped` (`42704`) and sends it through the same table. Different columns +always stop, because resuming would mix rows of two shapes in one target, even +when the history-loss policy is `resnapshot`. On a signal, the SDK will finish the in-flight batch, commit, close the cursor, and disconnect. @@ -391,25 +479,66 @@ Each view's `RETAIN HISTORY` counts from its own upper, so a lagging view can leave the stored cut readable on one view and compacted on another. The retention-margin check will run per member. +Members must share the default `EpochMilliseconds` timeline, because timestamps +from different timelines are not comparable and the cut would be unsound. Only +sources with `ENVELOPE MATERIALIZE`, which is behind a flag, get another +timeline (`src/sql/src/plan/statement/ddl.rs:1000-1004`), and objects built on +them inherit it. The SDK will reject a member that is such a source or depends +on one, using `mz_sources.envelope_type` and +`mz_internal.mz_object_transitive_dependencies`. + A multi-view subscription will hold one connection per member. Any member error will end all members, and recovery will resume every member at the stored cut. +If one member's history no longer covers the cut and the history-loss policy is +`resnapshot`, the SDK will follow the recipe in #38468. It will re-snapshot each +expired member at a new timestamp `t_i` while the others resume from the cut, +re-establish the cut at `t*`, the largest `t_i`, and buffer every stream until +its progress passes `t*`. It will then apply in one step each re-snapshotted +member's snapshot, replacing its rows through a generation sweep, and every +member's changes through `t*`. The cost is buffering. The live members must hold +every change from the old cut to `t*`, which spans the whole outage, and the +re-snapshotted members hold their changes between `t_i` and `t*`. Targets with +generations will stage these changes in the target and make them visible at +`t*`, and other targets buffer them in memory. The catch-up can still hit +`max_result_size` (see "Buffering limits"). ### Typed errors | Error | Detected by | Default | | --- | --- | --- | -| `Transient` | no SQLSTATE (connection-level) | reconnect and resume | -| `FellBehind` | `53200` | resume from the last commit, report a metric | -| `HistoryLost` | `22000` plus the timestamp-selection message | history-loss policy | -| `ObjectDropped` | `42704` | `refollow` policy, else stop | -| `StreamPoisoned` | dataflow error | stop | -| `SchemaMismatch` | checkpoint fingerprint | history-loss policy | +| `Transient` | no SQLSTATE, `08000`, `08001`, `08003`, `08004`, `08006` | reconnect through the connection factory and resume, within the retry budget, then stall | +| `CredentialsExpired` | `28000` | one reconnect through the connection factory, then `Fatal` | +| `Canceled` | `57014` (a cancel or a timeout) | stop | +| `FellBehind` | `53200` | resume from the last commit with backoff, report a metric | +| `ResultTooLarge` | `XX000` plus the "exceeds max size" message | stop with remedy | +| `HistoryLost` | `22000` plus the timestamp-selection message, after the index re-check | history-loss policy | +| `ObjectDropped` | `42704` | see "Object identity" | +| `SchemaMismatch` | checkpoint fingerprint | see "Object identity" | +| `StreamPoisoned` | `XX000` with any other message, which includes genuine internal errors | stop | | `Fenced` | checkpoint commit | stop | -| `IndexedTarget` | startup check | stop with remedy | +| `IndexedTarget` | startup check, and the re-check before `HistoryLost` | stop with remedy | | `Fatal` | anything else (auth, TLS, SQL) | stop | -The `HistoryLost` text match will be the only message matching in the SDK, kept -in one function and gated on server version until a dedicated SQLSTATE exists. +Reconnects go through the connection factory. balancerd refuses connections +with `08004` while `environmentd` restarts (`src/balancerd/src/lib.rs:964-1004`), +so that is retried. The same code also means an unsupported protocol version, +which is why every reconnect counts against the retry budget and a deterministic +failure stalls instead of looping. `08P01` protocol errors are `Fatal`. A session +whose credentials expired ends with `28000` "authentication expired" +(`src/pgwire/src/protocol.rs:667`). Bad credentials share that code, so the SDK +reconnects once with credentials from the factory and treats a second `28000` +as `Fatal`. `57014` covers an operator's cancel and statement timeouts, and +retrying it would make a sink impossible to stop. During +a subscribe, `53200` means `FellBehind`, because the subscribe's own size error +takes the `XX000` path. Until dedicated SQLSTATEs exist, two errors need message +text: `HistoryLost` and `ResultTooLarge`. All message matching will live in one +function, gated on server version, with conformance vectors pinned to each +message. The workstream asks for dedicated codes so this can go away. + +Durable subscriptions add three failures: attaching to an expired subscription +and an `AS OF` outside its window both map to `HistoryLost`, and an attach by a +newer reader, which ends the older reader's stream, maps to `Fenced`. Their +detection waits on the error codes #38468 settles on. ### Connections @@ -433,10 +562,17 @@ snapshot, resume, retry, fencing, and the history-loss policy. Keyed-state targets (caches, search indexes, tables) will read the upsert envelope. When they re-snapshot, the SDK will write the snapshot under a new generation and, at the end, call a sweep that removes keys from older -generations. That removes keys deleted while the sink was offline, which a plain +generations. The generation is the snapshot's `AS OF`, and writes after the +snapshot carry it too. It grows across attempts, so a re-snapshot that crashed +and was retried never reuses a generation, and keys only the failed attempt +wrote are swept. That removes keys deleted while the sink was offline, which a plain re-snapshot leaves behind. A target that needs the snapshot to appear at once will make the new generation visible only at the sweep. +The upsert envelope emits `key_violation` rows when a key has more than one +value. The SDK will pass them to `reject` by default, and a sink can choose +`stall` instead, so a key violation never crashes the process. + Event targets (webhooks, queues, notifications) will read the diff envelope under `at_least_once`, with the idempotency key above. They will declare what a retraction means: ignore it, or call a compensating function. Compensation will @@ -448,69 +584,135 @@ later retraction. The first sink will keep turbopuffer namespaces equal to views. It answers the search-index use case and will run against our internal context graph, which -has no Kafka, so the existing Kafka-based sink does not fit there. +has no Kafka, so the existing Kafka-based sink does not fit there. The sink will +write only to namespaces it creates, so every document carries `mz_timestamp`. turbopuffer's documentation states that one write request to one namespace is applied atomically and is durable on return, and that there are no transactions -across namespaces. Conditional writes compare the stored document with the -incoming one (`$ref_new`) and silently skip rows whose condition fails. An -upsert to a document that does not exist is applied unconditionally, and a -delete's condition sees `null` for every `$ref_new` attribute. - -Conditional writes alone therefore do not make a replay safe. If a later batch +across namespaces. Conditional writes compare each stored document with the +incoming one (`$ref_new`) and silently skip a row whose condition fails. For +patches, the response counts only the rows whose condition held. An upsert to a +document that does not exist is applied unconditionally, a patch to one is +skipped, and in a namespace with vector attributes every upserted document must +carry every vector. + +Plain deletes therefore cannot make writes safe to repeat. If a later batch deleted a key, replaying an earlier upsert of that key finds no document and -recreates it. The sink will instead write a checkpoint document into each -namespace in the same request as that namespace's data, so a namespace's data -and its frontier commit atomically. On restart the sink will resume from the -lowest checkpoint across its namespaces, and for each namespace it will drop -every change below that namespace's own checkpoint. Batch boundaries move across -a resume, so the comparison is per change, using its `mz_timestamp`. That gives -each namespace exactly-once state. - -Upserts will also be conditional on the stored `mz_timestamp` being older, and -deletes on it being below the batch frontier, passed as a literal because a -delete's `$ref_new` values are `null`. A stale worker that slips past fencing -writes the same rows at the same timestamps as the live worker, because -Materialize output is deterministic per timestamp. The condition only has to -stop its older rows from replacing newer ones. - -Writes across namespaces are not atomic. Between the writes of one cut, a reader -can see one namespace at the new cut and another at the old one. Each namespace -converges to every cut, and a reader that needs a joint view filters on -`mz_timestamp`. The sink's docs will state this. +recreates it. A stale worker can do the same before it learns it was fenced, +because conditions are evaluated per document, so its data writes cannot be +conditioned on the epoch in a checkpoint. + +The sink will write tombstones in place of deletes. A delete will upsert the +key's document with `deleted = true`, the delete's `mz_timestamp`, a wall-clock +`written_at`, and a placeholder vector where the namespace requires one. Every +upsert, of data or of a tombstone, will be conditional on the stored +`mz_timestamp` being older than the new one. A replayed or stale write of an +older change then meets a newer document or tombstone and is skipped. Before each +request, the sink will net its changes to the latest change per key, because a +batch spans timestamps and turbopuffer rejects a request that names one id +twice. The conditions only ever let a newer timestamp win, so dropping the older +change from a request is safe. Every write is then +safe to repeat and ordered by timestamp, so each namespace reaches exactly-once +state under `at_least_once` delivery, and a batch can be split across requests. + +A re-snapshot's generation sweep will turn documents from older generations into +tombstones instead of deleting them, so the same protection holds after a +re-snapshot. A patch by filter has no `$ref_new`, so it will set constants: +`deleted = true`, `mz_timestamp` to the snapshot's `AS OF` `t_s`, and +`written_at` to the current time, on documents matching +`generation < t_s AND deleted = false`. The patch re-evaluates its filter before +applying (turbopuffer's guarantees page), so a document a live write moves into +the new generation meanwhile is left alone. + +Searches will filter on `deleted = false`. A sweep will remove a tombstone only +when its `written_at` is older than a grace period and its `mz_timestamp` is +below the sink's committed frontier. The grace period runs from when the +tombstone was written, not from its `mz_timestamp`, because a sink that is +catching up writes tombstones for old timestamps. It has to exceed how long a +fenced worker can keep writing before its next checkpoint commit fails: +`commit_interval`, the retry budget, and the longest pause a sink process can +survive. A worker paused for longer could still recreate a swept key, so the +grace period bounds that risk without removing it. Filter operations are capped +per call (5 million rows for a delete by filter, 50 thousand for a patch by +filter), so both sweeps loop. + +The checkpoint will live in a separate checkpoint namespace, one document per +sink, created once when the sink is set up. Workers will change it only with +patches, which never create a document, so a stale worker cannot recreate a +deleted checkpoint. A worker will take the next epoch with a patch conditional +on the stored epoch being strictly lower than its own, so of two workers starting +at once only one wins. It will commit a frontier with a patch conditional on the +stored epoch being equal to its own, and a patch count of zero means it was +fenced. The data namespaces hold no reserved documents. + +Writes across namespaces are not atomic, and turbopuffer keeps one version of +each document. Between the writes of one cut, a reader can see one namespace at +the new cut and another at the old one. Filtering on `mz_timestamp` cannot +rebuild an earlier cut, because an upsert replaces the earlier version. Each +namespace converges to every cut, but the sink does not give a consistent view +across namespaces, so for turbopuffer R4 holds per namespace only. The sink's +docs will state this, and open question 12 asks whether that is enough. Embedding cost will follow the existing sink's transform model: a transform declares the columns it reads, and runs only for rows where those columns -changed. +changed. The sink will store a hash of each transform's source +columns on the document. For each netted change it will first make a patch of +the attributes without vectors, conditional on the stored `mz_timestamp` being +older, every stored source hash being equal to the new one, and the document not +being a tombstone. The patch keeps the stored vectors, which are then known to +match the source columns. If the patch count shows it was not applied, because a +source column changed, the document is missing or a tombstone, or a newer +version exists, the sink will compute the document's vectors and make the +conditional upsert. Comparing with the stored hash, not with the previous change +in the stream, stays correct when netting drops intermediate changes and after a +replay. Tombstones skip transforms. A replay re-runs transforms for the replayed rows, +which costs embedding calls but does not affect correctness. ### Durable subscriptions Durable subscriptions (#38468) give the server a per-consumer hold advanced by -`ACKNOWLEDGE`, with a wall-clock deadline in place of a window measured from the -upper. When they land, the SDK will attach with -`SUBSCRIBE USING DURABLE SUBSCRIPTION` in place of `AS OF F - 1`, and will -acknowledge only after the target commit. The target-side checkpoint will stay -mandatory, because the server alone is at-least-once. `START AT` will migrate an -existing sink at its stored frontier without a gap. A consumer that resumes by -filtering on its own position will check the opening progress message, because -recreating the subscription can move the start past that position without an -error. Stateless workers become possible at that point, because the server owns -the resume point. - -Durable subscriptions will remove the retention sizing problem, the -`AS OF F - 1` arithmetic on the default path, and most of the retention-margin -check, because the server holds history for each consumer. The rest of the -protocol core stays. Rows still need decoding and release at progress, the -server can resend data the target already committed, a multi-view cut still -spans several subscriptions, and fencing, retry, and dead-lettering do not -depend on where history lives. +`ACKNOWLEDGE`, with a wall-clock deadline (`ACKNOWLEDGE WITHIN`) in place of a +window measured from the upper. When they land, the SDK will attach with +`SUBSCRIBE USING DURABLE SUBSCRIPTION` and still pass `AS OF F - 1`, positioning +the read at its own checkpoint. If a recreate or `RESET` has moved the +subscription's hold past the checkpoint, that attach fails loudly. Resuming from +the server's position and filtering below the checkpoint would hide the same gap +whenever the opening progress check is missed. The SDK will acknowledge only +after the target commit, including for empty batches. The target-side checkpoint +will stay mandatory, because the server alone is at-least-once. `START AT` will +migrate an existing sink at its stored frontier without a gap. + +Subscriptions are provisioned per logical consumer, because creating one is a +catalog transaction (#38468). The SDK will attach to an existing subscription by +name and will not create one on start. Creating it, and choosing its deadline, +is a deploy step. + +A blue/green cutover needs a new subscription on the new object, created +`START AT` the sink's checkpoint, which works only if the new object's history +covers it. Dropping the old object needs `CASCADE` while the old subscription +exists. Deploy tooling will own both steps, and `refollow` then attaches to the +new subscription. + +Durable subscriptions will remove the retention sizing problem and turn the +retention-margin check into a deadline check (see "History loss and retention +margin"). Stateless workers become possible, because the server holds history +for each consumer. `IndexedTarget` should no longer apply if the durable attach +reads storage, which is still an open question in #38468. The rest of the +protocol core stays. Rows still need decoding and +release at progress, the server can resend data the target already committed, a +multi-view cut still spans several subscriptions, and fencing, retry, and +dead-lettering do not depend on where history lives. ### Security The SDK will hold Materialize credentials and target credentials in the user's process. The docs will recommend a dedicated role with `SELECT` on the object and -`USAGE` on the subscribe cluster, and nothing else. The protocol core performs no -network I/O, so it adds no network surface of its own. +`USAGE` on the subscribe cluster. The Materialize-table checkpoint store adds +`SELECT`, `INSERT`, and `UPDATE` on its checkpoint table, and durable +subscriptions add the privileges #38468 defines for attaching and acknowledging. +Creating a subscription needs more, which is one more reason it is a deploy step +and not something the SDK does on start. The protocol core performs no network +I/O, so it adds no network surface of its own. A resume token is not secret but is also not authenticated. Anyone who can write the checkpoint can move the frontier forward and make the sink skip data, @@ -572,7 +774,8 @@ will ship type stubs generated from the PyO3 module. Type mapping lives in the protocol core: it decodes each Materialize type into one documented value model, and each binding converts that model to the language's native types (for example `numeric` to `Decimal` in Python), with the conversions covered by -the vectors. Each package will hand-write only the transport over its driver, +the vectors. `mz_timestamp` is a u64, so the Node package will expose it as a +`BigInt`, which holds every u64 value where a JavaScript `number` does not. Each package will hand-write only the transport over its driver, the idiomatic API surface (iterators in Python, async iterators in Node), and the sink modules. @@ -626,14 +829,19 @@ The conformance suite is the contract, whatever the generation strategy. One runner will drive a thin adapter per language over stdin and stdout. Scenarios: snapshot then stream, resume after a kill, resume after a server restart, history loss, dropped object, poisoned dataflow, fell-behind, - indexed object refused, bounded stream. A server change that breaks the SDK, + result over `max_result_size`, indexed object refused at startup and after an + index appears mid-run, a name swap, a replacement materialized view, and a + bounded stream. A server change that + breaks the SDK, such as a changed error message, then breaks a nightly before it reaches a customer. 3. Convergence test: kill a sink at random points (mid-batch, between write and commit, mid-snapshot), restart, repeat, then compare the target with - `SELECT ... AS OF` the committed frontier, for both guarantees. + `SELECT ... AS OF F - 1` for the committed frontier `F`, for both guarantees. 4. Fencing test: two workers share one checkpoint, the stale one is refused, and - the target stays consistent. + the target stays consistent. For turbopuffer, the stale worker also writes + an older upsert of a key the live worker has deleted, and the key must stay + deleted. 5. User test kit: the recorded-stream player from the vectors will ship to users, so sink authors can test their targets offline (R11). @@ -654,15 +862,24 @@ snapshot, apply, and commit. Where correct behavior needs the database, the database change is part of this program. -1. Dedicated SQLSTATEs for history loss and for dataflow errors. +1. Dedicated SQLSTATEs for history loss, dataflow errors, a subscribe that fell + behind, and a result over `max_result_size`. Today these share codes with + other errors (see "Errors"). 2. SQL-528: a final progress row clamped to `UP TO`. -3. Durable subscriptions (#38468). -4. Snapshot elision for a projection and filter without a temporal predicate, +3. Server-side chunking: pass the updates between two progress messages on in + bounded pieces, so a snapshot or catch-up larger than `max_result_size` can be + delivered. Durable subscriptions do not change this, because a snapshot is one + timestamp. +4. In #38468, a way for a sink on a slow-moving object to keep its subscription + alive while healthy, for example an acknowledgement at the current position + that refreshes the `ACKNOWLEDGE WITHIN` deadline. +5. Durable subscriptions (#38468). +6. Snapshot elision for a projection and filter without a temporal predicate, which #38468 also needs. -5. Non-poisoning subscribe errors (database-issues#5182). -6. Docs that cross-link the durable-subscriptions pattern from every client page +7. Non-poisoning subscribe errors (database-issues#5182). +8. Docs that cross-link the durable-subscriptions pattern from every client page now, and lead with the SDK once it ships. -7. A stable WebSocket `SUBSCRIBE`, the gate for browser and function transports. +9. A stable WebSocket `SUBSCRIBE`, the gate for browser and function transports. The server-side buffering bound (#37905) has landed and needs no further work. @@ -674,8 +891,8 @@ graph. It tests the three riskiest claims: - that a protocol core with no I/O binds into a language package without packaging or performance problems, -- that per-namespace checkpoint documents give each turbopuffer namespace - exactly-once state under random kills, +- that tombstones and conditional writes give each turbopuffer namespace + exactly-once state under random kills and a stale worker, - that the nightly end-to-end suite catches server changes that break the SDK. ## Delivery plan @@ -732,7 +949,7 @@ right, and this design keeps it. The parts that change: | `name` defaults to the class name, rows keyed by list index | Required names, rows tagged by view name | Shared checkpoints fence each other, reordering remaps rows | | `commit_interval` in progress messages | Minimum time between commits | Progress cadence is not a user-facing unit | | Snapshot handling left to the user | Partial snapshot chunks, generation sweep, history-loss policy | Large snapshots and orphan keys are the common failure | -| Conditional turbopuffer writes for replay safety | A checkpoint document per namespace, written with the data | A replayed upsert recreates a key a later batch deleted | +| Conditional turbopuffer writes for replay safety | Tombstones plus conditional writes, with the checkpoint in its own namespace | A replayed or stale upsert recreates a key a later batch deleted | ### Build only on durable subscriptions @@ -774,18 +991,19 @@ produce. package. 3. Generation strategy: is a protocol core with bindings acceptable for packaging and support, or do we start with B plus A? -4. turbopuffer checkpoint document: the per-namespace checkpoint document is what - makes replay safe, but every query must exclude it. The alternative is - tombstones in place of deletes, swept later, with the checkpoint in a - Materialize table. Which cost is easier for users? +4. turbopuffer tombstones: searches must filter on `deleted = false`, tombstones + need a placeholder vector in namespaces with vectors, and a sweep removes them + after a grace period measured from when they were written. What default grace + period is safe, and is a placeholder vector acceptable to users? 5. Retention margin check: can a sink role read the catalog state it needs (the object's readable frontier and the cluster's indexes) without extra grants? 6. `mzq`: a transport under this SDK, or a separate path? -7. Object swaps: when a blue/green deploy swaps a view a sink reads, should the - sink follow the new view (`refollow`), re-snapshot the target, or stop for an - operator? The new view's history starts at its creation, so `refollow` works - only if that history covers the sink's checkpoint. Can deploy tooling - guarantee that, for example with `RETAIN HISTORY` on staged views? +7. Object swaps: should `refollow` be the default for a blue/green name swap? + The new view's history starts at its creation, so `refollow` works only if + that history covers the sink's checkpoint. Can deploy tooling guarantee that, + for example with `RETAIN HISTORY` on staged views? For a replacement + materialized view, does a running subscribe end at the switch, and does the + new catalog id's readable frontier cover a checkpoint taken before it? 8. Egress cost: data leaves through `environmentd`. What does a sink with a large snapshot cost a customer, and does the docs story need a sizing page? 9. Stateless workers before durable subscriptions: is there demand that cannot @@ -795,3 +1013,6 @@ produce. the protocol core, vectors, and end-to-end suite here? 11. Future-work modules: are any of them worth their maintenance cost, and what evidence should trigger one? +12. turbopuffer across namespaces: is per-namespace convergence enough for + search, or do readers need a consistent view across namespaces? That would + need versioned documents and a visible-cut pointer that readers filter on. From 7a8343c1b44225db28acbc7ce17d7bd2b1f16af6 Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Thu, 8 Oct 2026 17:54:17 +0300 Subject: [PATCH 09/19] design: Make Rust the first SDK package and write the turbopuffer sink in Rust --- .../design/20260706_materialize_sdk.md | 49 ++++++++++++------- 1 file changed, 32 insertions(+), 17 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index ddc180dc0f81e..8bed3a1b0a6f8 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -201,8 +201,9 @@ will expose two modules: `subscribe` for live and durable consumption, and `sink` for writing to targets. Every change will keep its `mz_timestamp`. Durable consumption will store its checkpoint in the target, fenced by an epoch, and will read one storage collection (a table, materialized view, or source) -plus an optional projection, filter, and envelope. The first sink will be -turbopuffer. The spec, conformance vectors, and an end-to-end +plus an optional projection, filter, and envelope. The Rust package comes first, +with Python and Node to follow, and the first sink will be turbopuffer, written +in Rust. The spec, conformance vectors, and an end-to-end suite will live in this repository and run in the nightlies. ### Naming @@ -587,6 +588,12 @@ search-index use case and will run against our internal context graph, which has no Kafka, so the existing Kafka-based sink does not fit there. The sink will write only to namespaces it creates, so every document carries `mz_timestamp`. +The October sink will be written in Rust on the Rust package and will call +turbopuffer's HTTP API directly. Its transforms will be Rust functions that call +an embedding provider over HTTP, so the existing sink's Python transforms are not +reused. The tombstone, condition, and checkpoint rules below do not depend on +the language, so a Python version can follow the same design. + turbopuffer's documentation states that one write request to one namespace is applied atomically and is durable on return, and that there are no transactions across namespaces. Conditional writes compare each stored document with the @@ -885,27 +892,34 @@ The server-side buffering bound (#37905) has landed and needs no further work. ## Minimal Viable Prototype -The prototype will be the protocol core extracted from the Rust package, one -language package, and the turbopuffer sink running against the internal context -graph. It tests the three riskiest claims: +The prototype will be the protocol core extracted from the Rust prototype, the +Rust package on top of it, and the turbopuffer sink, written in Rust, running +against the internal context graph. Python and Node packages are stretch goals +for the same period. It tests three of the riskiest claims: -- that a protocol core with no I/O binds into a language package without - packaging or performance problems, +- that a protocol core with no I/O carries a full SDK and a real sink with only a + thin transport around it, - that tombstones and conditional writes give each turbopuffer namespace exactly-once state under random kills and a stale worker, - that the nightly end-to-end suite catches server changes that break the SDK. +A Python or Node package, if it lands in the same period, also tests that the +protocol core binds into another language without packaging or performance +problems. + ## Delivery plan -October: the prototype above, with the end-to-end suite in the nightlies. Exit -criteria are a passing convergence test, a passing fencing test, and every typed -error reproduced in the end-to-end suite. +October: the prototype above, with the end-to-end suite in the nightlies, and +the Python and Node packages as stretch goals. Exit criteria are a passing +convergence test, a passing fencing test, and every typed error reproduced in +the end-to-end suite. -November and December: the second language, durable subscriptions as they land, -and a Redis or Postgres reference sink chosen by demand. +November and December: the Python and Node packages if they did not land in +October, durable subscriptions as they land, and a Redis or Postgres reference +sink chosen by demand. GA requires dedicated SQLSTATEs released, the client docs rewritten on the SDK, -and both languages passing the same vectors and end-to-end suite. +and every published language passing the same vectors and end-to-end suite. ## Future work @@ -985,10 +999,11 @@ produce. 1. Name: Materialize SDK, Subscribe SDK, or Sink SDK? The reasoning is under "Naming". This is cheap to change now and expensive after the first publish. -2. October language: Python matches the existing turbopuffer sink, the - turbopuffer client, and the field team's tooling. Rust is where the protocol - core already lives. Recommendation: the Rust protocol core plus the Python - package. +2. First language: Rust is decided, with Python and Node as stretch goals, so the + October turbopuffer sink is written in Rust against turbopuffer's HTTP API. + The existing turbopuffer sink and its transforms are Python. Is a Rust + turbopuffer sink acceptable for October, or should it wait for the Python + package? 3. Generation strategy: is a protocol core with bindings acceptable for packaging and support, or do we start with B plus A? 4. turbopuffer tombstones: searches must filter on `deleted = false`, tombstones From 7eec9f72662ec5ab0b28e16b6c25e6b38524b84f Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Thu, 8 Oct 2026 18:47:32 +0300 Subject: [PATCH 10/19] design: Address second review on replacements, swap race, sweep and commit ordering --- .../design/20260706_materialize_sdk.md | 71 ++++++++++++------- 1 file changed, 44 insertions(+), 27 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 8bed3a1b0a6f8..ece9ba75b01aa 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -425,14 +425,11 @@ fails every resume, so the SDK will refuse to start durable consumption when the object's retention is below the commit lag. Under durable subscriptions the margin becomes the time since the last -acknowledgement against the subscription's `ACKNOWLEDGE WITHIN` deadline. -Acknowledgements advance only when the frontier does. A `REFRESH EVERY` view -between refreshes, a paused source, or a cluster with no replicas sends no -progress, and #38468 treats an acknowledgement at the current position as a -no-op, so the deadline keeps running. The deadline must therefore exceed the -retry budget, an operator's response to a stall, and the longest time the -object's frontier can stand still. The SDK will warn at startup when it does -not. +acknowledgement against the subscription's `ACKNOWLEDGE WITHIN` deadline. A +subscription that has caught up does not age (#38468), so a healthy sink on a +`REFRESH EVERY` view, a paused source, or a cluster with no replicas does not +expire. The deadline must exceed the retry budget plus an operator's response to +a stall, and the SDK will warn at startup when it does not. ### Object identity @@ -441,7 +438,7 @@ The checkpoint fingerprint lets the SDK tell three kinds of change apart: | Change | Example | Default | With `refollow` | | --- | --- | --- | --- | | Different output columns | the view's definition changed | stop with `SchemaMismatch` | stop | -| Same columns, same storage shard, new catalog id | `ALTER MATERIALIZED VIEW ... APPLY REPLACEMENT` | resume | resume | +| Same columns, same storage shard, new catalog id | `ALTER MATERIALIZED VIEW ... APPLY REPLACEMENT` | re-run the retention-margin check, then resume | same | | Same columns, new storage shard | `ALTER SCHEMA ... SWAP` in a blue/green deploy | stop | resume if the new object's history covers the checkpoint, else the history-loss policy | | The name resolves to nothing | the object was dropped while the sink was offline | stop with `ObjectDropped` | stop | @@ -450,9 +447,22 @@ that lookup, a name that resolves to nothing fails planning with `XX000`, which would be misread as `StreamPoisoned` (`src/adapter/src/error.rs:1003`). `42704` arrives only for a drop during a running stream. +The lookup and the `SUBSCRIBE` are two statements, so a swap between them would +subscribe to the new object after the check passed. Right after `DECLARE`, the +SDK will compare `referenced_object_ids` in `mz_internal.mz_subscriptions` with +the fingerprint, and on a mismatch close the cursor and classify the change +through the table above. + A replacement keeps the storage shard and gives the view a new catalog id -(`src/adapter/src/catalog/transact.rs:1334`), so its history is continuous. A -name swap moves the name to a different object. The running stream keeps +(`src/adapter/src/catalog/transact.rs:1334`), so its data history is continuous. +It takes its retention and refresh schedule from the replacement object, though +(`MaterializedView::apply_replacement` in `src/catalog/src/memory/objects.rs`). +A replacement created without `RETAIN HISTORY` drops the shard to the +one-second default, so the SDK will re-run the retention-margin check after it +sees a new catalog id, and deploy tooling has to carry `RETAIN HISTORY` onto the +replacement. After a successful resume, the SDK will store the new catalog id in +the checkpoint, so later resumes compare against it. A name swap moves the name +to a different object. The running stream keeps reading the old object until the deploy drops it, which ends the stream with `ObjectDropped` (`42704`) and sends it through the same table. Different columns always stop, because resuming would mix rows of two shapes in one target, even @@ -627,9 +637,13 @@ tombstones instead of deleting them, so the same protection holds after a re-snapshot. A patch by filter has no `$ref_new`, so it will set constants: `deleted = true`, `mz_timestamp` to the snapshot's `AS OF` `t_s`, and `written_at` to the current time, on documents matching -`generation < t_s AND deleted = false`. The patch re-evaluates its filter before -applying (turbopuffer's guarantees page), so a document a live write moves into -the new generation meanwhile is left alone. +`generation < t_s AND deleted = false AND mz_timestamp < t_s`. The last clause +keeps the sweep from moving a timestamp backwards: a stale worker from before +the re-snapshot may have written a key at a time after `t_s`, the snapshot's +older upsert of that key is then skipped, and without the clause the sweep would +tombstone a correct document. The patch re-evaluates its filter before applying +(turbopuffer's guarantees page), so a document a live write moves into the new +generation meanwhile is left alone. Searches will filter on `deleted = false`. A sweep will remove a tombstone only when its `written_at` is older than a grace period and its `mz_timestamp` is @@ -649,8 +663,13 @@ patches, which never create a document, so a stale worker cannot recreate a deleted checkpoint. A worker will take the next epoch with a patch conditional on the stored epoch being strictly lower than its own, so of two workers starting at once only one wins. It will commit a frontier with a patch conditional on the -stored epoch being equal to its own, and a patch count of zero means it was -fenced. The data namespaces hold no reserved documents. +stored epoch being equal to its own and the stored frontier being lower, so a +delayed retry of an older commit cannot move the frontier back. The tombstone +sweep reads that frontier, so it must only move forward. A patch count of zero +then has two causes, and the worker reads the checkpoint back to tell them apart: +a different epoch means it was fenced, and the same epoch means the frontier was +already at or past this commit, which counts as committed. The data namespaces +hold no reserved documents. Writes across namespaces are not atomic, and turbopuffer keeps one version of each document. Between the writes of one cut, a reader can see one namespace at @@ -877,16 +896,13 @@ program. bounded pieces, so a snapshot or catch-up larger than `max_result_size` can be delivered. Durable subscriptions do not change this, because a snapshot is one timestamp. -4. In #38468, a way for a sink on a slow-moving object to keep its subscription - alive while healthy, for example an acknowledgement at the current position - that refreshes the `ACKNOWLEDGE WITHIN` deadline. -5. Durable subscriptions (#38468). -6. Snapshot elision for a projection and filter without a temporal predicate, +4. Durable subscriptions (#38468). +5. Snapshot elision for a projection and filter without a temporal predicate, which #38468 also needs. -7. Non-poisoning subscribe errors (database-issues#5182). -8. Docs that cross-link the durable-subscriptions pattern from every client page +6. Non-poisoning subscribe errors (database-issues#5182). +7. Docs that cross-link the durable-subscriptions pattern from every client page now, and lead with the SDK once it ships. -9. A stable WebSocket `SUBSCRIBE`, the gate for browser and function transports. +8. A stable WebSocket `SUBSCRIBE`, the gate for browser and function transports. The server-side buffering bound (#37905) has landed and needs no further work. @@ -1016,9 +1032,10 @@ produce. 7. Object swaps: should `refollow` be the default for a blue/green name swap? The new view's history starts at its creation, so `refollow` works only if that history covers the sink's checkpoint. Can deploy tooling guarantee that, - for example with `RETAIN HISTORY` on staged views? For a replacement - materialized view, does a running subscribe end at the switch, and does the - new catalog id's readable frontier cover a checkpoint taken before it? + for example with `RETAIN HISTORY` on staged views and on replacement + materialized views, which take their retention from the replacement? For a + replacement, does a running subscribe end at the switch, and does the new + catalog id's readable frontier cover a checkpoint taken before it? 8. Egress cost: data leaves through `environmentd`. What does a sink with a large snapshot cost a customer, and does the docs story need a sizing page? 9. Stateless workers before durable subscriptions: is there demand that cannot From 01263d3c04e6377975551ea96027434506cd4b07 Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Thu, 8 Oct 2026 19:34:22 +0300 Subject: [PATCH 11/19] design: Tighten swaps, generations, commits, and end-of-stream detection --- .../design/20260706_materialize_sdk.md | 228 +++++++++++------- 1 file changed, 146 insertions(+), 82 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index ece9ba75b01aa..a058d59372bde 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -80,7 +80,8 @@ dataflow imports the index instead of the persist shard index's `since` (`src/adapter/src/coord/timestamp_selection.rs:277-297`). An index's default window is one second (`src/adapter-types/src/compaction.rs:20`), and the view's `RETAIN HISTORY` does not carry over to its indexes. So -`AS OF F - 1` fails on an indexed view, however long the view's retention. An +`AS OF F - 1` fails on an indexed view once the checkpoint is more than about a +second behind, however long the view's retention. An index on a different cluster does not count. Index-level `RETAIN HISTORY` requires `enable_index_options`, which is off by default. @@ -165,9 +166,10 @@ this document adds. | R11 | Sink authors can test without a live environment | Design | R4 is met in full by targets that can write the whole cut atomically, such as -one Postgres transaction. A target split into parts with no transaction across -them, such as several turbopuffer namespaces, gets R4 per part only (see -"turbopuffer sink"). +one Postgres transaction. A target with no transaction over the whole cut, such +as turbopuffer, where only one write request is atomic, gets convergence: after +each completed batch it equals the cut, but readers can see part of a batch +while it is written (see "turbopuffer sink"). A user following the happy path cannot commit an unclosed timestamp, cannot get the `AS OF` arithmetic wrong, and cannot silently lose or duplicate updates @@ -352,22 +354,33 @@ A frontier advance with no data will yield an empty batch with a fresh token, so checkpoints keep moving through quiet hours. `UP TO` will be supported. On server versions without the SQL-528 fix, the SDK -will release everything below `UP TO` when a bounded stream ends without error, -so no closed data is lost. Gating on the server version switches the workaround -off once the fix ships. +will release everything below `UP TO` once a bounded stream has ended, so no +closed data is lost. An empty `FETCH` with a timeout does not prove the end, +because a timeout and an exhausted cursor return the same empty result +(`src/pgwire/src/protocol.rs:2607-2624`). So after an empty timed fetch on a +bounded stream, the SDK will issue a `FETCH` without a timeout, which waits for +data or the end of the stream (`src/sql/src/plan/statement/scl.rs:262`), and +only an empty result there counts as the end. Gating on the server version +switches the workaround off once the fix ships. ### Checkpoints and fencing -A checkpoint store will have `load(name)` and `commit(name, epoch, frontier)`. -The target is the recommended store, because writing data and checkpoint in -one transaction gives exactly-once state (R1, R5). A Materialize-table store and -a local file store will ship for targets with no transaction. Neither commits -atomically with the target, so both give `at_least_once` only. - -Each worker start will take the next epoch. A commit will be conditional on the -stored epoch not being newer, and a stale worker will get a typed `Fenced` -error. Without the fence, a second instance of one sink would overwrite the -first one's progress. +A checkpoint store will have `load(name)` and +`commit(name, epoch, expected, frontier)`. The target is the recommended store, +because writing data and checkpoint in one transaction gives exactly-once state +(R1, R5). A Materialize-table store and a local file store will ship for targets +with no transaction. Neither commits atomically with the target, so both give +`at_least_once` only. + +Each worker start will take the next epoch. A commit will succeed only if the +stored epoch equals the worker's and the stored frontier equals `expected`, the +frontier the batch started from. A stale worker gets a typed `Fenced` error. +Without the fence, a second instance of one sink would overwrite the first one's +progress. Under `transactional`, this check runs in the same transaction as the +data, so a batch whose frontier is already committed rolls back as a whole. That +matters when a commit succeeds but its response is lost: the retry finds the +stored frontier already past `expected`, rolls back, and the SDK reads the +checkpoint and treats the batch as committed instead of applying it twice. Every durable subscription will require an explicit name, which is its checkpoint identity. A default derived from a class name would give two @@ -420,9 +433,10 @@ below a configured threshold. A connected subscribe holds history only up to what it has emitted, not up to what the sink has committed (see "Retention window"). The margin therefore has to cover the commit lag (the client buffer, `commit_interval`, and the target's write time) as well as the retry budget and -an operator's response to a stall. An object with the default one-second window -fails every resume, so the SDK will refuse to start durable consumption when the -object's retention is below the commit lag. +an operator's response to a stall. A resume fails once `F - 1` falls below the +object's `since`. With the default one-second window that happens after about a +second of commit lag or downtime, so the SDK will refuse to start durable +consumption when the object's retention is below the commit lag. Under durable subscriptions the margin becomes the time since the last acknowledgement against the subscription's `ACKNOWLEDGE WITHIN` deadline. A @@ -439,7 +453,7 @@ The checkpoint fingerprint lets the SDK tell three kinds of change apart: | --- | --- | --- | --- | | Different output columns | the view's definition changed | stop with `SchemaMismatch` | stop | | Same columns, same storage shard, new catalog id | `ALTER MATERIALIZED VIEW ... APPLY REPLACEMENT` | re-run the retention-margin check, then resume | same | -| Same columns, new storage shard | `ALTER SCHEMA ... SWAP` in a blue/green deploy | stop | resume if the new object's history covers the checkpoint, else the history-loss policy | +| Same columns, new storage shard | `ALTER SCHEMA ... SWAP` in a blue/green deploy | stop | re-snapshot from the new object, into a new namespace for targets without transactions | | The name resolves to nothing | the object was dropped while the sink was offline | stop with `ObjectDropped` | stop | The SDK will look the name up in the catalog before each `SUBSCRIBE`. Without @@ -447,11 +461,28 @@ that lookup, a name that resolves to nothing fails planning with `XX000`, which would be misread as `StreamPoisoned` (`src/adapter/src/error.rs:1003`). `42704` arrives only for a drop during a running stream. +A name swap needs a re-snapshot even when the new object's history covers the +checkpoint. The target holds the old object's state at `F - 1`, and resuming the +new object without a snapshot delivers only its changes after that point. A key +the old object had and the new one never had would then survive forever. The +generation sweep after the re-snapshot removes such keys from a target that +writes data and checkpoint in one transaction, because the epoch check in that +transaction also rolls back a stale worker still reading the old object. A +target without such a transaction cannot fence those writes per document: a +stale worker could write a key the new object never has, at a timestamp after +the snapshot, and nothing would remove it. For such targets, turbopuffer among +them, `refollow` will write the new object into a new namespace and switch +readers when the snapshot completes. + The lookup and the `SUBSCRIBE` are two statements, so a swap between them would -subscribe to the new object after the check passed. Right after `DECLARE`, the -SDK will compare `referenced_object_ids` in `mz_internal.mz_subscriptions` with -the fingerprint, and on a mismatch close the cursor and classify the change -through the table above. +subscribe to the new object after the check passed. The SDK will therefore +subscribe by the catalog id it fingerprinted, with the bracket syntax the catalog +uses for stored definitions: `SUBSCRIBE [u123 AS "db"."schema"."name"]`. That +resolves by id, not by name, and fails with `InvalidId` once the id is gone +(`src/sql/src/names.rs:1585-1608`). A swap leaves ids unchanged, so the stream +keeps reading the object that was checked, and a failed resolution sends the SDK +back to the name lookup and the table above. The syntax is not documented for +users, so open question 13 asks whether the SDK can rely on it. A replacement keeps the storage shard and gives the view a new catalog id (`src/adapter/src/catalog/transact.rs:1334`), so its data history is continuous. @@ -495,8 +526,10 @@ from different timelines are not comparable and the cut would be unsound. Only sources with `ENVELOPE MATERIALIZE`, which is behind a flag, get another timeline (`src/sql/src/plan/statement/ddl.rs:1000-1004`), and objects built on them inherit it. The SDK will reject a member that is such a source or depends -on one, using `mz_sources.envelope_type` and -`mz_internal.mz_object_transitive_dependencies`. +on one, using `mz_sources.envelope_type`, the envelope of source tables in +`mz_internal.mz_kafka_source_tables.envelope_type` (a table created +`FROM SOURCE ... ENVELOPE MATERIALIZE` carries the envelope, not its source), +and `mz_internal.mz_object_transitive_dependencies`. A multi-view subscription will hold one connection per member. Any member error will end all members, and recovery will resume every member at the stored cut. @@ -573,10 +606,13 @@ snapshot, resume, retry, fencing, and the history-loss policy. Keyed-state targets (caches, search indexes, tables) will read the upsert envelope. When they re-snapshot, the SDK will write the snapshot under a new generation and, at the end, call a sweep that removes keys from older -generations. The generation is the snapshot's `AS OF`, and writes after the -snapshot carry it too. It grows across attempts, so a re-snapshot that crashed -and was retried never reuses a generation, and keys only the failed attempt -wrote are swept. That removes keys deleted while the sink was offline, which a plain +generations. The generation is a counter stored in the checkpoint, raised by a +conditional commit before each re-snapshot attempt starts, and writes after the +snapshot carry it too. A snapshot's `AS OF` cannot serve as the generation, +because a retry can pick the same `AS OF` and a new object after a swap can +have an older one. With a counter, a re-snapshot that crashed and was retried +never reuses a generation, and keys only the failed attempt wrote are swept. +That removes keys deleted while the sink was offline, which a plain re-snapshot leaves behind. A target that needs the snapshot to appear at once will make the new generation visible only at the sweep. @@ -628,7 +664,7 @@ older change then meets a newer document or tombstone and is skipped. Before eac request, the sink will net its changes to the latest change per key, because a batch spans timestamps and turbopuffer rejects a request that names one id twice. The conditions only ever let a newer timestamp win, so dropping the older -change from a request is safe. Every write is then +change from a request is safe. While tombstones are kept, every write is then safe to repeat and ordered by timestamp, so each namespace reaches exactly-once state under `at_least_once` delivery, and a batch can be split across requests. @@ -637,7 +673,8 @@ tombstones instead of deleting them, so the same protection holds after a re-snapshot. A patch by filter has no `$ref_new`, so it will set constants: `deleted = true`, `mz_timestamp` to the snapshot's `AS OF` `t_s`, and `written_at` to the current time, on documents matching -`generation < t_s AND deleted = false AND mz_timestamp < t_s`. The last clause +`generation < g AND deleted = false AND mz_timestamp < t_s`, where `g` is the +attempt's generation. The last clause keeps the sweep from moving a timestamp backwards: a stale worker from before the re-snapshot may have written a key at a time after `t_s`, the snapshot's older upsert of that key is then skipped, and without the clause the sweep would @@ -645,39 +682,54 @@ tombstone a correct document. The patch re-evaluates its filter before applying (turbopuffer's guarantees page), so a document a live write moves into the new generation meanwhile is left alone. -Searches will filter on `deleted = false`. A sweep will remove a tombstone only -when its `written_at` is older than a grace period and its `mz_timestamp` is -below the sink's committed frontier. The grace period runs from when the -tombstone was written, not from its `mz_timestamp`, because a sink that is -catching up writes tombstones for old timestamps. It has to exceed how long a -fenced worker can keep writing before its next checkpoint commit fails: -`commit_interval`, the retry budget, and the longest pause a sink process can -survive. A worker paused for longer could still recreate a swept key, so the -grace period bounds that risk without removing it. Filter operations are capped -per call (5 million rows for a delete by filter, 50 thousand for a patch by -filter), so both sweeps loop. +Searches will filter on `deleted = false`. By default the sink will keep +tombstones forever, because removing one reopens the hole: a fenced worker that +was paused before an older upsert, and resumes after the tombstone is gone, +recreates the key, and nothing deletes it again until the key changes upstream. +A sink can opt into a sweep that removes a tombstone only when its `written_at` +is older than a grace period and its `mz_timestamp` is below the sink's +committed frontier. The grace period runs from when the tombstone was written, +not from its `mz_timestamp`, because a sink that is catching up writes tombstones +for old timestamps. With the sweep on, exactly-once state holds only if no +worker writes later than the grace period after it was fenced. The SDK cannot +enforce that bound, because a process can pause between checking its epoch and +sending a write, so the sink's docs will state the condition and the grace +period has to exceed `commit_interval`, the retry budget, and the longest pause +the deployment can see. Filter operations are capped per call (5 million rows +for a delete by filter, 50 thousand for a patch by filter), so sweeps loop. The checkpoint will live in a separate checkpoint namespace, one document per sink, created once when the sink is set up. Workers will change it only with patches, which never create a document, so a stale worker cannot recreate a deleted checkpoint. A worker will take the next epoch with a patch conditional on the stored epoch being strictly lower than its own, so of two workers starting -at once only one wins. It will commit a frontier with a patch conditional on the -stored epoch being equal to its own and the stored frontier being lower, so a -delayed retry of an older commit cannot move the frontier back. The tombstone -sweep reads that frontier, so it must only move forward. A patch count of zero -then has two causes, and the worker reads the checkpoint back to tell them apart: -a different epoch means it was fenced, and the same epoch means the frontier was -already at or past this commit, which counts as committed. The data namespaces -hold no reserved documents. - -Writes across namespaces are not atomic, and turbopuffer keeps one version of -each document. Between the writes of one cut, a reader can see one namespace at -the new cut and another at the old one. Filtering on `mz_timestamp` cannot -rebuild an earlier cut, because an upsert replaces the earlier version. Each -namespace converges to every cut, but the sink does not give a consistent view -across namespaces, so for turbopuffer R4 holds per namespace only. The sink's -docs will state this, and open question 12 asks whether that is enough. +at once only one wins. A commit follows the general contract (see "Checkpoints +and fencing"): a patch conditional on the stored epoch being equal to its own and +the stored frontier being equal to `expected`. So a delayed retry of an older +commit cannot move the frontier back, which matters because the tombstone sweep +reads that frontier. On a patch count of zero the worker reads the checkpoint +back: a different epoch means `Fenced`, the same epoch with a frontier at or past +the requested one means the commit already happened, and anything else is a +conflict that sends the worker back to the stored checkpoint. The data +namespaces hold no reserved documents. + +Only one write request is atomic, and turbopuffer keeps one version of each +document. A batch can span several requests, even within one namespace, and a +cut spans one request per namespace at least. Between those requests a reader +can see part of a batch: two keys that changed at the same timestamp, one new +and one old, or one namespace at the new cut and another at the old one. +Filtering on `mz_timestamp` cannot rebuild an earlier cut, because an upsert +replaces the earlier version. The guarantee for turbopuffer is therefore +convergence: after each completed batch, every namespace equals the cut. The +sink does not give readers a consistent view while a batch is being written. The +sink's docs will state this, and open question 12 asks whether that is enough. + +After a name swap, the sink writes the new object into new namespaces and +switches readers when the snapshot completes (see "Object identity"). A stale +worker still reading the old object keeps writing only to the old namespaces, +which readers no longer use. The sink deletes them after a grace period. turbopuffer +creates a namespace on its first write, so a stale write after that recreates an +unused namespace, which the next cleanup removes. Embedding cost will follow the existing sink's transform model: a transform declares the columns it reads, and runs only for rows where those columns @@ -713,11 +765,11 @@ catalog transaction (#38468). The SDK will attach to an existing subscription by name and will not create one on start. Creating it, and choosing its deadline, is a deploy step. -A blue/green cutover needs a new subscription on the new object, created -`START AT` the sink's checkpoint, which works only if the new object's history -covers it. Dropping the old object needs `CASCADE` while the old subscription -exists. Deploy tooling will own both steps, and `refollow` then attaches to the -new subscription. +A blue/green cutover needs a new subscription on the new object, and the sink +re-snapshots from it, for the reason "Object identity" gives. Dropping the old +object needs `CASCADE` while the old subscription exists. Deploy tooling will own +both steps, and `refollow` then attaches to the new subscription with a +snapshot. Durable subscriptions will remove the retention sizing problem and turn the retention-margin check into a deadline check (see "History loss and retention @@ -801,7 +853,8 @@ protocol core: it decodes each Materialize type into one documented value model, and each binding converts that model to the language's native types (for example `numeric` to `Decimal` in Python), with the conversions covered by the vectors. `mz_timestamp` is a u64, so the Node package will expose it as a -`BigInt`, which holds every u64 value where a JavaScript `number` does not. Each package will hand-write only the transport over its driver, +`BigInt`, which holds every u64 value where a JavaScript `number` does not. +Each package will hand-write only the transport over its driver, the idiomatic API surface (iterators in Python, async iterators in Node), and the sink modules. @@ -866,9 +919,13 @@ The conformance suite is the contract, whatever the generation strategy. `SELECT ... AS OF F - 1` for the committed frontier `F`, for both guarantees. 4. Fencing test: two workers share one checkpoint, the stale one is refused, and the target stays consistent. For turbopuffer, the stale worker also writes - an older upsert of a key the live worker has deleted, and the key must stay - deleted. -5. User test kit: the recorded-stream player from the vectors will ship to users, + an older upsert of a key the live worker has deleted, and with tombstones + kept the key must stay deleted. With the sweep on, the same holds when the + stale write comes within the grace period. +5. Lost-response test: a transactional commit succeeds but its response is + dropped, and the retry must find the frontier already committed and apply + nothing twice. +6. User test kit: the recorded-stream player from the vectors will ship to users, so sink authors can test their targets offline (R11). Keeping the vectors and the end-to-end suite next to the server is what lets a @@ -1023,19 +1080,21 @@ produce. 3. Generation strategy: is a protocol core with bindings acceptable for packaging and support, or do we start with B plus A? 4. turbopuffer tombstones: searches must filter on `deleted = false`, tombstones - need a placeholder vector in namespaces with vectors, and a sweep removes them - after a grace period measured from when they were written. What default grace - period is safe, and is a placeholder vector acceptable to users? + need a placeholder vector in namespaces with vectors, and by default they are + kept forever. Is that storage cost acceptable, and is a placeholder vector + acceptable to users? For sinks that opt into the sweep, what grace period + should the docs recommend? 5. Retention margin check: can a sink role read the catalog state it needs (the object's readable frontier and the cluster's indexes) without extra grants? 6. `mzq`: a transport under this SDK, or a separate path? -7. Object swaps: should `refollow` be the default for a blue/green name swap? - The new view's history starts at its creation, so `refollow` works only if - that history covers the sink's checkpoint. Can deploy tooling guarantee that, - for example with `RETAIN HISTORY` on staged views and on replacement - materialized views, which take their retention from the replacement? For a - replacement, does a running subscribe end at the switch, and does the new - catalog id's readable frontier cover a checkpoint taken before it? +7. Object swaps: should `refollow`, which re-snapshots from the new object, be + the default for a blue/green name swap? A re-snapshot after every deploy + costs a full snapshot. Could deploy tooling confirm that the old and new + objects are equal at the handoff, so the sink can skip it? Deploy tooling also + has to carry `RETAIN HISTORY` onto replacement materialized views, which take + their retention from the replacement. For a replacement, does a running + subscribe end at the switch, and does the new catalog id's readable frontier + cover a checkpoint taken before it? 8. Egress cost: data leaves through `environmentd`. What does a sink with a large snapshot cost a customer, and does the docs story need a sizing page? 9. Stateless workers before durable subscriptions: is there demand that cannot @@ -1045,6 +1104,11 @@ produce. the protocol core, vectors, and end-to-end suite here? 11. Future-work modules: are any of them worth their maintenance cost, and what evidence should trigger one? -12. turbopuffer across namespaces: is per-namespace convergence enough for - search, or do readers need a consistent view across namespaces? That would - need versioned documents and a visible-cut pointer that readers filter on. +12. turbopuffer visibility: is convergence after each batch enough for search, + or do readers need a consistent view while a batch is written, within one + namespace or across several? That would need versioned documents and a + visible-cut pointer that readers filter on. +13. Subscribing by id: can the SDK rely on the `[u123 AS "db"."schema"."name"]` + reference syntax in `SUBSCRIBE`, which the catalog uses for stored + definitions but the user docs do not cover? It is what makes the identity + check race-free. From aba1fd381290ae062a32602984573d8c2f0e562e Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Thu, 8 Oct 2026 22:11:40 +0300 Subject: [PATCH 12/19] design: Limit sinks to SUBSCRIBE on a table, source, or materialized view --- .../design/20260706_materialize_sdk.md | 62 +++++++++++-------- 1 file changed, 36 insertions(+), 26 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index a058d59372bde..0f63b8e049876 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -201,9 +201,10 @@ which performs no I/O and is compiled into a thin package per language. Each package will use that language's own database driver for the connection and will expose two modules: `subscribe` for live and durable consumption, and `sink` for writing to targets. Every change will keep its `mz_timestamp`. -Durable consumption will store its checkpoint in the target, fenced by an epoch, -and will read one storage collection (a table, materialized view, or source) -plus an optional projection, filter, and envelope. The Rust package comes first, +Durable consumption, which powers sinks, will store its checkpoint in the target, +fenced by an epoch, and will read one table, source, or materialized view +directly with `SUBSCRIBE `, the same objects `CREATE SINK` accepts. Live +streams that never resume will accept any query. The Rust package comes first, with Python and Node to follow, and the first sink will be turbopuffer, written in Rust. The spec, conformance vectors, and an end-to-end suite will live in this repository and run in the nightlies. @@ -300,23 +301,30 @@ sequenceDiagram ### Subscription scope -A durable subscription, one that checkpoints and resumes, will read one storage -collection (a table, materialized view, or source) plus an optional projection, -filter, and envelope. The filter must not call `mz_now()`. Plain views, indexes -as targets, and temporal filters will be rejected. Arbitrary SQL will be accepted -only for live streams that never resume. - -Resuming a general query rehydrates its whole dataflow and needs history on -every input of the query. A plain view is inlined, so resuming it costs the -same, and a temporal filter needs the snapshot on every resume. Durable -subscriptions (#38468) accept exactly this narrower surface, so starting with it -makes the later move a transport change, not an API break. Widening later is -compatible, and narrowing is not. Structured input also lets the SDK build the -statement and place `ENVELOPE` before `WITH` without parsing user SQL. - -Projections and filters still read the snapshot on resume (see "Snapshot -elision"). The docs will recommend the plain collection form for large -collections, and the SDK will log the resume cost at startup. +The SDK will separate two uses of `SUBSCRIBE`: + +- A live subscribe, for a UI, a cache warmer, or an agent watching results, + never resumes. It will accept any query. +- A sink, or any durable subscription that checkpoints and resumes, will read + one table, source, or materialized view directly, as `SUBSCRIBE ` with + an optional envelope. It accepts no query, projection, or filter. These are the + objects `CREATE SINK` accepts (`plan_create_sink` in + `src/sql/src/plan/statement/ddl.rs`), so sinks built on the SDK follow the same + rule as Kafka sinks. + +A sink that needs a projection, a filter, or a join will read a materialized view +that computes it. Resuming a query rehydrates its whole dataflow and needs +history on every input of the query, and a plain view is inlined, so resuming it +costs the same. Only the plain object form skips the snapshot on resume (see +"Snapshot elision"), so with this rule a sink's resume never reads the snapshot. +Reading an object backed by a storage shard also leaves room for server-side +optimizations aimed at that case. + +Durable subscriptions (#38468) accept one storage collection with an optional +projection and filter, a wider surface than this. Starting narrow keeps the later +move a transport change, because widening later is compatible and narrowing is +not. Building the statement from an object and an envelope also lets the SDK +place `ENVELOPE` before `WITH` without parsing user SQL. The SDK will refuse an indexed object at startup. It will check whether the subscribed object has an index on the subscribing cluster and fail with an error @@ -345,7 +353,8 @@ arrive as chunks marked partial, with the token on the closing chunk. Chunks bound client memory only. A snapshot or catch-up larger than `max_result_size` fails on the server before the first chunk arrives (see "Buffering limits"), and the SDK will report it as `ResultTooLarge` with the remedies: subscribe to a -narrower projection, or have an administrator raise `max_result_size`. Server-side +smaller materialized view, for example one with fewer columns, or have an +administrator raise `max_result_size`. Server-side chunking would remove the limit (see "Materialize-side workstream"). The sink module's generations (see "Sink module") give atomic snapshot visibility to targets that need it. @@ -390,8 +399,8 @@ when the list is reordered. The checkpoint will also record a fingerprint of the subscription: the object's name, catalog id, and storage shard (from `mz_internal.mz_storage_shards`, joined -through `mz_internal.mz_object_global_ids`), and -the projection, filter, envelope, and output column types. A resume compares it +through `mz_internal.mz_object_global_ids`), and the envelope and output column +types. A resume compares it with the object the name resolves to now, and "Object identity" decides what each difference means. The shard matters because blue/green deploys use `ALTER SCHEMA ... SWAP`, which moves names and leaves ids unchanged. After a swap @@ -954,8 +963,9 @@ program. delivered. Durable subscriptions do not change this, because a snapshot is one timestamp. 4. Durable subscriptions (#38468). -5. Snapshot elision for a projection and filter without a temporal predicate, - which #38468 also needs. +5. Snapshot elision for a projection and filter without a temporal predicate. + #38468 needs it, and it would let sinks accept a projection and filter later + without a snapshot read on every resume. 6. Non-poisoning subscribe errors (database-issues#5182). 7. Docs that cross-link the durable-subscriptions pattern from every client page now, and lead with the SDK once it ships. @@ -1029,7 +1039,7 @@ right, and this design keeps it. The parts that change: | Strawman | This design | Why | | --- | --- | --- | | `update()` commits and fetches | `commit()` inside the transaction, `next()` outside | A target transaction must not wait on Materialize | -| `queries` as SQL strings, split at `ENVELOPE` | Object, projection, filter, envelope as structured input | Resume cost, durable subscription compatibility, no string splicing | +| `queries` as SQL strings, split at `ENVELOPE` | One table, source, or materialized view plus an envelope, as `CREATE SINK` takes | Resume never reads the snapshot, durable subscription compatibility, no string splicing | | One cursor per query | One `AS OF` for all members, release at the minimum frontier | Independent cursors produce torn cuts | | Rows stamped with the batch frontier | Every change keeps its `mz_timestamp` | Batches span timestamps | | Retractions come first | Order by timestamp, net per key within a timestamp | Not true by default | From 77968ad037696ae1719f7475365d8c273a65c6fa Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Fri, 9 Oct 2026 14:25:39 +0300 Subject: [PATCH 13/19] design: Specify the turbopuffer reader switch and classify errors by statement --- .../design/20260706_materialize_sdk.md | 39 ++++++++++++------- 1 file changed, 26 insertions(+), 13 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 0f63b8e049876..6f271111ea0b2 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -465,10 +465,14 @@ The checkpoint fingerprint lets the SDK tell three kinds of change apart: | Same columns, new storage shard | `ALTER SCHEMA ... SWAP` in a blue/green deploy | stop | re-snapshot from the new object, into a new namespace for targets without transactions | | The name resolves to nothing | the object was dropped while the sink was offline | stop with `ObjectDropped` | stop | -The SDK will look the name up in the catalog before each `SUBSCRIBE`. Without -that lookup, a name that resolves to nothing fails planning with `XX000`, which -would be misread as `StreamPoisoned` (`src/adapter/src/error.rs:1003`). `42704` -arrives only for a drop during a running stream. +The SDK will look the name up in the catalog before each `SUBSCRIBE`, and it will +classify errors by the statement that raised them. A name that resolves to +nothing, and an id that no longer exists, both fail planning at `DECLARE` with +`XX000` (`src/adapter/src/error.rs:1003`), the same code a dataflow error has. So +any failure of `DECLARE` sends the SDK back to the name lookup and the table +above, and only errors raised during `FETCH` are classified as `StreamPoisoned` +or by the rest of "Typed errors". `42704` arrives only for a drop during a +running stream. A name swap needs a re-snapshot even when the new object's history covers the checkpoint. The target holds the old object's state at `F - 1`, and resuming the @@ -489,8 +493,8 @@ subscribe by the catalog id it fingerprinted, with the bracket syntax the catalo uses for stored definitions: `SUBSCRIBE [u123 AS "db"."schema"."name"]`. That resolves by id, not by name, and fails with `InvalidId` once the id is gone (`src/sql/src/names.rs:1585-1608`). A swap leaves ids unchanged, so the stream -keeps reading the object that was checked, and a failed resolution sends the SDK -back to the name lookup and the table above. The syntax is not documented for +keeps reading the object that was checked, and an `InvalidId` failure at +`DECLARE` follows the rule above. The syntax is not documented for users, so open question 13 asks whether the SDK can rely on it. A replacement keeps the storage shard and gives the view a new catalog id @@ -567,7 +571,7 @@ generations will stage these changes in the target and make them visible at | `HistoryLost` | `22000` plus the timestamp-selection message, after the index re-check | history-loss policy | | `ObjectDropped` | `42704` | see "Object identity" | | `SchemaMismatch` | checkpoint fingerprint | see "Object identity" | -| `StreamPoisoned` | `XX000` with any other message, which includes genuine internal errors | stop | +| `StreamPoisoned` | `XX000` during `FETCH` with any other message, which includes genuine internal errors | stop | | `Fenced` | checkpoint commit | stop | | `IndexedTarget` | startup check, and the re-check before `HistoryLost` | stop with remedy | | `Fatal` | anything else (auth, TLS, SQL) | stop | @@ -733,12 +737,21 @@ convergence: after each completed batch, every namespace equals the cut. The sink does not give readers a consistent view while a batch is being written. The sink's docs will state this, and open question 12 asks whether that is enough. -After a name swap, the sink writes the new object into new namespaces and -switches readers when the snapshot completes (see "Object identity"). A stale -worker still reading the old object keeps writing only to the old namespaces, -which readers no longer use. The sink deletes them after a grace period. turbopuffer -creates a namespace on its first write, so a stale write after that recreates an -unused namespace, which the next cleanup removes. +After a name swap, the sink writes the new object into new namespaces (see +"Object identity"). Namespace names carry an incarnation number, for example +`articles__i3`, and the checkpoint document records the active incarnation and +the list of retired ones. The checkpoint document is the reader contract: a +reader looks up the active incarnation there, with a small helper the sink +package ships, and derives every namespace name from it. When the snapshot +completes, the sink switches readers by patching the active incarnation, one +write to one document, so the switch is atomic for every namespace at once for +any reader that looks it up again. Until then, readers of the old namespaces see +data frozen at the swap, and the sink's docs will say so. A stale worker still +reading the old object keeps writing only to the retired namespaces. The sink +deletes them after a grace period, using the durable list in the checkpoint +document. turbopuffer creates a namespace on its first write, so a stale write +after that recreates an unused namespace. Retired incarnations therefore stay on +the list, and every later cleanup removes their namespaces again if they exist. Embedding cost will follow the existing sink's transform model: a transform declares the columns it reads, and runs only for rows where those columns From 866a544fb23024a41abd6959637f29defbcd404b Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Fri, 9 Oct 2026 15:42:54 +0300 Subject: [PATCH 14/19] design: Resolve or defer each open question with an owner --- .../design/20260706_materialize_sdk.md | 108 +++++++++++------- 1 file changed, 64 insertions(+), 44 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 6f271111ea0b2..3ea924bfd2e08 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -495,7 +495,7 @@ resolves by id, not by name, and fails with `InvalidId` once the id is gone (`src/sql/src/names.rs:1585-1608`). A swap leaves ids unchanged, so the stream keeps reading the object that was checked, and an `InvalidId` failure at `DECLARE` follows the rule above. The syntax is not documented for -users, so open question 13 asks whether the SDK can rely on it. +users, so deferred question 2 asks whether the SDK can rely on it. A replacement keeps the storage shard and gives the view a new catalog id (`src/adapter/src/catalog/transact.rs:1334`), so its data history is continuous. @@ -735,7 +735,7 @@ Filtering on `mz_timestamp` cannot rebuild an earlier cut, because an upsert replaces the earlier version. The guarantee for turbopuffer is therefore convergence: after each completed batch, every namespace equals the cut. The sink does not give readers a consistent view while a batch is being written. The -sink's docs will state this, and open question 12 asks whether that is enough. +sink's docs will state this, and deferred question 5 asks whether that is enough. After a name swap, the sink writes the new object into new namespaces (see "Object identity"). Namespace names carry an incarnation number, for example @@ -1093,45 +1093,65 @@ produce. ## Open questions -1. Name: Materialize SDK, Subscribe SDK, or Sink SDK? The reasoning is under - "Naming". This is cheap to change now and expensive after the first publish. -2. First language: Rust is decided, with Python and Node as stretch goals, so the - October turbopuffer sink is written in Rust against turbopuffer's HTTP API. - The existing turbopuffer sink and its transforms are Python. Is a Rust - turbopuffer sink acceptable for October, or should it wait for the Python - package? -3. Generation strategy: is a protocol core with bindings acceptable for - packaging and support, or do we start with B plus A? -4. turbopuffer tombstones: searches must filter on `deleted = false`, tombstones - need a placeholder vector in namespaces with vectors, and by default they are - kept forever. Is that storage cost acceptable, and is a placeholder vector - acceptable to users? For sinks that opt into the sweep, what grace period - should the docs recommend? -5. Retention margin check: can a sink role read the catalog state it needs (the - object's readable frontier and the cluster's indexes) without extra grants? -6. `mzq`: a transport under this SDK, or a separate path? -7. Object swaps: should `refollow`, which re-snapshots from the new object, be - the default for a blue/green name swap? A re-snapshot after every deploy - costs a full snapshot. Could deploy tooling confirm that the old and new - objects are equal at the handoff, so the sink can skip it? Deploy tooling also - has to carry `RETAIN HISTORY` onto replacement materialized views, which take - their retention from the replacement. For a replacement, does a running - subscribe end at the switch, and does the new catalog id's readable frontier - cover a checkpoint taken before it? -8. Egress cost: data leaves through `environmentd`. What does a sink with a large - snapshot cost a customer, and does the docs story need a sizing page? -9. Stateless workers before durable subscriptions: is there demand that cannot - wait? -10. Repository home: should the packages stay under `misc/` in this repository - for good, or move to their own repository once the spec settles, keeping - the protocol core, vectors, and end-to-end suite here? -11. Future-work modules: are any of them worth their maintenance cost, and what - evidence should trigger one? -12. turbopuffer visibility: is convergence after each batch enough for search, - or do readers need a consistent view while a batch is written, within one - namespace or across several? That would need versioned documents and a - visible-cut pointer that readers filter on. -13. Subscribing by id: can the SDK rely on the `[u123 AS "db"."schema"."name"]` - reference syntax in `SUBSCRIBE`, which the catalog uses for stored - definitions but the user docs do not cover? It is what makes the identity - check race-free. +Each question is either resolved by the review or deferred with an owner and a +point by which it has to be answered. None of the deferred ones blocks the +October work, because the doc picks a safe default for each. + +### Resolved during review + +- First language: the Rust package comes first, with Python and Node as stretch + goals, and the October turbopuffer sink is written in Rust against + turbopuffer's HTTP API. +- Generation strategy: one protocol core with bindings (option C). If native + packaging proves too costly when the second language lands, the fallback is B + plus A. +- Catalog access for the startup checks: every relation the SDK reads is + readable by `PUBLIC`: `mz_internal.mz_frontiers` (the readable frontier), + `mz_catalog.mz_indexes`, `mz_internal.mz_storage_shards`, + `mz_internal.mz_object_global_ids`, + `mz_internal.mz_object_transitive_dependencies`, and + `mz_internal.mz_kafka_source_tables`. A sink role needs no extra grants for + them. +- Sink scope: sinks read one table, source, or materialized view with + `SUBSCRIBE `, the same objects `CREATE SINK` accepts. + +### Deferred + +1. Name. The doc proposes Materialize SDK (see "Naming"). Owner: product. + Answer before the first package is published, since renaming after that is + expensive. +2. Subscribing by id. Resolving `[u123 AS "db"."schema"."name"]` by catalog id, + with the name used only as an alias, is confirmed in review and covered for + user queries by `test/sqllogictest/id.slt`. Owner: database team. Confirm it + is a supported contract for `SUBSCRIBE` before the Rust package ships, since + it is what makes the identity check race-free. +3. Object swaps. Until answered, a name swap stops the sink by default, and + `refollow` re-snapshots into new namespaces, which is safe but costs a full + snapshot per deploy. Owner: deploy tooling with the database team. Answer + before `refollow` ships: can deploy tooling confirm that the old and new + objects are equal at the handoff, so the sink can skip the snapshot, and + carry `RETAIN HISTORY` onto staged views and replacement materialized views? + For a replacement, does a running subscribe end at the switch, and does the + new catalog id's readable frontier cover a checkpoint taken before it? +4. turbopuffer tombstones. Tombstones are kept forever by default and carry a + placeholder vector in namespaces with vectors. Owner: DevEx, with the first + turbopuffer users, during the sink build. Is the storage cost and the + placeholder vector acceptable, and what grace period should the docs + recommend for sinks that opt into the sweep? +5. turbopuffer visibility. The sink documents convergence after each batch. + Owner: DevEx, with the first search users. Do readers need a consistent view + while a batch is written? That would need versioned documents and a + visible-cut pointer that readers filter on. +6. Repository home. The doc proposes keeping the packages under `misc/` in this + repository. Owner: DevEx with the engineering leads, once the spec settles: + do the packages move to their own repository, keeping the protocol core, + vectors, and end-to-end suite here? +7. `mzq` (PRD-87). Owner: product. Is it a transport under this SDK, or a + separate path? +8. Egress cost. Owner: product with the cloud team, before the GA docs. What + does a sink with a large snapshot cost a customer, and do the docs need a + sizing page? +9. Stateless workers. Out of scope until durable subscriptions land. Owner: + product. Is there demand that cannot wait? +10. Future-work modules. The rule for adding one is in "Future work". Owner: + DevEx, per module, when evidence appears. From 16ab79ad613f084c544d90b52a880b6e8b5beec4 Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Fri, 9 Oct 2026 17:34:43 +0300 Subject: [PATCH 15/19] design: Address review on re-snapshot isolation, sink ids, commit size, and mz-sink-sdk --- .../design/20260706_materialize_sdk.md | 143 +++++++++++++----- 1 file changed, 103 insertions(+), 40 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 3ea924bfd2e08..715c55ff7c0a4 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -9,7 +9,7 @@ - SQL-528: `SUBSCRIBE ... WITH (PROGRESS) ... UP TO` emits no final progress row. - [database-issues#5182](https://github.com/MaterializeInc/database-issues/issues/5182): dataflow errors poison a subscribe. - [Durable subscriptions pattern](https://materialize.com/docs/transform-data/patterns/durable-subscriptions/): the manual protocol this SDK packages. - - Prior art: [mz-redis-sync](https://github.com/MaterializeIncLabs/mz-redis-sync), [novu-materialize-integration](https://github.com/MaterializeIncLabs/novu-materialize-integration), [mz-turbopuffer-sink](https://github.com/MaterializeInc/mz-turbopuffer-sink). + - Prior art: [mz-sink-sdk](https://github.com/MaterializeIncLabs/mz-sink-sdk) and its private-preview [build your own sink](https://materialize.com/docs/export-data/build-your-own-sink/) guide (#38872), [mz-redis-sync](https://github.com/MaterializeIncLabs/mz-redis-sync), [novu-materialize-integration](https://github.com/MaterializeIncLabs/novu-materialize-integration), [mz-turbopuffer-sink](https://github.com/MaterializeInc/mz-turbopuffer-sink). ## The Problem @@ -64,7 +64,11 @@ upgrade the upper catches up and the window slides with it, so a one-hour window can expire during a one-hour outage. A `REFRESH EVERY` view's upper jumps to the next refresh, so a sink that is down across a refresh can lose history however short the downtime. Retention belongs to the view's owner, and -any `ALTER` changes every sink's margin without telling it. +any `ALTER` changes every sink's margin without telling it. `RETAIN HISTORY` on +materialized views is also in private preview +(`doc/user/data/examples/create_materialized_view.yml`). Until durable +subscriptions land, the SDK depends on a preview feature, and sinks built on it +will raise demand for that feature. A connected subscribe holds its input readable up to what it has emitted, not up to what the consumer has committed. The data between those two points needs @@ -304,7 +308,8 @@ sequenceDiagram The SDK will separate two uses of `SUBSCRIBE`: - A live subscribe, for a UI, a cache warmer, or an agent watching results, - never resumes. It will accept any query. + never resumes. It will accept any query, including a temporal filter with + `mz_now()`. - A sink, or any durable subscription that checkpoints and resumes, will read one table, source, or materialized view directly, as `SUBSCRIBE ` with an optional envelope. It accepts no query, projection, or filter. These are the @@ -374,8 +379,8 @@ switches the workaround off once the fix ships. ### Checkpoints and fencing -A checkpoint store will have `load(name)` and -`commit(name, epoch, expected, frontier)`. The target is the recommended store, +A checkpoint store will have `load(sink_id)` and +`commit(sink_id, epoch, expected, frontier)`. The target is the recommended store, because writing data and checkpoint in one transaction gives exactly-once state (R1, R5). A Materialize-table store and a local file store will ship for targets with no transaction. Neither commits atomically with the target, so both give @@ -391,9 +396,11 @@ matters when a commit succeeds but its response is lost: the retry finds the stored frontier already past `expected`, rolls back, and the SDK reads the checkpoint and treats the batch as committed instead of applying it twice. -Every durable subscription will require an explicit name, which is its -checkpoint identity. A default derived from a class name would give two -deployments one checkpoint, and they would fence each other. Rows from several +Every durable subscription will require an explicit `sink_id`, which is its +checkpoint identity. It is chosen by the user and stays stable across restarts, +like a consumer group id, and it is not the name of the subscribed object. A +default derived from a class name would give two deployments one checkpoint, and +they would fence each other. Rows from several views will be tagged by view name, because tagging by list position remaps rows when the list is reordered. @@ -412,8 +419,7 @@ the shard a reconnect would switch objects silently. The user will choose one of two guarantees. Under `transactional`, the target commits the batch and the checkpoint together, so state is exactly once. Under `at_least_once`, effects run first and the checkpoint commits after. The SDK -will provide an idempotency key built from the subscription name, -`mz_timestamp`, the key columns, and an ordinal within the timestamp. An ordinal +will provide an idempotency key built from the `sink_id`, `mz_timestamp`, the key columns, and an ordinal within the timestamp. An ordinal within the batch would break deduplication, because batch boundaries move across a resume. @@ -427,7 +433,12 @@ will always be logged. A target that cannot represent a single row will call rest of the batch will commit. `commit_interval` will be the minimum time between commits, for targets that -prefer fewer, larger writes. +prefer fewer, larger writes. `max_commit_bytes` will bound one commit. A batch +after a catch-up can be large, and every timestamp in it is closed on its own, +so the SDK will split a batch at timestamp boundaries and commit the frontier at +each split. One timestamp larger than the limit, such as a snapshot, cannot be +split for a target that has to apply it atomically, so it goes through a +generation or a new incarnation instead (see "Sink module"). ### History loss and retention margin @@ -484,8 +495,8 @@ transaction also rolls back a stale worker still reading the old object. A target without such a transaction cannot fence those writes per document: a stale worker could write a key the new object never has, at a timestamp after the snapshot, and nothing would remove it. For such targets, turbopuffer among -them, `refollow` will write the new object into a new namespace and switch -readers when the snapshot completes. +them, every re-snapshot goes into a new incarnation of the target and switches +readers when it completes (see "Sink module"). The lookup and the `SUBSCRIBE` are two statements, so a swap between them would subscribe to the new object after the check passed. The SDK will therefore @@ -544,6 +555,14 @@ on one, using `mz_sources.envelope_type`, the envelope of source tables in `FROM SOURCE ... ENVELOPE MATERIALIZE` carries the envelope, not its source), and `mz_internal.mz_object_transitive_dependencies`. +Each member's snapshot is collected in `environmentd` up to `max_result_size` +(see "Buffering limits"), so N members starting together can hold N times that +at once. The SDK will start members one at a time at the shared `AS OF`, and +start the next only after the previous member's snapshot has drained, so at most +one snapshot sits in `environmentd`. Each member's retention has to cover the +time until it starts. Server-side chunking removes the limit itself (see +"Materialize-side workstream"). + A multi-view subscription will hold one connection per member. Any member error will end all members, and recovery will resume every member at the stored cut. If one member's history no longer covers the cut and the history-loss policy is @@ -551,7 +570,8 @@ If one member's history no longer covers the cut and the history-loss policy is expired member at a new timestamp `t_i` while the others resume from the cut, re-establish the cut at `t*`, the largest `t_i`, and buffer every stream until its progress passes `t*`. It will then apply in one step each re-snapshotted -member's snapshot, replacing its rows through a generation sweep, and every +member's snapshot, replacing its rows through a generation sweep or a new +incarnation (see "Sink module"), and every member's changes through `t*`. The cost is buffering. The live members must hold every change from the old cut to `t*`, which spans the whole outage, and the re-snapshotted members hold their changes between `t_i` and `t*`. Targets with @@ -617,9 +637,13 @@ call `commit(frontier)` inside the target's transaction, and optionally snapshot, resume, retry, fencing, and the history-loss policy. Keyed-state targets (caches, search indexes, tables) will read the upsert -envelope. When they re-snapshot, the SDK will write the snapshot under a new -generation and, at the end, call a sweep that removes keys from older -generations. The generation is a counter stored in the checkpoint, raised by a +envelope. How they re-snapshot depends on whether the target commits data and +checkpoint in one transaction. + +A target with such a transaction will re-snapshot in place: the SDK writes the +snapshot under a new generation and, at the end, calls a sweep that removes keys +from older generations. The epoch check in each transaction rolls back a stale +worker's writes, so nothing outside the snapshot can reach the target. The generation is a counter stored in the checkpoint, raised by a conditional commit before each re-snapshot attempt starts, and writes after the snapshot carry it too. A snapshot's `AS OF` cannot serve as the generation, because a retry can pick the same `AS OF` and a new object after a swap can @@ -629,6 +653,14 @@ That removes keys deleted while the sink was offline, which a plain re-snapshot leaves behind. A target that needs the snapshot to appear at once will make the new generation visible only at the sweep. +A target without such a transaction will re-snapshot into a new incarnation, a +fresh namespace, keyspace, or table, and switch readers when the snapshot +completes. In place, a key that was deleted before the snapshot's `AS OF` never +reaches the target, so nothing marks it deleted, and a stale worker's older write +of that key would create it with nothing left to remove it. Per-document +conditions cannot stop that write, because there is no document to compare +against. A new incarnation keeps every stale write in the retired one. + The upsert envelope emits `key_violation` rows when a key has more than one value. The SDK will pass them to `reject` by default, and a sink can choose `stall` instead, so a key violation never crashes the process. @@ -681,19 +713,12 @@ change from a request is safe. While tombstones are kept, every write is then safe to repeat and ordered by timestamp, so each namespace reaches exactly-once state under `at_least_once` delivery, and a batch can be split across requests. -A re-snapshot's generation sweep will turn documents from older generations into -tombstones instead of deleting them, so the same protection holds after a -re-snapshot. A patch by filter has no `$ref_new`, so it will set constants: -`deleted = true`, `mz_timestamp` to the snapshot's `AS OF` `t_s`, and -`written_at` to the current time, on documents matching -`generation < g AND deleted = false AND mz_timestamp < t_s`, where `g` is the -attempt's generation. The last clause -keeps the sweep from moving a timestamp backwards: a stale worker from before -the re-snapshot may have written a key at a time after `t_s`, the snapshot's -older upsert of that key is then skipped, and without the clause the sweep would -tombstone a correct document. The patch re-evaluates its filter before applying -(turbopuffer's guarantees page), so a document a live write moves into the new -generation meanwhile is left alone. +turbopuffer has no transaction around data and checkpoint, so the sink never +re-snapshots in place. Tombstones only protect keys the namespace has seen: a +key deleted before a re-snapshot's `AS OF` has no document and no tombstone, so +a stale worker's older upsert of it would be applied unconditionally. Every +re-snapshot, after history loss or a name swap, therefore writes into new +namespaces, described below. Searches will filter on `deleted = false`. By default the sink will keep tombstones forever, because removing one reopens the hole: a fenced worker that @@ -708,8 +733,8 @@ worker writes later than the grace period after it was fenced. The SDK cannot enforce that bound, because a process can pause between checking its epoch and sending a write, so the sink's docs will state the condition and the grace period has to exceed `commit_interval`, the retry budget, and the longest pause -the deployment can see. Filter operations are capped per call (5 million rows -for a delete by filter, 50 thousand for a patch by filter), so sweeps loop. +the deployment can see. A delete by filter is capped at 5 million rows per call, +so the sweep loops. The checkpoint will live in a separate checkpoint namespace, one document per sink, created once when the sink is set up. Workers will change it only with @@ -737,8 +762,8 @@ convergence: after each completed batch, every namespace equals the cut. The sink does not give readers a consistent view while a batch is being written. The sink's docs will state this, and deferred question 5 asks whether that is enough. -After a name swap, the sink writes the new object into new namespaces (see -"Object identity"). Namespace names carry an incarnation number, for example +Every re-snapshot, after history loss or a name swap (see "Object identity"), +writes into new namespaces. Namespace names carry an incarnation number, for example `articles__i3`, and the checkpoint document records the active incarnation and the list of retired ones. The checkpoint document is the reader contract: a reader looks up the active incarnation there, with a small helper the sink @@ -746,8 +771,8 @@ package ships, and derives every namespace name from it. When the snapshot completes, the sink switches readers by patching the active incarnation, one write to one document, so the switch is atomic for every namespace at once for any reader that looks it up again. Until then, readers of the old namespaces see -data frozen at the swap, and the sink's docs will say so. A stale worker still -reading the old object keeps writing only to the retired namespaces. The sink +data frozen at the start of the re-snapshot, and the sink's docs will say so. A +stale worker keeps writing only to the retired namespaces. The sink deletes them after a grace period, using the durable list in the checkpoint document. turbopuffer creates a namespace on its first write, so a stale write after that recreates an unused namespace. Retired incarnations therefore stay on @@ -943,7 +968,10 @@ The conformance suite is the contract, whatever the generation strategy. the target stays consistent. For turbopuffer, the stale worker also writes an older upsert of a key the live worker has deleted, and with tombstones kept the key must stay deleted. With the sweep on, the same holds when the - stale write comes within the grace period. + stale write comes within the grace period. A third case: the stale worker + pauses before an older upsert of a key, the key is deleted upstream, a + history-loss re-snapshot completes, and the stale worker resumes. The key must + not appear in the active namespaces. 5. Lost-response test: a transactional commit succeeds but its response is dropped, and the retry must find the frontier already committed and apply nothing twice. @@ -974,12 +1002,18 @@ program. 3. Server-side chunking: pass the updates between two progress messages on in bounded pieces, so a snapshot or catch-up larger than `max_result_size` can be delivered. Durable subscriptions do not change this, because a snapshot is one - timestamp. + timestamp. #38789 is a proof of concept that supports larger results and moves + snapshot processing out of `environmentd`. Running subscribes on clusters with + direct reads from persist would also suit sinks, which read storage-backed + objects only. 4. Durable subscriptions (#38468). 5. Snapshot elision for a projection and filter without a temporal predicate. #38468 needs it, and it would let sinks accept a projection and filter later without a snapshot read on every resume. -6. Non-poisoning subscribe errors (database-issues#5182). +6. Non-poisoning subscribe errors (database-issues#5182). #38916 proposes + letting a subscribe ignore errors. A subscribe that emitted both its data and + its errors would let a sink dead-letter bad rows (R3) instead of stopping on + `StreamPoisoned`. 7. Docs that cross-link the durable-subscriptions pattern from every client page now, and lead with the SDK once it ships. 8. A stable WebSocket `SUBSCRIBE`, the gate for browser and function transports. @@ -1056,11 +1090,32 @@ right, and this design keeps it. The parts that change: | One cursor per query | One `AS OF` for all members, release at the minimum frontier | Independent cursors produce torn cuts | | Rows stamped with the batch frontier | Every change keeps its `mz_timestamp` | Batches span timestamps | | Retractions come first | Order by timestamp, net per key within a timestamp | Not true by default | -| `name` defaults to the class name, rows keyed by list index | Required names, rows tagged by view name | Shared checkpoints fence each other, reordering remaps rows | +| `name` defaults to the class name, rows keyed by list index | Required `sink_id`, rows tagged by view name | Shared checkpoints fence each other, reordering remaps rows | | `commit_interval` in progress messages | Minimum time between commits | Progress cadence is not a user-facing unit | | Snapshot handling left to the user | Partial snapshot chunks, generation sweep, history-loss policy | Large snapshots and orphan keys are the common failure | | Conditional turbopuffer writes for replay safety | Tombstones plus conditional writes, with the checkpoint in its own namespace | A replayed or stale upsert recreates a key a later batch deleted | +### Extend mz-sink-sdk + +[mz-sink-sdk](https://github.com/MaterializeIncLabs/mz-sink-sdk) is a Python sink +SDK that the private-preview "build your own sink" guide already documents. It +shares most of this design's choices: a `sink_id` per deployment, a committed +frontier, fencing of a stale process, refusing to start when the query or schema +changed, a commit-interval policy, and disk spools for bounded memory. It ships +Postgres, Kafka, Iceberg, and Redis sinks. + +It differs in four ways that drive this design. It is Python only, with no +shared core for other languages. It subscribes to any `SELECT`, so a resume +rehydrates the query and reads its snapshot again. Its sinks must commit data and +checkpoint together, which leaves out targets such as turbopuffer and webhooks. +And it reads one subscription, with no consistent cut across several views. + +The two should converge rather than compete. The proposed direction is that the +Python package of this SDK takes over mz-sink-sdk's transactional sinks and its +sink interface, on top of the protocol core, and the guide moves to the SDK once +it ships. Deferred question 12 records this for agreement with the owners of +mz-sink-sdk. + ### Build only on durable subscriptions The semantics would be simplest, but nothing ships until the server does, and @@ -1155,3 +1210,11 @@ October work, because the doc picks a safe default for each. product. Is there demand that cannot wait? 10. Future-work modules. The rule for adding one is in "Future work". Owner: DevEx, per module, when evidence appears. +11. Schema changes. A change in output columns stops the sink by default. + Owner: DevEx, with the first users. Should the SDK offer a policy that + accepts added columns, dropping them on the client or passing rows through a + transform between the subscribe and sink modules, instead of stopping? +12. mz-sink-sdk. Owner: DevEx with the owners of mz-sink-sdk, before the Python + package starts. Do they agree that the Python package takes over its + transactional sinks and sink interface, and that the "build your own sink" + guide moves to the SDK (see "Extend mz-sink-sdk")? From 637caa24056f1e14a868b30c6f2318d12a93e541 Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Fri, 9 Oct 2026 22:01:59 +0300 Subject: [PATCH 16/19] design: Address review on many-sink memory, native embedding, and a Postgres sink before GA --- .../design/20260706_materialize_sdk.md | 86 +++++++++++++------ 1 file changed, 59 insertions(+), 27 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index 715c55ff7c0a4..bc7f78a9ab320 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -128,6 +128,12 @@ after long downtime can hit the same limit when the frontier advances in one step. The failure is deterministic, so a retry repeats it, and nothing the client does with chunks can avoid it. +Both limits apply per subscribe, and nothing bounds their sum +(`src/adapter/src/coord/message_handler.rs`, the backlog check per active +subscribe). Every sink holds its own subscribe, so many sinks on one environment +multiply the worst case: each can hold up to `max_result_size` while it takes a +snapshot, and up to `subscribe_max_buffered_bytes` if its client falls behind. + ### Errors | Condition | SQLSTATE | Message | @@ -473,7 +479,7 @@ The checkpoint fingerprint lets the SDK tell three kinds of change apart: | --- | --- | --- | --- | | Different output columns | the view's definition changed | stop with `SchemaMismatch` | stop | | Same columns, same storage shard, new catalog id | `ALTER MATERIALIZED VIEW ... APPLY REPLACEMENT` | re-run the retention-margin check, then resume | same | -| Same columns, new storage shard | `ALTER SCHEMA ... SWAP` in a blue/green deploy | stop | re-snapshot from the new object, into a new namespace for targets without transactions | +| Same columns, new storage shard | `ALTER SCHEMA ... SWAP` in a blue/green deploy | stop | re-snapshot from the new object, into a new incarnation (a new namespace in turbopuffer) for targets without transactions | | The name resolves to nothing | the object was dropped while the sink was offline | stop with `ObjectDropped` | stop | The SDK will look the name up in the catalog before each `SUBSCRIBE`, and it will @@ -678,12 +684,18 @@ The first sink will keep turbopuffer namespaces equal to views. It answers the search-index use case and will run against our internal context graph, which has no Kafka, so the existing Kafka-based sink does not fit there. The sink will write only to namespaces it creates, so every document carries `mz_timestamp`. +A turbopuffer namespace is a set of documents that are written and searched +together, roughly a table. + +The first version of the sink will be written in Rust on the Rust package and +will call turbopuffer's HTTP API directly. The tombstone, condition, and +checkpoint rules below do not depend on the language, so a Python version can +follow the same design. -The October sink will be written in Rust on the Rust package and will call -turbopuffer's HTTP API directly. Its transforms will be Rust functions that call -an embedding provider over HTTP, so the existing sink's Python transforms are not -reused. The tombstone, condition, and checkpoint rules below do not depend on -the language, so a Python version can follow the same design. +The sink will read the upsert envelope keyed by the document id. The server +emits a delete for a key only when no row has that key, so the sink needs no +multiplicity count, and a key with several rows arrives as `key_violation`, +which goes to `reject`. turbopuffer's documentation states that one write request to one namespace is applied atomically and is durable on return, and that there are no transactions @@ -778,20 +790,31 @@ document. turbopuffer creates a namespace on its first write, so a stale write after that recreates an unused namespace. Retired incarnations therefore stay on the list, and every later cleanup removes their namespaces again if they exist. -Embedding cost will follow the existing sink's transform model: a transform -declares the columns it reads, and runs only for rows where those columns -changed. The sink will store a hash of each transform's source +turbopuffer can compute embeddings itself. Its native embedding takes an `embed` +option on a string attribute, names one of the models turbopuffer hosts, bills +per token, and recomputes the vector on every write of that attribute. The sink +will use native embedding by default, so neither the sink nor Materialize runs +embedding code or holds model credentials. The cost follows writes of the text +attribute, so the sink will leave the text out of a write when it has not +changed, through the patch path below. A sink that needs a model turbopuffer does +not host can use a transform instead: Rust code that calls the embedding +provider and supplies the vector. + +Either way, embedding cost follows the existing sink's transform model: an +embedding declares the columns it reads, and is computed only for rows where +those columns changed. The sink will store a hash of each embedding's source columns on the document. For each netted change it will first make a patch of -the attributes without vectors, conditional on the stored `mz_timestamp` being -older, every stored source hash being equal to the new one, and the document not -being a tombstone. The patch keeps the stored vectors, which are then known to -match the source columns. If the patch count shows it was not applied, because a -source column changed, the document is missing or a tombstone, or a newer -version exists, the sink will compute the document's vectors and make the -conditional upsert. Comparing with the stored hash, not with the previous change -in the stream, stays correct when netting drops intermediate changes and after a -replay. Tombstones skip transforms. A replay re-runs transforms for the replayed rows, -which costs embedding calls but does not affect correctness. +the other attributes, leaving out vectors and the text that native embedding +reads, conditional on the stored `mz_timestamp` being older, every stored source +hash being equal to the new one, and the document not being a tombstone. The +patch keeps the stored vectors, which are then known to match the source +columns. If the patch count shows it was not applied, because a source column +changed, the document is missing or a tombstone, or a newer version exists, the +sink will make the conditional upsert with the text for native embedding, or with +vectors from a transform. Comparing with the stored hash, not with the previous +change in the stream, stays correct when netting drops intermediate changes and +after a replay. Tombstones skip embedding. A replay re-embeds the replayed rows, +which costs embedding but does not affect correctness. ### Durable subscriptions @@ -1045,11 +1068,17 @@ convergence test, a passing fencing test, and every typed error reproduced in the end-to-end suite. November and December: the Python and Node packages if they did not land in -October, durable subscriptions as they land, and a Redis or Postgres reference -sink chosen by demand. +October, durable subscriptions as they land, and a Postgres reference sink. The +turbopuffer sink exercises only the path for targets without transactions, so +Postgres is the sink that proves the transactional path: data and checkpoint in +one transaction, the expected-frontier check, and re-snapshot in place. MySQL +follows by demand. GA requires dedicated SQLSTATEs released, the client docs rewritten on the SDK, -and every published language passing the same vectors and end-to-end suite. +every published language passing the same vectors and end-to-end suite, the +Postgres sink passing the convergence and lost-response tests, and an answer to +snapshot size: either server-side chunking, or a documented per-sink size limit +with sizing guidance for environments that run many sinks. ## Future work @@ -1188,11 +1217,14 @@ October work, because the doc picks a safe default for each. carry `RETAIN HISTORY` onto staged views and replacement materialized views? For a replacement, does a running subscribe end at the switch, and does the new catalog id's readable frontier cover a checkpoint taken before it? -4. turbopuffer tombstones. Tombstones are kept forever by default and carry a - placeholder vector in namespaces with vectors. Owner: DevEx, with the first - turbopuffer users, during the sink build. Is the storage cost and the - placeholder vector acceptable, and what grace period should the docs - recommend for sinks that opt into the sweep? +4. turbopuffer tombstones and native embedding. Tombstones are kept forever by + default and carry a placeholder vector in namespaces with vectors. Owner: + DevEx, with the first turbopuffer users, during the sink build. Is the storage + cost and the placeholder vector acceptable, and what grace period should the + docs recommend for sinks that opt into the sweep? turbopuffer's docs do not + say how native embedding treats a patch that leaves the text attribute out, or + a document without text, such as a tombstone, so the sink build has to test + both first. 5. turbopuffer visibility. The sink documents convergence after each batch. Owner: DevEx, with the first search users. Do readers need a consistent view while a batch is written? That would need versioned documents and a From 74211d757207e9c39e1885d0ad59c89843e3865d Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Sat, 10 Oct 2026 00:07:29 +0300 Subject: [PATCH 17/19] design: Require a unique sink key, since the subscribe upsert envelope is stateless --- .../design/20260706_materialize_sdk.md | 35 +++++++++++++++---- 1 file changed, 28 insertions(+), 7 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index bc7f78a9ab320..ff529a5b4f024 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -667,9 +667,27 @@ of that key would create it with nothing left to remove it. Per-document conditions cannot stop that write, because there is no document to compare against. A new incarnation keeps every stale write in the retired one. -The upsert envelope emits `key_violation` rows when a key has more than one -value. The SDK will pass them to `reject` by default, and a sink can choose -`stall` instead, so a key violation never crashes the process. +The upsert envelope of `SUBSCRIBE` is stateless per timestamp. It turns one +timestamp's changes for a key into `upsert`, `delete`, or `key_violation` +without knowing the key's other rows (`process_response` in +`src/adapter/src/active_compute_sink.rs`). If the key is not unique in the +object, a removed row becomes a `delete` even while another row with that key +remains, and an added duplicate becomes an `upsert` that replaces the first +value. A `key_violation` appears only when a key's changes inside one timestamp +do not fit an insert, update, or delete. So a sink's key must be unique in the +object. `CREATE SINK` enforces this: the key must match a unique key Materialize +infers for the object, or the user writes `KEY (...) NOT ENFORCED` +(`src/sql/src/plan/statement/ddl.rs:3429-3445`). `SUBSCRIBE` only checks that +the key columns exist (`src/sql/src/plan/statement/dml.rs:1727-1750`). + +The SDK will apply the `CREATE SINK` rule. The catalog does not expose the +unique keys Materialize infers for a user object, so the check needs a server +change (see "Materialize-side workstream"). Until it lands, a sink must declare +its key as not enforced, the same acknowledgement `CREATE SINK` asks for, and +the docs will state that a key that is not unique makes deletes wrong. + +The SDK will pass `key_violation` rows to `reject` by default, and a sink can +choose `stall` instead, so a key violation never crashes the process. Event targets (webhooks, queues, notifications) will read the diff envelope under `at_least_once`, with the idempotency key above. They will declare what a @@ -692,10 +710,9 @@ will call turbopuffer's HTTP API directly. The tombstone, condition, and checkpoint rules below do not depend on the language, so a Python version can follow the same design. -The sink will read the upsert envelope keyed by the document id. The server -emits a delete for a key only when no row has that key, so the sink needs no -multiplicity count, and a key with several rows arrives as `key_violation`, -which goes to `reject`. +The sink will read the upsert envelope keyed by the document id, which must be +unique in the object (see "Sink module"). With a unique key, one document per +key is enough and the sink needs no multiplicity count. turbopuffer's documentation states that one write request to one namespace is applied atomically and is durable on return, and that there are no transactions @@ -1040,6 +1057,10 @@ program. 7. Docs that cross-link the durable-subscriptions pattern from every client page now, and lead with the SDK once it ships. 8. A stable WebSocket `SUBSCRIBE`, the gate for browser and function transports. +9. Upsert key validation for `SUBSCRIBE`: check the `ENVELOPE UPSERT` key + against the unique keys Materialize infers, with the same `NOT ENFORCED` + escape as `CREATE SINK`, or expose those keys in the catalog so the SDK can + check them. Until then, a sink's key is unchecked. The server-side buffering bound (#37905) has landed and needs no further work. From d6e09dbb710ec5ba518df6d3f40077efef07be3a Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Sat, 10 Oct 2026 13:43:19 +0300 Subject: [PATCH 18/19] design: Fix incarnation allocation, recovery, commit splitting, start errors, and embedding limits --- .../design/20260706_materialize_sdk.md | 232 ++++++++++++------ 1 file changed, 158 insertions(+), 74 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index ff529a5b4f024..e8a916c9e4f74 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -146,10 +146,10 @@ snapshot, and up to `subscribe_max_buffered_bytes` if its client falls behind. An error raised inside a running subscribe reaches the client as an unstructured adapter error, which maps to `XX000` -(`src/adapter/src/active_compute_sink.rs:283`, `src/adapter/src/error.rs:1066`). +(`src/adapter/src/active_compute_sink.rs:283`, `src/adapter/src/error.rs:1063`). So a dataflow error, a result-size failure, and a genuine internal error share one code. `53200` covers both `SubscribeFellBehind` and the adapter's -result-size error (`error.rs:1017-1018`), but a subscribe's size error takes the +result-size error (`error.rs:1014-1015`), but a subscribe's size error takes the `XX000` path, so during a subscribe `53200` means the client fell behind. The history-loss message has changed once already (#34712), which broke the one client that matched its @@ -177,15 +177,19 @@ this document adds. R4 is met in full by targets that can write the whole cut atomically, such as one Postgres transaction. A target with no transaction over the whole cut, such -as turbopuffer, where only one write request is atomic, gets convergence: after -each completed batch it equals the cut, but readers can see part of a batch -while it is written (see "turbopuffer sink"). +as turbopuffer, where only one write request is atomic, gets convergence: once +the sink has caught up past every write any worker made, the target equals the +cut at the committed frontier. Before that, readers can see part of a batch, and +after a crash they can see state newer than the checkpoint (see "turbopuffer +sink"). A user following the happy path cannot commit an unclosed timestamp, cannot get the `AS OF` arithmetic wrong, and cannot silently lose or duplicate updates -across a restart. The check is a sink that, after being killed at random points, -holds exactly `SELECT ... AS OF F - 1`, where `F` is its committed frontier. -The frontier is exclusive, so `AS OF F` would also include changes at `F`. +across a restart. The check kills a sink at random points. A target that commits +data and checkpoint together must then hold exactly `SELECT ... AS OF F - 1`, +where `F` is its committed frontier. The frontier is exclusive, so `AS OF F` +would also include changes at `F`. A target without that transaction must hold +the same once the sink has caught up and stopped writing. ## Out of Scope @@ -439,12 +443,17 @@ will always be logged. A target that cannot represent a single row will call rest of the batch will commit. `commit_interval` will be the minimum time between commits, for targets that -prefer fewer, larger writes. `max_commit_bytes` will bound one commit. A batch -after a catch-up can be large, and every timestamp in it is closed on its own, -so the SDK will split a batch at timestamp boundaries and commit the frontier at -each split. One timestamp larger than the limit, such as a snapshot, cannot be -split for a target that has to apply it atomically, so it goes through a -generation or a new incarnation instead (see "Sink module"). +prefer fewer, larger writes. `max_commit_bytes` will be a soft bound on one +commit. A batch after a catch-up can be large, and every timestamp in it is +closed on its own, so the SDK will split a batch at timestamp boundaries and +commit the frontier at each split. Netting per key happens only within one +piece: a piece that netted in a change from a later piece would commit a +frontier past data it has not written. The SDK never splits inside a timestamp. +A timestamp larger than the limit commits as one unit, in one transaction where +the target has one, or across several requests followed by one commit where it +does not. Only a snapshot is staged instead, through a generation or a new +incarnation (see "Sink module"), because staging replaces the whole state, which +an incremental timestamp must not do. ### History loss and retention margin @@ -483,13 +492,18 @@ The checkpoint fingerprint lets the SDK tell three kinds of change apart: | The name resolves to nothing | the object was dropped while the sink was offline | stop with `ObjectDropped` | stop | The SDK will look the name up in the catalog before each `SUBSCRIBE`, and it will -classify errors by the statement that raised them. A name that resolves to -nothing, and an id that no longer exists, both fail planning at `DECLARE` with -`XX000` (`src/adapter/src/error.rs:1003`), the same code a dataflow error has. So -any failure of `DECLARE` sends the SDK back to the name lookup and the table -above, and only errors raised during `FETCH` are classified as `StreamPoisoned` -or by the rest of "Typed errors". `42704` arrives only for a drop during a -running stream. +classify errors by when they happen. A cursor's statement is planned at +`DECLARE` and checked against the catalog again at its first `FETCH`, which +plans it again if the catalog changed in between (`verify_portal` in +`src/adapter/src/coord/command_handler.rs`). So a name that resolves to nothing, +or an id that no longer exists, fails with `XX000` +(`src/adapter/src/error.rs:1001`) at either point, the same code a dataflow +error has, and changed columns fail at the first `FETCH` with `0A000` +(`ChangedPlan`, `src/adapter/src/error.rs:942`). Any error before the cursor's +first progress row is therefore a start error: it sends the SDK back to the +name lookup and the table above. The one exception is `HistoryLost`, which keeps +its own classification. Errors after the first progress row are classified by +"Typed errors". `42704` arrives only for a drop during a running stream. A name swap needs a re-snapshot even when the new object's history covers the checkpoint. The target holds the old object's state at `F - 1`, and resuming the @@ -565,44 +579,58 @@ Each member's snapshot is collected in `environmentd` up to `max_result_size` (see "Buffering limits"), so N members starting together can hold N times that at once. The SDK will start members one at a time at the shared `AS OF`, and start the next only after the previous member's snapshot has drained, so at most -one snapshot sits in `environmentd`. Each member's retention has to cover the -time until it starts. Server-side chunking removes the limit itself (see -"Materialize-side workstream"). +one snapshot sits in `environmentd`. A member that has not started counts as +being at the shared `AS OF` in the joint minimum, so nothing is released past a +point it has not reached. Before the first member starts, the SDK checks that +every member is readable at the shared `AS OF`, and its retention check includes +the time the staggered start takes. The same one-at-a-time rule applies to every +catch-up, including recovery, not only the first start. Members that have +started can still each hold up to `subscribe_max_buffered_bytes`. Server-side +chunking removes the limit itself (see "Materialize-side workstream"). A multi-view subscription will hold one connection per member. Any member error will end all members, and recovery will resume every member at the stored cut. If one member's history no longer covers the cut and the history-loss policy is -`resnapshot`, the SDK will follow the recipe in #38468. It will re-snapshot each -expired member at a new timestamp `t_i` while the others resume from the cut, -re-establish the cut at `t*`, the largest `t_i`, and buffer every stream until -its progress passes `t*`. It will then apply in one step each re-snapshotted -member's snapshot, replacing its rows through a generation sweep or a new -incarnation (see "Sink module"), and every -member's changes through `t*`. The cost is buffering. The live members must hold -every change from the old cut to `t*`, which spans the whole outage, and the -re-snapshotted members hold their changes between `t_i` and `t*`. Targets with -generations will stage these changes in the target and make them visible at -`t*`, and other targets buffer them in memory. The catch-up can still hit -`max_result_size` (see "Buffering limits"). +`resnapshot`, recovery depends on the target. A target that commits data and +checkpoint in one transaction follows the recipe in #38468. The SDK +re-snapshots each expired member at a new timestamp `t_i` while the others +resume from the cut, re-establishes the cut at `t*`, the largest `t_i`, and +stages every stream in the target under a new generation until its progress +passes `t*`. It then makes in one step each re-snapshotted member's snapshot and +every member's changes through `t*` visible, and sweeps older generations. The +live members stage every change from the old cut to `t*`, which spans the whole +outage, and the catch-up can still hit `max_result_size` (see "Buffering +limits"). + +A target without that transaction re-snapshots every member, not only the +expired ones, at one new `AS OF`, into a new incarnation. Readers resolve one +incarnation for all of a sink's namespaces (see "turbopuffer sink"), so +re-snapshotting only some members would leave the others behind in namespaces +readers no longer use. Before durable subscriptions the SDK chooses that `AS OF` +itself, so the spread between `t_i` values does not arise. ### Typed errors | Error | Detected by | Default | | --- | --- | --- | -| `Transient` | no SQLSTATE, `08000`, `08001`, `08003`, `08004`, `08006` | reconnect through the connection factory and resume, within the retry budget, then stall | +| `Transient` | no SQLSTATE, `08000`, `08001`, `08003`, `08004`, `08006`, `25P03` | reconnect through the connection factory and resume, within the retry budget, then stall | | `CredentialsExpired` | `28000` | one reconnect through the connection factory, then `Fatal` | | `Canceled` | `57014` (a cancel or a timeout) | stop | | `FellBehind` | `53200` | resume from the last commit with backoff, report a metric | | `ResultTooLarge` | `XX000` plus the "exceeds max size" message | stop with remedy | | `HistoryLost` | `22000` plus the timestamp-selection message, after the index re-check | history-loss policy | | `ObjectDropped` | `42704` | see "Object identity" | -| `SchemaMismatch` | checkpoint fingerprint | see "Object identity" | -| `StreamPoisoned` | `XX000` during `FETCH` with any other message, which includes genuine internal errors | stop | +| `SchemaMismatch` | checkpoint fingerprint, or `0A000` before the first progress row | see "Object identity" | +| `StreamPoisoned` | `XX000` after the first progress row with any other message, which includes genuine internal errors | stop | | `Fenced` | checkpoint commit | stop | | `IndexedTarget` | startup check, and the re-check before `HistoryLost` | stop with remedy | | `Fatal` | anything else (auth, TLS, SQL) | stop | -Reconnects go through the connection factory. balancerd refuses connections +A session that sits idle inside its transaction for longer than +`idle_in_transaction_session_timeout` (two minutes by default, +`src/sql/src/session/vars/definitions.rs:391-393`) is ended with `25P03`. The +fetch loop never idles that long, so this means the process itself was paused, +and the SDK reconnects. Reconnects go through the connection factory. balancerd refuses connections with `08004` while `environmentd` restarts (`src/balancerd/src/lib.rs:964-1004`), so that is retried. The same code also means an unsupported protocol version, which is why every reconnect counts against the retry budget and a deterministic @@ -665,7 +693,9 @@ completes. In place, a key that was deleted before the snapshot's `AS OF` never reaches the target, so nothing marks it deleted, and a stale worker's older write of that key would create it with nothing left to remove it. Per-document conditions cannot stop that write, because there is no document to compare -against. A new incarnation keeps every stale write in the retired one. +against. A new incarnation keeps every stale write in the retired one. Each +incarnation belongs to one snapshot `AS OF`, recorded in the checkpoint before +its first write (see "turbopuffer sink"). The upsert envelope of `SUBSCRIBE` is stateless per timestamp. It turns one timestamp's changes for a key into `upsert`, `delete`, or `key_violation` @@ -675,14 +705,15 @@ object, a removed row becomes a `delete` even while another row with that key remains, and an added duplicate becomes an `upsert` that replaces the first value. A `key_violation` appears only when a key's changes inside one timestamp do not fit an insert, update, or delete. So a sink's key must be unique in the -object. `CREATE SINK` enforces this: the key must match a unique key Materialize -infers for the object, or the user writes `KEY (...) NOT ENFORCED` +object. `CREATE SINK` enforces this: the key must contain a unique key +Materialize infers for the object, or the user writes `KEY (...) NOT ENFORCED` (`src/sql/src/plan/statement/ddl.rs:3429-3445`). `SUBSCRIBE` only checks that the key columns exist (`src/sql/src/plan/statement/dml.rs:1727-1750`). The SDK will apply the `CREATE SINK` rule. The catalog does not expose the -unique keys Materialize infers for a user object, so the check needs a server -change (see "Materialize-side workstream"). Until it lands, a sink must declare +unique keys Materialize infers for a user object. `EXPLAIN ... WITH (keys)` +shows them, but its output is written for people and has no stable format, so +the check needs a server change (see "Materialize-side workstream"). Until it lands, a sink must declare its key as not enforced, the same acknowledgement `CREATE SINK` asks for, and the docs will state that a key that is not unique makes deletes wrong. @@ -786,37 +817,78 @@ cut spans one request per namespace at least. Between those requests a reader can see part of a batch: two keys that changed at the same timestamp, one new and one old, or one namespace at the new cut and another at the old one. Filtering on `mz_timestamp` cannot rebuild an earlier cut, because an upsert -replaces the earlier version. The guarantee for turbopuffer is therefore -convergence: after each completed batch, every namespace equals the cut. The -sink does not give readers a consistent view while a batch is being written. The -sink's docs will state this, and deferred question 5 asks whether that is enough. +replaces the earlier version. After a crash, a namespace can also be ahead of the +checkpoint: a worker may have written changes past the frontier it committed, +and the next worker skips them as already applied. The checkpoint is a replay +position, not a statement that the namespace equals one cut. The guarantee for +turbopuffer is therefore convergence: once the sink has caught up past every +write any worker made, and turbopuffer has indexed those writes, every namespace +equals the cut at the committed frontier. The sink does not give readers a +consistent view before that. The sink's docs will state this, and deferred +question 5 asks whether that is enough. Every re-snapshot, after history loss or a name swap (see "Object identity"), writes into new namespaces. Namespace names carry an incarnation number, for example `articles__i3`, and the checkpoint document records the active incarnation and the list of retired ones. The checkpoint document is the reader contract: a reader looks up the active incarnation there, with a small helper the sink -package ships, and derives every namespace name from it. When the snapshot -completes, the sink switches readers by patching the active incarnation, one -write to one document, so the switch is atomic for every namespace at once for -any reader that looks it up again. Until then, readers of the old namespaces see -data frozen at the start of the re-snapshot, and the sink's docs will say so. A -stale worker keeps writing only to the retired namespaces. The sink -deletes them after a grace period, using the durable list in the checkpoint -document. turbopuffer creates a namespace on its first write, so a stale write -after that recreates an unused namespace. Retired incarnations therefore stay on -the list, and every later cleanup removes their namespaces again if they exist. +package ships, and derives every namespace name from it. + +An incarnation belongs to one snapshot `AS OF`. Before its first write, a worker +records the new incarnation number and its `AS OF` in the checkpoint with a +patch conditional on its epoch, and a number is never given to a different +`AS OF`. A crashed attempt may continue in the same incarnation at the same +`AS OF` while that `AS OF` is readable, because the retry writes every change a +stale worker of that attempt could write, so the timestamp conditions skip the +stale worker's older writes. A retry at another `AS OF` allocates a new +incarnation, and the abandoned one joins the retired list. The first snapshot +follows the same rule. + +When the snapshot completes, the sink waits until turbopuffer has indexed the +new namespaces, because a namespace with more than 128 MiB of unindexed writes +hides later writes from queries for up to about an hour (turbopuffer's +guarantees page). It then switches readers with one patch that sets the active +incarnation, the committed frontier, and the retired list together, conditional +on the epoch and the expected frontier. The switch is atomic for every namespace +at once. The reader helper resolves the active incarnation for each query, or +caches it for less than the retirement grace period, and resolves it again when a +namespace is not found. Until the switch, readers of the old namespaces see data +frozen at the start of the re-snapshot, and the sink's docs will say so. A stale +worker keeps writing only to retired namespaces. The sink deletes them after a +grace period, using the durable list in the checkpoint document. turbopuffer +creates a namespace on its first write, so a stale write after that recreates an +unused namespace. Retired incarnations therefore stay on the list, and every +later cleanup removes their namespaces again if they exist. + +While a snapshot is being written, the sink writes changes after its `AS OF` as +they arrive instead of buffering them, since the timestamp conditions make the +order of writes irrelevant, and it commits the frontier only once the snapshot +is complete. A snapshot can take hours at turbopuffer's embedding limits (see +below), and buffering those changes would overflow the client buffer and +restart the snapshot. turbopuffer's `branch_from_namespace` makes an instant +copy of a namespace, which could let a re-snapshot reuse unchanged vectors +instead of embedding everything again. That needs its own correctness review and +is left for after the first version. turbopuffer can compute embeddings itself. Its native embedding takes an `embed` option on a string attribute, names one of the models turbopuffer hosts, bills per token, and recomputes the vector on every write of that attribute. The sink will use native embedding by default, so neither the sink nor Materialize runs -embedding code or holds model credentials. The cost follows writes of the text -attribute, so the sink will leave the text out of a write when it has not -changed, through the patch path below. A sink that needs a model turbopuffer does -not host can use a transform instead: Rust code that calls the embedding +embedding code or holds model credentials. A sink that needs a model turbopuffer +does not host can use a transform instead: Rust code that calls the embedding provider and supplies the vector. +Native embedding has limits the sink is built around. A write with embedded +attributes takes at most 256 rows, a namespace can have at most 4 embedded +attributes (turbopuffer's limits page), and a new organization gets 1,024 +requests and 2 million tokens per minute for each model (turbopuffer's embedding +page). A rate-limit response (HTTP 429) counts as throttling, not as a failure: +the sink paces its writes by tokens, and throttling does not use the retry +budget, so a long snapshot does not stall. turbopuffer's docs do not say whether +a patch that leaves the text out keeps the stored vector, or how a tombstone +without text is treated. The patch path below assumes the first, and the sink +build tests both before relying on them (deferred question 4). + Either way, embedding cost follows the existing sink's transform model: an embedding declares the columns it reads, and is computed only for rows where those columns changed. The sink will store a hash of each embedding's source @@ -825,13 +897,16 @@ the other attributes, leaving out vectors and the text that native embedding reads, conditional on the stored `mz_timestamp` being older, every stored source hash being equal to the new one, and the document not being a tombstone. The patch keeps the stored vectors, which are then known to match the source -columns. If the patch count shows it was not applied, because a source column -changed, the document is missing or a tombstone, or a newer version exists, the -sink will make the conditional upsert with the text for native embedding, or with -vectors from a transform. Comparing with the stored hash, not with the previous -change in the stream, stays correct when netting drops intermediate changes and -after a replay. Tombstones skip embedding. A replay re-embeds the replayed rows, -which costs embedding but does not affect correctness. +columns. The patch asks for `return_affected_ids`, so the response lists which +documents were patched. For the others, because a source column changed, the +document is missing or a tombstone, or a newer version exists, the sink will +make the conditional upsert with the text for native embedding, or with vectors +from a transform. Comparing with the stored hash, not with the previous change +in the stream, stays correct when netting drops intermediate changes and after a +replay. Snapshot writes into a new incarnation skip the patch path, because the +namespace starts empty and a patch costs an extra query per request. Tombstones +carry no text to embed. A replay re-embeds the replayed rows, which costs +embedding but does not affect correctness. ### Durable subscriptions @@ -1004,6 +1079,9 @@ The conformance suite is the contract, whatever the generation strategy. 3. Convergence test: kill a sink at random points (mid-batch, between write and commit, mid-snapshot), restart, repeat, then compare the target with `SELECT ... AS OF F - 1` for the committed frontier `F`, for both guarantees. + A target without a transaction around data and checkpoint is compared after + the sink has caught up and stopped writing, since until then it can be ahead + of its checkpoint. 4. Fencing test: two workers share one checkpoint, the stale one is refused, and the target stays consistent. For turbopuffer, the stale worker also writes an older upsert of a key the live worker has deleted, and with tombstones @@ -1097,9 +1175,10 @@ follows by demand. GA requires dedicated SQLSTATEs released, the client docs rewritten on the SDK, every published language passing the same vectors and end-to-end suite, the -Postgres sink passing the convergence and lost-response tests, and an answer to -snapshot size: either server-side chunking, or a documented per-sink size limit -with sizing guidance for environments that run many sinks. +Postgres sink passing the convergence and lost-response tests, upsert key +validation (workstream item 9) or a duplicate-key check in each keyed sink, and +an answer to snapshot size: either server-side chunking, or a documented +per-sink size limit with sizing guidance for environments that run many sinks. ## Future work @@ -1158,7 +1237,12 @@ It differs in four ways that drive this design. It is Python only, with no shared core for other languages. It subscribes to any `SELECT`, so a resume rehydrates the query and reads its snapshot again. Its sinks must commit data and checkpoint together, which leaves out targets such as turbopuffer and webhooks. -And it reads one subscription, with no consistent cut across several views. +And it gets a consistent cut across several views only by folding them into one +query, for example a `UNION` with a column naming each view, which reads the +whole snapshot again on every resume. Its runner also bounds memory by streaming +writes into an open target transaction, where this design keeps the target +transaction short and separate from fetching (see "Batches"). Deferred question +12 covers reconciling the two. The two should converge rather than compete. The proposed direction is that the Python package of this SDK takes over mz-sink-sdk's transactional sinks and its From e5ea2e163e7f711965c71020216733dfd60125e5 Mon Sep 17 00:00:00 2001 From: Bobby Iliev Date: Sat, 10 Oct 2026 13:45:40 +0300 Subject: [PATCH 19/19] design: Scope recovery sweeps and narrow start errors --- .../design/20260706_materialize_sdk.md | 33 +++++++++++-------- 1 file changed, 20 insertions(+), 13 deletions(-) diff --git a/doc/developer/design/20260706_materialize_sdk.md b/doc/developer/design/20260706_materialize_sdk.md index e8a916c9e4f74..f95219280b2f7 100644 --- a/doc/developer/design/20260706_materialize_sdk.md +++ b/doc/developer/design/20260706_materialize_sdk.md @@ -499,11 +499,15 @@ plans it again if the catalog changed in between (`verify_portal` in or an id that no longer exists, fails with `XX000` (`src/adapter/src/error.rs:1001`) at either point, the same code a dataflow error has, and changed columns fail at the first `FETCH` with `0A000` -(`ChangedPlan`, `src/adapter/src/error.rs:942`). Any error before the cursor's -first progress row is therefore a start error: it sends the SDK back to the -name lookup and the table above. The one exception is `HistoryLost`, which keeps -its own classification. Errors after the first progress row are classified by -"Typed errors". `42704` arrives only for a drop during a running stream. +(`ChangedPlan`, `src/adapter/src/error.rs:942`). So an `XX000` other than the +result-size error, or a `0A000`, before the cursor's first progress row is a +start error: the SDK repeats the name lookup and classifies the change through +the table above. If the object is unchanged, the original error keeps its own +classification. Every other error before the first progress row, such as a +dropped connection, a cancel, expired credentials, or `HistoryLost`, is +classified by "Typed errors" as it would be later. After the first progress row, +an `XX000` is a dataflow error. `42704` arrives only for a drop during a running +stream. A name swap needs a re-snapshot even when the new object's history covers the checkpoint. The target holds the old object's state at `F - 1`, and resuming the @@ -594,10 +598,12 @@ If one member's history no longer covers the cut and the history-loss policy is `resnapshot`, recovery depends on the target. A target that commits data and checkpoint in one transaction follows the recipe in #38468. The SDK re-snapshots each expired member at a new timestamp `t_i` while the others -resume from the cut, re-establishes the cut at `t*`, the largest `t_i`, and -stages every stream in the target under a new generation until its progress -passes `t*`. It then makes in one step each re-snapshotted member's snapshot and -every member's changes through `t*` visible, and sweeps older generations. The +resume from the cut, and re-establishes the cut at `t*`, the largest `t_i`. +Each re-snapshotted member writes its snapshot under a new generation scoped to +that member. The other members stage their changes as ordinary changes, not as +replacement state. In one step at `t*`, the SDK makes the new snapshots and every +member's changes through `t*` visible and sweeps older generations of the +re-snapshotted members only, so unchanged rows of the other members stay. The live members stage every change from the old cut to `t*`, which spans the whole outage, and the catch-up can still hit `max_result_size` (see "Buffering limits"). @@ -1330,10 +1336,11 @@ October work, because the doc picks a safe default for each. say how native embedding treats a patch that leaves the text attribute out, or a document without text, such as a tombstone, so the sink build has to test both first. -5. turbopuffer visibility. The sink documents convergence after each batch. - Owner: DevEx, with the first search users. Do readers need a consistent view - while a batch is written? That would need versioned documents and a - visible-cut pointer that readers filter on. +5. turbopuffer visibility. The sink documents convergence: once it has caught up + past every write and turbopuffer has indexed them, each namespace equals the + cut at the committed frontier. Owner: DevEx, with the first search users. Do + readers need a consistent view before that? That would need versioned + documents and a visible-cut pointer that readers filter on. 6. Repository home. The doc proposes keeping the packages under `misc/` in this repository. Owner: DevEx with the engineering leads, once the spec settles: do the packages move to their own repository, keeping the protocol core,