Repository navigation
applier: flush drained batches column-wise with the unique-move fallback - #149
Conversation
Add Flusher, the write half of pkg/applier. Flush applies one drained Batch in one bounded read-write transaction under the table lock session, guarded as the copier's chunk inserts are (owner role, catalog-only search_path, ACCESS SHARE on both relations, lock confirmation, relation-OID check). The flush first completes every image Batch.CompleteFirst names: from the old key's shadow row, read FOR UPDATE so the batch's own delete of that key cannot race it; from the source row under Key when the old key was reused or its shadow row is absent; an image the source has since removed is skipped (D13). It then writes column-wise inside a savepoint: a delete marker as a DELETE, a whole image as an upsert, a marker-bearing image of an unmoved row as an UPDATE of only its present columns, where zero rows updated is a CO-8 refusal rather than an insert that invents the omitted value. On SQLSTATE 23505 the flush rolls back to the savepoint and reapplies the batch whole-row (CO-6): every remaining marker completed from the shadow row under its own key, every key in the batch deleted in one statement, every surviving image inserted plain. A second 23505 returns ErrBatchDeferred with the shadow untouched; Buffer.Requeue puts the batch back for a later drain and refuses, fail closed, if any of its keys is held. The flush pins extra_float_digits so a float lands as the value the verifier will hash. Result counts images, deletes, completions by source, and skips, and reports whether the fallback ran. Integration tests cover the column-wise path, the omitted-column UPDATE, the seats unique exchange converging in one flush, completion from the shadow and from the source (including a reused old key), a moved row that is gone, a batch a unique index still refuses, the owner role, a replaced shadow, and a lost lock. SAFETY.md, the design's package map, and CO-4/CO-6/CO-8 now describe the flush.
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
🤖 1/2: adversarial correctness review of 2 blocking, 3 non-blocking. The column-wise path, the savepoint fallback and BlockingB1. A skipped image leaves the deleted predecessor's row under When a key-moving UPDATE lands on a key whose row was deleted earlier in the same window,
The fix is local. Write a skipped image as a delete at Test that fails on e93aa1d and passes with the fixfunc everyKeyLanded() copier.Position {
return copier.Position{Watermark: copier.NewWatermark(math.MaxInt64), Cut: copier.NewWatermark(math.MaxInt64)}
}
func moveEvent(lsn decode.LSN, from, to int64, cols ...decode.Column) decode.ChangeEvent {
return decode.ChangeEvent{Kind: decode.Update, LSN: lsn, Key: to, OldKey: &from, Columns: cols}
}
// Row 3 moves onto key 2 after the row there was deleted, then on to key
// 10; the copier read key 3's chunk after the first move. The first flush
// has nothing to complete key 2 from, and the second still completes the
// move to key 10 with row 3's own document, not the deleted row's.
func TestFlushSkippedMoveLeavesNoPredecessorRow(t *testing.T) {
f := newFlushFixture(t)
f.createSeats(t, 3)
p := f.prepare(t, "seats")
doc3, _ := f.text(t, "seats", "doc", 3)
f.exec(t, `DELETE FROM %s.`+p.shadow.ShadowTable()+` WHERE id = 3`)
f.exec(t, `DELETE FROM %s.seats WHERE id = 2`)
f.exec(t, `UPDATE %s.seats SET id = 2 WHERE id = 3`)
f.exec(t, `UPDATE %s.seats SET id = 10 WHERE id = 2`)
buffer := applier.NewBuffer()
require.NoError(t, buffer.Add(decode.ChangeEvent{Kind: decode.Delete, LSN: 1, Key: 2}))
require.NoError(t, buffer.Add(moveEvent(2, 3, 2, col("slot", "C"), marker("doc"), col("note", "seat 3"))))
result, err := p.flush.Flush(t.Context(), f.pool, buffer.Drain(everyKeyLanded()))
require.NoError(t, err)
assert.Equal(t, 1, result.Skipped)
require.NoError(t, buffer.Add(moveEvent(3, 2, 10, col("slot", "C"), marker("doc"), col("note", "seat 3"))))
_, err = p.flush.Flush(t.Context(), f.pool, buffer.Drain(everyKeyLanded()))
require.NoError(t, err)
got, _ := f.text(t, p.shadow.ShadowTable(), "doc", 10)
assert.Equal(t, *doc3, *got, "the moved row keeps its own document")
f.assertConverged(t, p.shadow)
}
B2. A source completion can take another row's value, and a later flush spreads it. The source row under The reuse flag has the same blind spot on the shadow side. D13 relies on "the buffer sees the reuse" (L373). But the copier reads chunks live. It can copy the old key's chunk after another row has moved in, while the stream has not yet delivered that row's event. This is a D13 decision, so I am not offering a patch. The gate needs this property: a completion read counts only once the buffer holds every event the read could have seen. One shape is to read Test that fails on e93aa1d (fix is a D13 decision)// Row 3 moves to key 10 and on to 11, and a new row is inserted at key 10,
// before the first flush reads the source. Every key ends holding its own
// row's document.
func TestFlushSourceCompletionNeverTakesAnotherRowsValue(t *testing.T) {
f := newFlushFixture(t)
f.createSeats(t, 3)
p := f.prepare(t, "seats")
doc3, _ := f.text(t, "seats", "doc", 3)
f.exec(t, `DELETE FROM %s.`+p.shadow.ShadowTable()+` WHERE id = 3`)
f.exec(t, `UPDATE %s.seats SET id = 10 WHERE id = 3`)
f.exec(t, `UPDATE %s.seats SET id = 11 WHERE id = 10`)
f.exec(t, `INSERT INTO %s.seats (id, slot, doc, note) VALUES (10, 'Z', 'the other row''s document', 'seat 10')`)
buffer := applier.NewBuffer()
require.NoError(t, buffer.Add(moveEvent(1, 3, 10, col("slot", "C"), marker("doc"), col("note", "seat 3"))))
_, err := p.flush.Flush(t.Context(), f.pool, buffer.Drain(everyKeyLanded()))
require.NoError(t, err)
require.NoError(t, buffer.Add(moveEvent(2, 10, 11, col("slot", "C"), marker("doc"), col("note", "seat 3"))))
_, err = p.flush.Flush(t.Context(), f.pool, buffer.Drain(everyKeyLanded()))
require.NoError(t, err)
require.NoError(t, buffer.Add(decode.ChangeEvent{Kind: decode.Insert, LSN: 3, Key: 10,
Columns: []decode.Column{col("slot", "Z"), col("doc", "the other row's document"), col("note", "seat 10")}}))
_, err = p.flush.Flush(t.Context(), f.pool, buffer.Drain(everyKeyLanded()))
require.NoError(t, err)
got, _ := f.text(t, p.shadow.ShadowTable(), "doc", 11)
assert.Equal(t, *doc3, *got, "the moved row keeps its own document")
f.assertConverged(t, p.shadow)
}
Non-blockingN1. Every move to a lower key on a table with a unique index takes the fallback. The column-wise pass writes in key order. A row that moves 3→0 keeps its slot, so the upsert at 0 collides with the row still under 3, which the delete marker at 3 has not removed yet. That costs the whole batch the fallback: a completion read Test that fails on e93aa1d and passes with deletes firstfunc TestFlushMoveToALowerKeyStaysColumnWise(t *testing.T) {
f := newFlushFixture(t)
f.createSeats(t, 3)
p := f.prepare(t, "seats")
doc3, _ := f.text(t, "seats", "doc", 3)
f.exec(t, `UPDATE %s.seats SET id = 0 WHERE id = 3`)
result, err := p.flush.Flush(t.Context(), f.pool, batch(
moved(3, 0, col("slot", "C"), col("doc", *doc3)),
deleted(3),
))
require.NoError(t, err)
assert.False(t, result.Fallback, "only the row's own old key held its slot")
f.assertConverged(t, p.shadow)
}
N2. An exclusion constraint has the CO-6 problem but gets neither the fallback nor Preflight admits a table with Test that fails on e93aa1d and passes with 23P01 in the fallbackfunc TestFlushConvergesAnExclusionExchange(t *testing.T) {
f := newFlushFixture(t)
f.exec(t, `CREATE TABLE %s.seats (id bigint PRIMARY KEY, slot text, doc text, note text, EXCLUDE USING btree (slot WITH =))`)
f.exec(t, `INSERT INTO %s.seats VALUES (1, 'A', 'd1', 'n1'), (2, 'B', 'd2', 'n2')`)
p := f.prepare(t, "seats")
f.exec(t, `UPDATE %s.seats SET slot = NULL WHERE id = 1`)
f.exec(t, `UPDATE %s.seats SET slot = 'A' WHERE id = 2`)
f.exec(t, `UPDATE %s.seats SET slot = 'B' WHERE id = 1`)
result, err := p.flush.Flush(t.Context(), f.pool, batch(
image(1, col("slot", "B"), col("doc", "d1")),
image(2, col("slot", "A"), col("doc", "d2")),
))
require.NoError(t, err)
assert.True(t, result.Fallback)
f.assertConverged(t, p.shadow)
}
N3. Several branches of the copied guard have no flusher test, so nothing pins them in this package's copy. Each of these mutants survives the whole
The veto list keeps the guard per package so that each copy is reviewed where it lives. That works only if each copy carries its own tests. Porting the copier's search_path and replaced-shadow-after-claim tests would close most of the list. Verified
Mutants ran against
This review was generated by Claude Code (claude-opus-5-5). |
|
🤖 2/2: OSS adoption and integration ease, at For adoption, this is the first piece of pg-sprite that writes change data, and it reads well. Each SQL shape is a small builder pinned against its literal statement. The savepoint fallback is ten lines that a reader can check against CO-6. The integration tests state each source change as SQL next to the batch the buffer would drain, so a newcomer can see what every vector means. The "Decisions to veto" list made this review faster: B1 and B2 in 1/2 are both answers to one of its items. For SchemaBot and other importers, nothing calls 1. Nothing bounds a batch, and a batch is one transaction with one round trip per entry.
2. The completions a host should trust least are reported as counts, without keys.
3. The guard now has three copies, and this one is the least tested. The veto list keeps a guard per package, so that a change to one package's session setup is reviewed in that package. For someone maintaining a fork, three hand-kept copies of a security-relevant sequence (role, catalog-only This review was generated by Claude Code (claude-opus-5-5). |
aparajon
left a comment
There was a problem hiding this comment.
🤖 Stamping with comments: 2 blocking, 6 non-blocking (see the two review comments above). Please address B1 (write a skipped moved image as a delete at its key) and B2 (gate source completion so it cannot take another row's value, and correct D13's "overlays it" claim) before merging.
This stamp was left by Claude Code (claude-opus-5-5).
morgo
left a comment
There was a problem hiding this comment.
🤖 Automated adversarial review, posted on Morgan Tocker's behalf.
Approving. The write half is carefully ordered, and the two orderings that matter most are both right: completing a moved image from the old key's shadow row before anything in the batch deletes that key, and deleting every key in one statement before inserting any of them so the batch cannot collide with itself on a unique index. Refusing a zero-row UPDATE as a CO-8 violation rather than converting it to an insert is the right instinct — inventing a value for an omitted TOAST column is exactly the failure that would be invisible until a verifier run.
Verified rather than assumed:
extra_float_digits = 3is the one rendering setting that genuinely needs pinning, and the others genuinely do not.rowSQLreads columns::textand writes the text straight back, so output lossiness matters — and float output is the only case where the output function can lose information the input function cannot recover.DateStyle,IntervalStyleandbytea_outputchange the form of the text but not its information content, and output and input happen in the same session with the same settings, so they round-trip exactly. Pinning one GUC and not the others looks arbitrary at first read; it is correct.- The
$1::bigint[]cast indeleteAllSQLmatches the copier's established convention, not a new assumption.pkg/copier/chunker.godeclares its parameters bigint "whatever the key's integer type" with the operator-family reasoning written out, so the applier is consistent with it. - The fallback completes markers before it deletes. Reading
FOR UPDATEunder the table lock makes the read-then-delete sequence safe without depending on the isolation level, which the transaction does not set. Requeuerefusing when the buffer already holds one of the batch's keys is the right invariant: it makes "nothing mayAddbetweenDrainand the flush's outcome" checkable rather than conventional.
A deferred batch has no bound and no way to tell "waiting on the copier" from "stuck forever"
ErrBatchDeferred → Requeue → drain again. The reasoning given is that the caller "drains again once the copier has landed the chunk the rows wait on", which is a sound account of the expected cause and the only cause the code can recover from. Nothing distinguishes that case from one that will never resolve.
Concretely: the shadow holds a row under key 3 with slot = 'C' from an earlier chunk copy. A batch wants key 1 to take slot = 'C'. The column-wise UPDATE hits 23505; the whole-row fallback deletes only the batch's own keys, so row 3 survives and the insert hits 23505 again; the batch defers. If the event that would have freed 'C' was skipped — and the flush has a Skipped counter, so skipping is a real outcome — or if the shadow has drifted for any other reason, no later chunk and no later batch will clear it. The stream stops advancing and nothing raises an error. Buffer.OldestPending quietly stops moving.
The scheduling loop is explicitly a later PR, and that is the right place for the detection, so this is not a change I would hold this PR for. But it should be named now while the reasoning is fresh: a deferral count on the batch, or a check that OldestPending advances within some bound, turns a silent stall into a reportable condition. Without one, the difference between "the copier is behind" and "the shadow has drifted" is invisible from the outside, and those need opposite responses.
Non-blocking
split treats an unnamed column and a value-less column as the same thing, and they are not. A value-less column is a legitimate unchanged-TOAST marker: the correct handling is to leave that column alone. An unnamed column is a decoder defect. Folding the second into the first means a decoding bug does not surface as an error — it surfaces as a column silently retaining its old value in the shadow, which is precisely the drift that the deferral above cannot recover from, discovered later by a verifier with no way to trace it back. Given how much care the rest of this package takes to refuse rather than guess (zero-row UPDATE, absent shadow row in the fallback, a missing source row on completion), an unnamed column reaching split reads like it belongs in the same category: refuse it.
deleteAllSQL's comment says "the bigint array" without the reasoning the copier's equivalent carries. pkg/copier/chunker.go explains that parameters are declared bigint whatever the key's integer type and why the operator family makes that safe. A reader of pkg/applier alone sees a hardcoded ::bigint[] next to four other builders that bind $1 with no cast at all, and cannot tell whether that is a deliberate widening or an assumption nobody checked. A one-line cross-reference would settle it.
The key-only upsert writes a real row version. DO UPDATE SET pk = EXCLUDED.pk makes the row exist either way, which is the point, but on the conflict branch it is a genuine update: a new tuple, WAL, index maintenance, and any triggers on the shadow. For a key-only image whose row already exists with the right key, nothing needed to change. DO NOTHING would not be equivalent if the intent is to lock or touch the row, so if the write is deliberate it is worth saying so in the comment on upsertSQL; if it is not, this is a free saving on a path the applier takes often.
…read A moved, marker-bearing image is completed from the old key's shadow row or the source row under its key. Neither row is known to be the stream's: both are live, so a later source transaction can already have put another row's value there. The flush no longer writes such a completion. It reads the value together with pg_current_wal_insert_lsn() in the same statement and returns the image in Result.Held; Buffer.Hold keeps it buffered with the completion pending, any event that replaces the key or reuses its old key drops the completion, an UPDATE keeps it pending behind its own values, and Buffer.Release(passed) fills the markers still left once the stream has delivered every change committed through the read. The next Drain flushes the whole image (D13, CO-8). A moved image neither row can complete is written as a delete marker at its key, named in Result.Skipped, instead of being left out: the source holds no row under that key, so a row the shadow may hold there is stale. The column-wise pass writes every delete marker before any image, so a row that moved to a lower key while keeping a unique value finds its old row gone before its upsert. The fallback now also catches SQLSTATE 23P01, since an exclusion constraint over a moved value fails like a unique index. CompleteFirst returns copies of the batch's entries, never pointers into them, and the guard tests prove the shadow proof check, the catalog-only search_path, the extra_float_digits pin on a completion read, and that an empty batch opens no transaction. Docs: D13, the pkg/applier package-map row, CO-4, CO-6, CO-8, doc.go.
|
🤖 Adversarial review response — created by Kiran's code review agent (Amp, Claude Opus 4.6) — pull/149, follow-up commit Both blocking findings are fixed by one change to D13: a moved image's completion is held in the buffer with the WAL position of its read and written only once the stream has passed it; a skipped image is a delete at its key. N1, N2, and L2 are fixed; N3 is fixed for the branches that matter to correctness and the rest, with L1 and L3, are tracked as internal follow-ups.
Source: block/pg-sprite#149, review comments 6031869738 and 6031870515 and review 5438186038 at head |
…bm/cs7-pgoutput * origin/kiran01bm/cs7-slot: decode: prove the server, verify the publication's shape, type the refusals schemachange: re-derive the swapped-table proof from the catalog for the post-swap resume (#146) migrate: let only the caller's cancellation end the unknown-outcome test's lock wait (#145) applier: flush drained batches column-wise with the unique-move fallback (#149) # Conflicts: # SAFETY.md # docs/copy-and-swap-design.md # docs/invariants.md # pkg/decode/doc.go
Why
pkg/appliercould merge and drain but not write.Buffer.Drainhands back aBatchof landed entries — delete markers, whole images, and marker-bearing images whose unchanged-TOAST columns still need completing, withCompleteFirstnaming the moved ones — and nothing consumed it. This PR adds the write half:Flusher, which applies one batch in one guarded transaction under the table lock, column-wise where it can (CO-5, CO-8) and whole-row when a unique index or exclusion constraint refuses the column-wise form (CO-6, D13), plusBuffer.Requeuefor the batch a constraint still refuses andBuffer.Hold/Buffer.Releasefor the moved images whose completion the flush read from a row that can be newer than the stream.The scheduling that makes a chunk copy and a backlog flush mutually exclusive, and the stream owner that calls
Drain→Flush→Hold→Releasein a loop, are the next PRs. This one is the SQL and its transaction.Before / after
seats(id PRIMARY KEY, slot UNIQUE, doc EXTERNAL)holding(1,'A',…),(2,'B',…)in the shadow; the source swaps the slots in one transaction and leavesdocuntouched, so both decoded images carrydocas a marker:A moved image,
UPDATE seats SET id = 10 WHERE id = 2captured withdocas a marker. The old key's shadow row was copied live and the source row is live, so either can already hold a change the stream has not delivered — a different row'sdoc, if a later source transaction put another row at the key. The flush reads the completion but does not write it; the buffer holds the image until the stream has passed the read:What
pkg/applier/flusher.go—Options{LockTimeout, StatementTimeout}(defaults; sub-millisecond refused asErrInvalidOptions),Result{Images, Deletes, Completed, Held []HeldImage, Skipped []int64, Fallback},ErrBatchDeferred.NewFlusher(target, shadow, lock, opts)runs the copier's proof and shadow checks (ST-6).Flush(ctx, pool, batch)is a no-op for an empty batch; otherwise binds the lock session's context, begins one read-write transaction with the guard, reads a completion for everyCompleteFirstimage intoResult.Held(or writes a delete at its key intoResult.Skipped), applies the rest column-wise in a savepoint — every delete marker before any image — falls back whole-row on23505or23P01, commits, and reports a lost lock over any other error.pkg/applier/hold.go—HeldImage{Entry, Completed, FromSource, ReadLSN};Buffer.Hold(held)puts each image back with its completion pending (a key the buffer holds isErrInvariantViolation, CO-5, buffer unchanged);Buffer.Release(passed) intfills the markers still left on every image whose read is at or belowpassed;Buffer.Held(). AnyAddthat replaces a held key's entry or reuses a held image's old key drops the pending completion;Drainkeeps a held image buffered and counts it inBatch.Held.pkg/applier/complete.go—completeFromreads the named row's columns astexttogether withpg_current_wal_insert_lsn()in the same statement; the shadow read isFOR UPDATE, the source read plain.pkg/applier/drain.go—CompleteFirstreturns copies of the batch's entries;Batch.Held.pkg/applier/flush_sql.go— the statements as free builders over sanitized identifiers:upsertSQL(present columns; a key-only image becomesDO UPDATE SET pk = EXCLUDED.pkso the row exists either way),updateSQL(present columns,WHERE pk = $1),insertSQL(whole row, no conflict clause),deleteSQL,deleteAllSQL(pk = ANY($1::bigint[])),rowSQL(each column::text, plus the WAL position);splitputs an image's columns in the shadow's column order and treats an unnamed or value-less column as a marker.pkg/applier/guard.go— the package's copy of the copier's guard (owner role, catalog-onlysearch_path, bounded timeouts,extra_float_digits = 3,ACCESS SHAREon both relations, lock confirmation, relation-OID confirmation, lock-lost reporting), now with its own tests: a shadow that is not the target's is refused atNewFlusher, a decoy=operator on the session'ssearch_pathis ignored, a connection configured atextra_float_digits = 0still completes afloat8[]marker exactly, and an empty batch takes no connection from the pool.pkg/applier/requeue.go—Buffer.Requeue(batch)puts a deferred batch back as the flush left it; a key the buffer already holds isErrInvariantViolation(CO-5) and the buffer is unchanged.SAFETY.mdapplier row; design D13 (the hold gate replaces the "a later event overlays it" reasoning; deletes first;23P01; a skipped image is a delete) and package map;docs/invariants.mdCO-4 (what is enforced now vs. the mutual-exclusion scheduling still planned), CO-6 (Enforced), CO-8 (Enforced);pkg/applier/doc.go.Decisions to veto
pg_current_wal_insert_lsn()bounds every commit the read could have seen; the image waits in the buffer until the stream owner has delivered through that position and callsRelease. An event for the key meanwhile wins: an UPDATE overlays its own values and the completion fills only the markers left; anything that replaces the key's entry or reuses the old key drops the completion. The alternative — write the read value and let a later event overlay it — is unsound, since a later move away from the key reads the shadow row instead of overlaying it.copier.Positionas it is. A per-chunk read position can be added toPositionlater without changing the applier's contract.Key. The source holds no row under that key, so whatever the shadow holds there is stale, and the key-moving UPDATE that landed on the key may have replaced the delete marker that was the only record of a row to remove. Deleting is right whatever the shadow holds; a later move away from the key finds no row and falls through to the source. The batch's own entry is left as it was soRequeuecan put the image back.23P01is a collision like23505. Preflight admits exclusion constraints, and an exclusion over a moved value fails the same way a unique index does; the fallback converges it. ADEFERRABLE INITIALLY DEFERREDunique constraint reports atCOMMITand would surface as a commit error rather thanErrBatchDeferred; tracked as an internal follow-up.Result.HeldandResult.Skippedcarry the entries, not counts. The stream owner needs the held images to callHold; an operator triaging a later verifier mismatch needs the skipped keys.pkg/applier, not shared. Same convention aspkg/checksum: each writing package carries its own guard so a change to one package's session setup is reviewed in that package, and this copy now has the tests that pin it. A shared internal package besidechunksqlwould save ~100 lines and couple three packages' transaction shape; tracked as an internal follow-up.textand writes the text back. The completed value is bound as atextparameter and the server casts it to the column's type on write, the same round trip the decoder's text-mode values take. Reading typed values would need per-type scanning for no gain.extra_float_digits = 3is pinned on the flush session too. The verifier hashes the text rendering; a float the flush rendered at the session's default could land one digit off what the decoder saw and read as a mismatch. The completion read runs in the same transaction, so a heldfloat8value carries every digit.UPDATE, not an upsert. The marker stands for a value only a present shadow row holds; an upsert would insert the row with the omitted column invented. Zero rows updated is a CO-8 refusal.ErrBatchDeferred, not a retry and not a refusal. With every key in the batch deleted first, a remaining violation means the colliding row is outside the batch — a key an in-flight chunk is about to land or a deferred entry still holds. The copier moving resolves it; the caller requeues and drains later. A retry inside the flush would spin on the same state.RequeueandHoldrefuse a held key and run before anyAddafter theDrain. The stream owner's loop is drain → flush → (requeue | hold + confirm) → add; a key both the batch and the buffer hold means that order broke, and merging the two would pick an arbitrary winner.Batch.OldestFirstLSNof a batch whose flush committed andBuffer.OldestPendingof what has not, which covers held images; the stream owner that confirms is the next PR's.Flush, which is one transaction with one round trip per entry. The bound belongs with the stream owner (Drain(pos, limit)keeping pairs together) and set-based statements are the natural next step for the round trips; tracked as an internal follow-up.23505when a flush landed a row carrying a unique value a chunk is about to copy; the copier'sON CONFLICT (pk) DO NOTHINGdoes not cover a secondary unique index. The mutual-exclusion scheduling is where that is decided. Tracked as an internal follow-up.Verification
make lint0 issues;go vet ./pkg/applier/;go build ./...;SKIP_INTEGRATION=1 go test -race ./pkg/applier/green.go test -race -count=1 ./pkg/applier/green on PostgreSQL 14, 16, and 18 (~50s each);scripts/test-flaky.sh8/8 on the hold, skip, and empty-batch tests.doc(SET STORAGE EXTERNAL,ToastBytes() > 0proven) is left in place byte-for-byte; a marker for a row the shadow lacks isErrInvariantViolation (CO-8); theseatsexchange{1→'B', 2→'A'}converges in one flush withFallbackset, and so does the same exchange underEXCLUDE USING btree (slot WITH =); a move to a lower key keeping its slot stays column-wise; the fallback completes markers from the shadow; a batch colliding with a row outside it returnsErrBatchDeferred,Requeuetakes it back, and the next flush converges once the colliding row is in the batch; an image the fallback cannot complete is refused.docand converges afterRelease; a move whose old key never landed is held with the source's; a reused old key completes from the source; a move whose row is gone writes a delete and names the key inSkipped; a row that moves onto a deleted key and on again keeps its owndoc(the skipped delete removes the predecessor); a row that moves 3→10→11 while another row is inserted at 10 before the first flush reads the source — the first hold carries the other row'sdoc, the stream's move away from 10 drops it, and key 11 ends with row 3's own; a held shadow completion is dropped when the stream delivers a reuse of the old key; an UPDATE of a held key keeps the completion pending andReleasefills only the markers it left.NewFlusherrefuses a shadow that is not the target's (nil lock, other table's lock); a decoy=operator on the session'ssearch_pathis ignored; a connection atextra_float_digits = 0still completes afloat8[]marker holding0.1 + 0.2with every digit; an empty batch takes no connection; the writes run as the table owner; a replaced shadow is refused (ST-6); a gone lock is reported as lost (LK-1). Removing thesearch_pathpin or theextra_float_digitspin fromguard.gofails the corresponding test.Holdkeeps the image until the stream passes the read, releases only the completions at or belowpassed, refuses a held key, fills only markers, drops on delete/move/reuse, copies the image and its completion, no-op on empty;CompleteFirstreturns copies; the SQL builders against literal expected statements;splitfollows the shadow's column order;Requeueround-trips a batch, refuses a held key, no-op on empty; option defaults and refusals;23505and23P01matched by SQLSTATE only.