Skip to content

migrate: let only the caller's cancellation end the unknown-outcome test's lock wait - #145

Merged
Kiran01bm merged 5 commits into
mainfrom
kiran01bm/accept-blocking-cancel-test
Oct 8, 2026
Merged

Kiran01bm merged 5 commits into
mainfrom
kiran01bm/accept-blocking-cancel-test

Conversation

@Kiran01bm

@Kiran01bm Kiran01bm commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

TestRunAcceptBlockingReportsAnUnknownOutcomeHonestly could observe a statement-budget failure instead of the unknown outcome it asserts, and when it did, the test hung on pool close until the package deadline instead of failing. Both budgets the test relies on are now widened, and the pool is closed from a cleanup registered before the lock holder's. Review round 1 then closed the product edge the flake exposed: BlockingBudget requires the statement bound to be longer than the lock bound. Review round 2 closed what that rule could not: the executor now takes the acknowledged table's lock as a statement of its own before the DDL, so a statement-budget cancellation on the accepted path is never an ungranted lock wait.

Why

The test holds ACCESS EXCLUSIVE on orders, runs DROP INDEX through RunAcceptBlocking, waits for pg_stat_activity to show the statement waiting on the lock, then cancels the context and asserts *BlockingOutcomeUnknownError. It widens lock_timeout to 30s so that only the caller's cancellation can end the wait. statement_timeout also counts the lock wait, and it was left at the fixture's 5s default. Whenever the poll took more than 5s to see the waiting statement (a loaded runner is enough), the server cancelled the statement first; the context was not yet cancelled, so acceptedBlockingStatementError classified 57014 as BudgetError{CauseStatement} and require.ErrorAs failed.

That failure then became a 10-minute timeout rather than a red assertion: defer pool.Close() ran before the holder's t.Cleanup rolled back the lock transaction, so Close waited for a connection that would only be returned after Close finished. Every test in pkg/migrate for that PostgreSQL version was lost with it.

What

  • opts.Budget.Brief.StatementTimeout widened alongside the existing LockTimeout line (and kept the longer of the two); the comment states why both budgets have to be widened.
  • defer pool.Close() → t.Cleanup(pool.Close) registered before the holder's cleanup, so cleanup LIFO order rolls back the holder first and the pool closes without waiting; a failed assertion now fails the test.

Review round 1

The flake was reachable in production too: an operator running --accept-blocking with an explicit --statement-timeout shorter than --lock-timeout (3s default) against a contended table gets budget-statement-exceeded / exit 1 for a lock that was never granted, where the design promises a lock-budget refusal / exit 2.

  • BlockingBudget.validate refuses a statement bound that is not longer than the lock bound (ErrInvalidBlockingBudget, before a session is acquired). Round 2 reframes this as a consistency check — a pair in the other order could never let the lock budget apply — rather than the guarantee it was first described as.
  • --accept-blocking help, AB-1 in docs/invariants.md, the budgets and failure-semantics sections of docs/lock-budgeted-passthrough.md, and the invalid-blocking-budget row in docs/execution-model.md state the ordering rule.
  • New TestRunAcceptBlockingRefusesAStatementBudgetThatCouldEndTheLockWait: lock 10s, statement 1s, table locked by another session → ErrInvalidBlockingBudget, zero verdict, index stands. TestBlockingBudgetValidation gains equal / shorter / one-millisecond-longer cases. The statement-cutoff subtest moves its bound from 100ms to 200ms so it stays above the fixture's 100ms lock budget while the 500ms rebuild still overruns it.

Review round 2

lock_timeout bounds each lock acquisition on its own; statement_timeout counts them all together. Every admitted statement acquires more than one lock in sequence — DROP INDEX the table then the index, REINDEX INDEX the table then the index, REINDEX TABLE the table then each index — so the ordering rule did not close the window: a first wait granted late followed by a second one is cut off by the statement bound before the second wait's own lock bound, and the 57014 reads as work that ran.

  • ExecuteAcceptedBlocking takes the schema-qualified table the caller resolved from the catalog (pgx.Identifier) and runs LOCK TABLE <table> IN ACCESS EXCLUSIVE MODE as its own statement after SET LOCAL and before the DDL. Every way that wait can end arrives while nothing has been submitted: 55P03, and a 57014 of the request itself (reachable only when it recurses to partitions), are the lock-budget refusal; the caller's context ending is cancelled-by-caller with a verdict that can say nothing was submitted, not the unknown outcome. Holding the table's ACCESS EXCLUSIVE lock means no other session holds a conflicting lock on any of its indexes, so the DDL's own requests are granted at once and its statement budget measures work.
  • migrate.lockedTable returns the table's relkind alongside its name; a materialized view, which LOCK TABLE cannot name, is the one target passed as nil, and the statement acquires its own locks there. The acceptBlocking comment, AB-1, AB-2, the session / budgets / failure-semantics sections of the passthrough design, and the cancelled-by-caller row of docs/execution-model.md describe the lock request and stop claiming the ordering rule is what keeps a lock wait from being reported as work.
  • Tests. TestExecuteAcceptedBlockingLocksTheTableBeforeTheStatement is the reviewer's scenario: REINDEX INDEX behind a writer that releases the table after 1s and a reader whose open transaction keeps the index locked, budgets lock 1.5s / statement 1.6s — with the pre-lock a 55P03 lock refusal at 1.5s, without it (run once with nil, not committed) a CauseStatement failure at 1.6s. …ReportsACancelledLockWaitAsNothingSubmitted at both layers: cancel during the LOCK TABLE wait → ErrCancelledByCaller, code cancelled-by-caller, index stands. TestRunAcceptBlockingReportsAnUnknownOutcomeHonestly now cancels a running slow REINDEX INDEX (the only place the unknown outcome can still arise) instead of a lock wait. REINDEX TABLE on a materialized view commits through both layers without a pre-lock. The remaining inline holder rollbacks carry a one-line comment on why the cleanup's second rollback is harmless.

Decisions to veto

  • ACCESS EXCLUSIVE for every admitted form, including REINDEX, whose own table lock is SHARE. In practice REINDEX already blocks every new query on the table — planning opens every index of a table under a lock the index's ACCESS EXCLUSIVE conflicts with — so the stronger table lock adds nothing an operator who acknowledged the table has not accepted, and it is what makes the index requests free. Taking the statement's own mode instead would leave REINDEX TABLE with one lock wait per index inside the statement budget.
  • No ONLY: on a partitioned table the lock request recurses to the partitions, which DROP INDEX of a partitioned index locks in turn anyway, and a 57014 of that request is classified as the lock refusal with the statement budget named as the one that cut it off.
  • A context cancelled during the lock wait is a failed verdict with code cancelled-by-caller (the existing code for the caller's own cancellation), not a refusal; the detail says nothing was committed.
  • The ordering rule is enforced in the executor's budget validation only; the CLI does not duplicate the check.
  • Two review-round-1 suggestions are tracked as internal follow-ups rather than taken here: a testutil.NewPool that registers t.Cleanup(pool.Close); and splitting blocking-outcome-unknown into "lost before COMMIT" and a genuinely ambiguous commit failure, which changes AB-5 and is a design decision of its own.

Before / after

Round 2, REINDEX INDEX with a writer on the table (releases at 1s) and a reader on the index, lock 1.5s / statement 1.6s:

Before                                               After
──────                                               ─────
BEGIN; SET LOCAL lock_timeout/statement_timeout      BEGIN; SET LOCAL lock_timeout/statement_timeout
REINDEX INDEX i                                      LOCK TABLE t IN ACCESS EXCLUSIVE MODE
  t SHARE ........ waits 1.0s (writer)  granted        t AEL ....... waits (writer, reader)
  i AEL   ........ waits (reader)                      t+1.5s  55P03  → budget-lock-exceeded, exit 2
  t+1.6s  57014  → budget-statement-exceeded, exit 1   REINDEX INDEX i   never submitted
          "work ran" — nothing did

Round 1 (unchanged), slow pg_stat_activity poll before cancel():

Before                                            After
──────                                            ─────
t+5s  server fires 57014 (statement_timeout)      cancel()
      → BudgetError{CauseStatement}                     → BlockingOutcomeUnknownError
      require.ErrorAs fails                             test passes
      defer pool.Close() waits on holder conn
t+10m package deadline; panic, all tests lost

Tests

  • PG_VERSION=16 scripts/test-flaky.sh TestExecuteAcceptedBlockingLocksTheTableBeforeTheStatement 5 ./pkg/executor/ — 5/5; the same test with the pre-lock disabled fails with CauseStatement at 1.6s.
  • PG_VERSION=16 scripts/test-flaky.sh TestExecuteAcceptedBlockingReportsACancelledLockWaitAsNothingSubmitted 5 ./pkg/executor/ and PG_VERSION=16 scripts/test-flaky.sh 'TestRunAcceptBlockingReportsAnUnknownOutcomeHonestly|TestRunAcceptBlockingReportsACancelledLockWaitAsNothingSubmitted' 5 ./pkg/migrate/ — 5/5 each.
  • PG_VERSION=16 go test ./pkg/executor/ ./pkg/migrate/ ./internal/cli/ -race -count=1, SKIP_INTEGRATION=1 go test ./..., make lint — clean.

…est's lock wait

TestRunAcceptBlockingReportsAnUnknownOutcomeHonestly widens lock_timeout
so that cancelling the context is the only thing that can end the DROP
INDEX's wait behind the held lock. statement_timeout counts the lock
wait too, and it was left at the fixture's 5s default. When the
pg_stat_activity poll took longer than that to see the waiting
statement, the server cancelled it first and the test observed a
statement-budget failure instead of the unknown outcome it asserts.
Widen statement_timeout with lock_timeout.

Register the pool close as a cleanup before the lock holder's cleanup
rather than as a defer. The holder keeps a connection until its cleanup
rolls it back; a deferred Close ran first and waited on that connection,
so a failed assertion parked the test until the package deadline instead
of failing it.
@Kiran01bm
Kiran01bm marked this pull request as ready for review October 5, 2026 03:21
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@aparajon

aparajon commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

🤖 1/2: adversarial correctness review of 99e9549. I read the diff and the code the test drives: executeAcceptedBlocking and acceptedBlockingStatementError in pkg/executor, the verdict mapping in migrate.acceptBlocking, and testutil.NewSchema. I checked them against AB-1, AB-2 and AB-5 and the design's failure semantics. Every run below used real PostgreSQL 16.

0 blocking, 2 non-blocking.

The fix is right, and it fixes the root cause rather than buying time. The test is about the caller's cancellation. A 5s statement budget gave the server a second way to end the wait, and the test never meant to allow one. Widening it states the test's intent; nothing in the system under test gets more time to be slow. I reproduced both halves of the description with a 6s sleep injected before cancel():

1f46fb9 (base) 99e9549 (head)
6s slow poll panic: test timed out after 1m0s, parked in pgxpool.(*Pool).Close ok (7.27s)
no injection, -count=10 ok ok

The test still pins the safety property. The verdict is failed, never committed. It never carries the "nothing committed" wording. Its code is blocking-outcome-unknown. Each assertion still kills a production mutant (table under Verified). AB-5 is upheld. The test is named in AB-5's Enforced line and keeps its name, so the citation still resolves.

The cleanup reorder fixes a second thing as well. On base, NewSchema's DROP SCHEMA cleanup ran after defer pool.Close() and logged drop throwaway schema t_51089_1: closed pool, so on a shared PG_DSN server this test leaked its schema. At head the drop runs before the close and succeeds.

Non-blocking

1. The flake was a real product edge, and nothing pins it now: a lock that is never granted, cut off by a shorter statement budget, is reported as a statement failure, not a refusal. accept_blocking.go:103-107, accepted_blocking.go:139-151, lock-budgeted-passthrough.md:360-363

This PR is right to take the test off that path: the test is about cancellation. The path itself is still reachable, though. statement_timeout counts the lock wait, and BlockingBudget.validate does not require the statement budget to exceed the lock budget. Take an operator who runs --accept-blocking with an explicit --statement-timeout 2s (a plain DROP INDEX is quick) and leaves --lock-timeout at its 3s default. If the table is contended for longer than 2s, the server cancels the statement with 57014 before the lock is granted. That is classified CauseStatement, so the run fails with exit 1 and budget-statement-exceeded.

The design says otherwise: "failure before the statement starts … remains a typed refusal with exit code 2 because nothing ran". The comment at L104-106 says "the lock was granted", and in this case it was not. This is fail-closed, since nothing executed and nothing is reported as committed, so it is not blocking. The cost is that the operator is told the work needs a bigger statement budget when the table was just busy.

The smallest fix is in AB-1's validation: refuse a statement budget that is not longer than the lock budget, before a session is acquired. Then statement_timeout cannot fire during a single lock wait. The test below asserts the refusal form. If validation is the fix, its assertion becomes require.ErrorIs(t, err, executor.ErrInvalidBlockingBudget).

Test that fails on 99e9549
// An accepted statement whose lock is never granted executes nothing,
// whichever bound ends the wait: when the statement budget is the shorter
// of the two it is still a lock that was never granted (AB-2, AB-5).
func TestRunAcceptBlockingReportsAnUngrantedLockAsARefusalUnderEitherBound(t *testing.T) {
	url := testutil.StartPostgres(t)
	pool, err := dbconn.NewPool(t.Context(), dbconn.Config{URL: url})
	require.NoError(t, err)
	t.Cleanup(pool.Close)
	schema := testutil.NewSchema(t, pool)
	_, err = pool.Exec(t.Context(), fmt.Sprintf(`
		CREATE TABLE %[1]s.orders (id int PRIMARY KEY);
		CREATE INDEX orders_id_idx ON %[1]s.orders (id)`, schema))
	require.NoError(t, err)
	holder, err := pool.Begin(t.Context())
	require.NoError(t, err)
	t.Cleanup(func() { _ = holder.Rollback(context.WithoutCancel(t.Context())) })
	_, err = holder.Exec(t.Context(), fmt.Sprintf("LOCK TABLE %s.orders IN ACCESS EXCLUSIVE MODE", schema))
	require.NoError(t, err)

	opts := acceptBlockingOptions(schema + ".orders")
	opts.Budget.Brief.LockTimeout = 10 * time.Second
	opts.Budget.Brief.StatementTimeout = time.Second
	v, err := migrate.Run(t.Context(), pool,
		parseOne(t, fmt.Sprintf("DROP INDEX %s.orders_id_idx", schema)), opts)

	require.NoError(t, err)
	assert.Equal(t, verdict.OutcomeRefused, v.Outcome)
	assert.Equal(t, verdict.CauseLockBudget, v.Cause)
	require.NoError(t, holder.Rollback(t.Context()))
	assert.True(t, indexExists(t, pool, schema, "orders_id_idx"))
}

--- FAIL: TestRunAcceptBlockingReportsAnUngrantedLockAsARefusalUnderEitherBound (1.47s) on 99e9549, with outcome=failed code="budget-statement-exceeded" and the error execution exceeded its statement budget (1s) and was cancelled. The lock was never granted.

2. The cleanup-order fix applies to many other tests, which still leak their schemas on a shared server. 71 tests call defer pool.Close() and then testutil.NewSchema(t, pool) at the top level: 35 in the CLI package, and the rest in pkg/schemadiff, pkg/diffplan, pkg/preflight and pkg/migrate. Their DROP SCHEMA cleanups run on a closed pool, the same way this test's did on base. None of them holds a connection in a cleanup, so none of them can hang. They do leave their schemas behind on CI's PG_DSN matrix server and on a long-lived make test-db database. A testutil.NewPool(t, url) that registers t.Cleanup(pool.Close) would fix all of them at once, and new tests could not get the order wrong. This belongs outside this PR.

Verified

  • Not hiding a backend wait. executeAcceptedBlocking returns BlockingOutcomeUnknownError straight from the cancelled Exec. It never follows the backend, so the cancelled run returns in about 1s, and widening the budget does not mask a slow return.
  • Not a timeout bump under AGENTS.md. The widened value is the bound under test, not a wait for the system to finish. require.Eventually and the server budget are both 30s, so a poll slow enough to race the server fails Eventually first.

Each mutant was run against the head test. All edits were restored with git checkout.

Mutant Result
T1: 6s slow poll, StatementTimeout line removed, cleanup fix kept FAIL in 7.2s (expected: *executor.BlockingOutcomeUnknownError): fails red instead of hanging
T2: 6s slow poll, defer pool.Close() restored, budget fix kept ok: the budget fix alone removes the flake; the cleanup fix only changes how a failure ends
P1: drop the ctx.Err() check in acceptedBlockingStatementError FAIL (BlockingOutcomeUnknownError not in chain)
P2: drop the unknown-outcome Detail override in acceptBlocking FAIL (detail contains "nothing committed")
P3: drop the CodeBlockingOutcomeUnknown mapping FAIL (execution-failed)
P4: report the cancelled statement as success (return nil) FAIL (An error is expected but got nil)
P5: drop the deferred tx.Rollback killed only by the package deadline (pool Close waits on the leaked transaction)

This review was generated by Claude Code (claude-opus-5-5).

@aparajon

aparajon commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

🤖 2/2: OSS adoption and integration ease, at 99e9549. These are lenses, not correctness findings. 0 blocking, 1 non-blocking.

For adoption, this is the kind of test fix contributors should copy. The comment explains why both budgets have to move and why the pool close has to be registered before the holder, so the next person who adds a lock-holding test can see the trap. A suite that hangs for 10 minutes and takes every other test in the package down with it is one of the fastest ways to lose an outside contributor's trust in CI. This PR turns that hang back into a red assertion. 1/2 non-blocking 2 is the same lesson applied to the rest of the suite.

1. An importer would get more from "lost before COMMIT" than from "unknown". accepted_blocking.go:113-119, lock-budgeted-passthrough.md:365-368

This test pins blocking-outcome-unknown for a cancel that lands while the statement is still in Exec. At that point the client has not sent COMMIT. PostgreSQL commits an explicit transaction block only on COMMIT, so the transaction rolls back whatever the backend does next. The test shows it: once the holder rolls back, the index is still there. The design doc draws the same line. It says a statement cancellation rolls back, and only "a dropped connection at the commit boundary is ambiguous". AB-5's wording ("lost at or after submission") is broader, and the code follows AB-5.

For an orchestrator such as SchemaBot the difference matters. blocking-outcome-unknown means "stop and send an operator to the catalog". A cancel before COMMIT could instead mean "nothing committed, safe to retry". Splitting the two would leave the genuinely ambiguous case (tx.Commit failing, L117-119) as the only unknown outcome. That would be a change to AB-5 and to this test's subject, so it is a design decision for its own PR, not something for this one. Keeping it conservative is fail-closed and correct as it stands.

This review was generated by Claude Code (claude-opus-5-5).

@aparajon aparajon left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Stamping with comments: 0 blocking, 3 non-blocking (see the two review comments above).

This stamp was left by Claude Code (claude-opus-5-5).

… the lock wait

statement_timeout counts the lock wait, so a statement budget that is not
longer than the lock budget could cancel an accepted statement whose lock
was never granted. That cancellation arrives as SQLSTATE 57014 and the
accepted path classifies it as a statement failure (exit 1,
budget-statement-exceeded), although nothing ran and the design promises
a lock-budget refusal (exit 2) for an ungranted lock.

BlockingBudget.validate now requires the statement bound to be longer than
the lock bound, so lock_timeout is the only bound that can end a lock wait
and a statement-budget cancellation on this path always means the work had
started. The pair is refused before a session is acquired, as
ErrInvalidBlockingBudget, like every other invalid bound.

The CLI help, AB-1, the passthrough design doc and the outcome-code table
state the ordering rule. The tests that set budgets keep the statement
bound the longer one, and a new integration test holds the table's lock
and proves the inverted pair is refused with no verdict and no DROP.
@Kiran01bm

Copy link
Copy Markdown
Collaborator Author

🤖 Adversarial review response — created by Kiran's code review agent (Amp, Claude Opus 4.6) — pull/145, follow-up commit

One of the three non-blocking findings is fixed in this PR with the reviewer's smallest fix; the two that the review itself placed outside this PR are tracked as internal follow-ups. The reviewer's test was taken in the validation form the review proposed. PG_VERSION=16 go test ./pkg/executor/ ./pkg/migrate/ ./internal/cli/ -race -count=1 is green, the new and changed pkg/migrate tests pass 5/5 under scripts/test-flaky.sh, and SKIP_INTEGRATION=1 go test ./... and make lint are clean.

# Finding Status Explanation
C1-F1 A lock never granted but cut off by a shorter statement_timeout is reported as budget-statement-exceeded / exit 1, where the design promises a lock-budget refusal / exit 2 fixed Taken as the review proposed, in AB-1's validation: BlockingBudget.validate refuses a statement bound that is not longer than the lock bound as ErrInvalidBlockingBudget, before a session is acquired. With statement_timeout > lock_timeout, lock_timeout is the only bound that can end a lock wait, so a 57014 on this path always means the lock was granted and the work had started — which is what the acceptBlocking comment now says, instead of asserting it. TestRunAcceptBlockingRefusesAStatementBudgetThatCouldEndTheLockWait is the reviewer's test with its assertion in the validation form: lock 10s, statement 1s, table locked by another session → ErrInvalidBlockingBudget, zero verdict, index stands. TestBlockingBudgetValidation gains equal / shorter / one-millisecond-longer cases. The --accept-blocking help, AB-1, the budgets and failure-semantics sections of the passthrough doc, and the invalid-blocking-budget row in execution-model.md state the rule. The existing tests that set budgets keep the statement bound the longer one: this PR's own test moves to 30s / 60s, and the statement-cutoff subtest moves from 100ms to 200ms against the fixture's 100ms lock budget, still well under its 500ms rebuild. The CLI does not duplicate the check.
C1-F2 Many other tests still defer pool.Close() before testutil.NewSchema and leak their schemas on a shared PG_DSN server deferred As the review says, this belongs outside this PR: a testutil.NewPool(t, url) that registers t.Cleanup(pool.Close) touches every package's tests and is its own change. Tracked as an internal follow-up and named in the PR body. The one test this PR touches keeps the cleanup-registered close; the new test uses the same shape.
C2-F1 An importer would get more from "lost before COMMIT, safe to retry" than from blocking-outcome-unknown deferred Agreed on the direction and on the review's own placement: splitting the cancelled-in-Exec case from the failed-tx.Commit case changes AB-5 and this test's subject, so it is a design decision for its own PR. Tracked as an internal follow-up and named in the PR body; the conservative mapping stands here.

Verified — no action: AB-5 upheld as the review found; the mutant table's results are unchanged by this commit, and the cleanup-order fix the review confirmed is kept.

Fixes refer to the two review comments at 99e9549: 5987651487 (1/2 correctness), 5987652135 (2/2 OSS adoption and integration); approving review 5409771888.

@morgo morgo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🤖 Automated adversarial review, posted on Morgan Tocker's behalf.

Approving. Turning a flake investigation into a product fix is the right instinct, and the round-1 finding is a real one: a 57014 on this path genuinely could mean a lock that was never granted, and that misreports a refusal as work that ran. The fix is in the right place (executor validation, before a session is acquired), the docs and AB-1 were updated together rather than left to drift, and the test for the new rule asserts the whole refusal contract — zero verdict, nothing submitted, index still standing with the table locked by someone else.

One note below is load-bearing: the ordering rule does not actually establish the invariant AB-1 now claims for it. The rest is minor.

Verified rather than assumed:

  • The cleanup reordering is correct, and for the stated reason. Deferred calls in a test body run before the testing framework invokes t.Cleanup functions, so the old defer pool.Close() really did run ahead of the holder's rollback and really would park on a connection that could only come back afterwards. Registering pool.Close as the first cleanup puts it last in LIFO order, behind both the holder's rollback and the schema drop that testutil.NewSchema registers after it — and the schema drop still has a live pool when it runs.
  • The widened budgets are self-consistent. 60s statement against 30s lock satisfies the new rule, and the statement-cutoff subtest's move to 200ms clears the fixture's 100ms lock budget while the ~500ms rebuild still overruns it.
  • The default flag pair is not broken by the new validation. --lock-timeout defaults to 3s and --statement-timeout to 30s, so an operator who overrides neither still validates comfortably. This was my first concern — a validation added below a pair of defaults that happened to be equal would have failed every invocation — and it does not apply.

The ordering rule does not make lock_timeout the only bound that can end a lock wait

This is the claim the invariant, the doc section and the new acceptBlocking comment all now rest on:

Requiring statement_timeout > lock_timeout makes lock_timeout the only bound that can end a lock wait, so an ungranted lock is always the lock-budget refusal

lock_timeout applies separately to each lock acquisition attempt, not to the statement's total time spent waiting. statement_timeout is cumulative from when the backend receives the statement. So the implication only holds for a statement that acquires exactly one lock.

All three admitted forms acquire more than one. DROP INDEX takes ACCESS EXCLUSIVE on the index's owning table and on the index itself; REINDEX INDEX the same; REINDEX TABLE takes the table and then each of its indexes.

Concrete failure, using the pair the new test explicitly blesses as valid (lock + 1ms):

  • lock_timeout = 10s, statement_timeout = 10.001s, operator runs DROP INDEX orders_id_idx.
  • The backend waits 9s for ACCESS EXCLUSIVE on orders — under the 10s lock bound, so no 55P03 — and is granted it.
  • It then begins waiting for the lock on the index.
  • At t = 10.001s statement_timeout fires. The index lock wait has been running ~1s, nowhere near its own bound.
  • 57014 arrives with a lock still ungranted. acceptBlocking classifies it CauseStatement, and the operator gets budget-statement-exceeded and exit 1 for a statement that never started its work — the exact outcome round 1 set out to eliminate, and a violation of the exit-2 refusal contract in lock-budgeted-passthrough.md.

The requirement the invariant actually needs is statement_timeout > N × lock_timeout, where N is the number of lock acquisitions the statement makes — not knowable from the parsed statement, since REINDEX TABLE's N depends on how many indexes the table has.

To be clear about blame: this gap predates the PR, and requiring the statement bound to be the longer one strictly reduces the window. What is new is the assertion that the window is closed. AB-1, the failure-semantics section and the comment in acceptBlocking all now tell the next reader that a 57014 here is proof the lock was granted, which means nobody re-derives it.

Two ways out, ranked:

  1. Take the lock explicitly, first, in the engine-owned transaction. LOCK TABLE <table> IN ACCESS EXCLUSIVE MODE as its own statement, then the DDL. A 55P03 on the LOCK is unambiguously a lock refusal; a 57014 on the DDL is unambiguously work that ran. No inference from budget ordering at all, and the acknowledgement already names exactly the table to lock. It also collapses N: once the transaction holds ACCESS EXCLUSIVE on the owning table, nothing else can hold a conflicting lock on its indexes, so the remaining acquisitions are uncontended. This makes the invariant true rather than approximately true.
  2. Keep the ordering rule but require a real margin — some multiple of the lock bound rather than one millisecond — and state in AB-1 that it is a heuristic that narrows the window rather than a proof. Cheaper, but it leaves REINDEX TABLE on a many-index table unbounded in N, so the margin is a guess.

If neither is worth doing now, the honest minimum is softening the three places that claim the window is closed, so the classification stays visibly approximate.

Minor

The new test rolls the holder back twice. require.NoError(t, holder.Rollback(t.Context())) runs inline, and the t.Cleanup registered above rolls back again. The cleanup discards its error so nothing fails, but the inline call is what the assertion depends on and the cleanup is what makes it safe on an early failure — a reader meeting both wonders which is load-bearing. A one-line comment, or dropping the inline rollback and asserting indexExists through a separate connection, removes the question.

ErrInvalidBlockingBudget now covers three distinct operator mistakes — a bound that is absent or sub-millisecond, one that is unrepresentable, and a pair in the wrong order. The first two are "you did not give me a usable number"; the third is "both numbers are fine but their relationship is wrong", and it is the only one whose remedy is to change the other flag. The message text distinguishes them well, and invalid-blocking-budget is documented as covering all three, so this is only worth noting if the exit-code table ever grows a consumer that keys on the code to suggest a remedy.

lock_timeout bounds each lock acquisition on its own; statement_timeout
counts them all together. The admitted statements acquire more than one
lock — DROP INDEX the table then the index, REINDEX INDEX the table then
the index, REINDEX TABLE the table then each index — so a statement bound
longer than the lock bound did not keep statement_timeout from ending a
lock wait: a first wait granted late followed by a second one was cut off
as 57014 and reported as budget-statement-exceeded, work that never ran.

ExecuteAcceptedBlocking now takes the schema-qualified table the caller
resolved from the catalog and requests its ACCESS EXCLUSIVE lock as a
statement of its own before the DDL. Every way that wait can end arrives
while nothing has been submitted: 55P03 and a 57014 of the request itself
are the lock-budget refusal, and the caller's context ending is
cancelled-by-caller rather than an unknown outcome. Holding the table's
ACCESS EXCLUSIVE lock means no other session holds a conflicting lock on
any of its indexes, so the DDL's own requests are granted at once and its
statement budget measures work. The request recurses to partitions, which
the DDL would lock in turn. A materialized view, which LOCK TABLE cannot
name, is the one target the statement still locks itself; migrate passes
nil for it.

The statement-longer-than-lock rule stays as a consistency check — a pair
in the other order could never let the lock budget apply — and the
invariants, passthrough design, execution-model row, and comments stop
claiming it is what keeps a lock wait from being reported as work.

Tests: a REINDEX INDEX behind a writer that releases the table after a
second and a reader that keeps the index locked, with budgets that fit
each wait but not both, is a lock refusal at the lock bound (without the
pre-lock it is a statement failure at the statement bound); a context
cancelled during the table lock wait is cancelled-by-caller with the
index standing, at both the executor and migrate layers; the
unknown-outcome test cancels a running slow REINDEX instead of a lock
wait; REINDEX TABLE on a materialized view commits without a pre-lock.
The inline holder rollbacks in the remaining tests say why the cleanup's
second rollback is harmless.
@Kiran01bm

Copy link
Copy Markdown
Collaborator Author

🤖 Adversarial review response — created by Kiran's code review agent (Amp, Claude Opus 4.6) — pull/145, follow-up commit

Round 2: 3 findings — 2 fixed, 1 no action. Round 1's three findings were answered in the earlier response (1 fixed, 2 tracked as internal follow-ups) and are repeated here for one consolidated view.

# Finding Status Explanation
C1-F1 Statement budget shorter than the lock budget turns an ungranted lock into budget-statement-exceeded fixed (in 5c871ac) BlockingBudget.validate refuses a statement bound that is not longer than the lock bound before any session is acquired; docs, help text, and AB-1 state the rule. Round 2 reframes it as a consistency check — see C3-F1 for why it was never sufficient on its own.
C1-F2 A testutil.NewPool that registers t.Cleanup(pool.Close) so the suite cannot repeat the cleanup-order mistake deferred Mechanical migration of dozens of tests across five packages; tracked as an internal follow-up, unchanged from the round-1 response.
C2-F1 Split blocking-outcome-unknown into "lost before COMMIT, rolled back" and the genuinely ambiguous commit failure deferred Changes AB-5 and trades fail-closed conservatism for a retry signal; a design decision of its own, tracked as an internal follow-up, unchanged from the round-1 response.
C3-F1 lock_timeout is per acquisition, statement_timeout is cumulative; every admitted statement takes more than one lock, so the ordering rule leaves a window where a late-granted first wait plus a second wait is cut off as 57014 and reported as work fixed Taken as proposed: ExecuteAcceptedBlocking now receives the schema-qualified table the caller resolved from the catalog and runs LOCK TABLE <table> IN ACCESS EXCLUSIVE MODE as its own statement after SET LOCAL and before the DDL. Every way that wait can end arrives with nothing submitted: 55P03, and a 57014 of the request itself (reachable only when it recurses to partitions), are the lock-budget refusal; the caller's context ending is cancelled-by-caller with a verdict that says nothing was submitted, not the unknown outcome. Holding the table's ACCESS EXCLUSIVE lock means no other session holds a conflicting lock on any of its indexes, so the DDL's own requests are granted at once and its statement budget measures work. Two decisions worth a look: the mode is ACCESS EXCLUSIVE for REINDEX too (its own table lock is SHARE, but REINDEX already blocks every new query on the table, and the stronger lock is what makes REINDEX TABLE's per-index requests free instead of one wait each inside the statement budget); and a materialized view, which LOCK TABLE cannot name, is the one target passed without a pre-lock, so the statement acquires its own locks there and the docs say so. Your scenario is the new test: REINDEX INDEX behind a writer that releases the table at 1s and a reader keeping the index locked, lock 1.5s / statement 1.6s — a 55P03 lock refusal at 1.5s with the pre-lock, a CauseStatement failure at 1.6s without it. The ordering rule stays as a consistency check (a pair in the other order could never let the lock budget apply); AB-1/AB-2, the passthrough design, the execution-model row, and the comments no longer claim it is what keeps a lock wait from being reported as work.
C3-F2 Double holder rollback in the new test (inline plus t.Cleanup) fixed The unknown-outcome test is restructured — it now cancels a running slow REINDEX INDEX, the only place that outcome can still arise, rather than a lock wait — and the two remaining inline holder.Rollback sites carry a one-line comment on why the cleanup's second rollback is harmless (pgx returns ErrTxClosed, which the cleanup ignores by design).
C3-F3 ErrInvalidBlockingBudget covers three distinct mistakes (disabled, out of range, wrong order) no action The three cases produce distinct messages under one sentinel; nothing keys on the sub-cause today — migrate maps the sentinel to one verdict code and one exit code, and an adapter has no reason to branch between them. If an orchestrator ever needs to distinguish "fix your ordering" from "fix your range", a typed cause on the error is a small additive change; splitting the sentinel now would widen the error taxonomy for no consumer.

Source: comment 5987651487 · comment 5987652135 · review 5409771888 · review 5420614804 — reviewed at head 5c871ac.

@Kiran01bm
Kiran01bm enabled auto-merge (squash) October 7, 2026 05:28
@Kiran01bm
Kiran01bm merged commit 15679ba into main Oct 8, 2026
16 checks passed
@Kiran01bm
Kiran01bm deleted the kiran01bm/accept-blocking-cancel-test branch October 8, 2026 01:16
Kiran01bm added a commit that referenced this pull request Oct 8, 2026
* origin/main:
  migrate: let only the caller's cancellation end the unknown-outcome test's lock wait (#145)
  applier: flush drained batches column-wise with the unique-move fallback (#149)

# Conflicts:
#	docs/copy-and-swap-design.md
Kiran01bm added a commit that referenced this pull request Oct 8, 2026
…bm/cs7-pgoutput

* origin/kiran01bm/cs7-slot:
  decode: prove the server, verify the publication's shape, type the refusals
  schemachange: re-derive the swapped-table proof from the catalog for the post-swap resume (#146)
  migrate: let only the caller's cancellation end the unknown-outcome test's lock wait (#145)
  applier: flush drained batches column-wise with the unique-move fallback (#149)

# Conflicts:
#	SAFETY.md
#	docs/copy-and-swap-design.md
#	docs/invariants.md
#	pkg/decode/doc.go
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants