Repository navigation
Fix the TimescaleDB deadlock between first writes of new series and chunk creation - #47
Merged
Merged
Conversation
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
… that it is not limited to one series Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…n TimescaleDB Creating a chunk (a new week, or the second hash partition of a week) takes a ShareUpdateExclusiveLock on the hypertable and then, to copy the foreign key onto the chunk, a ShareRowExclusiveLock on `sensors`, which waits for every transaction that inserted a sensor and has not committed. A batch that registered a new sensor and then needs a chunk takes the two in the other order: PostgreSQL reports a deadlock and a request fails with a 500. It is not limited to one series: two writers of different series meet too, for example on a new week. A batch with sensors to register now locks the hypertables it writes to (ONLY, in alphabetical order, one statement) before the first insert into `sensors`, so the order is the same for everybody. One transaction as before: a failed batch leaves no sensor behind. Batches of known series take no lock. Tests: eight writers of different new series, mixing four types, on a time range without storage (fails 3 times out of 3 without the lock, 0 deadlocks in 30 runs with it), and the same with two writers meeting two types. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…a deadlock The hypertables are locked in alphabetical order before the sensors of a batch are registered, but the inserts meet them in the order of the code (integers, numerics, floats, strings, then the other types), so a writer that registers boolean and string series and a writer that creates chunks of the two can still wait for each other. PostgreSQL aborts one of them with SQLSTATE 40P01 and documents that the application tries again. The transaction was rolled back whole, so a failed attempt leaves nothing behind; the batch is tried up to 5 times, after a random pause of 10 to 60 ms for a deadlock. The foreign-key retry (a sensor deleted by another instance) shares the loop. Test: a connection plays the chunk creator, deterministically. It fails with one attempt and passes with five. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…caleDB The rounds where eight writers register the same new series at the same moment were skipped on TimescaleDB because of the first-write deadlock. Alone, the test failed 4 times out of 4 before the lock order fix, and passes 10 times out of 10 now, with no deadlock in the server log. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…sured cost The lock is `sensors`, taken by the chunk creation, and the cycle is a lock-order inversion that does not need the same series. The fix and its tests are described in done/, together with the numbers: no change on a normal write, 2.4 times slower for parallel bulk loads of new series (follow-up in ideas/). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A deferred transaction that read before it writes fails at once with "database is locked" in WAL mode when another writer committed in between, whatever the busy timeout, so concurrent first writes of different series returned a 500 on SQLite (found by the new concurrent first-write tests on CI). IMMEDIATE takes the write lock at the start: the writers wait for each other. Both tests and the concurrent deduplication test pass 5 times out of 5, the SQLite suite passes, and a write of 3000 series is unchanged (0.39 s new, 0.14 s again). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
On TimescaleDB, concurrent writers could deadlock (
deadlock detected, HTTP 500) when one registered a new series and another created a chunk. Reproduced with eight writers of the same new series, and with writers of different series.Stacked on #46 (
dedup-at-ingestion, not merged yet): the reproducing test lives there, so the diff againstmainincludes its commits until it merges. The commits of this PR are the last five.Cause
The value hypertables are partitioned by 7-day range and by
by_hash('sensor_id', 2). The first sample in a chunk that does not exist yet creates it:ShareUpdateExclusiveLockon the hypertable, thenShareRowExclusiveLockonsensors(the foreign key is copied onto the chunk). That waits for every transaction holding an uncommitted sensor insert. A batch that registered a new sensor and then needs a chunk takes the two locks in the opposite order. Not limited to the same series; confirmed withpg_locks, the server log and two hand-driven sessions.Fix
LOCK TABLE ONLY <its hypertables> IN SHARE UPDATE EXCLUSIVE MODE(alphabetical), then inserts the sensors. Still one transaction, so a failed batch leaves no sensor behind. Batches of known series take no lock. The PostgreSQL backend passes an empty list.publishretries up to 5 times ondeadlock_detected(SQLSTATE 40P01) with a 10-60 ms random pause. The foreign-key retry shares the loop.Rejected: advisory lock per uuid (does not cover different series), a first committed transaction (breaks the no-leftover-sensor guarantee), retry alone (54 deadlocks and 26 s in the extreme case, still failing).
Tests
cargo test --no-default-features --features timescaledb(243 unit, 292 integration) passes;cargo clippy --all-targets --features all-storage -- -D warningsandcargo fmt --checkare clean. The PostgreSQL suite was not run (no container).Cost (release binaries before/after,
tests/perf/scale.sh 3000, 3 runs each)ideas/timescaledb-parallel-first-writes.md.Not covered
Deduplication on plus concurrent registration under load (a cycle through the series advisory lock looks possible by reasoning; the retry should recover it), and large batches holding the lock while others wait.
Details:
done/timescaledb-first-write-deadlock.md.🤖 Generated with Claude Code