Opt-in deduplication of samples at ingestion - #46
Merged
Merged
Conversation
…e numbers A benchmark of the write patterns that deduplication changes (append, retry, replay of old samples, small requests) and its numbers on the five backends without deduplication. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…B (experiment) Behind SENSAPP_DEDUPLICATE_ON_INGEST (off by default): the INSERT leaves out the rows that are stored already, bounded to the time window of the batch so that the planner reads the window once instead of probing the BRIN index for every row. A transaction-level advisory lock per series, taken in order, makes it hold with several writers or instances: a read before the insert cannot see the rows another transaction has not committed (the concurrent test stored 401 rows where 51 were expected before the lock). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ndow New samples are newer than the statistics, so the window was estimated at one row and the planner chose a nested loop (seconds for 20 000 samples), and a table that was never analyzed got a sequential scan. The deduplicating transaction now turns both off for its inserts. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ment) SQLite: the multi-row INSERT becomes WITH u AS (VALUES ..) INSERT .. SELECT DISTINCT .. WHERE NOT EXISTS, answered by the (sensor_id, timestamp_us) index. DuckDB: the appenders write to a temporary table and one statement copies the new rows, looking only at the time window of the batch. Both have a single writer, so nothing can slip in between the check and the insert. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
One advisory lock per series made 10 concurrent requests of 8000 series each fail with out of shared memory (the lock table is shared by all transactions). A fixed number of buckets bounds the entries whatever the batch. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…escaleDB deadlock Eight writers that register the same new series at once deadlock on TimescaleDB with or without the deduplication (8 runs out of 8 with it off). The test registers the series first on that backend; the bug is written down in ideas/. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…re it is unsupported SENSAPP_DEDUPLICATE_ON_INGEST / deduplicate_on_ingest (off by default) replaces the environment read of the factory; the server stops with a clear message on a backend that cannot do it. A PostgreSQL migration sets autosummarize on the BRIN indexes so the probe does not read the ranges written since the last vacuum. Documented in CONFIGURATION.md and DATA_LIFECYCLE.md, chart value added. Tests: repeated and overlapping requests through the HTTP API, compressed TimescaleDB chunks, which backends accept the switch, the migration. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
SENSAPP_DEDUPLICATE_ON_INGEST=true(ordeduplicate_on_ingest = true), off by default, makes the write itself leave out the samples that are stored already, and write the repeated samples of one request once. The rule is the vacuum's: same series, timestamp and value (coordinates for a location). Two different values at one timestamp are both kept, nothing is rejected, the request succeeds. It covers independent, repeated and overlapping requests, not only retries. The vacuum stays for duplicates written before it is turned on, and for ClickHouse.How
src/storage/pg_samples.rs, shared): theunnestinsert becomesINSERT .. SELECT DISTINCT .. WHERE NOT EXISTSagainst the stored rows inside the time window of the batch (a per-row probe of the BRIN index cost 0.36 ms a row), with a transaction-local planner guard (no nested loop / seq scan: new samples are newer than the statistics, which gave 3 to 5 s plans for 20 000 samples).WITH u AS (VALUES ..) INSERT .. WHERE NOT EXISTS. DuckDB: the appenders write a temp table, one statement copies the new rows.autosummarizeon the BRIN indexes.settings.toml, Helm value,docs/CONFIGURATION.md,docs/DATA_LIFECYCLE.md.For reviewers
out of shared memory, HTTP 500); hence the fixed 1 024 buckets. Consequence: concurrent large batches take turns.tests/perf/dedup.sh, one run, indicative): no visible difference for a 20 000 sample request on SQLite, DuckDB, TimescaleDB; PostgreSQL 0.14 s vs 0.19 s once summarized. Small writes: SQLite +0, DuckDB +1.5 ms, TimescaleDB +0.5 to 1.5 ms, PostgreSQL +1 to 3 ms. Re-sending known samples is as fast or faster. All numbers and the baseline are indone/ingestion-deduplication.md.tests/integration/deduplication.rs): every value type, repeats inside a request, exact duplicates only, partial overlap, switching off, concurrent writers, repeated/overlapping requests through the HTTP API, compressed TimescaleDB chunks, which backends accept the switch, the migration. Full suites on the five backends andcargo clippy --all-targets --features all-storage -- -D warningspass (the DuckDB doctest only fails locally for the knownlibduckdb.dylibreason).ideas/timescaledb-concurrent-first-write-deadlock.md, to fix separately.done/ingestion-deduplication.md.🤖 Generated with Claude Code