Skip to content

Opt-in deduplication of samples at ingestion - #46

Merged
fungiboletus merged 12 commits into
mainfrom
dedup-at-ingestion
Oct 3, 2026
Merged

fungiboletus merged 12 commits into
mainfrom
dedup-at-ingestion

Conversation

@fungiboletus

Copy link
Copy Markdown
Member

What

SENSAPP_DEDUPLICATE_ON_INGEST=true (or deduplicate_on_ingest = true), off by default, makes the write itself leave out the samples that are stored already, and write the repeated samples of one request once. The rule is the vacuum's: same series, timestamp and value (coordinates for a location). Two different values at one timestamp are both kept, nothing is rejected, the request succeeds. It covers independent, repeated and overlapping requests, not only retries. The vacuum stays for duplicates written before it is turned on, and for ClickHouse.

Backend Support Several writers / instances
PostgreSQL, TimescaleDB yes exact: advisory lock per series, hashed into 1 024 buckets, held until the transaction ends
SQLite, DuckDB yes exact: single writer / single process
ClickHouse, BigQuery, RRDCached no the server refuses to start with the flag set

How

  • PostgreSQL / TimescaleDB (src/storage/pg_samples.rs, shared): the unnest insert becomes INSERT .. SELECT DISTINCT .. WHERE NOT EXISTS against the stored rows inside the time window of the batch (a per-row probe of the BRIN index cost 0.36 ms a row), with a transaction-local planner guard (no nested loop / seq scan: new samples are newer than the statistics, which gave 3 to 5 s plans for 20 000 samples).
  • SQLite: WITH u AS (VALUES ..) INSERT .. WHERE NOT EXISTS. DuckDB: the appenders write a temp table, one statement copies the new rows.
  • PostgreSQL migration: autosummarize on the BRIN indexes.
  • Config entry, settings.toml, Helm value, docs/CONFIGURATION.md, docs/DATA_LIFECYCLE.md.

For reviewers

  • A read before the insert is not enough with several instances. A concurrent-writers test first stored 401 rows where 51 were expected; the advisory locks fix it. A stress test (10 concurrent requests of 8 000 series) also showed one lock per series exhausting PostgreSQL's shared lock table (out of shared memory, HTTP 500); hence the fixed 1 024 buckets. Consequence: concurrent large batches take turns.
  • PostgreSQL condition. The probe goes through BRIN, which returns every range not summarized yet in full. Right after a bulk load a small write took 64 ms (3.5 to 4 ms once autovacuum caught up; it did by itself in 90 s with default settings). A large import is slower with the flag on (10.5 s vs 6 s for 1 M samples). Documented, with the advice to turn it off for a clean one-time import.
  • Cost (release build, 1 M rows, tests/perf/dedup.sh, one run, indicative): no visible difference for a 20 000 sample request on SQLite, DuckDB, TimescaleDB; PostgreSQL 0.14 s vs 0.19 s once summarized. Small writes: SQLite +0, DuckDB +1.5 ms, TimescaleDB +0.5 to 1.5 ms, PostgreSQL +1 to 3 ms. Re-sending known samples is as fast or faster. All numbers and the baseline are in done/ingestion-deduplication.md.
  • Tests (backend-generic, tests/integration/deduplication.rs): every value type, repeats inside a request, exact duplicates only, partial overlap, switching off, concurrent writers, repeated/overlapping requests through the HTTP API, compressed TimescaleDB chunks, which backends accept the switch, the migration. Full suites on the five backends and cargo clippy --all-targets --features all-storage -- -D warnings pass (the DuckDB doctest only fails locally for the known libduckdb.dylib reason).
  • Found on the way, not caused by this PR: on TimescaleDB, concurrent first writes of the same new series deadlock (8 of 8 runs with the flag off). The concurrent test skips that half of its rounds on TimescaleDB; written up in ideas/timescaledb-concurrent-first-write-deadlock.md, to fix separately.
  • Not done: ClickHouse (no exact mechanism at insert time), throughput of many concurrent writers of overlapping series, DuckDB staging refinements. See the end of done/ingestion-deduplication.md.

🤖 Generated with Claude Code

fungiboletus and others added 12 commits October 3, 2026 12:25
…e numbers

A benchmark of the write patterns that deduplication changes (append, retry,
replay of old samples, small requests) and its numbers on the five backends
without deduplication.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…B (experiment)

Behind SENSAPP_DEDUPLICATE_ON_INGEST (off by default): the INSERT leaves out
the rows that are stored already, bounded to the time window of the batch so
that the planner reads the window once instead of probing the BRIN index for
every row. A transaction-level advisory lock per series, taken in order, makes
it hold with several writers or instances: a read before the insert cannot see
the rows another transaction has not committed (the concurrent test stored 401
rows where 51 were expected before the lock).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ndow

New samples are newer than the statistics, so the window was estimated at one
row and the planner chose a nested loop (seconds for 20 000 samples), and a
table that was never analyzed got a sequential scan. The deduplicating
transaction now turns both off for its inserts.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ment)

SQLite: the multi-row INSERT becomes WITH u AS (VALUES ..) INSERT .. SELECT
DISTINCT .. WHERE NOT EXISTS, answered by the (sensor_id, timestamp_us) index.
DuckDB: the appenders write to a temporary table and one statement copies the
new rows, looking only at the time window of the batch. Both have a single
writer, so nothing can slip in between the check and the insert.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
One advisory lock per series made 10 concurrent requests of 8000 series each
fail with out of shared memory (the lock table is shared by all transactions).
A fixed number of buckets bounds the entries whatever the batch.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…escaleDB deadlock

Eight writers that register the same new series at once deadlock on TimescaleDB
with or without the deduplication (8 runs out of 8 with it off). The test
registers the series first on that backend; the bug is written down in ideas/.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…re it is unsupported

SENSAPP_DEDUPLICATE_ON_INGEST / deduplicate_on_ingest (off by default) replaces
the environment read of the factory; the server stops with a clear message on a
backend that cannot do it. A PostgreSQL migration sets autosummarize on the BRIN
indexes so the probe does not read the ranges written since the last vacuum.
Documented in CONFIGURATION.md and DATA_LIFECYCLE.md, chart value added.
Tests: repeated and overlapping requests through the HTTP API, compressed
TimescaleDB chunks, which backends accept the switch, the migration.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@fungiboletus
fungiboletus merged commit 539c84b into main Oct 3, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant