Skip to content

BigQuery backend: back in sync with the storage interface, tested on a real dataset - #49

Merged
fungiboletus merged 10 commits into
mainfrom
bigquery-back-in-sync
Oct 4, 2026
Merged

fungiboletus merged 10 commits into
mainfrom
bigquery-back-in-sync

Conversation

@fungiboletus

Copy link
Copy Markdown
Member

What

Brings the BigQuery backend back in line with the current storage interface, and runs it for real. BigQuery stays an optional R&D backend (like RRDCached): simple, documented limits, no deduplication (like ClickHouse).

The old module only compiled because the storage trait has default methods. query_sensor_data returned no samples, label queries were a bail!, floats were stored as f32, JSON values as "", a query over 10 s returned zero rows silently, a rejected write was reported as a success, and every append took the process-wide write lock. Full audit in done/bigquery-back-in-sync.md.

Changes

  • Rewritten backend (src/storage/bigquery/): ids derived from the UUID and unit name (shared with ClickHouse, moved to storage::common), so two instances registering the same series insert identical rows instead of two ids. No dictionary tables, microsecond timestamps, monthly partitions, FLOAT64 as doubles, NUMERIC as text. Statements wait for their job and read every page, use parameters, and can carry max_bytes_billed. Writes check the error and row_errors of every answer.
  • Reads: paginated listing, label matchers (RE2), bulk selectors, latest sample, deletes, and aggregation in BigQuery (GROUP BY of buckets, single series and Prometheus selectors) with BigQuery's own AVG.
  • Connection string: optional key file (otherwise Application Default Credentials), location, max_bytes_billed, validated identifiers.
  • Removed: bigdecimal, its encoder, clru, tonic, and sinteflake with SENSAPP_INSTANCE_ID (only BigQuery used them). Breaking: that environment variable is gone.
  • Tests: 32 unit tests (SQL builders, rows checked against the migration, errors), integration tests in tests/integration/bigquery_integration.rs, BigQuery added to the backend lists of the generic suites. CI gets a bigquery-checks job (compile, clippy, unit tests, no cloud project needed).
  • Docs: docs/BIGQUERY.md (setup, no-key setup for organizations that forbid keys, cost safeguards, limits), docs/BACKENDS.md, docs/CONFIGURATION.md.

Tested

On a real BigQuery dataset (europe-north1): all the applicable integration tests pass, 293 in total (9 BigQuery-specific, 42 + 242 generic). 16.9k statements, 25 GiB billed, about $0.15.

Regression runs on the shared changes (storage::common, TypedSamples::truncate, config): ClickHouse 300 integration + 254 lib, TimescaleDB 292, SQLite 287, lib tests with every backend, clippy clean with --all-targets on default, bigquery and all-storage features.

The live suite needs a Google Cloud project and is run by hand (cargo make test-bigquery-live, wipes its dataset): CI only compiles and unit-tests BigQuery.

Things found by the live run (fixed)

🤖 Generated with Claude Code

fungiboletus and others added 10 commits October 3, 2026 22:13
…orage::common

BigQuery will use the same ids. Open the task on bringing BigQuery back in sync, with the
audit of the old module.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Nothing ran against Google Cloud yet (unit tests only).

- Schema: ids derived from the UUID and the unit name (shared with ClickHouse), no dictionaries,
  timestamps as microseconds, FLOAT64 written as doubles (were f32), NUMERIC as a decimal string,
  monthly partitions. Registration is idempotent: two writers insert identical rows, reads collapse them.
- Writes: one append per table through the default stream, all the answers checked (the `error` and
  `row_errors`, not only the status), no process-wide write lock, JSON values stored as JSON (were "").
- Reads: every statement waits for its job and reads all the pages (a job over 10 s was an empty
  result), parameters instead of formatted SQL, samples of the 8 types, paginated listing, label
  matchers, bulk selectors, latest sample, deletes, aggregation on the raw window.
- Connection string: optional key file (else Application Default Credentials), location,
  max_bytes_billed, identifiers validated.
- Remove the dependencies the old code needed (bigdecimal, its encoder, clru, tonic) and sinteflake
  with SENSAPP_INSTANCE_ID, which only BigQuery used.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Integration tests for what is specific to BigQuery, BigQuery added to the backend lists of the generic
suites, docs/BIGQUERY.md with the Google Cloud setup and the cost safeguards, a CI job that compiles and
unit tests the backend, and a cargo-make task for the live run.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A GROUP BY of buckets for one series and for the selectors of remote read, same buckets as the other
backends. The integration tests compare every numeric type and aggregation with the local reference.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…loud setup commands

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Found by running the integration tests on Google Cloud:
- install the rustls crypto provider in BigQueryStorage::connect: the gRPC client of the Storage Write API
  panicked when both ring and aws-lc-rs are compiled in and the caller did not install one
- averages are the sum over the count: BigQuery's AVG gives -5.5e-17 for integers that average to 0
- the cost cap test uses the INFORMATION_SCHEMA statement: rows still in the streaming buffer are billed 0 bytes

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Faster than a sum over a count, and the last bits of a floating point average do not matter in a data
warehouse (-5.5e-17 for integers that average to 0). The integration test compares averages with a tolerance.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nual Docker variant

One clippy step on the backends of the Docker image plus BigQuery replaces the 45 minute image build that
only ran on a manual dispatch and showed up as skipped on every pull request. The slim builder image of the
Dockerfile builds BigQuery without protoc (checked in a container).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@fungiboletus
fungiboletus merged commit 756000e into main Oct 4, 2026
17 checks passed
@fungiboletus
fungiboletus deleted the bigquery-back-in-sync branch October 4, 2026 10:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant