Repository navigation
BigQuery backend: back in sync with the storage interface, tested on a real dataset - #49
Merged
Merged
Conversation
…orage::common BigQuery will use the same ids. Open the task on bringing BigQuery back in sync, with the audit of the old module. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Nothing ran against Google Cloud yet (unit tests only). - Schema: ids derived from the UUID and the unit name (shared with ClickHouse), no dictionaries, timestamps as microseconds, FLOAT64 written as doubles (were f32), NUMERIC as a decimal string, monthly partitions. Registration is idempotent: two writers insert identical rows, reads collapse them. - Writes: one append per table through the default stream, all the answers checked (the `error` and `row_errors`, not only the status), no process-wide write lock, JSON values stored as JSON (were ""). - Reads: every statement waits for its job and reads all the pages (a job over 10 s was an empty result), parameters instead of formatted SQL, samples of the 8 types, paginated listing, label matchers, bulk selectors, latest sample, deletes, aggregation on the raw window. - Connection string: optional key file (else Application Default Credentials), location, max_bytes_billed, identifiers validated. - Remove the dependencies the old code needed (bigdecimal, its encoder, clru, tonic) and sinteflake with SENSAPP_INSTANCE_ID, which only BigQuery used. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Integration tests for what is specific to BigQuery, BigQuery added to the backend lists of the generic suites, docs/BIGQUERY.md with the Google Cloud setup and the cost safeguards, a CI job that compiles and unit tests the backend, and a cargo-make task for the live run. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
A GROUP BY of buckets for one series and for the selectors of remote read, same buckets as the other backends. The integration tests compare every numeric type and aggregation with the local reference. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…loud setup commands Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Found by running the integration tests on Google Cloud: - install the rustls crypto provider in BigQueryStorage::connect: the gRPC client of the Storage Write API panicked when both ring and aws-lc-rs are compiled in and the caller did not install one - averages are the sum over the count: BigQuery's AVG gives -5.5e-17 for integers that average to 0 - the cost cap test uses the INFORMATION_SCHEMA statement: rows still in the streaming buffer are billed 0 bytes Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Faster than a sum over a count, and the last bits of a floating point average do not matter in a data warehouse (-5.5e-17 for integers that average to 0). The integration test compares averages with a tolerance. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…nual Docker variant One clippy step on the backends of the Docker image plus BigQuery replaces the 45 minute image build that only ran on a manual dispatch and showed up as skipped on every pull request. The slim builder image of the Dockerfile builds BigQuery without protoc (checked in a container). Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Brings the BigQuery backend back in line with the current storage interface, and runs it for real. BigQuery stays an optional R&D backend (like RRDCached): simple, documented limits, no deduplication (like ClickHouse).
The old module only compiled because the storage trait has default methods.
query_sensor_datareturned no samples, label queries were abail!, floats were stored asf32, JSON values as"", a query over 10 s returned zero rows silently, a rejected write was reported as a success, and every append took the process-wide write lock. Full audit indone/bigquery-back-in-sync.md.Changes
src/storage/bigquery/): ids derived from the UUID and unit name (shared with ClickHouse, moved tostorage::common), so two instances registering the same series insert identical rows instead of two ids. No dictionary tables, microsecond timestamps, monthly partitions,FLOAT64as doubles,NUMERICas text. Statements wait for their job and read every page, use parameters, and can carrymax_bytes_billed. Writes check theerrorandrow_errorsof every answer.GROUP BYof buckets, single series and Prometheus selectors) with BigQuery's ownAVG.location,max_bytes_billed, validated identifiers.bigdecimal, its encoder,clru,tonic, andsinteflakewithSENSAPP_INSTANCE_ID(only BigQuery used them). Breaking: that environment variable is gone.tests/integration/bigquery_integration.rs, BigQuery added to the backend lists of the generic suites. CI gets abigquery-checksjob (compile, clippy, unit tests, no cloud project needed).docs/BIGQUERY.md(setup, no-key setup for organizations that forbid keys, cost safeguards, limits),docs/BACKENDS.md,docs/CONFIGURATION.md.Tested
On a real BigQuery dataset (europe-north1): all the applicable integration tests pass, 293 in total (9 BigQuery-specific, 42 + 242 generic). 16.9k statements, 25 GiB billed, about $0.15.
Regression runs on the shared changes (
storage::common,TypedSamples::truncate, config): ClickHouse 300 integration + 254 lib, TimescaleDB 292, SQLite 287, lib tests with every backend, clippy clean with--all-targetson default,bigqueryandall-storagefeatures.The live suite needs a Google Cloud project and is run by hand (
cargo make test-bigquery-live, wipes its dataset): CI only compiles and unit-tests BigQuery.Things found by the live run (fixed)
ringandaws-lc-rsboth compiled in).max_bytes_billedcannot be tested on them.f32because of agcp-bigquery-client0.22 bug (Float64declared as a protobuffloat, Unable to write double/f64 on float columns using the Storage Write API lquerel/gcp-bigquery-client#106), fixed since 0.26. Doubles are exact now, tested.🤖 Generated with Claude Code