Skip to content

Latest commit

 

History

History
242 lines (214 loc) · 21.5 KB

File metadata and controls

242 lines (214 loc) · 21.5 KB

Architecture

The map of the codebase: what the layers are, where each responsibility lives, and what exists today versus what each build phase adds. For the design rationale behind this shape start with high-level-design.md; for interfaces and lifecycle internals see low-level-design.md.

Table of contents

The three layers

pg-sprite is a decoupled planner → router → executor engine. The planner decides what changes, the router decides which strategy, interchangeable executors decide how.

The planner is itself a pipeline of five distinct stages. The two front-ends enter it at different points — an imperative migrate --alter already is DDL, so it goes straight to parse; a declarative diff --desired schema must first be compared against the live database to produce DDL — and the derived statements then re-enter the parse boundary like any hand-written statement, so both routes converge on the same parse → classify → lint tail and every operation is judged by the same rules regardless of how it arrived:

   user: migrate --alter "ALTER TABLE …"   user: diff --desired schema.sql
     (imperative: statements)               (declarative: whole schema)
                  │                                      │
        ╭─────────▼──────────────────────────────────────▼─────────╮
        │  CLI: migrate · pull · diff · fmt · lint · suggest ·     │
        │       capabilities · status                              │
        ╰─────────┬──────────────────────────────────────┬─────────╯
                  │                                      │
   ┌──────────────▼───── PLANNER (shared front-end) ─────▼──────────────┐
   │                                                                    │
   │  ╭─────────────────────╮            ╭─────────────────────────╮    │
   │  │ PARSE               │            │ INTROSPECT              │    │
   │  │ pkg/statement       │            │ pkg/schemadiff          │    │
   │  │ real PG grammar     │◀───────╮   │ read the live catalog;  │    │
   │  │ (go-pgquery) →      │        │   │ apply desired DDL to a  │    │
   │  │ typed per-operation │        │   │ scratch schema and      │    │
   │  │ descriptors         │        │   │ introspect that too     │    │
   │  ╰──────────┬──────────╯        │   ╰────────────┬────────────╯    │
   │             │           derived │                │ two schema      │
   │             │               DDL │                ▼ models          │
   │             │                   │   ╭─────────────────────────╮    │
   │             │                   │   │ DIFF                    │    │
   │             │                   │   │ pkg/schemadiff          │    │
   │             │                   ╰───┤ desired vs live →       │    │
   │             │                       │ ordered DDL statements  │    │
   │             │                       ╰─────────────────────────╯    │
   │             │ operations                                           │
   │             ╰──────────────────╮                                   │
   │                                ▼                                   │
   │                   ╭─────────────────────────╮                      │
   │                   │ CLASSIFY                │                      │
   │                   │ pkg/planner             │                      │
   │                   │ per op: native-safe ·   │                      │
   │                   │ needs-rewrite · refuse; │                      │
   │                   │ emit the safer native   │                      │
   │                   │ sequence where one      │                      │
   │                   │ exists                  │                      │
   │                   ╰────────────┬────────────╯                      │
   │                                ▼                                   │
   │                   ╭─────────────────────────╮                      │
   │                   │ LINT                    │                      │
   │                   │ pkg/lint                │                      │
   │                   │ reject unsafe or        │                      │
   │                   │ unsupported operations  │                      │
   │                   │ before any write        │                      │
   │                   ╰────────────┬────────────╯                      │
   │                                │ Plan (ordered, classified steps)  │
   └────────────────────────────────┼───────────────────────────────────┘
                                    ▼
                            ╭───────────────╮
                            │    ROUTER     │      policy + cluster facts
                            ╰───────┬───────╯
           ╭────────────────────────┼───────────────┬────────────────╮
           ▼                        ▼               ▼                ▼
      native DDL             copy-and-swap    expand/contract    refuse with
      CONCURRENTLY           (later phase)    (reversible,       not-native-safe
      NOT VALID …                              later)            verdict
           ╰────────────────────────┴───────────────╯
                                    │ cross-cutting: connection mgmt,
                                    │ lock bounding, lag-aware throttling
                                    ▼
                         ╭─────────────────────╮
                         │     PostgreSQL      │
                         ╰─────────────────────╯

The five front-end stages

Stage Package Live DB? Input → output Why it is a separate stage
Parse pkg/statement No — the grammar is embedded (Wasm libpg_query) SQL text → typed per-operation descriptors One parse boundary using the real PostgreSQL grammar — no hand-parsing anywhere else; a parse failure is an error surfaced to the caller, never a guess
Introspect pkg/schemadiff Yes — reads the target's catalog; executes desired DDL in a rolled-back scratch schema live catalog (and desired DDL applied to a scratch schema) → schema models The classifier and diff need facts, not text: column types, defaults, and constraint state come from PostgreSQL's own catalog, not a reimplementation of its semantics
Diff pkg/schemadiff No — pure model-vs-model comparison, but both its inputs come from Introspect desired model vs live model → ordered DDL statements Declarative mode is a front-end that produces statements; its output re-enters the parse boundary and flows through the same classify → lint tail as hand-written DDL, so both modes get identical safety treatment
Classify pkg/planner No — a pure function over parsed operations and facts gathered earlier each operation + introspected facts → native-safe · needs-rewrite · refuse, with the safer native sequence where one exists The safety decision lives in one pure, testable place — PostgreSQL's missing ALGORITHM=/LOCK= declaration (design-principles.md)
Lint pkg/lint No — pure over classified operations classified operations → typed findings (errors refuse, warnings advise) Policy-level rejection of unsafe or unsupported changes before any write — separate from the mechanical can-this-run-online judgment

Only Introspect touches the database. That is why lint, suggest, and fmt run fully offline against SQL text alone, while diff and dry-run require a connection for Introspect's facts — and migrate requires one for the facts and then to execute the change.

All five stages exist today (Phases 1–2.5); see the package map for per-package status.

Why introspect — ask PostgreSQL, don't reimplement it?

The engine constantly needs two facts it must not get wrong: what the live table actually looks like, and what a desired-state file actually means. There are two ways to answer the second one. The tempting way is AST transformation: parse the DDL and apply it to an in-memory schema model in Go. But that model is only correct if it reproduces PostgreSQL's own semantics exactly — type-alias resolution (varchar is character varying, int4 is integer), default-expression normalization, implicit constraint and index naming, identity and sequence wiring, and every place those rules shift across server versions 14 → 18. Each divergence is a wrong diff, and a wrong diff becomes wrong DDL against someone's production table.

pg-sprite refuses to maintain that reimplementation. Instead it executes and introspects: hand the DDL to the one authority that cannot disagree with PostgreSQL — PostgreSQL — and read back what it built. Step by step, for a diff against a desired-state file:

  1. Parse for admission. The file goes through the pkg/statement parse boundary (the real PostgreSQL grammar); anything unparseable is an error, never a guess.
  2. Create a scratch schema. Inside a transaction on the target connection, a throwaway schema is created — invisible to every other session.
  3. Execute the desired DDL there. PostgreSQL applies its own semantics: it resolves type aliases, normalizes defaults, names implicit constraints and indexes, and wires up identity sequences — on the exact server version being targeted.
  4. Introspect the result. The catalogs (pg_attribute, pg_constraint, pg_index, pg_attrdef, …) are read back into a normalized schema model.
  5. Roll back. The transaction is discarded; nothing persists and the live schema was never touched.
  6. Introspect the live table into the same model shape, canonicalize both, and diff them — the comparison is model vs model, never text vs text.
  7. Re-enter the pipeline. The derived statements go back through the parse boundary and into classify → lint, exactly as if a user had submitted them with --alter — declarative mode earns no shortcut past the safety rules.

The result is a diff computed from what PostgreSQL means, not from what a Go reimplementation thinks it means. One honest limit: an empty scratch table reveals nothing about the size or lock cost of the real one, so classification still predicts rewrite and lock behaviour — execute-and-introspect complements the classifier, never replaces it. Mechanism details and the recorded decision live in low-level-design.md.

The planner's verdicts are requests, not permissions — executors re-verify their own preconditions. Which components are safety-critical (and the stricter rules inside that boundary) is defined in ../SAFETY.md.

What may a consumer depend on?

A consumer deciding between shelling out to the CLI and importing a package is making an architectural decision; this is the permission slip. Two integration surfaces exist, at different levels of commitment:

  • The CLI and its JSON output — the intended seam for orchestrators. diff and --dry-run emit machine-readable verdicts and plans; the plan report is a single versioned contract (format_version, plan.FormatVersion; additive-only changes within a version, a consumer rejects a version it does not understand), documented in plan-report.md. The suggest report carries its own format_version on the same rule (suggest-report.md). capabilities --json is deliberately not versioned separately: its version is the binary version string, because the matrix is committed with and tested against one binary, so consumers pin the binary rather than a matrix version and ignore fields they do not recognize (capabilities-contract.md).
  • The Go packages — everything under pkg/ is importable, and the front-end seams (pkg/statement, pkg/schemadiff, pkg/planner, pkg/router, pkg/verdict) are each designed as a standalone entry point; internal/ is unimportable by construction. Before a v1 module tag the Go API carries no compatibility promise: the JSON contract is the stability boundary, the Go API follows at v1. (The SAFETY.md core/periphery split answers a different question — review bar, not import permission.)

Package map

Package Role Status
cmd/pg-sprite CLI entry point (Kong): migrate · pull · diff · fmt · lint · suggest · capabilities · status all eight exist
internal/cli Command tree and flag handling (including migrate --dry-run) all eight exist
internal/testutil Test harness: containerized PostgreSQL, throwaway schemas exists
pkg/dbconn Pool with bounded session timeouts, retries, RDS/Aurora auto-TLS (embedded CA bundle), terminate-blockers; advisory-lock mutual exclusion lands here exists
pkg/statement go-pgquery (Wasm libpg_query) parse boundary, typed per-operation descriptors, and advisory rewrites (never hand-parse SQL); shadow DDL is validated by executing the retargeted statement on the empty shadow, and fingerprints come from pkg/schemadiff's transaction-scoped scratch schema — execute-and-introspect, never AST surgery exists
pkg/preflight Precondition verification and refusals before any write: target facts + table-size guard, tiered privilege checks (a refusal carries the exact provisioning GRANT), partitioned-table support gates, the copy-and-swap shape proof and its cluster/volume environment check exists
pkg/verdict Structured outcome contract (executed / refused / failed + reason, stable executor code, and safer idiom), rendering, exit codes exists (Phase 1)
pkg/schemadiff Execute-and-introspect desired state, introspect the live catalog, and produce an ordered declarative diff exists
pkg/planner Classify typed operations and emit safer native SQL exists
pkg/lint Offline lint findings with typed codes: unsupported operations are errors; blocking idioms, rewrites, and destructive drops are warnings exists (Phase 2.5)
pkg/suggest Offline advisory surface: maps risky DDL to the safer native form the engine would run, with typed caveats and manual-path guidance; emits the versioned suggest report (suggest-report.md) exists
pkg/plan Versioned machine-readable dry-run plan report — the one JSON contract both front doors emit and an orchestrator consumes exists (Phase 2.5)
pkg/diffplan The declarative front door as a library: desired schema in, routed plan.Report out — the CLI diff and embedding orchestrators share this one pipeline exists
pkg/migrate The imperative front door as a library: one parsed statement in — gate, resolve, classify, route, execute — one verdict.Verdict out; the CLI migrate and embedding orchestrators share this one pipeline. Also the desired-state execution loop: RunDesired derives the convergence plan (diffplan.Plan), admits it as a whole (existence, destructive guard, dispositions, optional fingerprint pin), and runs each planned statement back through Run — per-statement verdicts, committed-prefix semantics exists
pkg/router Route classified statements to native / copy-and-swap / refuse dispositions; copy-and-swap reports unavailable until that backend lands exists (Phase 2.4)
pkg/executor Native backend with stable outcome codes: the bounded optimistic attempt, the concurrent index build and its invalid-index recovery, the autocommit safer-sequence runner, the greenfield CREATE TABLE path, and the accepted-blocking passthrough primitive; the full Executor contract (Plan/Execute/Status/Abort) arrives with the copy-and-swap backend native execution exists
pkg/progress Strategy-wide, pollable progress snapshots: native phase/elapsed time, sequence position, retry attempt, and server-reported concurrent-index work, and engine-measured work polled from the copier and the checksum verifier (WorkSource) native progress and the copy and checksum work sources exist
pkg/copier PK-range chunker over one integer-family primary key with dynamic time-based sizing (produces Chunk and Watermark; composite keys refused in v1), and the parallel chunked copy into the shadow table (never overwrites; reports the cut frontier, in-flight chunks, and landed watermark for the applier, and fills the progress tracker's copy counters while it runs) — there is no separate chunker package chunker, copy loop, and progress fillers exist
pkg/checksum The mandatory correctness gate: the chunk verifier (single-snapshot per-chunk digests through the shadow's types, reporting the chunks that differ up to the landed watermark), Check under an explicit divergence policy (abort, or repair each differing chunk with the copier's statement and reread it), and the proof constructors (VerifiedShadow, CleanWatermark) minted only by a pass that found nothing and repaired nothing; the pass reports its chunks compared, rows hashed, mismatches, and repairs to the progress tracker as its WorkSource; continuous checker to follow verifier, policy, repair, proofs, and progress counters exist; checker planned
pkg/decode Logical-decoding change capture, LSN accounting, slot lifecycle Phase 6, 8
pkg/applier Change apply onto the shadow (always wins), buffer/dedup, flush scheduling Phase 6
pkg/schemachange Orchestrator: lifecycle, cutover swap + fidelity gate, checkpoint/resume Phase 7–8
pkg/checkpoint Durable resume state: one row per (schema, table) in the target database Phase 8
pkg/throttler Chunk-time targeting and the hard slot-lag ceiling; replica-lag throttling deferred Phase 8

The copy-and-swap lifecycle

In a later phase, when the in-house heavy path is available and the router picks it:

 create shadow table  ─▶  start change capture   ─▶  bulk-copy existing rows
 with the new schema      (logical decoding,         in parallel chunks
                           off the WAL)                     │
                                                            ▼
        cut over  ◀──  CHECKSUM GATE  ◀──  drain the captured-change backlog
   (brief ACCESS           (must prove          onto the shadow
    EXCLUSIVE swap,         shadow == source
    bounded + retried)      before cutover)

The interleaving rules that make the copier and applier converge — and the invariant IDs every component must uphold — are registered in invariants.md; the trust boundary and domain-type design are in tcb-model.md.

Where to read more