Skip to content
Merged
Show file tree
Hide file tree
Changes from 13 commits
Commits
Show all changes
14 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,8 @@ Index of the design notes for the harper core: one line per note, grouped by the
- [`Table.ts` — section map](resources/DESIGN.md#tablets--section-map) — Section markers for the 4.7K-line `makeTable()` factory.
- ["Where is X" cheat sheet](resources/DESIGN.md#where-is-x-cheat-sheet) — Symbol lookup for the read/write path, audit, subscriptions and schema.
- [Full-text declarations and reader snapshots](resources/DESIGN.md#full-text-declarations-and-reader-snapshots) — Declaration names stay separate from stored attributes; a query retains one native reader across every page.
- [Audit retention floor](resources/DESIGN.md#audit-retention-floor) — A saved audit cursor below the floor must resync; the floor is internal, and `Table.commit` skips the out-of-order walk below it.
- [Audit retention floor](resources/DESIGN.md#audit-retention-floor) — The floor records what pruning removed; it is internal, and `Table.commit` skips the out-of-order walk below it.
- [Database generation and resumable positions](resources/DESIGN.md#database-generation-and-resumable-positions) — Every copy path stamps a new generation before the copy is readable; a position resumes only if it names it and sits at or above its resume floor.
- [Path routing & parameterised routes](resources/DESIGN.md#path-routing--parameterised-routes) — How resource paths and route parameters resolve.
- [Persisted relationship catalog](resources/DESIGN.md#persisted-relationship-catalog) — Where relationship definitions are stored and rebuilt.
- [Typed, discoverable resources (code-first schema + request contract)](resources/DESIGN.md#typed-discoverable-resources-code-first-schema--request-contract) — Declaring schema and request contracts from code.
Expand Down
43 changes: 38 additions & 5 deletions bin/copyDb.ts
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,12 @@ import OpenEnvironmentObject from '../utility/lmdb/OpenEnvironmentObject.ts';
import { OpenDBIObject } from '../utility/lmdb/OpenDBIObject.ts';
import { INTERNAL_DBIS_NAME, AUDIT_STORE_NAME } from '../utility/lmdb/terms.ts';
import { CONFIG_PARAMS, DATABASES_DIR_NAME, MIGRATING_DIR_SUFFIX } from '../utility/hdbTerms.ts';
import { AUDIT_STORE_OPTIONS, auditRetention } from '../resources/auditStore.ts';
import {
AUDIT_STORE_OPTIONS,
DATABASE_GENERATION_KEYS,
auditRetention,
stampDatabaseGeneration,
} from '../resources/auditStore.ts';
import { blobsReadmeContent, copyBlobRootsByIndex } from '../dataLayer/blobBackup.ts';
import { describeSchema } from '../dataLayer/schemaDescribe.ts';
import { updateConfigValue } from '../config/configUtils.ts';
Expand Down Expand Up @@ -323,7 +328,13 @@ export async function copyDb(
primaryStoresByDbi.set(table.primaryStore.name, table.primaryStore);
}
try {
await copyDbEnvironment(sourceDatabase, targetDatabasePath, rootStore, primaryStoresByDbi);
await copyDbEnvironment(
sourceDatabase,
targetDatabasePath,
rootStore,
primaryStoresByDbi,
blobDisposition === 'copy'
);
if (blobDisposition === 'copy') await copyDatabaseBlobs(sourceDatabase, targetDatabasePath, blobRoots);
} catch (error) {
// Every path removed here was created by this call — both targets are rejected above if they
Expand All @@ -339,7 +350,8 @@ async function copyDbEnvironment(
sourceDatabase: string,
targetDatabasePath: string,
rootStore,
primaryStoresByDbi: Map<string, any>
primaryStoresByDbi: Map<string, any>,
newHistory: boolean
) {
// this contains the list of all the dbis
const sourceDbisDb = rootStore.dbisDb;
Expand Down Expand Up @@ -389,7 +401,16 @@ async function copyDbEnvironment(
if (!sourceAuditDbi) throw new Error(`Could not open the audit store of ${sourceDatabase} to copy it`);
const targetAuditStore = (targetEnv as any).openDB(AUDIT_STORE_NAME, AUDIT_STORE_OPTIONS);
console.log('copying audit log for', sourceDatabase, 'to', targetDatabasePath);
await copyDbi(useRawBytes(sourceAuditDbi), useRawBytes(targetAuditStore), false, transaction);
// a copy beside its source is a separate history, and one cut short must not hold the source's
await copyDbi(
useRawBytes(sourceAuditDbi),
useRawBytes(targetAuditStore),
false,
transaction,
undefined,
newHistory ? DATABASE_GENERATION_KEYS : undefined
);
if (newHistory) stampDatabaseGeneration(targetAuditStore, { carriesLog: true });
}

/**
Expand Down Expand Up @@ -426,7 +447,14 @@ async function copyDbEnvironment(
}
}

async function copyDbi(sourceDbi, targetDbi, isPrimary, transaction, primaryStore?) {
async function copyDbi(
sourceDbi,
targetDbi,
isPrimary,
transaction,
primaryStore?,
skipKeys?: ReadonlySet<symbol>
) {
let recordsCopied = 0;
let bytesCopied = 0;
let skippedRecord = 0;
Expand All @@ -443,6 +471,7 @@ async function copyDbEnvironment(
)) {
try {
start = key;
if (skipKeys?.has(key)) continue;
// Drop a tombstone only once it is past audit retention, the point the runtime
// removes it too: dropping a live one loses the delete, letting a peer that
// missed it resurrect the record. A tombstone's body is a lone msgpack nil, so
Expand Down Expand Up @@ -983,6 +1012,10 @@ export async function copyDbToRocks(sourceRootStore, sourceDatabase: string, tar
targetRootStore.putSync(REMOTE_NODE_IDS_KEY, asBinary(idMappingBytes));
}

// flushed so the stamp is durable before the caller publishes the staging directory
stampDatabaseGeneration(targetRootStore, { carriesLog: false });
await targetRootStore.flush({ allowWriteStall: true });

console.log('migrated database ' + sourceDatabase + ' to RocksDB');
} finally {
endPendingMigrationBlobSaves();
Expand Down
5 changes: 5 additions & 0 deletions dataLayer/DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -149,6 +149,11 @@ Three non-obvious mechanics keep that safe:
calls `closeLoadedDatabases()` (`resources/databases.ts`) in its `finally`, closing every loaded
user database on that thread (the non-enumerable `system` DB is intentionally skipped), so an
exited job worker leaves no residual handle to be mistaken for a live holder.
- **A restore stamps a new database generation before `completeRestore`.** The restored files carry
the backup's generation, so both paths open the restored directory privately, stamp it and flush
(`stampDatabaseDirectory`, [database generation](../resources/DESIGN.md#database-generation-and-resumable-positions))
inside the destructive section: a failed stamp leaves the marker, and the rerun re-purges and
re-stamps. The stamp is as durable as the marker protocol it runs inside.
- **`dropDatabase` and `restore_backup` serialize on the same lock, not a check-then-act probe.**
A drop's `destroy()` interleaving with a restore's purge-and-copy on the same directory would gut
a "successful" restore (or vice versa). `dropDatabase` takes the restore lock for every RocksDB or
Expand Down
3 changes: 3 additions & 0 deletions dataLayer/rocksdbBackup.ts
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ import { setTimeout as delay } from 'node:timers/promises';
import { pack as tarPack, type Pack } from 'tar-stream';
import { RocksDatabase, backups, registryStatus, type BackupInfo } from '@harperfast/rocksdb-js';
import { getDatabases, resolveDatabasePath } from '../resources/databases.ts';
import { stampDatabaseDirectory } from '../resources/auditStore.ts';
import {
type BlobCaptureDisposition,
classifyBlobFileForCapture,
Expand Down Expand Up @@ -601,6 +602,7 @@ export async function restoreBackup(request: any) {
if (manifest.blobs) {
await restoreBlobSnapshot(backupDir, backupId, databaseName, getBlobPathsForDatabaseName(databaseName));
}
await stampDatabaseDirectory(databaseDir, { carriesLog: true });
} catch (error: any) {
// Leave the marker (so startup/rescan detection reports an incomplete restore until a rerun
// succeeds) when either the destructive purge has begun, OR this attempt was itself a recovery
Expand Down Expand Up @@ -1101,6 +1103,7 @@ export async function restoreBackupOffline(
getBlobPathsForDatabaseName(targetDatabase ?? databaseName)
);
}
await stampDatabaseDirectory(databaseDir, { carriesLog: true });
} catch (error: any) {
// Preserve the marker on a destructive failure or a recovery over a pre-existing marker (see
// the online restoreBackup for the rationale); otherwise clear the fresh marker so an intact,
Expand Down
99 changes: 45 additions & 54 deletions resources/DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -188,11 +188,10 @@ Index waits use the ordinary transaction timeout; they do not renew it. The adap
## Audit retention floor

`Table.subscribe`'s `startTime` replay just begins wherever the audit log now begins, so a consumer
resuming below the retention horizon is silently handed a short replay. The floor is the primitive
that makes that detectable (harper#2447). It is internal, with deliberately no public accessor, and **no
resume path consumes it yet**: harper#2448 is to put the check inside `Table.subscribe` itself — the same shape as
replication's `shouldForceBaseCopyForRetention`, and the only one where the floor cannot move between
being read and being acted on. Until then the short replay above is unchanged.
resuming below the retention horizon is silently handed a short replay. The floor records what
pruning removed (harper#2447); it is internal, with deliberately no public accessor. **A resume is
not checked against it** but against the database generation (next section): this floor also steers
reconciliation, survives a restore, and absorbs at `Infinity`.

**The one consumer today is not a resume**: `Table.commit`'s out-of-order reconciliation reads the
floor before entering the audit walk (harper#2642). The walk terminates at the incoming write only by
Expand Down Expand Up @@ -226,50 +225,31 @@ Things that are easy to get wrong here:
surviving entry would do, because they prune a database-wide time prefix. `Table.deleteHistory`
removes one table's entries from a database-scoped log, so a sibling's entry survives _below_ the
newest entry it removed, and a floor taken from that survivor certifies cursors over removed history.
- **The record's presence is the trust marker.** `Symbol.for('audit-floor')` is a different key from
`last-removed`, which is still live and still maintained by the LMDB retention loop (#2338 hardened
its write path and added tests for the retry-carry — do not remove it). They coexist because they
answer different questions: `last-removed` records where the LMDB loop got to, after the fact,
while the floor is written ahead of every one of the five prune paths and its commit is verified.
A value found under `last-removed` therefore cannot be told apart from one carrying those
guarantees, which is why the floor needs its own key rather than reusing it.
- **A store with no floor record is a store whose retention history we cannot account for.** That
includes the empty audit store an LMDB→RocksDB migration leaves behind, since `bin/copyDb.ts`
deliberately does not migrate it, and the audit-DBI-less result of a table-scoped backup taken
without `include_audit` — so `openAuditStore` stamps `max(Date.now(), newest retained key)` as a
one-time resync epoch. There is no permissive-baseline case: creating the audit DBI proves the
DBI was absent, not that the database is new.
- **That epoch is a guess, and it is recorded as one.** Its bound is surviving state, which cannot see
history a selective prune already removed: a legacy `deleteHistory` takes one table's entries out of
the shared log, so a table that held the newest entries can leave the newest _survivor_ older than
entries that are gone, and a clock rolled back between the two stamps a floor below them (#2458).
Refusing to stamp is worse — `AUDIT_FLOOR_UNKNOWN` is absorbing (`raiseAuditFloor` cannot lift it,
`establishAuditFloor` skips any existing record), so it would make every upgraded deployment fail
closed forever. So `establishAuditFloor` writes the epoch under `Symbol.for('audit-floor-bootstrap')`
first, then stamps the floor from what that record holds.

**The record's presence is the signal; comparing it against the floor is not.** A store carrying one
has an unverified pre-tracking window for as long as the record exists, however far the floor has
since moved — a prune raising the floor above the epoch certifies only what that prune removed, and
says nothing about history removed before tracking began, which may sit _above_ the epoch, since that
is precisely what the guess could not see. Worked example: a v4-era `deleteHistory` removes tableA up
to t=1000 while sibling tableB's newest survivor is 900; a rolled-back clock stamps bootstrap=900 and
floor=900; a later retention pass raises the floor to 950. A repair keyed on `floor > bootstrap` would
read 950 > 900, call it earned, and leave a consumer at cursor 970 certified over tableA's missing
950–1000. So the mark is retired by a database generation (#2451), never by a floor that climbed past
it; what the recorded _value_ is for is telling that repair how far the guess reached.

Two properties it does depend on. **Ordering:** the record is written first, so a crash between the
two writes leaves a record with no floor, which the next open retries because the early return tests
the _floor_. **Undecodable bytes are overwritten** rather than kept — unlike the floor, where a
present record may be a deliberate `AUDIT_FLOOR_UNKNOWN` and rewriting it would lower a floor.
Keeping torn bytes pinned the store to unknown _forever_: the resolver skipped the write because a
record existed, the read back failed identically on every later open, and no retry could succeed.

- **`getHistory` is not in the floor's time domain.** The floor is an audit-log key, which is what
`subscribe`'s events carry as `localTime`; `getHistory` reports each entry's origin `version` under
that same name, and a backdated or replicated write makes the two differ. A cursor saved from
`getHistory` cannot be compared against the floor.
- **The record's presence is the trust marker.** `Symbol.for('audit-floor')` is not `last-removed`,
which the LMDB retention loop still maintains (#2338 — do not remove it): that one records where the
loop got to, after the fact, so a value there cannot be told apart from a write-ahead, verified floor.
- **A store with no floor record is one whose retention history we cannot account for** (the empty
audit store a migration leaves, a table-scoped backup without `include_audit`), so `openAuditStore`
stamps `max(Date.now(), newest retained key)` as a one-time resync epoch. There is no
permissive-baseline case: creating the audit DBI proves only that it was absent.
- **That epoch is a guess, and it is recorded as one.** Surviving state cannot see history a legacy
selective prune removed, so a rolled-back clock can stamp it below entries that are gone (#2458).
Refusing to stamp is worse — `AUDIT_FLOOR_UNKNOWN` is absorbing — so `establishAuditFloor` writes the
epoch under `Symbol.for('audit-floor-bootstrap')` first, then stamps the floor from that record.

**The record's presence is the signal; comparing it against the floor is not.** Worked example: a
v4-era `deleteHistory` removes tableA up to t=1000 while tableB's newest survivor is 900; a
rolled-back clock stamps bootstrap=floor=900; a later pass raises the floor to 950 — and a cursor at
970 still sits over tableA's missing 950–1000. No timestamp can close that window, which is why
resumable positions are bound to a generation instead: a position naming one postdates tracking.

**Ordering:** the record is written first, so a crash between the two writes leaves a record with no
floor, which the next open retries. **Undecodable bytes are overwritten** (unlike the floor, where a
present record may be a deliberate unknown): keeping them pinned the store to unknown forever.

- **`getHistory` is not in the floor's time domain.** The floor is an audit-log key (`subscribe`'s
`localTime`); `getHistory` reports each entry's origin `version` under that name, which a backdated
or replicated write makes differ, so its cursors cannot be compared against the floor.
- **On RocksDB the floor tracks the configured retention horizon, not retained reality.** Whole-log-file
purge granularity means the branch cannot know which entries a purge will drop, and the floor is
written first, so each pass advances it to `Date.now() - auditRetention/(1+priority²)` whether a
Expand All @@ -284,11 +264,22 @@ Things that are easy to get wrong here:
- **Untrustworthy metadata resolves to `Infinity`, not to a number.** A wrong-length record, or eight
bytes decoding to NaN/negative, must not become a floor: `cursor < NaN` is false, so a consumer
spelling the check that way would read corrupt metadata as safe.
- **A restore is outside what the floor can see.** `restore_backup` reinstalls the backup's floor
along with everything else, so a cursor from after the backup point reads as safe against it. The
audit floor is one of three carriers of resumable state a restore rolls back (record versions and
per-node `Symbol.for('seq')` records are the others), so this wants a database-level generation
rather than a fix in this one field — harper#2451.
- **A copy keeps this floor honest for the log it carried.** A restore carries its log, so the floor
stands; a branch or migration carries none, so its stamp raises a finite floor to the copy's epoch.
Which history the database is belongs to the generation.

---

## Database generation and resumable positions

Every copy path gives the copy a generation of its own before anything can read it (harper#2451), so no live subscription or persisted position carries over from the source. **Invariant: no readable copy carries its source's generation; a position resumes only if it names the current generation, its cursor is finite, and no prune within the generation reached above it** (`isResumablePosition`).

- **Two records beside the floor:** the generation (a 16-byte random id and a float64 epoch, 0 for genesis) and the resume floor (highest prune cutoff in the generation). `raiseAuditFloor` raises both in one verified transaction; the resume floor is not absorbed by an unknown floor.
- **Stamped into the copy before publication:** restore (before `completeRestore`), branch and migration (before their renames), via a private open plus an engine flush — a root-store write is not power-loss durable and directory fsync is best-effort. `copydb` never copies the source's records and stamps the target; `compactOnStart` keeps them (same history).
- **An ordinary open never repairs:** it adopts the record, mints genesis as a compare-and-set on absence, or leaves the handle without one (every resume refused). The resume floor is read as never below a finite audit floor, since an older binary's prunes raise only the audit floor.
- **No scalar mode.** A position without an id is never resumable; bind an id only to a position established within the generation. Cursors must be progress-based: a snapshot's newest in-scope key can sit below a floor that retention advances on a quiet database, and would be refused forever.
- **Live subscriptions:** the per-path registry outlives the store, but a subscription from before a reopen can never deliver again (its commit listener and table stores belong to the closed handle). Every open first records its audit store as the only handle a registration on the path may use, then detaches the registry and ends each subscription: `DatabaseGenerationChangedError` for another or an unknown generation, the retryable `DatabaseClosingError` for the same one (its position still resumes). A registration through an earlier or closed handle, even from inside that close, is refused with the same pair of errors.
- **Not covered:** copies no generation-aware code made, a pre-generation binary pruning while the audit floor is unknown, keys reissued below a cursor after a clock rollback across a restart, cross-node identity (an id is per database per node), and subscription teardown and the handle check on a legacy LMDB `auditPath` root, which is reopened on every metadata read.

---

Expand Down
Loading
Loading