Operational reference for replays-fetcher — the Solid Stats ingest service. It discovers
new OCAP replay files from the external replay source, stores raw replay objects in
S3-compatible storage, and writes ingestion staging records that server-2 promotes into
canonical replay records and parse jobs.
This document is the technical companion to the front-door README. For the ownership
boundary with server-2, replay-parser-2, and web, see
integration-contract.md.
The repository contains the v1 TypeScript ingest path: package scripts, strict compiler
settings, config validation, a check command, discover --dry-run, discover --store-raw,
discover --store-raw --stage, run-once, tests, and integration-contract docs.
Use Node.js 25 for the current baseline. Tooling is intentionally pinned to the latest
starting point for new work: TypeScript 6, ESLint 10, @types/node 25, Vitest 4 with V8
coverage, and Prettier 3.
Install dependencies:
pnpm installValidate the repository:
pnpm run format
pnpm run lint
pnpm run typecheck
pnpm test
pnpm run test:integration
pnpm run test:coverage
pnpm run build
pnpm run verifypnpm run verify is the fast gate (format, lint, typecheck, unit tests, coverage, build,
depcruise, knip) — no Docker required, safe to run on every change.
pnpm run test:integration is a separate, slower pre-deploy gate that runs only on
master before deploy. It starts Docker-backed PostgreSQL and MinIO via Testcontainers and
runs the full *.integration.test.ts suite — including the golden end-to-end regression tests
below — and is intentionally NOT part of verify: it may take several minutes, and its
job is to catch regressions and bugs, not to be fast. Docker must be available to run it.
The golden integration tests (src/run/golden-e2e.integration.test.ts,
src/run/golden-watch.integration.test.ts) pin the full ingest pipeline's behavior against
real source pages + real replay bytes, replayed offline so CI never hammers the source. The
fixtures are captured by a human (the agent is denied live source access):
pnpm exec tsx scripts/capture-golden-fixtures.tsRun against a configured .env (real REPLAY_SOURCE_* creds/transport). It reuses the real
source/byte clients and the production URL/parse helpers, then writes a three-tier gzip corpus
under src/run/fixtures/golden/: manifest.json, list/page-*.html.gz (10 listing pages),
detail/<id>.html.gz (each replay's detail page), and bytes/<id>.ocap.gz (each replay's raw
bytes). Until the fixtures exist, the golden tests skip cleanly, so both verify and
test:integration stay green. Once the corpus is captured and committed, they run for real
inside the test:integration pre-deploy gate (never in verify).
Client-side git hooks are managed by lefthook and sourced from the
shared @solid-stats/ts-toolchain preset via extends (single source of truth — no hook
bodies live in this repo). pnpm install wires .git/hooks/pre-commit (oxfmt + oxlint over
staged files) and .git/hooks/pre-push (typecheck + unit tests), mirroring CI.
Bypass the hooks when needed: git commit --no-verify / git push --no-verify, or
LEFTHOOK=0 git ....
Expected v1 command shape:
# Validate config before ingest work
replays-fetcher check
# Dry-run discovery without writing S3 or staging records
replays-fetcher discover --dry-run
# Discover and store raw replay objects without staging records
replays-fetcher discover --store-raw
# Discover, store raw replay objects, and write pending staging rows
replays-fetcher discover --store-raw --stage
# Run one scheduled fetch cycle
replays-fetcher run-oncediscover --dry-run is implemented for non-mutating source inspection.
discover --store-raw is implemented for raw object storage.
discover --store-raw --stage is implemented for staging writes to ingest_staging_records.
run-once is implemented for scheduled v1 ingestion.
pnpm run checkreplays-fetcher check validates required configuration and then runs real read-only probes
for the external source, the configured S3-compatible bucket, and PostgreSQL staging access.
Successful full-config output includes concrete sourceConnectivity, s3Connectivity, and
stagingConnectivity objects; it must not contain not-implemented. Expected config or
connectivity failures emit structured JSON and set exit code 2.
pnpm exec tsx src/cli.ts discover --dry-runDry-run discovery reads the configured source and emits JSON to stdout. It does not write S3
objects, staging rows, parser artifacts, local replay-list files, or server-2 business
tables.
Top-level report fields:
ok— whether discovery completed without source-level errors.mode— currentlydry-run.sourceUrl— configured source URL used for the discovery pass.generatedAt— report timestamp.counts— candidate and diagnostic totals.candidates— normalized replay candidate evidence.diagnostics— structured source or candidate warnings/errors.
pnpm exec tsx src/cli.ts discover --store-rawdiscover --store-raw loads full source and S3 configuration, discovers candidates, fetches
each replay as opaque bytes, computes the SHA-256 checksum before deriving the final key, and
stores raw objects at raw/sha256/<sha256>.ocap.
The storage path performs HEAD before PUT for idempotency:
- Missing object: writes the raw bytes with
sha256object metadata and reportsstored. - Existing object with matching size and checksum metadata: does not rewrite and reports
skipped. - Existing object with mismatched evidence: does not overwrite and reports
conflict. - Source or storage failure: reports structured
failedevidence.
The raw storage command does not write staging rows, outbox rows, parser artifacts, local
replay-list files, or server-2 business tables.
pnpm exec tsx src/cli.ts discover --store-raw --stagediscover --store-raw --stage extends the raw storage path by writing pending rows to
server-2's existing ingest_staging_records table through DATABASE_URL. It uses
source_system, source_replay_id, object_key, checksum, size_bytes,
replay_timestamp, status, promotion_evidence, and conflict_details fields compatible
with server-2.
Staging idempotency follows the server-2 schema:
- Matching
source_system + source_replay_idwith matching object evidence reportsalready_staged. - Matching source identity with changed checksum/object key reports
conflict. - Matching checksum/object key under another source identity reports
conflictsoserver-2remains the owner of duplicate lineage decisions. - Raw storage failures are not staged and are counted as skipped staging items.
The staging command does not create canonical replays, does not create parse_jobs, does
not publish RabbitMQ messages, and does not write parser artifacts.
pnpm exec tsx src/cli.ts run-oncerun-once is the v1 command intended for cron, container schedules, or an external scheduler.
It executes exactly one bounded discovery → raw storage → staging cycle and exits.
Output streams:
- stdout — exactly one compact JSON document (
CompactRunSummary). The scheduler orserver-2operator should parse this. Usejqto extract fields. - stderr — per-page lifecycle NDJSON progress events (pino). Each line is a JSON object
with a stable
eventdiscriminator (run_start,page_complete,retry,page_failed,source_unavailable,run_complete,run_partial). Greppable without affecting stdout.
Exit codes:
0— the cycle completed without expected operational failures.2— configuration, source, fetch, storage, or staging completed as a structured expected failure.
Unexpected programmer errors still throw instead of being hidden as operational failures.
Compact stdout document fields:
runId— unique ID for the one-shot run.mode—run-once.startedAtandfinishedAt— ISO timestamps.ok— whether the run had no failure categories.sourceUrl— configured source URL when discovery reached source execution.discoveredRange— first and last completed page when at least one page completed.counts— totals fordiscovered,fetched,stored,staged,duplicate,conflict,failed,skipped, anddiagnostics.failureCategories— stable failure category values for operators and schedulers.status—complete,resumable,partial, orfailedwhen present.resumeInvocation— the resume command string when the run is resumable.
The per-candidate candidates, rawStorage, staging, and diagnostics arrays are not on
stdout. To persist full per-run evidence, use the opt-in flags below.
Opt-in evidence flags:
--emit-evidence— writes the full run evidence asruns/<runId>/evidence.jsonin the configured S3 bucket. Controlled byS3_EVIDENCE_PREFIX(defaultruns). Write failures are logged at warn and do not change the exit code. Bulk pruning ofruns/is delegated to infra-owned S3 lifecycle rules.--evidence-file <path>— also writes the full run evidence as a JSON file to<path>on local disk. Dev/debug convenience only. The operator owns the path and its cleanup.
Both flags are independent and non-exclusive. Neither is set by default.
Other flags:
--resume— resume from the last completed page using the source checkpoint.
Failure categories:
config_invalidsource_unavailablefetch_failedstorage_failedstorage_conflictstaging_failedstaging_conflictnot_stageable
Run output must not include S3 secrets, database credentials, SSH command secrets, raw replay
bytes, parser artifacts, or canonical server-2 business records.
The check command JSON and the single run-once compact stdout document are the structured
operational log surfaces for this service. They must not include S3 secrets, database
credentials, SSH command secrets, raw replay bytes, parser artifacts, canonical replay
records, parse jobs, parser results, identity records, stats rows, roles, requests, or
moderation data.
The check command requires these environment variables:
REPLAY_SOURCE_URLS3_ENDPOINTS3_REGIONS3_BUCKETS3_ACCESS_KEY_IDS3_SECRET_ACCESS_KEYDATABASE_URL
Optional:
S3_FORCE_PATH_STYLEdefaults totrue.S3_CHECKPOINT_PREFIXsets the S3 key prefix for run checkpoint objects. Defaults tocheckpoints. Checkpoint objects are written at<prefix>/<slug>.json.S3_EVIDENCE_PREFIXsets the S3 key prefix for opt-in evidence artifacts written by--emit-evidence. Defaults toruns. Evidence objects are written at<prefix>/<runId>/evidence.json. Bulk retention management is delegated to infra-owned S3 lifecycle rules.REPLAY_SOURCE_TRANSPORTdefaults todirect; set tosshto fetch source pages through an operator-managed SSH host.REPLAY_SOURCE_SSH_HOSTis required whenREPLAY_SOURCE_TRANSPORT=ssh.REPLAY_SOURCE_SSH_COMMANDdefaults tocurl -fsSL --max-time 30.REPLAY_SOURCE_CONCURRENCYbounds the per-page detail/byte/store/stage fan-out. Defaults to8; valid range1–32. Out-of-range or non-numeric values are rejected before any S3/PostgreSQL write.REPLAY_SOURCE_REQUEST_SPACING_MSis the minimum spacing applied between source requests (the pacing floor that replaces the old blanket per-request delay). Defaults to250; valid range0–5000. Out-of-range or non-numeric values are rejected before any write.REPLAY_SOURCE_MAX_PAGESis an optional safety-valve cap with no default. When unset, a fullrun-onceis unbounded and stops on the first empty source page (stop-on-empty governs the range). Set it only to cap partial runs or tests to a fixed number of pages; any operator or scheduler env that relied on the previousdefault(1)behavior to fetch only page 1 must now setREPLAY_SOURCE_MAX_PAGES=1explicitly. Must be a positive integer when set.
For operators whose local IP is blocked by Cloudflare, dry-run discovery can use an allowlisted SSH host:
REPLAY_SOURCE_TRANSPORT=ssh \
REPLAY_SOURCE_SSH_HOST=<allowlisted-host> \
pnpm exec tsx src/cli.ts discover --dry-runThe SSH path is an operator-managed source transport, not the old relay service.
For discover --store-raw, the same SSH transport can be used for source page reads and
replay byte downloads. The mutating storage path additionally requires valid S3-compatible
object storage settings:
S3_ENDPOINTS3_REGIONS3_BUCKETS3_ACCESS_KEY_IDS3_SECRET_ACCESS_KEYS3_FORCE_PATH_STYLEwhen the provider does not use the default path-style behavior.
For discover --store-raw --stage, the command also requires:
DATABASE_URLpointing at theserver-2PostgreSQL database that ownsingest_staging_records.