Skip to content

Special:GraphStores reports In sync after page projections have failed #1259

Description

@alistair3149

A store that fell behind during a backend outage keeps reporting In sync, so the page an operator would check to
find out reports the opposite of the truth.

Cause

GraphStoreStatusLookup::stateOf() derives the state from the last successful rebuild and whether the projection was
redefined since:

if ( $lastSuccessfulRun === null || $lastSuccessfulRun->failed > 0 ) {
	return StoreSyncState::NeverBuilt;
}

return $projectionChanged === null ? StoreSyncState::InSync : StoreSyncState::Stale;

Neither input moves when a save fails. FailureIsolatingGraphDatabasePlugin swallows the failure so the edit still
commits, and writes no run record, so the store keeps reporting the last rebuild's verdict indefinitely.

Reproduction

  1. Rebuild a store so it reads In sync.
  2. Point the backend at an unreachable host.
  3. Save some pages. Each edit commits; each projection fails.
  4. Check the status.

Measured on a dev stack:

after rebuild:  neo4j  state=in-sync  lastRebuild=20260806005636  processed=123  failed=0
after 3 saves:  neo4j  state=in-sync  lastRebuild=20260806005636  processed=123  failed=0

while the two sides had actually diverged:

MediaWiki:  126 pages, including RotProbe1, RotProbe2, RotProbe3
Neo4j:      123 pages, none of them matching RotProbe

Why it matters

Swallowing the failure is deliberate and correct — the graph is rebuildable, so an outage must not block editing. The
consequence is that the log line in FailureIsolatingGraphDatabasePlugin is the only signal that the store fell
behind, and MediaWiki routes no log channel anywhere by default, so on a stock install there is no signal at all.

It also undercuts the guidance in docs/operations/maintenance.md — "Once Neo4j is back, rebuild the graph" — because
nothing tells the operator a rebuild is due. The store reads healthy, and the drift is silent and permanent until
somebody happens to rebuild.

Possible directions

Not designed, listed only to frame the discussion:

  • Record failed projections per store (a counter, or a marker with a timestamp) and let the state reflect them, cleared
    by the next rebuild that reconciles every page.
  • Treat "a save failed for this store since the last rebuild" as a staleness input alongside projection change.
  • Leave the state as it is and surface a separate "projection failures since last rebuild" figure on the page, so the
    In sync verdict keeps its current meaning.

Found while working on #1258, which routes the log channel that currently carries the only evidence.

AI-authored — Claude Code, Opus 5 (1M context); investigation and verification in-session; reproduced on a dev stack against both the graph and the wiki database.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingneo4jTouches the Neo4j graph DB

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions