Skip to content

AOF: checkpoints (cut capture on saves, automatic trigger, health) #8410

Description

@romange

Part of #8404 (AOF MVP). Design: Checkpoint Protocol.

While AOF is on, every local save is a checkpoint.

Scope:

  • The cut: a global transaction. On each shard, without preemption, RegisterChangeListener,
    L_i = journal::GetLsn(), and AofStreamer::SealAndRotate(L_i).
  • aof-cut-lsn (replacing aof-preamble) and aof-cut-time aux fields in base files.
  • fsync each base file before its rename (snapshot files are not fsynced today). A failed fsync
    fails the save before the rename.
  • Manifest commit, then GC. Saves to cloud storage stay plain saves: no cut, no rotation, no
    manifest.
  • Automatic trigger (--aof_rewrite_min_size, --aof_rewrite_percentage), written as an AOF-owned
    base. It is dropped when a qualifying save is already running.
  • A sync failure schedules a checkpoint whose cut comes after the failure.
  • Checkpoint health in INFO persistence: aof_last_bgrewrite_status, aof_checkpoint_failures,
    aof_base_file, sizes. Retry with backoff.

Depends on: #8408, #8409.

Done when: tests cover a crash injected at each checkpoint step, a same-name save crashing
between its rename and the manifest commit, repeated checkpoint failures, and --snapshot_cron
keeping the log bounded.

Open edge cases to settle: DFS saves that reuse a name (files are renamed one by one); a base
deleted by the user at runtime; plain saves repeatedly delaying the automatic checkpoint.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions