Skip to content

Fastpath graceful shutdown not reliable on failure path #1217

Description

@LDiazN

When the fastpath process gets a sigterm, it stops receiving new measurements and waits for workers to finish all remaining work in the queue.

This approach works well for the usual use case: Everything is working properly and we want to reboot/restart/stop the Fastpath

However, when there's an issue with measurement processing (Clickhouse is unresponsive, for example) this approach presents some issues:

  1. Measurement burning: If workers are unable to process the measurement, they will error on every measurement they get from the queue, burning them forever. Measurements won't be stored in the DB, s3 or disk.
  2. sigkill escalation: If worker processes are too slow to process measurements on the queue, docker will kill them with a sigkill signal, interrupting the graceful shutdown and losing measurements stored in the in-memory queue

It's worth noting that since these measurements already left the ooniprobe service, they won't be backed up to s3 if they are lost from the queue

Impact

At the moment the maximum queue size is 500, nearly 10s worth of measurements at our current rate. So every time we're in this situation we can expect to lose ~10s worth of measurements

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working correctly

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions