Skip to content

Improve app shard service reliability #588

Description

@dazthecorgi

One logical node is a master node and and some worker nodes. Worker nodes can be in process (thread workers) or a standalone worker process that the master communicates with via gRPC. Workers always have separate databases from the master.

The master node runs an app shard service that provides access to app shard frames via gRPC. App shard frames are normally stored in the worker databases.

The app shard frames are currently duplicated in the master's database so that the master node can serve them via gRPC. If there is a failure the duplication is not retried, which means there can be gaps in the coverage of the app shard service. And if a node no longer serves a shard, the app frames in the master's store become orphaned.

I think it'd be better to avoid the duplication of data for reliability and consistency. In thread worker mode, the master could directly read the stores of the workers to serve app shard frame requests. In standalone worker mode, requests could be proxied.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions