Skip to content

Question about the release timeline and scope of the SA-Co training datasets #619

Description

@StargazerWang

Hi SAM 3 team,

Thank you for releasing SAM 3, the model checkpoints, code, and the SA-Co evaluation benchmarks.

I would like to ask about the release plan for the training datasets used by SAM 3, especially:

SA-Co/HQ
SA-Co/SYN
SA-Co/VIDEO

Could you please clarify:

Is there currently a plan to publicly release these training datasets?
If yes, is there an estimated release timeline?
How much of each dataset is expected to be released?
For example, will the full datasets be available, or only a subset?
Will the release include the original images/videos, annotations only, or image/video identifiers and download scripts?
Will the released annotations include the noun phrases, masks, hard negatives, and other metadata used during SAM 3 training?
Are there licensing or redistribution constraints that may prevent the full SA-Co training data from being released?

The paper provides useful statistics about the scale of SA-Co, but for researchers interested in reproducing or extending SAM 3, it would be very helpful to understand what portion of the training data will eventually become publicly accessible.

Even an approximate roadmap or clarification on which components are expected to remain private would be very useful.

Thank you!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions