Test set size and final submission selection

Two questions about the evaluation setup, for planning purposes rather than modelling.

  1. How many exams are in the test set, and how is it split between the public and the private leaderboard? Knowing the size helps judge how much of a public-leaderboard difference is meaningful rather than sampling noise.
  2. For the final private ranking, which submission counts: the best-scoring one on the private set, the most recent one, or one that participants select explicitly? If selection is explicit, how many submissions may be selected, and by when?

Thanks.

Hi @Adilours ,

  1. We do not disclose the number of exams in the test set or how they are split between public and private.
  2. For the final private ranking, the best-scoring submission on the private set counts. Participants do not have to make any explicit selections.

Thanks!