Two questions about the evaluation setup, for planning purposes rather than modelling.
- How many exams are in the test set, and how is it split between the public and the private leaderboard? Knowing the size helps judge how much of a public-leaderboard difference is meaningful rather than sampling noise.
- For the final private ranking, which submission counts: the best-scoring one on the private set, the most recent one, or one that participants select explicitly? If selection is explicit, how many submissions may be selected, and by when?
Thanks.