mine 0.28 cv to 0.30 lb
Mine is 0.2310 CV to 0.2586 LB
WOW!, any pretrained model be used or heavy long training?
Across 6 submissions with 5×5 repeated grouped CV, my CV was basically uncorrelated with the LB: best CV (0.2494) → worst LB (0.2817), worst CV (0.2606) → 0.2660. The CV–LB gap was stable within architectures (~0.012–0.015 for ResNet, ~0.023–0.032 for ConvNeXt), so I trust CV to rank variants of one model, but not to compare different families.
how are you setting up your folds? grouped fold or stratified grouped folds or something else?
i used stratified fold
Grouped, with the groups inferred rather than provided—there’s no patient or site ID, so I clustered the NIfTI headers (dimensions, voxel spacing, and field of view) into 10 groups as a proxy for acquisition site.
Plain GroupKFold, not stratified. I checked the grouping by rerunning it with the group labels shuffled while keeping the fold sizes unchanged: the shuffled version scored better, suggesting that the gap I’d attributed to site effects was mostly due to unequal folds.
The catch is that grouped CV assumes unseen test sites, and I don’t know whether that holds. If the split is random, I’ve discarded a useful prior; if it’s by site, using one would have hurt. I measured that downside at −0.013 and chose not to gamble.