# What correlation CV vs LB?

**URL:** <https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507>\
**Category:** DaT Parkinson's Prediction Challenge\
**Created:** [September 12, 2026, 9:24am UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507 "2026-09-12T09:24:39Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![bitguber](https://avatars.discourse-cdn.com/v4/letter/b/9dc877/32.png) [@bitguber](https://community.drivendata.org/u/bitguber)\
**Post date:** [September 12, 2026, 9:24am UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507/1 "2026-09-12T09:24:39Z")

</div>

mine 0.28 cv to 0.30 lb

---

<div class="post-metadata">

**Author:** ![MICADEE](https://avatars.discourse-cdn.com/v4/letter/m/d9b06d/32.png) [@MICADEE](https://community.drivendata.org/u/MICADEE)\
**Post date:** [September 12, 2026, 1:23pm UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507/2 "2026-09-12T13:23:56Z")

</div>

Mine is 0.2310 CV to 0.2586 LB

---

<div class="post-metadata">

**Author:** ![bitguber](https://avatars.discourse-cdn.com/v4/letter/b/9dc877/32.png) [@bitguber](https://community.drivendata.org/u/bitguber)\
**Post date:** [September 12, 2026, 1:58pm UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507/3 "2026-09-12T13:58:11Z")

</div>

WOW!, any pretrained model be used or heavy long training?

---

<div class="post-metadata">

**Author:** ![atatonal](https://avatars.discourse-cdn.com/v4/letter/a/e19adc/32.png) [@atatonal](https://community.drivendata.org/u/atatonal)\
**Post date:** [September 16, 2026, 8:30am UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507/4 "2026-09-16T08:30:13Z")

</div>

Across 6 submissions with 5×5 repeated grouped CV, my CV was basically uncorrelated with the LB: best CV (0.2494) → worst LB (0.2817), worst CV (0.2606) → 0.2660. The CV–LB gap was stable within architectures (~0.012–0.015 for ResNet, ~0.023–0.032 for ConvNeXt), so I trust CV to rank variants of one model, but not to compare different families.

---

<div class="post-metadata">

**Author:** ![juniorLeopard](https://avatars.discourse-cdn.com/v4/letter/j/7ba0ec/32.png) [@juniorLeopard](https://community.drivendata.org/u/juniorLeopard)\
**Post date:** [September 16, 2026, 12:20pm UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507/5 "2026-09-16T12:20:19Z")

</div>

how are you setting up your folds? grouped fold or stratified grouped folds or something else?

---

<div class="post-metadata">

**Author:** ![bitguber](https://avatars.discourse-cdn.com/v4/letter/b/9dc877/32.png) [@bitguber](https://community.drivendata.org/u/bitguber)\
**Post date:** [September 16, 2026, 12:25pm UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507/6 "2026-09-16T12:25:54Z")

</div>

i used stratified fold

---

<div class="post-metadata">

**Author:** ![atatonal](https://avatars.discourse-cdn.com/v4/letter/a/e19adc/32.png) [@atatonal](https://community.drivendata.org/u/atatonal)\
**Post date:** [September 16, 2026, 3:09pm UTC](https://community.drivendata.org/t/what-correlation-cv-vs-lb/11507/7 "2026-09-16T15:09:11Z")

</div>

Grouped, with the groups inferred rather than provided—there’s no patient or site ID, so I clustered the NIfTI headers (dimensions, voxel spacing, and field of view) into 10 groups as a proxy for acquisition site.

Plain GroupKFold, not stratified. I checked the grouping by rerunning it with the group labels shuffled while keeping the fold sizes unchanged: the shuffled version scored better, suggesting that the gap I’d attributed to site effects was mostly due to unequal folds.

The catch is that grouped CV assumes unseen test sites, and I don’t know whether that holds. If the split is random, I’ve discarded a useful prior; if it’s by site, using one would have hurt. I measured that downside at −0.013 and chose not to gamble.
