# Help with setting up cross validation

**URL:** <https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733>\
**Category:** Pover-T Tests: Predicting Poverty\
**Created:** [January 5, 2018, 2:57am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733 "2018-01-05T02:57:21Z")\
**Posts on this page:** 14\
**Page:** 1

<div class="post-metadata">

**Author:** ![2h2f](https://avatars.discourse-cdn.com/v4/letter/2/4bbf92/32.png) [@2h2f](https://community.drivendata.org/u/2h2f)\
**Post date:** [January 5, 2018, 2:57am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/1 "2018-01-05T02:57:22Z")

</div>

Hello,

I’m trying to set up my cross validation. I’m using stratified K-Fold and I wrote a mean log loss func.

I’m getting really different values in my CV and the leaderboard.

I’d appreciate any suggestions!

I’m not sure if my mean log loss is correct, but this is what I have written:

```
def a_mean_log_loss(y_true_a, y_pred_a, y_true_b, y_pred_b, y_true_c, y_pred_c):
    
    # log_loss is from sklearn.metrics 
    a_logloss = log_loss(y_true_a, y_pred_a)
    b_logloss = log_loss(y_true_b, y_pred_b)
    c_logloss = log_loss(y_true_c, y_pred_c)
    # average of each countries log loss
    return np.sum([a_logloss, b_logloss, c_logloss])/3
```

---

<div class="post-metadata">

**Author:** ![NxGTR](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/nxgtr/32/618_2.png) [@NxGTR](https://community.drivendata.org/u/NxGTR)\
**Post date:** [January 5, 2018, 5:24am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/2 "2018-01-05T05:24:23Z")

</div>

I am doing the same thing. Also got an interesting difference.

However, the difference in CV vs LB might or might not make sense… Since its logloss you cant really say… unless you know the points you are being evaluated.

I asked about this [here](http://community.drivendata.org/t/leaderboard-split/1718/1), but no response so far. So, no idea yet.

---

<div class="post-metadata">

**Author:** ![DenisVorotyntsev](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/denisvorotyntsev/32/622_2.png) [@DenisVorotyntsev](https://community.drivendata.org/u/DenisVorotyntsev)\
**Post date:** [January 5, 2018, 1:08pm UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/3 "2018-01-05T13:08:57Z")

</div>

You have to consider that datasets have a different number of rows, so you should use weighted mean.

---

<div class="post-metadata">

**Author:** ![2h2f](https://avatars.discourse-cdn.com/v4/letter/2/4bbf92/32.png) [@2h2f](https://community.drivendata.org/u/2h2f)\
**Post date:** [January 7, 2018, 1:44am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/4 "2018-01-07T01:44:21Z")

</div>

Thanks for the replies.

I tried weighted mean but that didn’t help. I still see very different numbers between CV and LB.

I also tried Adversarial validation but my classifiers could not distinguish between train and test samples, so according to what I’ve read on this, I think they are from similar distributions. Correct me if I’m wrong please.

Any other suggestions?

---

<div class="post-metadata">

**Author:** ![sagol](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/sagol/32/623_2.png) [@sagol](https://community.drivendata.org/u/sagol)\
**Post date:** [January 7, 2018, 8:21am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/5 "2018-01-07T08:21:09Z")

</div>

I suppose that public LB is imbalanced and private LB will be more close to CV.

---

<div class="post-metadata">

**Author:** ![sagol](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/sagol/32/623_2.png) [@sagol](https://community.drivendata.org/u/sagol)\
**Post date:** [January 7, 2018, 1:41pm UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/6 "2018-01-07T13:41:52Z")

</div>

I take my words back. I’v got very close CV results. I think that if the score is much less than CV, then most likely this is a result of overfiting due to imbalanced data for country B and C.

---

<div class="post-metadata">

**Author:** ![2h2f](https://avatars.discourse-cdn.com/v4/letter/2/4bbf92/32.png) [@2h2f](https://community.drivendata.org/u/2h2f)\
**Post date:** [January 7, 2018, 10:59pm UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/7 "2018-01-07T22:59:13Z")

</div>

I tried both sklearn cross\_val\_score and my own CV function using stratified K fold. With both, I get scores that are really different from LB, without a trend I can notice…  
For example: I tried a logistic regression with no feature engineering. On CV, I get ~ 0.3 but then I get ~3. on the LB. -------- here, CV is much lower than LB

However, I tried a lightgbm model and I get a CV score that is lower than the LB by ~0.1. --------CV is slightly lower than LB

Then, still using the same lightgbm parameters, I removed some features, transform, etc. and I saw little change in CV, but ~2-fold improvement on the leaderboard over the same model with all the features. It seems to be all over the place. -------- CV is much higher than LB.

If you could point out any problems with my approach, please let me know.  
My approach:

1. load data
2. Stratified KFold to split data
3. Preprocess/normalize etc on the current fold (I also tried doing this before splitting data, but wasn’t sure if that would case a data leak)
4. Train, Predict, and store those predictions
5. calculate LogLoss on those preds
6. take mean of all three country’s logloss to get final mean log loss (tried weighted average as well, but it isn’t a big difference)

---

<div class="post-metadata">

**Author:** ![sagol](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/sagol/32/623_2.png) [@sagol](https://community.drivendata.org/u/sagol)\
**Post date:** [January 8, 2018, 10:26am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/8 "2018-01-08T10:26:53Z")

</div>

1.First of all, you should pay attention to the imbalance of the classes and choose your own strategy. ([https://elitedatascience.com/imbalanced-classes](https://elitedatascience.com/imbalanced-classes)) Validation data should be chosen before any transformation.  
2. Instead of the np.mean, you should use np.average and calculate the weight of each prediction by the number of values in the resulting file (10% difference).  
3. Before any feature engineering try to create your own baseline prediction with very close CV and LB 😉

I hope this helps you.

---

<div class="post-metadata">

**Author:** ![2h2f](https://avatars.discourse-cdn.com/v4/letter/2/4bbf92/32.png) [@2h2f](https://community.drivendata.org/u/2h2f)\
**Post date:** [January 8, 2018, 2:08pm UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/9 "2018-01-08T14:08:21Z")

</div>

Thanks for the tips. I will try these out.

---

<div class="post-metadata">

**Author:** ![fbirarra](https://avatars.discourse-cdn.com/v4/letter/f/ed8c4c/32.png) [@fbirarra](https://community.drivendata.org/u/fbirarra)\
**Post date:** [January 13, 2018, 7:16pm UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/10 "2018-01-13T19:16:58Z")

</div>

Hi,

Did you manage to solve this?

Sagol, which one of the solutions proposed on elite data science would you recommend for this case?

---

<div class="post-metadata">

**Author:** ![2h2f](https://avatars.discourse-cdn.com/v4/letter/2/4bbf92/32.png) [@2h2f](https://community.drivendata.org/u/2h2f)\
**Post date:** [January 13, 2018, 9:43pm UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/11 "2018-01-13T21:43:36Z")

</div>

Yes, I think I got it working correctly now. I submitted several submissions to see how the LB scores match up to my CV scores, and it’s pretty close.

I think the key is to put the upsampling or downsampling inside the CV loop. Once I did that, I was getting closer scores between CV and LB.  
In addition to what sagol recommended, I looked at this site: [https://www.marcoaltini.com/blog/dealing-with-imbalanced-data-undersampling-oversampling-and-proper-cross-validation](https://www.marcoaltini.com/blog/dealing-with-imbalanced-data-undersampling-oversampling-and-proper-cross-validation) and this site: [http://www.alfredo.motta.name/cross-validation-done-wrong/](http://www.alfredo.motta.name/cross-validation-done-wrong/)

---

<div class="post-metadata">

**Author:** ![electricity](https://avatars.discourse-cdn.com/v4/letter/e/aca169/32.png) [@electricity](https://community.drivendata.org/u/electricity)\
**Post date:** [February 12, 2018, 10:21am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/12 "2018-02-12T10:21:42Z")

</div>

This link was helpful to me to get an approximate LogLoss: [r - How to incorporate logLoss in caret - Stack Overflow](https://stackoverflow.com/questions/29266804/how-to-incorporate-logloss-in-caret)

As per the link:

> summaryFunction=mnLogLoss

Extract LogLoss from, for example, the country A model:

> ```
> loglossA<-modelA$results$logLoss
> 
> ```

Then take a weighted average of the countries’ LogLoss values by their relative number of rows: 46% in A, 18% in B, and 36% in C.

---

<div class="post-metadata">

**Author:** ![ThanosT](https://avatars.discourse-cdn.com/v4/letter/t/a8b319/32.png) [@ThanosT](https://community.drivendata.org/u/ThanosT)\
**Post date:** [February 14, 2018, 1:27am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/13 "2018-02-14T01:27:16Z")

</div>

1. Before any feature engineering try to create your own baseline prediction with very close CV and LB

I suspect that the LB distribution is more close to 50/50 whereas training examples are 90/10 for countries B and C, so how we set up our cross\_validation to mimic this fact?  
Moreover, I tried a simple RandomOverSampler on the training set and i got better results than the benchmark solution(all things except the sampling method same), whereas when i applied a more sophisticated combined oversampling undersampling method, called SMOTEENN, the results were far worse( i saw that this method lowers the training examples).

---

<div class="post-metadata">

**Author:** ![payback](https://avatars.discourse-cdn.com/v4/letter/p/898d66/32.png) [@payback](https://community.drivendata.org/u/payback)\
**Post date:** [February 14, 2018, 11:12am UTC](https://community.drivendata.org/t/help-with-setting-up-cross-validation/1733/14 "2018-02-14T11:12:51Z")

</div>

I think you refer to SMOTE, with SMOTE you create syntetic data of minority class based on K neighbors. i Tried it and i didn’t have improvements too, i used other tecqniques to improve the LB score
