# Is your performance on training data quite different from that on testing data?

**URL:** <https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653>\
**Category:** Senior Data Science: Safe Aging with SPHERE\
**Created:** [June 15, 2016, 1:12am UTC](https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653 "2016-06-15T01:12:09Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![liuxixi](https://avatars.discourse-cdn.com/v4/letter/l/ad7895/32.png) [@liuxixi](https://community.drivendata.org/u/liuxixi)\
**Post date:** [June 15, 2016, 1:12am UTC](https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653/1 "2016-06-15T01:12:09Z")

</div>

I found that my performance on testing data given by the leaderboard is much lower than performance on training data.  
Is it because that the testing data is much more challenging than the training data?  
I just want to make sure whether the difference is due to the data itself, or due to my code.  
By the way, I was doing 10-fold cross validation on the training data, and the weighted Brier score looks good, but once I submitted to the leaderboard, the weighed Brier score becomes much much worse (increases with 0.08~0.09).  
Is there anyone who has the same problem as mine?  
Or your performance on training and testing data are quite close?

---

<div class="post-metadata">

**Author:** ![jgaines](https://avatars.discourse-cdn.com/v4/letter/j/bb73d2/32.png) [@jgaines](https://community.drivendata.org/u/jgaines)\
**Post date:** [June 16, 2016, 1:07am UTC](https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653/2 "2016-06-16T01:07:30Z")

</div>

My LB performance is worse than my CV results, but it’s always about the same amount worse.

This isn’t a huge dataset, and it’s probably easy to overfit.

---

<div class="post-metadata">

**Author:** ![liuxixi](https://avatars.discourse-cdn.com/v4/letter/l/ad7895/32.png) [@liuxixi](https://community.drivendata.org/u/liuxixi)\
**Post date:** [June 16, 2016, 9:15pm UTC](https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653/3 "2016-06-16T21:15:16Z")

</div>

My LB result is always much worse than CV results, and it’s not always the same amount worse. Sometimes, even though my CV results get better, the LB can still get worse, and I haven’t figured out why.  
It might be the problem of overfitting. Maybe it’s just because the training data and testing data are so different.

---

<div class="post-metadata">

**Author:** ![rpmcruz](https://avatars.discourse-cdn.com/v4/letter/r/cc9497/32.png) [@rpmcruz](https://community.drivendata.org/u/rpmcruz)\
**Post date:** [June 29, 2016, 12:39pm UTC](https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653/4 "2016-06-29T12:39:32Z")

</div>

One thing to realize when cross validating is that for whatever reason test data has less observations per file than training data. Training data has many minutes of observations, while testing has 5 to 30 seconds of observations. Depending on what you’re doing this may have an effect. You probably want to split each training file into small blocks in order to better simulate the testing files.

---

<div class="post-metadata">

**Author:** ![liuxixi](https://avatars.discourse-cdn.com/v4/letter/l/ad7895/32.png) [@liuxixi](https://community.drivendata.org/u/liuxixi)\
**Post date:** [July 7, 2016, 5:57pm UTC](https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653/5 "2016-07-07T17:57:23Z")

</div>

Is the performance given by the leaderboard for the whole data set? or just partial data set?

---

<div class="post-metadata">

**Author:** ![rpmcruz](https://avatars.discourse-cdn.com/v4/letter/r/cc9497/32.png) [@rpmcruz](https://community.drivendata.org/u/rpmcruz)\
**Post date:** [July 8, 2016, 9:04am UTC](https://community.drivendata.org/t/is-your-performance-on-training-data-quite-different-from-that-on-testing-data/653/6 "2016-07-08T09:04:28Z")

</div>

The way these websites work is that there is a **public leaderboard** and a **private leaderboard**.

You send your predictions. **Part** of them will be validated for the public leaderboard, and **part** of them will be validated by the private leaderboard. The private leaderboard is only revealed at the end. This is made to discourage overfit.
