# Data Quality in the Test Set

**URL:** <https://community.drivendata.org/t/data-quality-in-the-test-set/10263>\
**Category:** Kelp Wanted: Segmenting Kelp Forests\
**Created:** [February 2, 2024, 2:49pm UTC](https://community.drivendata.org/t/data-quality-in-the-test-set/10263 "2024-02-02T14:49:16Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![DavidEis](https://avatars.discourse-cdn.com/v4/letter/d/45deac/32.png) [@DavidEis](https://community.drivendata.org/u/DavidEis)\
**Post date:** [February 2, 2024, 2:49pm UTC](https://community.drivendata.org/t/data-quality-in-the-test-set/10263/1 "2024-02-02T14:49:16Z")

</div>

Hi,

I have a question on the Data Quality of the Test Set.  
The Problem Description Page mentions possible False Negatives in the Ground Truth Data, but that is not the only issue occurring in the Ground Truth of the Training Set.

Do we have to assume that the Test Set Kelp Ground Truth has the same structure as the Training Sets’ including all its shortcomings?

Thank you very much and best regards  
David

---

<div class="post-metadata">

**Author:** ![cszc](https://avatars.discourse-cdn.com/v4/letter/c/c68b51/32.png) [@cszc](https://community.drivendata.org/u/cszc)\
**Post date:** [February 2, 2024, 4:16pm UTC](https://community.drivendata.org/t/data-quality-in-the-test-set/10263/2 "2024-02-02T16:16:33Z")

</div>

Hi @DavidEis, thanks for the thoughtful question. Yes, you should assume that the test set has the same distribution as the training set, including errors.
