# Confusion on what to train the model on

**URL:** <https://community.drivendata.org/t/confusion-on-what-to-train-the-model-on/6662>\
**Category:** Image Similarity Challenge\
**Created:** [October 12, 2021, 2:06pm UTC](https://community.drivendata.org/t/confusion-on-what-to-train-the-model-on/6662 "2021-10-12T14:06:50Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![NawasNaziru](https://avatars.discourse-cdn.com/v4/letter/n/a3d4f5/32.png) [@NawasNaziru](https://community.drivendata.org/u/NawasNaziru)\
**Post date:** [October 12, 2021, 2:06pm UTC](https://community.drivendata.org/t/confusion-on-what-to-train-the-model-on/6662/1 "2021-10-12T14:06:50Z")

</div>

If it is not recommended to train on the reference data, then as per the model output requirement which needs the reference id, how can one connect the training data to the reference data since the model doesn’t know anything about the reference data?

---

<div class="post-metadata">

**Author:** ![mike-dd](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@mike-dd](https://community.drivendata.org/u/mike-dd)\
**Post date:** [October 13, 2021, 5:24pm UTC](https://community.drivendata.org/t/confusion-on-what-to-train-the-model-on/6662/2 "2021-10-13T17:24:58Z")

</div>

Hi @NawasNaziru,

I’d encourage you to check out the [“Getting Started” blog post](https://www.drivendata.co/blog/image-similarity-challenge/). The passage below might point you in the right direction:

> Unlike for a typical supervised machine learning competition, developing a model for this challenge is not as simple as just training a model on provided labeled training data. The query and reference sets are intended for evaluation and not for training.

Hope this helps.
