# Training on Reference images

**URL:** <https://community.drivendata.org/t/training-on-reference-images/6367>\
**Category:** Image Similarity Challenge\
**Created:** [August 4, 2021, 4:27pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367 "2021-08-04T16:27:27Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![gradient.machine](https://avatars.discourse-cdn.com/v4/letter/g/53a042/32.png) [@gradient.machine](https://community.drivendata.org/u/gradient.machine)\
**Post date:** [August 4, 2021, 4:27pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/1 "2021-08-04T16:27:27Z")

</div>

I’m a bit confused that can I train on the reference set ? Since we need to compute the similarity between the reference and query images pair, it is necessary to learn the embeddings of the reference set.

---

<div class="post-metadata">

**Author:** ![mike-dd](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@mike-dd](https://community.drivendata.org/u/mike-dd)\
**Post date:** [August 16, 2021, 5:04pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/2 "2021-08-16T17:04:00Z")

</div>

Hi gradient.machine. Thanks for the question.

Use of the reference images for training is discouraged but not prohibited. We would instead strongly encourage you to use the provided training set for training purposes. Check out the [“Getting started” blog post](https://www.drivendata.co/blog/image-similarity-challenge/) for some more information about how to use the training set.

Hope this helps. Good luck!

---

<div class="post-metadata">

**Author:** ![separate](https://avatars.discourse-cdn.com/v4/letter/s/b77776/32.png) [@separate](https://community.drivendata.org/u/separate)\
**Post date:** [August 17, 2021, 2:45am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/3 "2021-08-17T02:45:31Z")

</div>

I don’t understand.  
The rule says that “Submitted individual predictions may not take into account more than one query image or more than one reference image at a time.”  
Does the model trained with reference images is able to take into account only one reference image at each prediction? The model already saw all reference images and learned from them.

---

<div class="post-metadata">

**Author:** ![mike-dd](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@mike-dd](https://community.drivendata.org/u/mike-dd)\
**Post date:** [August 17, 2021, 12:50pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/4 "2021-08-17T12:50:25Z")

</div>

Hi separate! Thanks for the question.

The key point here is that predictions need to be independent of one another in the sense that they are not affected by the presence of other images. See the following section of the [Rules](https://www.drivendata.org/competitions/79/competition-image-similarity-1-dev/page/376/#datause):

> The goal of the competition is to reflect real world settings where new images are continuously added to the reference set, or old images get removed from it. Therefore, the used approach should have an equivalent result if all of the reference images are provided at once, or if they are provided in chunks over time. We restrict the focus here, and only allow approaches that treat each new image of the reference set independently and without any interaction with other reference images.

You should also check out the “Acceptable methods” section (6.5) of the [competition paper](https://arxiv.org/abs/2106.09672), which covers the same topic.

Please let us know if anything remains unclear.

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [August 17, 2021, 1:33pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/5 "2021-08-17T13:33:18Z")

</div>

I do NOT understand why you argue that training on the reference dataset is not prohibited. In the early reply, [Is legal for training on the reference dataset?](http://community.drivendata.org/t/is-legal-for-training-on-the-reference-dataset/6293) and the rule [Competition: Facebook AI Image Similarity Challenge: Matching Track](https://www.drivendata.org/competitions/79/competition-image-similarity-1-dev/page/376/#datause) have already pointed out "Use of augmented reference images for any other reason, including model training, is **prohibited**".

---

<div class="post-metadata">

**Author:** ![mike-dd](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@mike-dd](https://community.drivendata.org/u/mike-dd)\
**Post date:** [August 17, 2021, 6:34pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/6 "2021-08-17T18:34:55Z")

</div>

Hi wenhaowang.

Note in the sentence you have quoted that _ **augmenting** _ reference images for training purposes is prohibited. From the Rules section you can see that we define augmenting as applying a transformation to an image to generate a new image, such as the manner in which the query set images were derived from the reference set images.

Applying these transformations to the reference images and then training is prohibited. Training using the reference images themselves without augmentation is not prohibited; but again, participants are encouraged to use the training set for these purposes.

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [August 18, 2021, 4:04am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/7 "2021-08-18T04:04:04Z")

</div>

Ok, thanks for your reply.

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [August 18, 2021, 4:09am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/8 "2021-08-18T04:09:31Z")

</div>

By the way, is the resize operation, such as resizing an image to 512x512, is regarded as a transformation?

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [August 19, 2021, 5:00am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/9 "2021-08-19T05:00:04Z")

</div>

Or, whether RandomCrop is regarded as augmentation? Because for training, images with different shapes are difficult to process.

---

<div class="post-metadata">

**Author:** ![mike-dd](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@mike-dd](https://community.drivendata.org/u/mike-dd)\
**Post date:** [August 24, 2021, 1:49pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/10 "2021-08-24T13:49:29Z")

</div>

Thanks for the question, wenhaowang.

Resizing images once as part of preprocessing to get consistent image sizes is allowed, since this is a basic requirement for many CV techniques.

Any form of augmentation is allowed during inference. For example, you could run inference on multiple augmentations of a single image and aggregate the results for that image independently of other images.

Augmentation (including resizing or cropping) for training purposes is only allowed on the training set, not the reference or query sets.

I hope this helps.

From the [competition rules](https://www.drivendata.org/competitions/79/competition-image-similarity-1-dev/page/376/#datause) on augmenting:

- “Augmenting” images refers to applying a transformation to an image to generate a new image, such as the manner in which the query set images were derived from the reference set images.
- Augmenting reference images is permitted in the inference process of generating embeddings and matching scores, so long as each reference image is used independently without any interaction with other reference images. Use of augmented reference images for any other reason, including model training, is prohibited.

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [August 25, 2021, 4:35pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/11 "2021-08-25T16:35:02Z")

</div>

Thanks for your reply. It helps.

---

<div class="post-metadata">

**Author:** ![coin](https://avatars.discourse-cdn.com/v4/letter/c/779978/32.png) [@coin](https://community.drivendata.org/u/coin)\
**Post date:** [August 29, 2021, 4:38am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/12 "2021-08-29T04:38:38Z")

</div>

Hi,

Are we allowed to normalized reference images for training?

For example, we first divide the image by 255, and then submit mean and divide std of each channel. If we want to train on reference dataset, are we allowed to do this normalization please?

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [August 31, 2021, 6:48am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/14 "2021-08-31T06:48:04Z")

</div>

“Phase III: verification. (November) The organizers verify that the code submitted in advance reproduces the reported results. The participants are expected to reply to requests from the organizers to assist in this validation.  
The results will be announced at the NeurIPS competition workshop.”

---

<div class="post-metadata">

**Author:** ![coin](https://avatars.discourse-cdn.com/v4/letter/c/779978/32.png) [@coin](https://community.drivendata.org/u/coin)\
**Post date:** [September 1, 2021, 12:25am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/15 "2021-09-01T00:25:56Z")

</div>

But the organizer cannot catch collaborating between teams. For example, there are three people working together, but they registered the competition sperately as three teams. Thus they would have 3x submitting chances. The final rank would be inaccurate due to this. If they achieved rank-3 in phase-II, then rank-3 to rank-6 are occupied, the original rank-4 would be rank-7. Is this behavior allowed?

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [September 1, 2021, 4:07am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/16 "2021-09-01T04:07:34Z")

</div>

I suggest the organizer @mike-dd check the duplication of the final report and codes of the competition. I think in this way, the mentioned behavior could be avoided.

---

<div class="post-metadata">

**Author:** ![separate](https://avatars.discourse-cdn.com/v4/letter/s/b77776/32.png) [@separate](https://community.drivendata.org/u/separate)\
**Post date:** [September 4, 2021, 9:52am UTC](https://community.drivendata.org/t/training-on-reference-images/6367/17 "2021-09-04T09:52:34Z")

</div>

Hi @mike-dd,

> Resizing images once as part of preprocessing to get consistent image sizes is allowed, since this is a basic requirement for many CV techniques.

> Augmentation (including resizing or cropping) for training purposes is only allowed on the training set, not the reference or query sets.

Above two statements seem to contradict each other. Just to be clear, is ‘Resizing images once on the reference sets for training purposes’ allowed?

And another question is about the statement from the paper,

> This means that even if the dataset had a single query image and a single reference image, the  
> score of the image pair would be the same.

For one query-reference image pair A-B, one model that trained on all reference images and another model that trained only on one reference image(B) will have different score, which seems to break the above statement.  
So, can I understand that above statement(same score for different dataset size) is applied in ‘inference’ stage? And in ‘training’ stage, we are just okay to train using all reference images?

---

<div class="post-metadata">

**Author:** ![mike-dd](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@mike-dd](https://community.drivendata.org/u/mike-dd)\
**Post date:** [September 7, 2021, 2:49pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/18 "2021-09-07T14:49:15Z")

</div>

Hi separate,

Yes to all of your questions.

Resizing images once on the reference sets for training purposes is allowed.

The requirement that an image score be unchanged by the dataset size is applicable at the inference step. It is ok to train on all reference images, but as you know participants are encouraged to use the training set for training purposes.

I hope this helps.

---

<div class="post-metadata">

**Author:** ![separate](https://avatars.discourse-cdn.com/v4/letter/s/b77776/32.png) [@separate](https://community.drivendata.org/u/separate)\
**Post date:** [September 7, 2021, 4:30pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/19 "2021-09-07T16:30:57Z")

</div>

Thanks for answering repeated questions! Really appreciate it.

---

<div class="post-metadata">

**Author:** ![wenhaowang](https://avatars.discourse-cdn.com/v4/letter/w/b2d939/32.png) [@wenhaowang](https://community.drivendata.org/u/wenhaowang)\
**Post date:** [September 7, 2021, 4:54pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/20 "2021-09-07T16:54:11Z")

</div>

> [@separate](#):
>
> Augmentation (including resizing or cropping) for training purposes is only allowed on the training set, not the reference or query sets.

Sorry, I cannot understand. @mike-dd

> [@mike-dd](#):
>
> Resizing images once on the reference sets for training purposes is allowed.

I think they are incompatible.

---

<div class="post-metadata">

**Author:** ![mike-dd](https://avatars.discourse-cdn.com/v4/letter/m/f0a364/32.png) [@mike-dd](https://community.drivendata.org/u/mike-dd)\
**Post date:** [September 30, 2021, 3:41pm UTC](https://community.drivendata.org/t/training-on-reference-images/6367/21 "2021-09-30T15:41:33Z")

</div>

Hi wenhaowang,

Resizing images, if applied once to an image as part of preprocessing, is not considered augmentation. Augmentation for training purposes is only allowed on the training set, not the reference or query sets.

From the [competition rules](https://www.drivendata.org/competitions/79/competition-image-similarity-1-dev/page/376/#datause):

> “Augmenting” images refers to applying a transformation to an image to generate a new image, such as the manner in which the query set images were derived from the reference set images.

Again, participants are encouraged to use the training set for training purposes, and to use the ground truth labels as a test/validation set.
