# Official pre-trained models / external data thread

**URL:** <https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003>\
**Category:** Segmenting Buildings for Disaster Resilience\
**Created:** [December 18, 2019, 3:55pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003 "2019-12-18T15:55:20Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![glipstein](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/glipstein/32/920_2.png) [@glipstein](https://community.drivendata.org/u/glipstein)\
**Post date:** [December 18, 2019, 3:55pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/1 "2019-12-18T15:55:20Z")

</div>

Pre-trained models and external data are allowed in this competition as long as they can be released under an [Open Source License](https://opensource.org/licenses). We want to build on the best of what’s available. If you do use a pre-trained model or external data, please make sure to share in this thread for the rest of the community.

Thanks and good luck!

---

<div class="post-metadata">

**Author:** ![daveluo\_gfdrr](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@daveluo\_gfdrr](https://community.drivendata.org/u/daveluo_gfdrr)\
**Post date:** [December 20, 2019, 7:02pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/2 "2019-12-20T19:02:08Z")

</div>

To start us off, here are some great external datasets allowable for use upon sharing here and with the proper attributions:

- [SpaceNet datasets](https://spacenet.ai/datasets/), licensed as [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/):
  - also their [SpaceNet-pretrained models](https://solaris.readthedocs.io/en/latest/pretrained_models.html), available via [solaris](https://solaris.readthedocs.io/en/latest/index.html) library by [CosmiQ Works](https://www.cosmiqworks.org/)

- Any [OpenStreetMap](https://www.openstreetmap.org/) data, licensed as [ODbL 1.0](https://opendatacommons.org/licenses/odbl/summary/index.html). Can access via:
  - [HOTOSM’s export tool](https://export.hotosm.org/en/v3/)
  - [overpass turbo](https://overpass-turbo.eu/)

Generally speaking, allowable external data means that they’re publicly available and disclosed here for all participants to benefit from and licensed in a way that enables their use in models released under those [open source software licenses](https://opensource.org/licenses) mentioned above.

If you are wondering if a specific dataset or pre-trained model is allowed for use or not, please let the challenge organizers know in this thread and we’ll get back to you with an answer. Thank you!

Dave

---

<div class="post-metadata">

**Author:** ![johnowhitaker](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/johnowhitaker/32/874_2.png) [@johnowhitaker](https://community.drivendata.org/u/johnowhitaker)\
**Post date:** [December 21, 2019, 9:06am UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/3 "2019-12-21T09:06:08Z")

</div>

Do we have to share if we’re using a model trained on ImageNet (for eg most of the default models in fast.ai et al)?

---

<div class="post-metadata">

**Author:** ![daveluo\_gfdrr](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@daveluo\_gfdrr](https://community.drivendata.org/u/daveluo_gfdrr)\
**Post date:** [December 21, 2019, 3:51pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/4 "2019-12-21T15:51:08Z")

</div>

Well, you just did 🙂

For good measure, here’s a starter list of ImageNet pre-trained models:

- Fast.ai’s vision models zoo (thanks @johnowhitaker) : [https://docs.fast.ai/vision.models.html](https://docs.fast.ai/vision.models.html)
- PyTorch/torchvision pretrained models, [https://pytorch.org/docs/stable/torchvision/models.html](https://pytorch.org/docs/stable/torchvision/models.html)
- More PyTorch pretrained models from:
  - Cadene: [https://github.com/Cadene/pretrained-models.pytorch](https://github.com/Cadene/pretrained-models.pytorch)
  - Ross Wightman: [https://github.com/rwightman/pytorch-image-models](https://github.com/rwightman/pytorch-image-models)

- Tensorflow Keras pretrained models: [https://www.tensorflow.org/api\_docs/python/tf/keras/applications](https://www.tensorflow.org/api_docs/python/tf/keras/applications)

---

<div class="post-metadata">

**Author:** ![johnowhitaker](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/johnowhitaker/32/874_2.png) [@johnowhitaker](https://community.drivendata.org/u/johnowhitaker)\
**Post date:** [December 21, 2019, 3:55pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/5 "2019-12-21T15:55:13Z")

</div>

Many thanks @daveluo_gfdrr 🙂

---

<div class="post-metadata">

**Author:** ![Hasan\_N](https://avatars.discourse-cdn.com/v4/letter/h/c67d28/32.png) [@Hasan\_N](https://community.drivendata.org/u/Hasan_N)\
**Post date:** [January 18, 2020, 5:05pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/6 "2020-01-18T17:05:49Z")

</div>

Hello , i have recently participated in xview2 challenge, it is okay if i use their data also ?  
this is the link for the official xview2 challenge dataset website : [https://xview2.org/dataset](https://xview2.org/dataset)

and thank you!

---

<div class="post-metadata">

**Author:** ![daveluo\_gfdrr](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@daveluo\_gfdrr](https://community.drivendata.org/u/daveluo_gfdrr)\
**Post date:** [January 19, 2020, 10:24pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/7 "2020-01-19T22:24:51Z")

</div>

Hi @Hasan_N,

Welcome to the challenge and thank you for checking about using the xView2 dataset.

Unfortunately, that dataset can’t be used here. xView2 data is licensed as Creative Commons Attribution-Noncommercial-Sharealike 4.0 International (CC BY-NC-SA 4.0) and the NonCommercial-Sharealike part potentially limits the open sourcing of solutions developed in this challenge.

Dave

---

<div class="post-metadata">

**Author:** ![Hasan\_N](https://avatars.discourse-cdn.com/v4/letter/h/c67d28/32.png) [@Hasan\_N](https://community.drivendata.org/u/Hasan_N)\
**Post date:** [January 20, 2020, 5:45am UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/8 "2020-01-20T05:45:04Z")

</div>

Thank you @daveluo_gfdrr

---

<div class="post-metadata">

**Author:** ![akashintsev](https://avatars.discourse-cdn.com/v4/letter/a/58f4c7/32.png) [@akashintsev](https://community.drivendata.org/u/akashintsev)\
**Post date:** [February 14, 2020, 6:01am UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/9 "2020-02-14T06:01:59Z")

</div>

Open source deep learning code and pretrained models:  
[https://modelzoo.co/](https://modelzoo.co/)

---

<div class="post-metadata">

**Author:** ![azkalot1](https://avatars.discourse-cdn.com/v4/letter/a/50afbb/32.png) [@azkalot1](https://community.drivendata.org/u/azkalot1)\
**Post date:** [February 22, 2020, 11:00pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/10 "2020-02-22T23:00:24Z")

</div>

Any ideas if we can use this [https://github.com/dronedeploy/dd-ml-segmentation-benchmark](https://github.com/dronedeploy/dd-ml-segmentation-benchmark) dataset? Can’t really find the license.

---

<div class="post-metadata">

**Author:** ![azkalot1](https://avatars.discourse-cdn.com/v4/letter/a/50afbb/32.png) [@azkalot1](https://community.drivendata.org/u/azkalot1)\
**Post date:** [February 23, 2020, 3:57am UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/11 "2020-02-23T03:57:13Z")

</div>

[https://project.inria.fr/aerialimagelabeling/](https://project.inria.fr/aerialimagelabeling/) and this one

---

<div class="post-metadata">

**Author:** ![daveluo\_gfdrr](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@daveluo\_gfdrr](https://community.drivendata.org/u/daveluo_gfdrr)\
**Post date:** [February 24, 2020, 5:07pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/12 "2020-02-24T17:07:48Z")

</div>

Hi @azkalot1, thanks for sharing these resources and checking about their usage!

The [Inria Aerial Image Labeling dataset](https://project.inria.fr/aerialimagelabeling/) should be okay to use as all the imagery and labels are from public domain sources per their website:

> The dataset was constructed by combining public domain imagery and public domain official building footprints.

and the [paper](https://hal.inria.fr/hal-01468452/document) states their intention for the dataset to be open access:

> Let us first highlight the fact that we can only focus on regions where both the images and the reference data are available. In addition, we require the data to be open access in order to freely share our derived dataset with the community.

Re: the [DroneDeploy segmentation benchmark dataset](https://github.com/dronedeploy/dd-ml-segmentation-benchmark), there doesn’t seem to be any info anywhere about its license or permitted usage outside of benchmarking on Weights & Bias. I’m inquiring about it now and will update here if I hear back. Until updated otherwise here, the DroneDeploy dataset should not be used in this challenge.

---

<div class="post-metadata">

**Author:** ![akashintsev](https://avatars.discourse-cdn.com/v4/letter/a/58f4c7/32.png) [@akashintsev](https://community.drivendata.org/u/akashintsev)\
**Post date:** [February 26, 2020, 7:15am UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/13 "2020-02-26T07:15:33Z")

</div>

> **[GitHub - bonlime/keras-deeplab-v3-plus: Keras implementation of Deeplab v3+...](https://github.com/bonlime/keras-deeplab-v3-plus)**
>
> Keras implementation of Deeplab v3+ with pretrained weights - GitHub - bonlime/keras-deeplab-v3-plus: Keras implementation of Deeplab v3+ with pretrained weights

---

<div class="post-metadata">

**Author:** ![Agedev](https://avatars.discourse-cdn.com/v4/letter/a/dc4da7/32.png) [@Agedev](https://community.drivendata.org/u/Agedev)\
**Post date:** [March 1, 2020, 7:29am UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/14 "2020-03-01T07:29:56Z")

</div>

Im not sure about data from here:

[https://www.crowdai.org/challenges/mapping-challenge](https://www.crowdai.org/challenges/mapping-challenge)  
Mapping Challenge  
Building Missing Maps with Machine Learning  
License: This dataset is released under a [Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International](https://creativecommons.org/licenses/by-nc-sa/4.0/) - I presume this would be OK to use?

Also:

[https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/data](https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/data)  
[https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/rules](https://www.kaggle.com/c/dstl-satellite-imagery-feature-detection/rules)

Where the Dstl comp data comes from here:

[http://worldview3.digitalglobe.com/](http://worldview3.digitalglobe.com/)

But I cant find anywhere a specific note that data is open-source

---

<div class="post-metadata">

**Author:** ![qubvel](https://avatars.discourse-cdn.com/v4/letter/q/e36b37/32.png) [@qubvel](https://community.drivendata.org/u/qubvel)\
**Post date:** [March 1, 2020, 9:36pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/15 "2020-03-01T21:36:22Z")

</div>

> **[GitHub - qubvel/segmentation\_models.pytorch: Segmentation models with...](https://github.com/qubvel/segmentation_models.pytorch)**
>
> Segmentation models with pretrained backbones. PyTorch. - GitHub - qubvel/segmentation\_models.pytorch: Segmentation models with pretrained backbones. PyTorch.

---

<div class="post-metadata">

**Author:** ![daveluo\_gfdrr](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@daveluo\_gfdrr](https://community.drivendata.org/u/daveluo_gfdrr)\
**Post date:** [March 2, 2020, 6:25pm UTC](https://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/16 "2020-03-02T18:25:11Z")

</div>

Hi @Agedev, thanks for checking about these datasets!

RE: the CrowdAI Mapping Challenge dataset, because its license contains the NonCommercial clause, it’s not permissive enough and can’t be used in this challenge (same issue as with the [xView2 dataset inquired about earlier](http://community.drivendata.org/t/official-pre-trained-models-external-data-thread/4003/7) in thread).

In any case, I know that that CrowdAI dataset is a derived subset of the [SpaceNet 2 Buildings](https://spacenet.ai/spacenet-buildings-dataset-v2/) dataset. The original SpaceNet datasets are OK for use in this challenge as they all have a [Creative Commons Attribution-ShareAlike 4.0 International License](http://creativecommons.org/licenses/by-sa/4.0/).

Re: the DSTL Kaggle competition dataset, I can’t find any data licensing info either. If it’s WorldView imagery from DigitalGlobe/Maxar, those are usually not licensed for open usage. In general, if there’s no explicit and permissive open source license or public statement by the producers of how a dataset is intended for use by others, it can’t be used in our challenge. So for the DSTL dataset case, it’s not OK to be used in our challenge.
