# How to have access to the data

**URL:** <https://community.drivendata.org/t/how-to-have-access-to-the-data/8186>\
**Category:** The BioMassters\
**Created:** [November 14, 2022, 11:25pm UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186 "2022-11-14T23:25:53Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![hazhir\_bahrami](https://avatars.discourse-cdn.com/v4/letter/h/48db29/32.png) [@hazhir\_bahrami](https://community.drivendata.org/u/hazhir_bahrami)\
**Post date:** [November 14, 2022, 11:25pm UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/1 "2022-11-14T23:25:53Z")

</div>

Hi,  
I hope everything is going well with you. When I want to download data by command line, downloading data is interrupted. When I want to download data by passing the following link:  
s3://drivendata-competition-biomassters-public-us/train\_features/  
But, it needs “access key ID” and “secret access key.”  
I did not find these parameters in the .txt file. I would appreciate it if you help me in this regard.

---

<div class="post-metadata">

**Author:** ![nick.burns](https://avatars.discourse-cdn.com/v4/letter/n/47e85d/32.png) [@nick.burns](https://community.drivendata.org/u/nick.burns)\
**Post date:** [November 15, 2022, 5:56am UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/2 "2022-11-15T05:56:07Z")

</div>

Hi Hazhir - check out the “AWS CLI” section of the download instructions:

> ## AWS CLI
> 
> The easiest way to download data from AWS is using the AWS CLI:
> 
> ```
> https://docs.aws.amazon.com/cli/latest/userguide/cli-chap-welcome.html
> 
> ```
> 
> To download an individual data file to your local machine, the general structure is
> 
> ```
> aws s3 cp <S3 URI> <local path> --no-sign-request
> 
> ```
> 
> For example:
> 
> ```
> aws s3 cp s3://drivendata-competition-biomassters-public-us/train_features/001b0634_S1_00.tif ./ --no-sign-request
> 
> ```
> 
> The above downloads the file `001b0634_S1_00.tif` from the public bucket in the US region. Adding “–no-sign-request” allows data to be downloaded without configuring an AWS profile.
> 
> To download a directory rather than a file, use the `--recursive` flag. For example, to download all of the training data:
> 
> ```
> aws s3 cp s3://drivendata-competition-biomassters-public-us/train_features/ train_features/ --no-sign-request --recursive
> 
> ```

---

<div class="post-metadata">

**Author:** ![hazhir\_bahrami](https://avatars.discourse-cdn.com/v4/letter/h/48db29/32.png) [@hazhir\_bahrami](https://community.drivendata.org/u/hazhir_bahrami)\
**Post date:** [November 15, 2022, 6:39am UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/3 "2022-11-15T06:39:47Z")

</div>

Hi,  
Thank you so much for your reply. I did in the same way. But, the download was interrupted after about 30 minutes of downloading. When I am going to download data again using the following code:

aws s3 cp s3://drivendata-competition-biomassters-public-us/train\_features/ train\_features/ --no-sign-request --recursive

it will start with the first image. I am not able to continue downloading data from a specific image ID to end. Moreover, the number of images is a lot and it is not possible to download them one by one.

---

<div class="post-metadata">

**Author:** ![nick.burns](https://avatars.discourse-cdn.com/v4/letter/n/47e85d/32.png) [@nick.burns](https://community.drivendata.org/u/nick.burns)\
**Post date:** [November 15, 2022, 7:02am UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/4 "2022-11-15T07:02:12Z")

</div>

Ahhh, I understand! That worried me too, that it might error out during the download.

Could you perhaps write a script (like the python one below) to loop over the files and download them individually? I just checked and this works nicely. If you ran a few scripts in parallel, it wouldn’t be too slow. Sure, it’ll take a while ☹ But that’s usually one of the pain points working with good-sized data.

```auto
import os
import pandas as pd

metadata = pd.read_csv("features_metadata.csv")
for i, row in metadata.iterrows():
    
    if not os.path.exists(row.filename):
        cmd = f"aws s3 cp s3://drivendata-competition-biomassters-public-us/train_features/{row.filename} ./ --no-sign-request"
        os.system(cmd)
        
    if i > 5:
        break

```

---

<div class="post-metadata">

**Author:** ![hazhir\_bahrami](https://avatars.discourse-cdn.com/v4/letter/h/48db29/32.png) [@hazhir\_bahrami](https://community.drivendata.org/u/hazhir_bahrami)\
**Post date:** [November 15, 2022, 7:40am UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/5 "2022-11-15T07:40:01Z")

</div>

Thank you very much. I will do the same way you recommended.

---

<div class="post-metadata">

**Author:** ![ishashah](https://avatars.discourse-cdn.com/v4/letter/i/f475e1/32.png) [@ishashah](https://community.drivendata.org/u/ishashah)\
**Post date:** [November 15, 2022, 4:47pm UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/6 "2022-11-15T16:47:25Z")

</div>

Hi @hazhir_bahrami,

I’m sorry you’re having trouble downloading the data! If you are located outside of the US, you might find download speeds improve if you use another bucket. You might want to see if using the EU (`s3://drivendata-competition-biomassters-public-eu`) or Asia (`s3://drivendata-competition-biomassters-public-as`) buckets make the process faster.

Good luck, and let us know if you have any other questions.

---

<div class="post-metadata">

**Author:** ![hazhir\_bahrami](https://avatars.discourse-cdn.com/v4/letter/h/48db29/32.png) [@hazhir\_bahrami](https://community.drivendata.org/u/hazhir_bahrami)\
**Post date:** [November 17, 2022, 7:08am UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/7 "2022-11-17T07:08:26Z")

</div>

Hi,  
Thank you so much for the help.  
I started downloading data using python. The problem got solved.  
Best wishes.
