# How to have access to the data

**URL:** <https://community.drivendata.org/t/how-to-have-access-to-the-data/8186>\
**Category:** The BioMassters\
**Created:** [November 14, 2022, 11:25pm UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186 "2022-11-14T23:25:53Z")\
**Posts on this page:** 1\
**Showing post:** 4

<div class="post-metadata">

**Author:** ![nick.burns](https://avatars.discourse-cdn.com/v4/letter/n/47e85d/32.png) [@nick.burns](https://community.drivendata.org/u/nick.burns)\
**Post date:** [November 15, 2022, 7:02am UTC](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186/4 "2022-11-15T07:02:12Z")

</div>

Ahhh, I understand! That worried me too, that it might error out during the download.

Could you perhaps write a script (like the python one below) to loop over the files and download them individually? I just checked and this works nicely. If you ran a few scripts in parallel, it wouldn’t be too slow. Sure, it’ll take a while ☹ But that’s usually one of the pain points working with good-sized data.

```auto
import os
import pandas as pd

metadata = pd.read_csv("features_metadata.csv")
for i, row in metadata.iterrows():
    
    if not os.path.exists(row.filename):
        cmd = f"aws s3 cp s3://drivendata-competition-biomassters-public-us/train_features/{row.filename} ./ --no-sign-request"
        os.system(cmd)
        
    if i > 5:
        break

```

---

_[View the full topic](https://community.drivendata.org/t/how-to-have-access-to-the-data/8186)._
