# Smoke test data source

**URL:** https://community.drivendata.org/t/smoke-test-data-source/11358
**Category:** Children’s Speech Recognition Challenge
**Created:** [February 14, 2026, 12:43am UTC](https://community.drivendata.org/t/smoke-test-data-source/11358 "2026-02-14T00:43:21Z")
**Posts on this page:** 2
**Page:** 1

<div class="post-metadata">

### Author: ![ooousay](https://avatars.discourse-cdn.com/v4/letter/o/f0a364/32.png) [@ooousay](https://community.drivendata.org/u/ooousay)
#### Post date: [February 14, 2026, 12:43am UTC](https://community.drivendata.org/t/smoke-test-data-source/11358/1 "2026-02-14T00:43:21Z")

</div>

Good evening,  
I apologize if this is naive, but am I allowed to know if the smoke test data is a subset of the 77k test utterances or if its just a subset of training data.

Thanks,

Denis

---

<div class="post-metadata">

### Author: ![jayqi](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/jayqi/32/924_2.png) [@jayqi](https://community.drivendata.org/u/jayqi)
#### Post date: [February 14, 2026, 1:59am UTC](https://community.drivendata.org/t/smoke-test-data-source/11358/2 "2026-02-14T01:59:01Z")

</div>

Hi @ooousay,

You can find this information in the documentation of the smoke test data:

> In the smoke test runtime, `data/` contains 9,000 audio files from the training set.

— [Competition: On Top of Pasketti: Children’s Speech Recognition Challenge - Word Track](https://www.drivendata.org/competitions/308/childrens-word-asr/page/978/#smoke-tests)

> In the smoke test runtime, `data/` contains 3,000 audio files from the training set.

— [Competition: On Top of Pasketti: Children’s Speech Recognition Challenge - Phonetic Track](https://www.drivendata.org/competitions/309/childrens-phonetic-asr/page/981/#smoke-tests)

You can also take a look at the “Smoke test submission format” file on the data download page for each track for the specific utterance IDs for the smoke test audio files.
