# Scoring clarification: are known USGS/INGENIOUS faults masked when scoring, and are they in the Final Round label set?

**URL:** <https://community.drivendata.org/t/scoring-clarification-are-known-usgs-ingenious-faults-masked-when-scoring-and-are-they-in-the-final-round-label-set/11516>\
**Category:** GEMS Prize Challenge\
**Created:** [September 15, 2026, 1:22pm UTC](https://community.drivendata.org/t/scoring-clarification-are-known-usgs-ingenious-faults-masked-when-scoring-and-are-they-in-the-final-round-label-set/11516 "2026-09-15T13:22:56Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![exposed](https://avatars.discourse-cdn.com/v4/letter/e/e0b2c6/32.png) [@exposed](https://community.drivendata.org/u/exposed)\
**Post date:** [September 15, 2026, 1:22pm UTC](https://community.drivendata.org/t/scoring-clarification-are-known-usgs-ingenious-faults-masked-when-scoring-and-are-they-in-the-final-round-label-set/11516/1 "2026-09-15T13:22:56Z")

</div>

Hi organizers,

Two questions about how submissions are scored, since they change the optimal way to build a submission:

1. The problem description says the test set consists of _newly identified_ faults that are not in the USGS database, and the metric penalizes every predicted pixel of probability mass that is not within 300 m of a ground-truth pixel. If a submission predicts the known USGS/INGENIOUS fault traces (which we are asked to train on), do those predictions count as false positives, or are pixels near known faults masked out / excluded from the FP term when scoring against the new-fault labels?

2. For the Final Prize Round, is the “entire updated label set” the new-fault set (Initial Round labels + expert-verified additions), or does it also include the existing USGS/INGENIOUS faults?

Related: the description says predictions should cover “all faults in the region”. Should we read that as “predict the union of known and unknown faults”, or “predict the unknown ones”?

Thanks!

---

<div class="post-metadata">

**Author:** ![chrisk-dd](https://avatars.discourse-cdn.com/v4/letter/c/b5ac83/32.png) [@chrisk-dd](https://community.drivendata.org/u/chrisk-dd)\
**Post date:** [September 16, 2026, 6:58pm UTC](https://community.drivendata.org/t/scoring-clarification-are-known-usgs-ingenious-faults-masked-when-scoring-and-are-they-in-the-final-round-label-set/11516/2 "2026-09-16T18:58:28Z")

</div>

Good questions @exposed

1. Pixels corresponding to known USGS/INGENIOUS faults are masked / excluded from evaluation, so they do not count towards penalty terms.
2. Re-evaluation will also mask/exclude the existing USGS/INGENIOUS faults.

We’ll consider changing the description, but for scoring purposes it should not matter whether these known faults are included with predictions or not.

---

<div class="post-metadata">

**Author:** ![tarabird90](https://avatars.discourse-cdn.com/v4/letter/t/bbe5ce/32.png) [@tarabird90](https://community.drivendata.org/u/tarabird90)\
**Post date:** [September 20, 2026, 9:24pm UTC](https://community.drivendata.org/t/scoring-clarification-are-known-usgs-ingenious-faults-masked-when-scoring-and-are-they-in-the-final-round-label-set/11516/3 "2026-09-20T21:24:20Z")

</div>

Thanks for the earlier confirmation that pixels corresponding to known  
USGS/INGENIOUS faults are masked/excluded from evaluation in both rounds.  
I’d like to pin down the geometry of that mask, because the metric’s 300 m  
kernel means a pixel-exact mask and a buffered mask imply opposite  
submission strategies.

Three specific questions:

1. Is the mask pixel-exact — only the rasterized known-fault pixels  
themselves — or does it extend a buffer around them (for example the  
300 m / 3 px kernel support)?

2. Consider a predicted pixel that is not itself masked but lies 1-3 px  
from a known fault trace. Is its false-positive contribution computed  
normally, i.e. penalized at [1 - max\_g k(d)] against the new-fault  
ground truth only? Or is it also excluded?

3. Can a ground-truth pixel in the new-fault test set lie within 300 m of  
a known fault trace, or are such labels removed from the ground-truth  
set as well?

Why it matters: a lot of what an expert would add to an existing map sits  
just off the mapped line — splays, along-strike tip extensions,  
hanging-wall structures. Under a pixel-exact mask, predicting those  
corridors costs real false-positive mass. Under a buffered mask, the same  
predictions are free. I’d rather not spend a submission slot working out  
which regime applies.

Thanks!

---

<div class="post-metadata">

**Author:** ![chrisk-dd](https://avatars.discourse-cdn.com/v4/letter/c/b5ac83/32.png) [@chrisk-dd](https://community.drivendata.org/u/chrisk-dd)\
**Post date:** [September 21, 2026, 7:25am UTC](https://community.drivendata.org/t/scoring-clarification-are-known-usgs-ingenious-faults-masked-when-scoring-and-are-they-in-the-final-round-label-set/11516/4 "2026-09-21T07:25:04Z")

</div>

Hey @tarabird90

1. The mask is indeed pixel-exact - it is identical to the provided set of training fault labels.
2. Only new-fault ground truth is considered for scoring purposes. A predicted pixel that is near a known fault trace but far from a new-fault ground truth pixel will be fully penalized, i.e., the buffer does not apply to known faults.
3. A new-fault ground truth pixel can indeed lie within 300m of a known fault trace. Such pixels would constitute corrections or modifications to existing fault traces. Identifying these corrections is one outcome we are aiming for as part of this competition. Such corrections may already exist in the new-fault set, and may also exist in the final round evaluation set.
