# Testing Goodness of Fit of the Model

**URL:** <https://community.drivendata.org/t/testing-goodness-of-fit-of-the-model/530>\
**Category:** Warm Up: Predict Blood Donations\
**Created:** [April 11, 2016, 2:50pm UTC](https://community.drivendata.org/t/testing-goodness-of-fit-of-the-model/530 "2016-04-11T14:50:15Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![dkderden](https://avatars.discourse-cdn.com/v4/letter/d/65b543/32.png) [@dkderden](https://community.drivendata.org/u/dkderden)\
**Post date:** [April 11, 2016, 2:50pm UTC](https://community.drivendata.org/t/testing-goodness-of-fit-of-the-model/530/1 "2016-04-11T14:50:15Z")

</div>

Since we are doing logistic regression, I was wondering what you think the best test for “Goodness of Fit” you think is the best with your models.

---

<div class="post-metadata">

**Author:** ![isms](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/isms/32/9_2.png) [@isms](https://community.drivendata.org/u/isms)\
**Post date:** [April 11, 2016, 7:58pm UTC](https://community.drivendata.org/t/testing-goodness-of-fit-of-the-model/530/2 "2016-04-11T19:58:11Z")

</div>

Hi @dkderden,

The general approach for predictive modeling is to pick a metric in advance to tell us how well or poorly a given model is doing on the classification task at hand. This metric is what we ask our modeling tools to optimize in the course of fitting the model — for logistic regression this is where the optimal weights **β** are found — and then evaluated on a chunk of data for which the answers are withheld to see how the model does on new examples. (Nice explanation [here](https://class.coursera.org/ml-005/lecture/61).)

From the point of view of [the competition](https://www.drivendata.org/competitions/2/submissions/), the performance of a model is evaluated by a metric called [logarithmic loss](http://www.r-bloggers.com/making-sense-of-logarithmic-loss/). This is only one of many possible performance metrics for classification, but the twist with log loss is that it heavily penalizes classifications that are _very confident but wrong_.

Because we chose this particular metric for the competition, it’s probably the one you’ll want to use in your own exploration, but check out this [post on the Dato blog](http://blog.dato.com/how-to-evaluate-machine-learning-models-part-2a-classification-metrics) for a discussion of some other ways to do it.

Hope that was helpful, let me know if that was too basic or if it didn’t answer your question.
