# Deal with no dependent feature in test data.Interpret results with logloss

**URL:** <https://community.drivendata.org/t/deal-with-no-dependent-feature-in-test-data-interpret-results-with-logloss/1262>\
**Category:** Warm Up: Predict Blood Donations\
**Created:** [June 21, 2017, 7:13am UTC](https://community.drivendata.org/t/deal-with-no-dependent-feature-in-test-data-interpret-results-with-logloss/1262 "2017-06-21T07:13:23Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![ggeo](https://avatars.discourse-cdn.com/v4/letter/g/a88e4f/32.png) [@ggeo](https://community.drivendata.org/u/ggeo)\
**Post date:** [June 21, 2017, 7:13am UTC](https://community.drivendata.org/t/deal-with-no-dependent-feature-in-test-data-interpret-results-with-logloss/1262/1 "2017-06-21T07:13:24Z")

</div>

Hello,

I have build some models using logloss metric , I can see a result like this 🙂

```
nIter logLoss  
  11 0.5675282
  21 0.5544149
  31 0.5745408 

```

So, ok I am taking the smallest value.

I am using the  
`predict(mymodel, newdata=test_data)`

and I am receiving something like:

`no no no no no yes no no no no no no no no no no no no no no no no no no no no no no no no no no no ....`

hence, the predictions.

I am not sure how to interpret the results.  
My predictions are yes or no (I used that because the model demands a character and not a number (1 or 0)).  
The logloss is the result from the model.

How can I finilize the results?  
If the test\_data contained the dependent variable , I would use a confusion matrix.  
But again what about logloss?

Thanks!

---

<div class="post-metadata">

**Author:** ![enric1296](https://avatars.discourse-cdn.com/v4/letter/e/cab0a1/32.png) [@enric1296](https://community.drivendata.org/u/enric1296)\
**Post date:** [June 21, 2017, 3:39pm UTC](https://community.drivendata.org/t/deal-with-no-dependent-feature-in-test-data-interpret-results-with-logloss/1262/2 "2017-06-21T15:39:07Z")

</div>

Hi,

With that script you are predicting the total dataset.

What you really want to predict is the last column wich have a categorical feature (0/1 or no/yes ) so i recomend you to use this script in order to predict only that row with a probability that minimazes the log loss using your model.

prediction \<- predict(model, data.matrix(test[,-1]))

i hope it will help you, regards!

---

<div class="post-metadata">

**Author:** ![ggeo](https://avatars.discourse-cdn.com/v4/letter/g/a88e4f/32.png) [@ggeo](https://community.drivendata.org/u/ggeo)\
**Post date:** [June 22, 2017, 6:47am UTC](https://community.drivendata.org/t/deal-with-no-dependent-feature-in-test-data-interpret-results-with-logloss/1262/3 "2017-06-22T06:47:19Z")

</div>

Hi and thanks for the answer.

I don’t know what you are trying to do with using `test[,-1]`.If you just ommit the first column which is the ID’s, then ok, I have already dropped that when I use the test\_data.

I have figured how to interpret the results.  
You just need to add `type="prob"`:

`predict(model, test_data, type="prob")`

and you have the probabilities!

---

<div class="post-metadata">

**Author:** ![enric1296](https://avatars.discourse-cdn.com/v4/letter/e/cab0a1/32.png) [@enric1296](https://community.drivendata.org/u/enric1296)\
**Post date:** [June 22, 2017, 7:07am UTC](https://community.drivendata.org/t/deal-with-no-dependent-feature-in-test-data-interpret-results-with-logloss/1262/4 "2017-06-22T07:07:23Z")

</div>

Hi

I dropped the id before predicting. And whter adding `type="prob"` or not depends on your model. Im using xgboost but if you use randomforest the predictions are only 0 or 1 so you need to add that script.
