# What's your strategy?

**URL:** <https://community.drivendata.org/t/whats-your-strategy/23>\
**Category:** Warm Up: Predict Blood Donations\
**Created:** [December 24, 2014, 8:46pm UTC](https://community.drivendata.org/t/whats-your-strategy/23 "2014-12-24T20:46:58Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![bull](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/bull/32/8_2.png) [@bull](https://community.drivendata.org/u/bull)\
**Post date:** [December 24, 2014, 8:46pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/1 "2014-12-24T20:46:58Z")

</div>

We’d love to here what methods are working well in this competition! This competition is just for fun, so we want to treat it as a learning opportunity. Sharing your process and tools helps our community members that are launching their data science careers learn and improve.

- Do you use Python or R? Julia or Java? Stata or SAS?
- Are you preprocessing any of the features?
- Are you using an ensemble of methods or leaning on something standard?
- What features of the data help or hurt your solutions?

---

<div class="post-metadata">

**Author:** ![BKR](https://avatars.discourse-cdn.com/v4/letter/b/898d66/32.png) [@BKR](https://community.drivendata.org/u/BKR)\
**Post date:** [January 21, 2015, 4:38am UTC](https://community.drivendata.org/t/whats-your-strategy/23/2 "2015-01-21T04:38:49Z")

</div>

**Current Rank:** 22  
**Toolset:** R, nnet ensemble  
**Variables:** Use all variables except volume due to high correlation with number of donations. Derived one new variable - Average donations per donation period. Treat target as factor.  
**Preprocessing** : Scaled all numeric variables. Outlier removal doesn’t work for me (so far).

Any ideas on other derived variables? Other suggestions on how I can move from 0.4446 to 0.4223?

---

<div class="post-metadata">

**Author:** ![washier](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/washier/32/24_2.png) [@washier](https://community.drivendata.org/u/washier)\
**Post date:** [January 21, 2015, 8:40am UTC](https://community.drivendata.org/t/whats-your-strategy/23/3 "2015-01-21T08:40:30Z")

</div>

@BKR

Thanks for the info. I tried a GBM at first, but it seemed too keen on predicting everything as false.  
After reading your post I tried the avNNet package in R(powered by caret), which averages models built with the nnet package.

As for feature engineering : The volume feature is of no use - all donations are 250cc in size it seems. I also derived a donations per period feature, which is quite useful according to randomForest’s importance measures. Another feature I use is the ratio between the months since last donation and the months since first donation.

To summarize then :  
**Current Rank** : 33, score = 0.4492  
**Toolset** : R, avNNet package(via caret)  
**Feature engineering** : Drop volume, add donations per period and ratio between months since last and months since first donation.  
**Preprocessing** : Centered and scaled

---

<div class="post-metadata">

**Author:** ![dunkk157](https://avatars.discourse-cdn.com/v4/letter/d/bcef8e/32.png) [@dunkk157](https://community.drivendata.org/u/dunkk157)\
**Post date:** [January 22, 2015, 4:32am UTC](https://community.drivendata.org/t/whats-your-strategy/23/4 "2015-01-22T04:32:55Z")

</div>

I’ve tried h2o deep learning package and random forest.  
Since deep learning perform very well on training set but poorly on the test set , I switch to random forest [package.My](http://package.My) score improve a lot when I tune sampsize parameter.

**Current Rank** : 5, score = 0.4325  
**Toolset** : R, randomforest package  
**Feature engineering** : Drop volume, derived new variable = Tenure / Frequency  
**Preprocessing** : remove some outliers

---

<div class="post-metadata">

**Author:** ![erum123](https://avatars.discourse-cdn.com/v4/letter/e/edb3f5/32.png) [@erum123](https://community.drivendata.org/u/erum123)\
**Post date:** [January 29, 2015, 3:09pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/5 "2015-01-29T15:09:38Z")

</div>

Without doing any special preprocessing or feature engineering yet, l achieved results using generalized linear model (logistic regression). It worked better than random forest. Donated volume has strong correlation with No. of donations, hence volume is of no use.  
Current Rank: 31, score=0.4457  
Toolset: R, packages: glm, randomforest

---

<div class="post-metadata">

**Author:** ![kosta\_pal](https://avatars.discourse-cdn.com/v4/letter/k/51bf81/32.png) [@kosta\_pal](https://community.drivendata.org/u/kosta_pal)\
**Post date:** [February 9, 2015, 10:42am UTC](https://community.drivendata.org/t/whats-your-strategy/23/6 "2015-02-09T10:42:29Z")

</div>

My score moved down by 0.01 point by adding Principle Components to the data. Don’t know if that might help much.

Furthermore, average time between donations was also useful. I think further improvements might be achieved working a little bit on it.

---

<div class="post-metadata">

**Author:** ![george](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/george/32/228_2.png) [@george](https://community.drivendata.org/u/george)\
**Post date:** [November 22, 2015, 1:02am UTC](https://community.drivendata.org/t/whats-your-strategy/23/7 "2015-11-22T01:02:29Z")

</div>

# DrivenData’s _Predict Blood Donations_

Feature engineering did not improve scores in most cases. Scaling was used for algorithms that required it. Hyper-parameters were estimated by _GridSearchCV_, a brute-force stratified 10-fold cross-validated search.

**leaderboard\_score** is the contest score for predictions of the unknown test-set; lower is better. Camel-case model names refer to scikit-learn models; lower-case were hand-crafted in some way.

| model | leaderboard\_score |
| --- | --- |
| bagged\_nolearn | 0.4313 |
| ensemble of averages | 0.4370 |
| voting ensemble | 0.4396 |
| LogisticRegression | 0.4411 |
| bagged\_logit | 0.4442 |
| GradientBoostingClassifier | 0.4452 |
| LogisticRegressionCV | 0.4457 |
| bagged\_scikit\_nn | 0.4465 |
| bagged\_gbc | 0.4527 |
| nolearn | 0.4566 |
| ExtraTreesClassifier | 0.4729 |
| blending ensemble | 0.4834 |
| XGBClassifier | 0.4851 |
| BaggingClassifier | 0.4885 |
| scikit\_nn | 0.5020 |
| boosted\_svc | 0.5334 |
| SVC | 0.5336 |
| SGDClassifier | 0.5670 |
| cosine\_similarity | 0.5732 |
| boosted\_logit | 0.5891 |
| KMeans | 0.6289 |
| AdaBoostClassifier | 0.6642 |
| KNeighborsClassifier | 1.1870 |
| RandomForestClassifier | 1.7907 |

Simple logistic regression did quite well; it seems odd that bagging and boosting both reduced its performance. In general though, ensembling did improve performances.

* * *

A number of statistics were recorded for each model from 10-fold CV predictions of the training data:

- **accuracy** the proportion correctly predicted

- **logloss** the _sklearn.metrics.log\_loss_

- **AUC** the area under the ROC curve

- **f1** the weighted average of precision and recall

- **mu** the average over 100 cross-validated scores with permutations

- **std** the stdev over 100 cross-validated scores with permutations

Starting with all the variables, R’s _step_ function produced the following

```auto
Call:
lm(formula = leaderboard_score ~ mu + std, data = score_data,
    na.action = na.omit)

Residuals:
     Min 1Q Median 3Q Max
-0.18728 -0.05472 -0.03539 0.02082 0.42898

Coefficients:
            Estimate Std. Error t value Pr(>|t|)
(Intercept) 25.722 2.962 8.685 3.09e-07 ***
mu -33.089 3.897 -8.490 4.11e-07 ***
std -60.589 7.857 -7.711 1.35e-06 ***
---
Signif. codes: 0 ' ***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Residual standard error: 0.1499 on 15 degrees of freedom
  (8 observations deleted due to missingness)
Multiple R-squared: 0.8311,	Adjusted R-squared: 0.8086
F-statistic: 36.91 on 2 and 15 DF, p-value: 1.61e-06

```

Possibly **std** is a stand-in for statistical-learning’s **variance**.

* * *

The work is available on [GitHub](https://github.com/grfiv/predict-blood-donations) and [BitBucket](https://bitbucket.org/grfiv/predict-blood-donations/). (Only GitHub permits the viewing of IPython notebooks).

---

<div class="post-metadata">

**Author:** ![cj87holler](https://avatars.discourse-cdn.com/v4/letter/c/6a8cbe/32.png) [@cj87holler](https://community.drivendata.org/u/cj87holler)\
**Post date:** [May 4, 2016, 12:22pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/8 "2016-05-04T12:22:39Z")

</div>

I am very new to programming and Data Science, however, I am eager to learn! My background is in biology and education, but looking to make the shift to Data Science. I am currently taking a part time class and considering further full-time courses.

I actually plan on using this data set as a final project for my class.

**Toolset:** Python, statsmodel (to start off)

I plan on performing a simple logistic regression with one variable to start off. My problem is deciding which one. I would love to apply more variables, but again, just a beginner. The variable I plan on starting off with is total number of donations.

Any suggestions / advice are greatly appreciated!

---

<div class="post-metadata">

**Author:** ![Imbriglia](https://avatars.discourse-cdn.com/v4/letter/i/5f9b8f/32.png) [@Imbriglia](https://community.drivendata.org/u/Imbriglia)\
**Post date:** [September 8, 2016, 7:33am UTC](https://community.drivendata.org/t/whats-your-strategy/23/9 "2016-09-08T07:33:05Z")

</div>

Current Rank: 107 (0.4416)  
Toolset: R, glm, xgboost, caret  
Variables: All variables, some feature engeneering.  
Preprocessing: Tried to balance classes, but no improvement at all.

Linear models worked better for me than ensembled method (bagging or boosting). I’m surely missing something but not figured what yet.

---

<div class="post-metadata">

**Author:** ![zenzizenzi](https://avatars.discourse-cdn.com/v4/letter/z/2acd7d/32.png) [@zenzizenzi](https://community.drivendata.org/u/zenzizenzi)\
**Post date:** [September 28, 2016, 4:55am UTC](https://community.drivendata.org/t/whats-your-strategy/23/10 "2016-09-28T04:55:08Z")

</div>

**Score** : 0.4415  
**Rank** : 111  
**Toolset** : R, party::cforest() and glm()  
**Preprocessing** : Dropped volume and added average donation

I averaged results from random forest and regression function. I changed the values which were predicted negative to zero.

---

<div class="post-metadata">

**Author:** ![ghostintheshell](https://avatars.discourse-cdn.com/v4/letter/g/a88e57/32.png) [@ghostintheshell](https://community.drivendata.org/u/ghostintheshell)\
**Post date:** [December 13, 2016, 8:17am UTC](https://community.drivendata.org/t/whats-your-strategy/23/11 "2016-12-13T08:17:24Z")

</div>

Anyone using python? I encountered problem in using their log\_loss metric to select features/classifier. It seems to me their log loss function is scoring in way totally different from the one produced by R.

---

<div class="post-metadata">

**Author:** ![schliteur](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/schliteur/32/483_2.png) [@schliteur](https://community.drivendata.org/u/schliteur)\
**Post date:** [February 7, 2017, 1:22pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/12 "2017-02-07T13:22:45Z")

</div>

**Current Rank** : 105/2265 (top 5 %) , score = 0.4396  
**Toolset** : python, scikit-learn (LogisticRegression with CV)  
**Feature engineering** : New variable = log(“Months since First Donation”-“Months since Last Donation”), drop volume (perfectely linearly correlated to number of donations)  
**Preprocessing** : no outlier removal

---

<div class="post-metadata">

**Author:** ![schliteur](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/schliteur/32/483_2.png) [@schliteur](https://community.drivendata.org/u/schliteur)\
**Post date:** [February 7, 2017, 1:31pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/13 "2017-02-07T13:31:01Z")

</div>

Hi ,  
the best way to start with is to look at the correlation between the target and the single feature you want to include.  
To do so, just plot the distribution of each variable “matplotlib.pyplot.hist()” and set one color for a modality of the target variable and another color for the other one.  
The top feature would be the most seperable distribution regarding the 2 colors…

By way of an example, see below the histograms showing the distribution of each variable available in the **training** dataset:

 ![](https://canada1.discourse-cdn.com/flex028/uploads/drivendata1/original/1X/386aef88de50e9223c7f367c8b991cd782f96dc8.png)

The colors “blue” and “green” stand the modality 1 and 0 of the target variable (Made Donation in March 2007) respectively. We see that none of the variable is clearly separable in the 1D-plane AND it is generally the case in most of the “real life” datasets BUT it does not mean it is not seperable in higher dimensional plane !

A better idea may be to consider a combination of different features. Maybe you can try to divide the number of donations by the difference between months since first and last donation. See the distribution in the next post.

It is more separable, but still we are far from having two non overlapping bumps, far away from each other. You can try others combinations, the best one would be the one that gives the best result, so you can search for it by iteration on the training dataset (best done with cross-validation).

Mathieu

---

<div class="post-metadata">

**Author:** ![schliteur](https://yyz2.discourse-cdn.com/flex028/user_avatar/community.drivendata.org/schliteur/32/483_2.png) [@schliteur](https://community.drivendata.org/u/schliteur)\
**Post date:** [February 7, 2017, 4:38pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/14 "2017-02-07T16:38:46Z")

</div>

\<img src="//cdck-file-uploads-canada1.s3.dualstack.ca-central-1.amazonaws.com/flex028/uploads/drivendata1/original/1X/4ff5ea142c5171e7f54a3cf2c952946fbbafb13b.

---

<div class="post-metadata">

**Author:** ![ravibajpai](https://avatars.discourse-cdn.com/v4/letter/r/bcef8e/32.png) [@ravibajpai](https://community.drivendata.org/u/ravibajpai)\
**Post date:** [April 13, 2017, 7:15am UTC](https://community.drivendata.org/t/whats-your-strategy/23/15 "2017-04-13T07:15:45Z")

</div>

Hey, thanks for sharing the approach.  
I’ve implemented random forest too (via caret), but I don’t really understand how to tune sampsize here and the impact it can have, could you please help me with how to go about it?

---

<div class="post-metadata">

**Author:** ![ravibajpai](https://avatars.discourse-cdn.com/v4/letter/r/bcef8e/32.png) [@ravibajpai](https://community.drivendata.org/u/ravibajpai)\
**Post date:** [April 13, 2017, 7:21am UTC](https://community.drivendata.org/t/whats-your-strategy/23/16 "2017-04-13T07:21:28Z")

</div>

Hey, the approach really helped me, thanks!  
Could you please share your code please, as I’m relatively new to neural nets it would help me understand better?

---

<div class="post-metadata">

**Author:** ![Morack](https://avatars.discourse-cdn.com/v4/letter/m/a87d85/32.png) [@Morack](https://community.drivendata.org/u/Morack)\
**Post date:** [July 10, 2017, 7:29pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/17 "2017-07-10T19:29:56Z")

</div>

> [@schliteur](#):
>
> Toolset : python, scikit-learn (LogisticRegression with CV)

Hello. I am new to Data Science. I work with python and I am trying to solve the problem, but have not done good enough with the submissios. I even tried your method but couldn’ t do good enough. May be I have not done all the steps properly. This is my first real world problem.Can you please help me

Regards,

Mitul

---

<div class="post-metadata">

**Author:** ![regorhunt02052](https://avatars.discourse-cdn.com/v4/letter/r/48db29/32.png) [@regorhunt02052](https://community.drivendata.org/u/regorhunt02052)\
**Post date:** [July 24, 2017, 1:51pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/18 "2017-07-24T13:51:08Z")

</div>

Hi all, I’m a beginner. I tried out this code to do a simple test (it’s my first time doing this without DataCamp helping me along. Any comments are greatly appreciated!

import pandas as pd  
from sklearn.linear\_model import LogisticRegression  
from sklearn.model\_selection import train\_test\_split  
import numpy as np  
df = pd.read\_csv(‘C:\Users\Roger.Hunt\Downloads\BloodTRAINING.csv’)  
y = df[‘Made Donation in March 2007’]  
X = df[‘Months since First Donation’]  
X\_train, X\_test, y\_train, y\_test = train\_test\_split(X, y, test\_size=0.3)  
clf = LogisticRegression()  
clf.fit(X\_train, y\_train)  
print(clf.score(X\_test, y\_test))

---

<div class="post-metadata">

**Author:** ![regorhunt02052](https://avatars.discourse-cdn.com/v4/letter/r/48db29/32.png) [@regorhunt02052](https://community.drivendata.org/u/regorhunt02052)\
**Post date:** [July 24, 2017, 1:51pm UTC](https://community.drivendata.org/t/whats-your-strategy/23/19 "2017-07-24T13:51:43Z")

</div>

I forgot to mention that I was getting a whole bunch of errors…which is of course why I am posting 🙂

---

<div class="post-metadata">

**Author:** ![JarryJafery](https://avatars.discourse-cdn.com/v4/letter/j/4bbf92/32.png) [@JarryJafery](https://community.drivendata.org/u/JarryJafery)\
**Post date:** [July 25, 2017, 11:20am UTC](https://community.drivendata.org/t/whats-your-strategy/23/20 "2017-07-25T11:20:32Z")

</div>

I know its too late now but i am using python for this competition what problem is with your log loss function?

[Next page](https://community.drivendata.org/t/whats-your-strategy/23.md?page=2)
