# Top-ranked solutions

**URL:** <https://community.drivendata.org/t/top-ranked-solutions/5367>\
**Category:** Genetic Engineering Attribution\
**Created:** [October 21, 2020, 2:49am UTC](https://community.drivendata.org/t/top-ranked-solutions/5367 "2020-10-21T02:49:45Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![tarekhamdi](https://avatars.discourse-cdn.com/v4/letter/t/7cd45c/32.png) [@tarekhamdi](https://community.drivendata.org/u/tarekhamdi)\
**Post date:** [October 21, 2020, 2:49am UTC](https://community.drivendata.org/t/top-ranked-solutions/5367/1 "2020-10-21T02:49:45Z")

</div>

Is it possible that the Top-ranked competitors share their approaches for education purposes?

---

<div class="post-metadata">

**Author:** ![dexarsal](https://avatars.discourse-cdn.com/v4/letter/d/73ab20/32.png) [@dexarsal](https://community.drivendata.org/u/dexarsal)\
**Post date:** [October 21, 2020, 4:21pm UTC](https://community.drivendata.org/t/top-ranked-solutions/5367/2 "2020-10-21T16:21:37Z")

</div>

Think most of the people will do that after phase 2.

---

<div class="post-metadata">

**Author:** ![KieranLitschel](https://avatars.discourse-cdn.com/v4/letter/k/3da27b/32.png) [@KieranLitschel](https://community.drivendata.org/u/KieranLitschel)\
**Post date:** [October 24, 2020, 9:51am UTC](https://community.drivendata.org/t/top-ranked-solutions/5367/3 "2020-10-24T09:51:11Z")

</div>

Yep I think if you win it becomes property of Driven Data, but if not you can share it as you wish. So after phase 2 results are out (they said around late November) I’m planning to release a GitHub repo with a nice Readme explaining what I did and why along with my Jupyter notebook. I’ll defo add it to the “Share my work” section on here, and I’ll try and remember to post it here too. Excited to read what other people did too 🙂

---

<div class="post-metadata">

**Author:** ![corsair](https://avatars.discourse-cdn.com/v4/letter/c/c0e974/32.png) [@corsair](https://community.drivendata.org/u/corsair)\
**Post date:** [November 2, 2020, 7:57pm UTC](https://community.drivendata.org/t/top-ranked-solutions/5367/4 "2020-11-02T19:57:26Z")

</div>

not top-ranked (16th place), but simple, briefly:  
**basic estimator** :

- append reverse complement (GAT -\> GAT + NNN + ATC)
- TF-IDF (custom n-grams)
- TSVD (550 components)
- MLPClassifier (1 hidden, 800)

**main estimator** : ensemble (soft voting) of 3 basic estimators with different TF-IDF n-grams generator windows (custom analyzer parameter used):

```auto
    [0,1,2,3,4,5, 7]
    [0,1,2,3,4, 7]
    [0,1,2,3,4, 8]

```

---

<div class="post-metadata">

**Author:** ![KieranLitschel](https://avatars.discourse-cdn.com/v4/letter/k/3da27b/32.png) [@KieranLitschel](https://community.drivendata.org/u/KieranLitschel)\
**Post date:** [May 30, 2021, 11:27am UTC](https://community.drivendata.org/t/top-ranked-solutions/5367/5 "2021-05-30T11:27:23Z")

</div>

Much later than I planned, but here’s the repo of my final submission and report along with the Jupyter Notebooks of my model development along the way. Hope it helps 🙂

> **[GitHub - KieranLitschel/GeneticSWEM: A Simple Word-Embedding-based Model...](https://github.com/KieranLitschel/GeneticSWEM)**
>
> A Simple Word-Embedding-based Model (SWEM) for Genetic Engineering Attribution - GitHub - KieranLitschel/GeneticSWEM: A Simple Word-Embedding-based Model (SWEM) for Genetic Engineering Attribution

My report is a bit brief given the page limit for the report. I plan to extend it at some point in the future using the feedback I’ve received.
