Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The problem with semi-supervised learning (or rather the way it's used here) is that, if you don't label the low-confidence samples yourself, but just leave it alone instead, it might diverge and start producing worse and worse results as it thinks it knows it guessed correctly, but in fact didn't.

Basically, the problem is that you can't make a closed system learn from itself, without any outside feedback. The information has to come from somewhere.

It's a bit like someone giving you two Chinese phrases and their translation (without you knowing any Chinese beforehand), and then leaving you to translate a whole book. You will start guessing, based on what you already know, and by the end you'll have arrived to a (totally incorrect) interpretation of what you think each ideogram means.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: