Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"Relative newess" is a very shaky foundation for any argument. After all, in practice you could use all of the gradient descent methods we have today, you only need to hardcode a few derivatives. On the other hand biology suggests that activation functions are kind of irrelevant for complex networks beyond a certain size. So even in the case KANs slightly better, MLPs would win due to their simplicity in the implementation.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: