"Relative newess" is a very shaky foundation for any argument. After all, in practice you could use all of the gradient descent methods we have today, you only need to hardcode a few derivatives. On the other hand biology suggests that activation functions are kind of irrelevant for complex networks beyond a certain size. So even in the case KANs slightly better, MLPs would win due to their simplicity in the implementation.