Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The context extending methods still hurt perplexity/quality some. The longer the base model is, the more effective the context extending finetunes/post training tricks will be.


Sure it does, it's not magic. But the alternative is to start dropping out text out of context entirely, which is arguably far worse.

As someone else mentioned, this is probably more due to Llama 2 being already in training when this was figured out and it's not fully accepted yet, but I wouldn't be surprised if there was LLama 3 with out of the box dynamically scaled context at some point.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: