Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Why are these models able to reduce parameters but keep quality? I know the original intuition was scale data + params = quality but it looks like we have hit an s curve on improvements from pure scaling? Is this just because we are in a memory / data crunch? Are we learning how LLMs learn and effectively training better? Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc that isn't just brute force ablations?


> Do we have a way to derive the amount of intelligence an LLM will have based on size / training / etc

No? A large model obviously can be dumb, I don't think you can infer much other than by testing it.

These small models are almost certainly worse at some things than the big models. They prize is making them dumber at things no-one cares about while retaining the capabilities people do care about. A model probably does not need to be able to give me a political treatise on the late 19th century "silver question" to be able to write me code.


> Why are these models able to reduce parameters but keep quality?

That’s the thing…they aren’t. Well not in real life use anyway from my experience, but yeah in benchmarks they’re great at it.


Presumably there is distillation or similar being used to transfer from a larger model to a smaller one.


They innovated a lot.


Hardware constrains forced this?


Possible. But by looking at other industries Chinese don't seem to need to be forced to innovate. They just can and do. Unlike the West they seem to be on the way up and it seems like sky is the limit. In the West, the interests of the shareholders and other types of rent seekers seems to be the hard limit. Chinese have no qualms about making the cow obsolete before they milk it dry.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: