Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I suppose that current LLMs are incapable of answering such questions by saying "I don't know". The have no notion of facts, or any other epistemic categories.

They work basically by inventing a plausible-sounding continuation of a dialog, based on an extensive learning set. They will always find a plausible-sounding answer to a plausible-sounding question: so much learning material correlates to that.

Before epistemology is introduced explicitly into their architecture, language models will remain literary devices, so to say, unable to tell "truth" from "fiction". All they learn is basically "fiction", without a way to compare to any "facts", or the notion of "facts" or "logic".



No, that's a common misconception. They do what they are asked to do, and when they are asked to provide an answer they will provide an answer. If you ask them to provide an answer if they know, or tell you that they don't know if they don't know, they will comply with that quite well, and you'll hear a lot of "I don't know"s for questions it doesn't know the answer to.


I think the truth is somewhere in between, since I’ve seen both responses: “I don’t know” and something completely made up that was presented as facts.


They kind of do, since the predictions are well calibrated before they go through RLHF, so inside the model activations there is some notion of confidence.

Even with a RLHF model, you can say "is that correct?" and after an incorrect statement it is far more likely to correct itself than after a correct statement.


In my experience, GPT-4 answers "I don't know" fairly frequently.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: