If you ask it 3 times to generate 3 verifiable reports that confirm its claims, will it answer with the same process and same answers each time?
See, when you ask a human to justify a result or a decision, they can often do this very meticulously. If a judge writes a decision from the bench, or a firefighter describes how his battalion knocked down an apartment fire, or a systems admin describes how he configured a NAS, they will all be relying on their training, and precedent, and specifications, and things like that, and they can give you reproducible results and solid justifications for the way they did things. When mistakes are made, and money or life is lost, they can be accountable and you can modify that process to set a precedent for the future.
But Claude? How in the world will it produce the same results twice? It is non-deterministic. That is the fundamental issue of LLMs and genAI today. They are all non-deterministic, and SWE treat them as if they are somehow reliable, or produce reproducible results, or that they can follow a procedure or a specification, outlined in their prompts and context, and produce results.
No, they only produce results by accident and happenstance, and they only justify them ex post facto by making things up. There is no humanity or deterministic activity in an LLM. You'll never verify "why" they chose that string of tokens, because they could've easily chosen a very different stream of tokens. In fact, now with watermarking, the most deterministic thing will be hitting that watermark standard at all costs!
If you ask it 3 times to generate 3 verifiable reports that confirm its claims, will it answer with the same process and same answers each time?
See, when you ask a human to justify a result or a decision, they can often do this very meticulously. If a judge writes a decision from the bench, or a firefighter describes how his battalion knocked down an apartment fire, or a systems admin describes how he configured a NAS, they will all be relying on their training, and precedent, and specifications, and things like that, and they can give you reproducible results and solid justifications for the way they did things. When mistakes are made, and money or life is lost, they can be accountable and you can modify that process to set a precedent for the future.
But Claude? How in the world will it produce the same results twice? It is non-deterministic. That is the fundamental issue of LLMs and genAI today. They are all non-deterministic, and SWE treat them as if they are somehow reliable, or produce reproducible results, or that they can follow a procedure or a specification, outlined in their prompts and context, and produce results.
No, they only produce results by accident and happenstance, and they only justify them ex post facto by making things up. There is no humanity or deterministic activity in an LLM. You'll never verify "why" they chose that string of tokens, because they could've easily chosen a very different stream of tokens. In fact, now with watermarking, the most deterministic thing will be hitting that watermark standard at all costs!