What running the same prompt twice tells you
This update was drafted on a schedule by the AI I build with, from real project notes — part of the vibecoding experiment this blog documents.
Let me pull one thread out of an older post and actually turn it into a method, because I mentioned it in passing in confidence is not sufficiency and then never did anything with it. The thread is this: when I can't tell whether a model's answer is solid or made up — and the whole problem is that a confident answer and a guess look identical — one of the most useful things I can do is just run the same prompt again. Fresh context, same question. And then read the two answers against each other.
The idea is dead simple and the reason it works is worth saying out loud. A model isn't producing one true answer; it's sampling from a distribution. When the thing it's answering about is well-established — stuff it saw a thousand consistent versions of — that distribution is narrow, and two runs come back basically the same. When it's reaching, interpolating, or frankly inventing, the distribution is wide, and the two runs diverge. So the spread between runs is a readout of the model's own uncertainty — one it won't give you directly, because every single answer is delivered in the same confident voice regardless. Variance is the tell that confidence hides.
Here's how I actually use it, concretely, because "run it twice" on its own isn't a method.
I run it three times, not two. Two gives you agree-or-disagree, which is a coin that can land heads twice on a guess. Three lets me see the shape of the disagreement: all three agree, two-agree-one-drifts, or all three go somewhere different. Those are three genuinely different situations and I treat them differently. More than three I rarely bother with — the returns fall off fast and I'm not trying to run a study, I'm trying to decide whether to trust an answer in the next five minutes.
What I compare is not the wording — it's the load-bearing claims. Prose will vary every time; that's noise, ignore it. I'm looking at the parts an answer actually stands on: the specific number, the named statute or rule, the yes-or-no, the direction of the recommendation, the shape of the code's core logic. If all three runs independently land on the same specific claim, that claim is probably solid — three independent samples don't usually converge on the same wrong number by accident. If one run says the limit is 30 days and another says 45 and the third hedges, I've just learned the model doesn't actually know, no matter how confidently any single run stated it.
And the rule I actually act on: convergence on specifics means I can probably trust it; divergence on specifics means stop and go check the real source. That's it. Divergence isn't the model failing — it's the model being honest in the only language it has. It's pointing at the exact spot where its confidence and its correctness come apart, which is precisely the spot I'd otherwise walk straight past because the first answer sounded fine.
Now the honest limits, because this trick has a sharp edge and I've been fooled by it.
The big one: agreement is not truth, it's consistency. If the model learned the same popular misconception a thousand times, all three runs will cheerfully agree on the wrong thing, in unison, with total confidence. Confident consensus across runs feels like verification and it absolutely is not — it only tells me the belief is stable in the model, not that it's correct. For anything load-bearing, convergence lowers my suspicion but it never clears it; I still go bring my own answer key, the way I wrote about for any domain I'm new to. Variance-checking tells me where to look. It never tells me the answer is right.
It also does nothing for questions that are supposed to have many valid answers. Run "write me a tagline" three times and of course you get three different ones — that divergence is creativity, not uncertainty, and reading it as a warning sign is just misunderstanding what you asked for. The method only means something for questions with a fact of the matter, where the runs should agree if the model knows and will diverge if it's guessing. Point it at the wrong kind of question and it's noise.
So where it's earned a permanent spot is the narrow, high-stakes factual stuff — the specific legal rule, the one number a calculation hinges on, the "is this actually how the API behaves" question — where a single confident answer is exactly the thing I've learned not to trust on its own. It costs me two extra runs and thirty seconds. In exchange I get a rough map of where the model is standing on rock and where it's standing on air. That's not proof. It's just the cheapest uncertainty signal I know how to get, and on the days I remember to use it, it points me straight at the one claim I should have checked.