How I build enough competence to supervise the model
This update was drafted on a schedule by the AI I build with, from real project notes — part of the vibecoding experiment this blog documents.
There's a line I keep coming back to from nine projects, one operator: the model is most confident exactly where I can least check it. I've chewed on that danger from a couple of angles — confidence is not sufficiency was about the output looking identical whether the ground under it is solid or air. But I never answered the obvious next question, which is: okay, so what do you actually do about the domains you don't know? You're building across trading, credit, fragrance, social, community ops — you can't be an expert in all of those. So how do you catch the model when it's confidently wrong in a field you just walked into?
And the honest answer is that I build competence on purpose, fast, and only to a specific depth. Not expertise. Just enough to supervise. Those are different things and the difference is the whole post.
Expertise is being able to produce the right answer. Supervisory competence is being able to recognize a wrong one. The second is enormously cheaper than the first, and it turns out to be most of what I need. I don't have to be able to write the Metro 2 dispute logic from memory. I have to be able to look at what the model wrote and have the right thing itch. To know which claims are the load-bearing ones, where the field hides its landmines, and what "obviously wrong to anyone who actually knows this" looks like. That's a smaller target. I can hit a smaller target in a weekend.
Here's roughly how I do it. The first move is always to find the domain's failure modes before I find its facts. Every field has a short list of ways newcomers get it embarrassingly wrong — the confident mistake that makes an expert wince. In credit it's confusing what's legal to dispute with what's effective to dispute. In trading it's mistaking a backtest for proof, which is a hole I've fallen in and written my way out of. I go looking for those first, because they're exactly where the model is confident and I'm blind. The model will happily hand me the newcomer mistake in beautiful, well-structured prose. If I've already learned the three or four classic traps, I can smell it. If I haven't, it reads as authoritative, because it is authoritatively phrased. The phrasing is never the tell. That's the trap.
The second move is to find one real expert source and read it adversarially — not to absorb everything, but to build a list of questions I can ask the model that have known answers. This is the cheapest supervision trick I've got. I don't need to know the whole field. I need five or six questions where I already know the correct answer cold, so that when I ask the model and it fumbles one, I learn something true about where its confidence and its correctness come apart in this specific domain. Some fields the model is rock solid and I can relax a little. Some it's quietly reciting a popular misconception the internet is full of, and now I know to double-check everything adjacent to that. Same model, same confidence, totally different trust — and the only way to know which is to have brought your own answer key.
The third move is to make the model show its work in a form I can check against the world, not against itself. "Explain your reasoning" doesn't help much — a confidently wrong answer produces confidently wrong reasoning that hangs together fine. What helps is forcing it to cite something external and specific: the actual rule, the actual statute, the named standard, the primary source. Then I go check the citation exists and says what it claimed. Half the time the reasoning's fine and the citation's fine. Sometimes the prose is perfect and the citation is a plausible-sounding thing that doesn't exist, and that gap is the single most useful signal I get. A made-up citation under confident prose tells me I'm standing exactly on the spot the warning was about.
Now the honest limits, because this isn't a magic trick and I don't want to sell it as one.
It doesn't scale to true depth. Supervisory competence lets me catch the obvious-to-an-expert errors. It does not catch the subtle ones — the thing that's wrong in a way only ten years in the field would flag. For those I'm genuinely exposed, and the only real mitigations are to keep the stakes low until a real expert has looked, or to get a real expert to look. That's not a dodge; it's the actual boundary. I know the shape of what I can't see, which is better than not knowing, but it's not the same as seeing it.
And it takes discipline I don't always have. The whole point of building fast with a model is speed, and the supervision step is the part that's slow and boring and feels like it's slowing you down — because it is. The temptation, every time, is to skip the answer-key step because the output looks done. Looks-done is the enemy. The output always looks done. That's the one thing the model is unconditionally excellent at, regardless of whether it's right.
So "build competence fast enough to supervise" isn't really a study plan. It's a reflex I'm trying to make automatic: before I trust a confident answer in a field I'm new to, go spend the hour it takes to be able to disbelieve it. The model's confidence is free and constant. My competence is the only thing in the loop that's actually indexed to whether the thing is true — and in a new domain, that competence doesn't exist until I go and deliberately build the thin, specific slice of it that lets me say no, that's wrong, I've seen this mistake before.