

Most AI outputs are a mix of things the model knows and things it’s filling in. Nothing in the writing will tell you which is which.
AI gets a lot wrong on the first attempt. What makes it hard to catch on review is that a sourced fact and a guess present in exactly the same tone and formatting, sometimes woven into the same sentence.
We were reminded of this while testing Copilot last month. We asked it about its own features, including what it can do and which licence you need. It got details wrong often enough that we stopped assuming and started checking each claim. If a model can’t reliably describe its own product, its broader research outputs probably deserve a second look too.
So here’s a self-audit prompt that makes it separate the guesses from the facts.
The prompt: confidence check
This runs on an output you’ve already got, as a second pass after your first prompt.
Take the answer you just gave me and mark it up. For every factual claim, label it: [Certain] if you’re confident and could point to a source, [Likely] if it’s a reasonable inference, or [Unverified] if you’re unsure or filling a gap. Then list anything I’d need to check before acting on it. Don’t rewrite the answer to sound more confident, and if a claim is Unverified, say so plainly.
A few things are doing the work here. Forcing a label onto every claim stops the model skimming past the shaky ones. The three-way split gives it somewhere honest to put “reasonable inference” instead of rounding up to fact.
The last line is important because without it models tend to strengthen weaker claims and hand back something that reads better and checks out worse.
We ran it on our weekly AI news summary
Here are three sentences of a summary, the kind of thing you might send around to your team to keep everyone informed:
Anthropic shipped Claude Opus 4.5 this week, its strongest coding model to date, and cut Sonnet pricing at the same time. Gemini 3 now leads most public benchmarks. A 2026 Glean study found knowledge workers lose close to a day a week fixing AI outputs, and analysts expect agentic AI spend to triple by 2027.
Here’s what the confidence check returned:
[Certain] Claude Opus 4.5 exists and was positioned as Anthropic’s strongest coding model on release.
[Unverified] “this week.” I don’t have a reliable date for the week this summary covers against the release date.
[Unverified] “cut Sonnet pricing at the same time.” I have joined two announcements into one sentence. I can’t confirm they happened together, or that the second happened at all.
[Likely] “Gemini 3 leads most public benchmarks.” True of several leaderboards around its release. “Now” is doing work I can’t support — standings move weekly.
[Certain] Glean study, close to a day a week. Reported figure, attributable to the 2026 study.
[Unverified] “agentic AI spend to triple by 2027.” No analyst, no report, no baseline. This is the shape of a forecast rather than one I can point to.
Check before publishing: the Opus 4.5 release date; whether a Sonnet price change happened at all; current benchmark standings; the actual source for the 2027 forecast, or cut the line.
Six claims appeared in three sentences, and two survived.
The pricing claim was a really important catch. The model had welded two separate announcements into a single sentence that reads perfectly well and is completely wrong.
Nobody catches that by re-reading, and sharing unreliable claims like this one, even internally, can cause serious reputational damage.
Where to use the prompt
Run it on anything carrying numbers, dates, names, prices or product details. Run it on anything going to a client and on any research summary where you didn’t pull every source yourself.
Skip it when you’re drafting or rewriting your own words, because there are no facts to check in a tone edit.
Three things to remember
The check runs on the same model that wrote the answer, which makes it a self-report rather than verification. A model can be confidently wrong about being confident, so if stakes are high or you're still unsure about the output, try the prompt out with a different model.
A [Certain] label means the model believes it could point to a source, without having done so. The label tells you where to look, but it doesn’t always mean you should skip looking yourself.
The check is good at surfacing the joins - that is, the places where the model inferred or filled a gap. Those are the spots you would want to inspect, and re-reading alone will never find them.
Put the prompt to use today!
Save it somewhere you can reach in two seconds, like a pinned note or a text expander shortcut. We tested this one on a weekly news summary, and it has been a fixture in our content workflow ever since.
If you want hands-on practice with techniques like these, we run AI training for business teams at every level. And if you’d like a broader framework for writing your own prompts, check out this article.


