Judging output · Lesson 3
Check it against what you gave it
You paste in a forty page report and ask for a summary. What comes back is accurate about the four things you already knew, and wrong about one thing you did not. That one thing is the reason you asked.
The error is not random. The model had a gap where your document was thin, and it filled the gap with what usually goes there. Most reports of that kind say the renewal is annual, so the summary says the renewal is annual. Yours says quarterly, in a table on page 31.
This is the failure mode of every task where you supply the source: summarising, extracting, answering questions over your own documents. The answer is built partly from your text and partly from the shape of texts like yours, and nothing in the output tells you which parts came from where.
Fluency is not evidence
A sentence that reads well tells you the model has seen a lot of sentences. It tells you nothing about whether the claim inside it is in your document. Those are separate properties, and we are all bad at keeping them separate, because in human writing they usually travel together. Someone who knows a document well writes about it fluently. Someone who does not, hedges. That correlation is the thing you have spent your career reading, and it does not hold here.
Three tells, in rough order of how often they catch something.
Specificity the source cannot support. The summary says "roughly 12% of customers". Search your document for 12. If the number is not there, it came from somewhere else. This one is easy and people skip it.
Smoothness where the source was a mess. Real documents contradict themselves. A vendor contract says thirty days in one clause and sixty in an annex. If the summary reports one clean number, either it resolved the conflict without telling you or it never saw both. Both are worth knowing.
The answer matches the genre rather than the document. Ask for the risks in a project plan and you get scope creep, unclear ownership, and timeline pressure. Those are the risks in every project plan. If your plan named a specific dependency on a team that is about to reorganise, and that is not in the list, you got the genre and not the document.
Make the answer carry its evidence
The fix is mechanical and takes one line in the request. Do not ask for claims. Ask for claims with the text they came from.
For each finding, give me: the finding in one sentence, a verbatim quote from the document that supports it, and the section heading it appears under. If you cannot find a supporting quote, leave the finding out.
Now checking is a search rather than a reread. Take each quote, search the source for it, and anything that does not appear is gone. You are not verifying the reasoning, which would be slow. You are verifying that the words exist, which a text search does in seconds.
The rate at which quotes fail to appear is also the most useful signal you will get about a particular task. Two failures out of twenty is a task worth doing this way. Nine out of twenty means the model does not have what it needs and no amount of rewording will fix it.
Ask what the document does not say
The other half is absence. Models are far more willing to answer than to report that an answer is not there, so a question with no answer in your source produces a plausible one anyway.
Give it the option explicitly, and make the option cheap:
If the document does not address this, reply with "not addressed" and nothing else. That is a correct answer and I would rather have it than a guess.
Then test it. Ask something you know is absent. If it invents an answer for that, it has been inventing answers for the questions you could not check either, and you now know something important about the whole batch.
Where this doesn't help
Quotes prove a sentence exists. They do not prove it means what the summary says it means. A quote can be real, and lifted out of a clause that reverses it two lines later. For anything with stakes, the quote is where you start reading, not where you stop.
It also works badly when the answer is genuinely synthetic. "What is the overall tone of this feedback" has no sentence behind it, and demanding one produces a cherry-picked line standing in for a judgement. Use this on extraction and question answering, where there is a right passage. On synthesis, fall back to sampling: check three of the inputs yourself and see whether the summary survives them.
And a model can fabricate a quote. That is exactly why the check is a text search rather than a look at whether the quote seems plausible. Never grade the quote by eye.
The move
On any task where you supplied the source, ask for a verbatim quote per claim, then search the source for each quote. Before you trust a batch, ask one question whose answer is definitely not in the document, and see whether you get one anyway.
Exercise
You have a 40 page vendor report and you need the five findings that matter for a renewal decision. Write the request you would send, and then say what you would do with what comes back before you act on any of it.
How this gets marked
- 30%Asks for a verbatim quote or a locator alongside each claim, rather than claims alone.
- 30%Checks the quotes exist by searching the source rather than by judging whether they look right.
- 25%Gives the model a cheap way to say the source does not address something.
- 15%Tests the setup with something known to be absent, rather than trusting a run it cannot check.
The lesson is free and stays free. Marking is the part that costs us a model call, so it needs a name to record the score against.