Skip to content

Free lesson · about 12 minutes

When the right answer is in the pile

Ask a model for three answers instead of one and, more often, one of them is right. That doesn't mean your product shows the right one. Something still has to choose, and this lesson is about that choice.

No sign-up, no tracking, nothing stored. The 24 tasks are fictional maintenance notes I wrote by hand, and the scores are invented. The results show what each rule does with these cases. They are not a measured rate.

The misconception

“If one of the candidates is right, we've basically solved it.” The HumanEval paper that introduced Codex makes the distinction carefully. Its headline metric, pass@k, counts a problem as solved if any of k generated samples passes the unit tests. In a separate experiment, the authors can generate many samples but evaluate only one. Choosing by the model's own mean log-probability beats choosing at random, and both fall short of an oracle that already knows the tests.

In a product you rarely have the unit tests. You have candidates, maybe a ranking score, and whatever evidence you can check before showing something to a user. The gap between “one of these is right” and “we picked a right one” is where wrong answers get shipped.

Choose one yourself

A maintenance note gets turned into three fields: the component, the action and the day it's due. Below are three extractions of one note. Each quotes the part of the note it relied on and carries a ranking score. Whether each one is right stays hidden until you decide.

Note

Your call

The same 24 sets, four rules

20 of the 24 sets contain at least one right candidate. That count doesn't change with the rule; it's a property of the pile. What changes is what each rule does with it.

  • First shown. Take whichever candidate is listed first.
  • Highest score. Take the highest ranking score. Ties go to the lowest id.
  • Source check. Choose only if exactly one answer is backed by a verbatim quote. Otherwise send to review.
  • Answer key. Peeks at the labels. A ceiling for comparison, not something a product can do.

The cases were split into 12 for development and 12 held back, and the expected counts for both halves were written down before the rules ran. Both halves came out exactly as predicted.

Development set (12)

Scroll the table sideways to see every column.

Rule Right answers / all sets Wrong / answered automatically Answered automatically Sent to review (unresolved) Right answer missed / sets with one
First shown 6 / 12 6 / 12 12 / 12 0 4 / 10
Highest score 5 / 12 7 / 12 12 / 12 0 5 / 10
Source check 8 / 12 0 / 8 8 / 12 4 2 / 10
Answer key 10 / 12 0 / 10 10 / 12 2 0 / 10

Held-out set (12)

Scroll the table sideways to see every column.

Rule Right answers / all sets Wrong / answered automatically Answered automatically Sent to review (unresolved) Right answer missed / sets with one
First shown 6 / 12 6 / 12 12 / 12 0 4 / 10
Highest score 5 / 12 7 / 12 12 / 12 0 5 / 10
Source check 7 / 12 0 / 7 7 / 12 5 3 / 10
Answer key 10 / 12 0 / 10 10 / 12 2 0 / 10
Every set, every rule
Set Case First shownHighest scoreSource checkAnswer key
d01 One right answer Right c1 Right c1 Right c1 Right c1
d02 One right answer Wrong c1 Right c2 Right c2 Right c2
d03 Wrong answer scored highest Right c1 Wrong c2 Right c1 Right c1
d04 One right answer Wrong c1 Wrong c1 Right c2 Right c2
d05 No right answer Wrong c1 Wrong c1 Review no candidate passes the check Review answer key has no right answer
d06 Right words, wrong field Wrong c1 Wrong c1 Review 2 different answers pass the check Right c2
d07 Two right answers Right c1 Right c1 Right c1 Right c1
d08 Right answer, no quote Right c1 Wrong c2 Review no candidate passes the check Right c1
d09 Right answer, reworded Right c1 Right c1 Right c2 Right c1
d10 One right answer Right c1 Right c1 Right c1 Right c1
d11 Wrong answer scored highest Wrong c1 Wrong c1 Right c2 Right c2
d12 No right answer Wrong c1 Wrong c1 Review no candidate passes the check Review answer key has no right answer
h01 One right answer Right c1 Right c1 Right c1 Right c1
h02 One right answer Wrong c1 Right c2 Right c2 Right c2
h03 Wrong answer scored highest Right c1 Wrong c2 Right c1 Right c1
h04 One right answer Wrong c1 Wrong c1 Right c2 Right c2
h05 No right answer Wrong c1 Wrong c1 Review no candidate passes the check Review answer key has no right answer
h06 Right words, wrong field Wrong c1 Wrong c1 Review 3 different answers pass the check Right c2
h07 Two right answers Right c1 Right c2 Right c1 Right c1
h08 Right answer, no quote Right c1 Wrong c2 Review no candidate passes the check Right c1
h09 Right answer, reworded Right c1 Right c1 Right c2 Right c1
h10 One right answer Right c1 Right c1 Review 2 different answers pass the check Right c1
h11 Wrong answer scored highest Wrong c1 Wrong c1 Right c2 Right c2
h12 No right answer Wrong c1 Wrong c1 Review no candidate passes the check Review answer key has no right answer

What the sets show

Answering everything means answering wrong. First-shown and highest-score both answer all 24 sets. First-shown is wrong in 12 of them, highest-score in 14. Here the score does slightly worse than the display order. That's because I wrote fixtures where a confident wrong answer scores highest, and it's the point: a score is a ranking, not evidence.

Checking the source removes the wrong answers and leaves a queue. The source check answered 15 of 24 sets automatically and got none of those wrong. It sent the other 9 to a person. 4 of those had no right candidate at all, so review was the correct outcome. The rest had a right answer sitting in the pile that the rule couldn't safely pick: 5 of the 20 sets with one.

A check that passes is not a proof. In the generator set, one candidate swaps the component and the due day. Every word it reports is in its quote, so it passes. The check can see words, not which field they belong to. In the basement fan set, “replace the motor” is a substring of “replace the motor capacitor”, so a wrong action passes too. In both, the only thing that stopped a wrong answer was the rule's refusal to pick between two answers that both passed.

Deferral carries the load. In the tests, a variant that breaks ties by score instead of deferring starts shipping wrong answers. Remove the quote check as well and it becomes exactly highest-score.

Strict checks reject some right answers. A correct extraction that rewords “swap out” as “replace” fails the verbatim check, and so does a correct one that didn't quote the note. In these sets a literal candidate happened to be there as a fallback. In your product it may not be.

On the held-out half, the source check sent 5 sets to review (h05, h06, h08, h10, h12), one more than on the development half, because of the substring case. Both halves were predicted before running.

Check your reasoning

Two questions. Each answer explains itself, right or wrong.

1. A report says “a right answer was among the candidates in 83% of cases.” What does that tell you about users?
2. A rule makes no wrong automatic answers but sends 9 of 24 sets to review. How should you report it?

What to do on a real system

  1. Report the any-correct rate and the selected-correct rate side by side, with the same denominator. The difference is your selection gap.
  2. Count wrong automatic answers over automatic answers, and coverage over all cases. Keep deferrals as a separate, unresolved count; never score them as correct.
  3. Treat a ranking score as a ranking. Before trusting it, check how often the top-scored candidate is wrong.
  4. Decide what evidence a candidate must carry before it can be chosen automatically, and check it against the source, not against the candidate's own claims.
  5. Defer when more than one different answer passes. Build the review queue before you build the automation.
  6. Keep correctness labels out of anything the selector can read, and freeze your test split before tuning the rule.
  7. Include a broken control, such as a selector that guesses on ties, to prove your tests can see a wrong answer at all.

What this lesson does not show

  • No model. I wrote every candidate by hand to exercise one failure each, so the counts follow from that design. They aren't rates you should expect anywhere.
  • No real confidence scores. The scores are invented numbers chosen to make some wrong answers rank first.
  • No cost of review. A deferred set is counted as unresolved, not as solved, and the time it takes a person isn't modelled.
  • No semantic matching. The check is verbatim on purpose, which is why it rejects a correct paraphrase. A looser check would accept more right answers and more wrong ones; measure both.
  • No claim about any product, model or vendor.

The question to take back to your product

Generating more candidates is cheap. Choosing among them is the design problem. Before adding another sample, answer this: which evidence can this product actually obtain before it chooses? A quote from the source? A test that runs? A second system that agrees? A person? Whatever the answer is, it decides your coverage, your wrong answers and your review queue, far more than the size of the pile.

All 24 sets, the answer key, every rule's decision.

Questions or feedback?
I reply to every note.

Say hi