Free lesson · about 12 minutes
When the right answer is in the pile
Ask a model for three answers instead of one and, more often, one of them is right. That doesn't mean your product shows the right one. Something still has to choose, and this lesson is about that choice.
No sign-up, no tracking, nothing stored. The 24 tasks are fictional maintenance notes I wrote by hand, and the scores are invented. The results show what each rule does with these cases. They are not a measured rate.
The misconception
“If one of the candidates is right, we've basically solved it.” The HumanEval paper that introduced Codex makes the distinction carefully. Its headline metric, pass@k, counts a problem as solved if any of k generated samples passes the unit tests. In a separate experiment, the authors can generate many samples but evaluate only one. Choosing by the model's own mean log-probability beats choosing at random, and both fall short of an oracle that already knows the tests.
In a product you rarely have the unit tests. You have candidates, maybe a ranking score, and whatever evidence you can check before showing something to a user. The gap between “one of these is right” and “we picked a right one” is where wrong answers get shipped.
Choose one yourself
A maintenance note gets turned into three fields: the component, the action and the day it's due. Below are three extractions of one note. Each quotes the part of the note it relied on and carries a ranking score. Whether each one is right stays hidden until you decide.
Your call
The same 24 sets, four rules
20 of the 24 sets contain at least one right candidate. That count doesn't change with the rule; it's a property of the pile. What changes is what each rule does with it.
- First shown. Take whichever candidate is listed first.
- Highest score. Take the highest ranking score. Ties go to the lowest id.
- Source check. Choose only if exactly one answer is backed by a verbatim quote. Otherwise send to review.
- Answer key. Peeks at the labels. A ceiling for comparison, not something a product can do.
The cases were split into 12 for development and 12 held back, and the expected counts for both halves were written down before the rules ran. Both halves came out exactly as predicted.
Development set (12)
Scroll the table sideways to see every column.
| Rule | Right answers / all sets | Wrong / answered automatically | Answered automatically | Sent to review (unresolved) | Right answer missed / sets with one |
|---|---|---|---|---|---|
| First shown | 6 / 12 | 6 / 12 | 12 / 12 | 0 | 4 / 10 |
| Highest score | 5 / 12 | 7 / 12 | 12 / 12 | 0 | 5 / 10 |
| Source check | 8 / 12 | 0 / 8 | 8 / 12 | 4 | 2 / 10 |
| Answer key | 10 / 12 | 0 / 10 | 10 / 12 | 2 | 0 / 10 |
Held-out set (12)
Scroll the table sideways to see every column.
| Rule | Right answers / all sets | Wrong / answered automatically | Answered automatically | Sent to review (unresolved) | Right answer missed / sets with one |
|---|---|---|---|---|---|
| First shown | 6 / 12 | 6 / 12 | 12 / 12 | 0 | 4 / 10 |
| Highest score | 5 / 12 | 7 / 12 | 12 / 12 | 0 | 5 / 10 |
| Source check | 7 / 12 | 0 / 7 | 7 / 12 | 5 | 3 / 10 |
| Answer key | 10 / 12 | 0 / 10 | 10 / 12 | 2 | 0 / 10 |
Every set, every rule
| Set | Case | First shown | Highest score | Source check | Answer key |
|---|---|---|---|---|---|
| d01 | One right answer | Right c1 | Right c1 | Right c1 | Right c1 |
| d02 | One right answer | Wrong c1 | Right c2 | Right c2 | Right c2 |
| d03 | Wrong answer scored highest | Right c1 | Wrong c2 | Right c1 | Right c1 |
| d04 | One right answer | Wrong c1 | Wrong c1 | Right c2 | Right c2 |
| d05 | No right answer | Wrong c1 | Wrong c1 | Review no candidate passes the check | Review answer key has no right answer |
| d06 | Right words, wrong field | Wrong c1 | Wrong c1 | Review 2 different answers pass the check | Right c2 |
| d07 | Two right answers | Right c1 | Right c1 | Right c1 | Right c1 |
| d08 | Right answer, no quote | Right c1 | Wrong c2 | Review no candidate passes the check | Right c1 |
| d09 | Right answer, reworded | Right c1 | Right c1 | Right c2 | Right c1 |
| d10 | One right answer | Right c1 | Right c1 | Right c1 | Right c1 |
| d11 | Wrong answer scored highest | Wrong c1 | Wrong c1 | Right c2 | Right c2 |
| d12 | No right answer | Wrong c1 | Wrong c1 | Review no candidate passes the check | Review answer key has no right answer |
| h01 | One right answer | Right c1 | Right c1 | Right c1 | Right c1 |
| h02 | One right answer | Wrong c1 | Right c2 | Right c2 | Right c2 |
| h03 | Wrong answer scored highest | Right c1 | Wrong c2 | Right c1 | Right c1 |
| h04 | One right answer | Wrong c1 | Wrong c1 | Right c2 | Right c2 |
| h05 | No right answer | Wrong c1 | Wrong c1 | Review no candidate passes the check | Review answer key has no right answer |
| h06 | Right words, wrong field | Wrong c1 | Wrong c1 | Review 3 different answers pass the check | Right c2 |
| h07 | Two right answers | Right c1 | Right c2 | Right c1 | Right c1 |
| h08 | Right answer, no quote | Right c1 | Wrong c2 | Review no candidate passes the check | Right c1 |
| h09 | Right answer, reworded | Right c1 | Right c1 | Right c2 | Right c1 |
| h10 | One right answer | Right c1 | Right c1 | Review 2 different answers pass the check | Right c1 |
| h11 | Wrong answer scored highest | Wrong c1 | Wrong c1 | Right c2 | Right c2 |
| h12 | No right answer | Wrong c1 | Wrong c1 | Review no candidate passes the check | Review answer key has no right answer |
What the sets show
Answering everything means answering wrong. First-shown and highest-score both answer all 24 sets. First-shown is wrong in 12 of them, highest-score in 14. Here the score does slightly worse than the display order. That's because I wrote fixtures where a confident wrong answer scores highest, and it's the point: a score is a ranking, not evidence.
Checking the source removes the wrong answers and leaves a queue. The source check answered 15 of 24 sets automatically and got none of those wrong. It sent the other 9 to a person. 4 of those had no right candidate at all, so review was the correct outcome. The rest had a right answer sitting in the pile that the rule couldn't safely pick: 5 of the 20 sets with one.
A check that passes is not a proof. In the generator set, one candidate swaps the component and the due day. Every word it reports is in its quote, so it passes. The check can see words, not which field they belong to. In the basement fan set, “replace the motor” is a substring of “replace the motor capacitor”, so a wrong action passes too. In both, the only thing that stopped a wrong answer was the rule's refusal to pick between two answers that both passed.
Deferral carries the load. In the tests, a variant that breaks ties by score instead of deferring starts shipping wrong answers. Remove the quote check as well and it becomes exactly highest-score.
Strict checks reject some right answers. A correct extraction that rewords “swap out” as “replace” fails the verbatim check, and so does a correct one that didn't quote the note. In these sets a literal candidate happened to be there as a fallback. In your product it may not be.
On the held-out half, the source check sent 5 sets to review (h05, h06, h08, h10, h12), one more than on the development half, because of the substring case. Both halves were predicted before running.
Check your reasoning
Two questions. Each answer explains itself, right or wrong.
What to do on a real system
- Report the any-correct rate and the selected-correct rate side by side, with the same denominator. The difference is your selection gap.
- Count wrong automatic answers over automatic answers, and coverage over all cases. Keep deferrals as a separate, unresolved count; never score them as correct.
- Treat a ranking score as a ranking. Before trusting it, check how often the top-scored candidate is wrong.
- Decide what evidence a candidate must carry before it can be chosen automatically, and check it against the source, not against the candidate's own claims.
- Defer when more than one different answer passes. Build the review queue before you build the automation.
- Keep correctness labels out of anything the selector can read, and freeze your test split before tuning the rule.
- Include a broken control, such as a selector that guesses on ties, to prove your tests can see a wrong answer at all.
What this lesson does not show
- No model. I wrote every candidate by hand to exercise one failure each, so the counts follow from that design. They aren't rates you should expect anywhere.
- No real confidence scores. The scores are invented numbers chosen to make some wrong answers rank first.
- No cost of review. A deferred set is counted as unresolved, not as solved, and the time it takes a person isn't modelled.
- No semantic matching. The check is verbatim on purpose, which is why it rejects a correct paraphrase. A looser check would accept more right answers and more wrong ones; measure both.
- No claim about any product, model or vendor.
The question to take back to your product
Generating more candidates is cheap. Choosing among them is the design problem. Before adding another sample, answer this: which evidence can this product actually obtain before it chooses? A quote from the source? A test that runs? A second system that agrees? A person? Whatever the answer is, it decides your coverage, your wrong answers and your review queue, far more than the size of the pile.