Back to blog
Sep 23, 2026
20 min read

When to Trust a Decision Model

A confidence number is only worth what it lets you trust. I put nine decision models, open and hosted, through one test: keep only the answers each was 90% sure of, and score them, first on 111 hard AI-written questions and then on 1,599 questions labelled by 100 people each. Some earned the trust. Here is which, and under what conditions.

Part two of a series. Part one: A Model That Refuses to Talk.

I asked a model a question you can check on the back of an envelope. It told me it was 100% sure of the wrong answer.

Here is the question. A yearly service credit of 1,098 is prorated by active days in a 366-day leap year. The service ran February 28 to March 2 inclusive. A suspension removes February 29. How many credits are earned? Try it before you scroll.

One question, five answers

A yearly service credit is shared out evenly across the days the service was active.

  • The annual amount is 1,098 credits for 2028, which has 366 days.
  • The service activated February 28 and cancelled March 2.
  • A suspension covering February 29 removes that day from active days, but March 1 remains active.

How many credits?

The choices the systems had

What each system answered, and how sure it was
  • Qwen 27B, written-out12 the trap
    100% sure
  • Qwen 27B, probabilities12 the trap
    82% sure
  • Kev-9b12 the trap
    59% sure
  • Jev, hosted9 right answer
    55% sure
  • Decider-2b9 right answer
    41% sure

The bar is how much of its probability each system put on the answer it gave. A text tag marks the right answer and the planted wrong one, so the marking does not depend on colour.

Show the data
System691215Answered
Jev (hosted)0.1300.5500.2200.1009
Qwen 27B, logprobs0.0320.0870.8220.06012
Qwen 27B, logprobs, options reversed0.2030.5520.0660.1799
Qwen 27B, logprobs, Yes/No as words0.0320.0870.8220.06012
Qwen 27B, written-out probabilities0.0000.0001.0000.00012
Kev-9b0.1300.1800.5900.10012
Kev-4b0.3400.2100.2300.2206
Kev-0.8b0.0300.0300.7500.19012
Decider-2b0.3200.4070.1560.1179
Eve RLCD 0.6B0.3000.2920.2130.1956
Laya 421M, generalist0.2090.2490.3600.18212
Laya 421M, benchmark-tuned specialist0.2090.2290.3130.24912
Right answeryes9
Planted wrong answeryes12
Show the question as the systems got it

Classify the earned service credit.

An annual service credit is prorated by active calendar days in a leap year, including both activation and cancellation dates. The annual amount is 1,098 credits for 2028, which has 366 days. The service activated February 28 and cancelled March 2. A suspension covering February 29 removes that day from active days, but March 1 remains active. Compute exactly before classifying; no rounding is needed because the daily amount is integral.

Apply all stated timing, unit, inclusion, correction, and threshold rules; select the exact band.

REAL. One question from the hard test set (hard-sol-b-temporal_numeric-03, MIT licensed), answered by 5 systems. It was chosen because it is short and you can check it by hand. On 15 of the 111 hard questions the written-out arm was wrong and at least 90% sure on exactly the planted wrong answer. This question and its answer key were written by another AI model (GPT-5.6 Sol), not by a person. You can check this one by hand.

Four calendar days, minus the suspended leap day, leaves three. The yearly credit works out to 3 credits a day. So the answer is 9. The trap is 12, which is what you get if you forget the suspension.

That item is real. Its answer key was written by an AI model, GPT-5.6 Sol, and so was every other key in Part one.

On 15 of the 111 hard questions, the model that wrote out its own probabilities was wrong at 90% or higher on exactly the planted trap. On this one, both systems that got it right were unsure, and the one with no hesitation was wrong.

So this post runs one test on nine systems, first on 111 hard questions with AI-written keys, then on 1,599 questions labelled by 100 people each: hosted Jev, a Qwen 27B I run myself read two ways, and five open-weight models built in Jev’s style, on one GPU. Post one said the plain 27B was calibrated at least as well as Jev. That held on two easy datasets and does not hold here.

The short version

Right on 111 hard items, AI keyWrong answers it was 90%+ sure ofTrusting only answers at 90%+Same gate on 1,599 items, human key
Jev (hosted)72%2 of 31keeps 45, 96% rightkeeps 730, 72% right
Qwen 27B, read through its probabilities71 to 78% by prompt layout3 to 5keeps 49 to 57, 90 to 94% rightkeeps 628, 76% right
Same Qwen, writing its own probabilities73%23 of 30keeps 85, 73% rightkeeps 1,426, 63% right
Kev-9b (small open model, trained on the human set’s source)57%22 of 48keeps 72, 69% rightkeeps 1,355, 69% right

Read the last two columns against the first. The full nine-system table on the human key is in Part one; what the checks changed is Part two.

Part one: the models

The one concept: a confidence gate

A gate is the plainest thing you can do with a probability: accept the answers above some bar automatically, send the rest to a person or a slower model. So the test of a probability is this: keep only the answers where the system said it was at least 90% sure. How many did you keep, and how many of those were right? A system whose wrong answers also sit at 90% has a gate that passes the bad through with the good. So the rule this post scores everything by: a system earns trust when the answers it keeps are much more accurate than its answers overall.

Where you set the bar depends on what a wrong answer costs you; auto-tagging a ticket and auto-issuing a refund do not deserve the same bar. The 90% throughout this post is a reporting convention, not a recommendation.

Drag the bar and watch who keeps what.

Only trust it when it says it is at least this sure
50%99%
How many of the 111 answers are kept
Jev, hosted45 of 111Qwen 27B, probabilities57 of 111Qwen 27B, written-out85 of 111
How accurate the kept answers are
Jev, hosted96% (n=45)Qwen 27B, probabilities93% (n=57)Qwen 27B, written-out73% (n=85)
Systems:
Show the data
SystemKeptRightAccuracy95% interval
Jev (hosted) (read at 90%)45 of 1114395.6%85.2% to 98.8%
Qwen 27B, logprobs57 of 1115393.0%83.3% to 97.2%
Qwen 27B, logprobs, options reversed51 of 1114690.2%79.0% to 95.7%
Qwen 27B, logprobs, Yes/No as words49 of 1114693.9%83.5% to 97.9%
Qwen 27B, written-out probabilities85 of 1116272.9%62.7% to 81.2%
Kev-9b72 of 1115069.4%58.0% to 78.9%
Kev-4b70 of 1113955.7%44.1% to 66.8%
Kev-0.8b42 of 1111842.9%29.1% to 57.8%
Decider-2b45 of 1112555.6%41.2% to 69.1%
Eve RLCD 0.6B11 of 111436.4%15.2% to 64.6%
Laya 421M, generalist2 of 11100.0%0.0% to 65.8%
Laya 421M, benchmark-tuned specialist0 of 1110n/an/a
REAL. The same 111 hard questions for every system. Labels were written by AI models, not people. Whiskers are 95% Wilson intervals worked out from the counts shown, so they widen as fewer answers survive. Jev is published as 20 confidence bins rather than per question, so its cut snaps down to the5% step at or below the bar you set.

The familiar calibration score, ECE, asks whether a system that says 80% is right about 80% of the time. It is easy to flatter: Eve scored a tidy 0.056 on my second test set while getting 47.9% right, below the 52% from always guessing the most common answer for each question type. The gate is harder to flatter, as long as you also report how many answers it kept.

Same model, two ways to read its confidence

This is the cleanest result I have. One Qwen 27B, same weights, same server, same 111 hard questions, two ways of asking how sure it was. Asked to write its probabilities out as JSON, it held most of its wrong answers at high confidence. Read from its own token probabilities, which I will call reading from its probabilities, it held almost none there. One readout gave me a gate I could use. The other did not. The prompt and output format differ between the two as well, so this compares two pipelines, not one switch.

How sure each system was when it was wrong
Qwen 27B, written-out73% right, 23 of 30 wrong at 90% or moreQwen 27B, probabilities78% right, 4 of 24 wrong at 90% or moreQwen 27B, probabilities, reversed71% right, 5 of 32 wrong at 90% or moreQwen 27B, probabilities, Yes/No78% right, 3 of 24 wrong at 90% or moreJev, hosted (binned)72% right, 2 of 31 wrong at 90% or moreKev-0.8b32% right, 24 of 75 wrong at 90% or moreKev-4b48% right, 31 of 58 wrong at 90% or moreKev-9b57% right, 22 of 48 wrong at 90% or moreDecider-2b45% right, 20 of 61 wrong at 90% or moreEve 0.6B34% right, 7 of 73 wrong at 90% or moreLaya, generalist34% right, 2 of 73 wrong at 90% or moreLaya, benchmark-tuned28% right, 0 of 80 wrong at 90% or more0%25%50%75%90%100%how sure it said it was, for the answers it got wrong
one wrong answera bin of wrong answers, height is how manythe 90% mark

Marks piled against the right edge are answers the system got wrong while saying it was nearly certain. Those are the ones that survive if you only trust it when it says it is at least 90% sure.

Show the data
SystemRight on the hard setWrong answersWrong and at least 90% sureShare of its wrong answers
Qwen 27B, written-out probabilities73.0%302376.7%
Qwen 27B, logprobs78.4%24416.7%
Qwen 27B, logprobs, options reversed71.2%32515.6%
Qwen 27B, logprobs, Yes/No as words78.4%24312.5%
Jev (hosted) (binned)72.1%3126.5%
Kev-0.8b32.4%752432.0%
Kev-4b47.7%583153.4%
Kev-9b56.8%482245.8%
Decider-2b45.0%612032.8%
Eve RLCD 0.6B34.2%7379.6%
Laya 421M, generalist34.2%7322.7%
Laya 421M, benchmark-tuned specialist27.9%8000.0%
REAL. Wrong answers only, from the same 111 hard questions. Each mark is one wrong answer, placed at how sure the system was of it. A mark at zero is an answer that failed to parse, not a measured zero. Jev is published as 20 confidence bins, so its row is drawn as binned blocks and labelled that way. Counts are printed on every row.

One caveat first. The probability I read is conditional on the option letters I captured, and on 21 of the 111 questions less than 90% of the model’s probability sat on any letter at all.

Option order matters to Qwen in a way accuracy hides. Listing the options in reverse changed its top answer on 36 of the 105 reorderable questions when it wrote its probabilities out, and on 26 of 105 read through them, while accuracy moved only 2 and 7 points, because flips went both ways. Jev is not deterministic, and reversing its options changed its answer (8 of 105) about as often as asking twice did (7 of 105), so I found no order effect in Jev beyond its own noise. One run per layout, and I do not know the reason.

This is not a discovery. Kadavath and colleagues in 2022 and Xiong and colleagues in 2023 point the same way, and Tian and colleagues in 2023, whom I cited in post one, found the opposite on TriviaQA, SciQ and TruthfulQA.

So the claim is local: on these items, with this model, on my server. It held again when 100 people wrote the key, by a rule I set before running it; see Scored by 100 people.

Five open models: what each earned

Post one promised this run: the open-weight models built in this style, on a single GH200 I run myself. Each of them is good at something I measured, and each has a limit I measured in the same session. Both belong in the same paragraph.

Decider-2b (Mark Marosi) is the fastest single answer here: 5 ms for one question, 166 ms for fifty. On the human-labelled set its gate is respectable, 593 answers kept at 90% and 75% of them right, though it trained on that set’s source. On the hard AI-keyed questions it is 45% right and holds 20 of its 61 wrong answers at 90%, so there it is a fast answer, not a gate.

Eve RLCD (Anthony Maio) is the model whose confidence you can act on. It is trained with a reward that pays for honest probabilities, a proper scoring rule, and it behaves that way: on the hard questions it holds 7 of 73 wrong answers at 90%, on human labels 12 of 673, and it sits closest of the nine to the human label distribution. The cost is coverage. It is right on only 34% of the hard questions and keeps 115 of 1,599 human-labelled answers at 90%. When it is sure, believe it; it is rarely sure.

Laya (Nandakishor, Convai Innovations) has the flattest latency line I measured, 119 ms for one question and 113 for fifty, and its author’s published figures reproduced on my machine. Its gate does little: on the hard questions it is 34% right and keeps only 2 answers at 90%; on human labels it keeps 890 at 66% right, five points above just answering everything. It reads the first 512 tokens by design.

Kev (Jared Palmer) speaks the hosted service’s API, so my harness ran against it unchanged, and its accuracy climbs with size: 32%, 48%, 57% on the hard questions at 0.8B, 4B and 9B, and 56%, 62%, 66% on human labels, where the 9B is the most accurate system in this post (it trained on that set’s source). What does not climb is its honesty about being wrong. The share of its wrong hard answers held at 90% is 24 of 75, 31 of 58, 22 of 48; at 9B nearly half of what it got wrong it was sure of, and on human labels its gate adds about 3 points over just answering, at every size. Bigger got righter.

JevBench (Florian Standhartinger) is the benchmark all of this runs on.

So the two failure modes are different, and a small model can have either. Kev and Decider on hard questions are often wrong and often sure. Eve and Laya are often wrong and almost never sure. “Small models are overconfident” is false for two of the five.

Some of these numbers may differ from what the authors would get: Kev ran on a newer torch than its author pins, and Decider’s accuracy run had its CUDA graphs off to fit beside my own servers.

Scored by 100 people, the gap held

Everything so far is agreement with an answer key that an AI model argued its way to. So before publishing I ran the whole table against people: ChaosNLI (Nie, Zhou and Bansal, 2020), 1,599 premise-and-hypothesis pairs, each labelled by 100 people, asking whether the hypothesis follows from the premise, contradicts it, or neither. Its authors chose items people disagreed on, so only 57 of the 1,599 have a 90% human majority. That makes it a hard set for any gate, which is the point. “Right” means matching the human majority.

I committed the rule before any system saw an item (commit 3dfff8c, 19:15 Central, September 21): the finding holds only if the whole 95% interval for the gate gap between the two readouts sits above 5 points. Then I ran all nine systems on all 1,599. Read the third column against the first: a gate that beats its own overall accuracy is doing something.

RightKept at 90%+ sureRight among keptWrong answers held at 90%+
Qwen 27B, writing its own probabilities63%1,426 of 1,59963%524 of 596
Same Qwen, read through its probabilities59%62876%149 of 651
Jev (hosted)62%73072%206 of 614
Kev-9b, trained on MNLI66%1,35569%417 of 545
Kev-4b, trained on MNLI62%1,31765%467 of 608
Kev-0.8b, trained on MNLI56%1,16960%473 of 701
Decider-2b, trained on MNLI62%59375%146 of 605
Eve RLCD58%11590%12 of 673
Laya, generalist61%89066%304 of 624

Kev and Decider list MNLI among their training sets, and these items come from MNLI’s development split, so their rows may reflect memory. Laya’s page does not say; Eve’s author says it never trained on MNLI.

The 90% gate, scored by people
Qwen 27B, written-out
kept 1,426
63% right
Qwen 27B, probabilities
kept 628
76% right
Jev, hosted
kept 730
72% right
Kev-9b (trained on MNLI)
kept 1,355
69% right
Kev-4b (trained on MNLI)
kept 1,317
65% right
Kev-0.8b (trained on MNLI)
kept 1,169
60% right
Decider-2b (trained on MNLI)
kept 593
75% right
Eve 0.6B
kept 115
90% right
Laya, generalist (training data unknown; reports XNLI scores)
kept 890
66% right
Show the data
SystemRight overallKept at 90%+Right among keptWrong answers held at 90%+Distance to the human distribution
Qwen 27B, written-out62.7%1,42663.3%524 of 5960.40
Qwen 27B, probabilities59.3%62876.3%149 of 6510.32
Jev, hosted61.6%73071.8%206 of 6140.32
Kev-9b (trained on MNLI)65.9%1,35569.2%417 of 5450.40
Kev-4b (trained on MNLI)62.0%1,31764.5%467 of 6080.42
Kev-0.8b (trained on MNLI)56.2%1,16959.5%473 of 7010.42
Decider-2b (trained on MNLI)62.2%59375.4%146 of 6050.31
Eve 0.6B57.9%11589.6%12 of 6730.30
Laya, generalist (training data unknown; reports XNLI scores)61.0%89065.8%304 of 6240.36
REAL. ChaosNLI, 1,599 items, 100 people labelled each one; right means matching the human majority. Only 57 items have a 90% human majority, so this is a hard set for any gate. Grey: share of the 1,599 answers the system held at 90% sure or higher. Purple: how many of those were right. Flags are training-data overlap as stated in each project's own documentation.

Asked to write down its confidence, Qwen kept 89% of its answers and gained nothing by it: the kept set was exactly as accurate as answering every question. The probability reading kept 40% and was 13 points better on what it kept. The paired interval for that gap is +10.2 to +15.8 points, so by the rule I set, the finding stands on human labels.

Then the check I had promised myself: options in reverse. The gap shrank to +3.5 to +8.8 points, inconclusive under my rule, mostly because the written-out version held 310 fewer answers at 90% and got better at what it kept. So the honest sentence is: reading the probabilities gave the better gate under both layouts, and the size of the advantage depends on how the options are listed.

Jev held a third of its wrong answers at 90% or more here, against 2 of 31 on the AI-keyed questions. On items where 100 people split, though, the median human share on those confident wrong answers was 38%, so “wrong” is a harder judgment than the table makes it look.

Against the full human distribution rather than the majority, Eve, Decider, the probability reading and Jev were closest, within 0.02 of each other, and the written-out version and the Kevs were furthest. Whether Qwen, Jev or Laya have seen MNLI, I cannot say.

If you gate on a written-out confidence today, this is the experiment to run this week: same model, both readouts, two hundred of your own items.

Latency: how many questions per call

Jev and Laya answer fifty questions about as fast as they answer one. Nothing else here does. My own inference servers were down for about 23 minutes so each local model could be measured alone on the GPU (Jev is hosted; the one-structured-request Qwen row was measured separately with my servers running). Median response time is in the chart, in milliseconds.

Typical time for one call (p50)
510501005001s3s151050questions in one calltime, log scale27B, one per question27B, one requestKev-9bKev-4bEve 0.6BJev, hostedDecider-2bKev-0.8bLaya 421M

The two thick lines with square markers are the same general purpose 27B model twice: once as one request per question, and once as a single request that returns all the answers. Everything else takes many questions in one call, so the line stays close to flat.

Show the data
Deployment1 question5 questions10 questions50 questions
Jev, hosted156 ms154 ms165 ms190 ms
Decider-2b5.0 ms25 ms47 ms166 ms
Kev-0.8b80 ms85 ms89 ms146 ms
Kev-4b89 ms95 ms102 ms297 ms
Kev-9b62 ms105 ms120 ms396 ms
Eve 0.6B140 ms127 ms144 ms196 ms
Laya 421M119 ms92 ms95 ms113 ms
the same 27B model, one request per question120 ms331 ms656 ms2.77 s
the same 27B model, one request that returns all the answers181 ms302 ms460 ms1.84 s
REAL. My own timings of 9 deployments at 1, 5, 10, 50 questions in one call, median (p50) and 90th percentile (p90). Jev was timed from Austin over the internet on a connection that stays open. Every other system was timed on the machine itself, so this is not like for like.

For a builder the useful reading is the slope. Everything but Jev and Laya grows from one question to fifty, from Eve’s 40% to Decider’s thirty-fold, and the general LLM most in absolute terms. If one input carries twenty questions, a flat system is a different component from one that grows with every question. Tails matter too: at fifty questions Decider’s 90th percentile opens to 749 ms. Medians are what people quote; tails are what page you at night.

These are timings of these deployments on a synthetic workload. Loopback against a hosted API across the internet is not like for like, and I cannot subtract my network time to recover anyone’s compute time. My general LLM ran at its production setting of four concurrent sequences. One oddity I have not explained: Kev-9b answered a single question faster than Kev-4b. Rerun that before leaning on it.

Part two: the checks

Everything above is about when a number can be trusted. Here is when mine could not.

Three checks that changed the numbers

Prompts. My wrapper for Eve dropped the fine print on its yes/no questions. Fixed and rerun, Eve’s accuracy landed within two questions on each benchmark tier and a point lower on the second set, but its confidence changed: it was never 90% sure on a hard question, against 11 times before, 7 of them wrong. I report the first run throughout. The option-order control moved Qwen’s probability reading by 7 points, as large as its apparent lead over Jev, so that lead is not reported as one. And on my second test set, asking yes or no questions as lettered options instead of the words yes and no cost 16 points on the 600 yes or no decisions, 61% against 77%.

The same model, three ways of laying out the question
The hard test set, 111 questions
60%65%70%75%80%Jev 72.1%options as written78.4%options reversed71.2%down 7.2 pointsYes and No as words78.4%no changethe axis starts at 60%, not 0, because this figure is about small moves

Same model, same questions, same way of reading how sure it is.

The second test set, 2,000 decisions
60%65%70%75%80%Jev 73.5%options as written64.7%options reversed70.6%up 5.9 pointsYes and No as words69.6%up 4.9 pointsthe axis starts at 60%, not 0, because this figure is about small moves

Scored as agreement with a teacher model, not against a human answer.

How often the top answer changed
Jev, hosted
7 of 105
8 of 105
Qwen 27B, written-out
0 of 105
36 of 105
Qwen 27B, probabilities
0 of 105
26 of 105

Bars show the share of the 105 hard questions whose options can be reordered where the top answer stayed the same; the count beside each bar is how many changed. Grey: the identical request sent a second time. Purple: the same request with the options in reverse. One run each.

Laying the options out in reverse cost 7.2 points on the hard questions. On the second test set the ordering flipped: options reversed came last on the hard questions and first here. On these prompts I would not read a gap smaller than that.

Show the data
Prompt layoutHard set, 111 questionsChangeSecond set, 2000 decisionsChange
options as written78.4%baseline64.7%baseline
options reversed71.2%down 7.2 points70.6%up 5.9 points
Yes and No as words78.4%no change69.6%up 4.9 points
Jev, its own request format72.1%reference73.5%reference
Always guessing the most common answern/a52.5%floor
Show the data
SystemIdentical request sent twiceOptions listed in reverse
Jev, hosted7 of 1058 of 105
Qwen 27B, written-out0 of 10536 of 105
Qwen 27B, probabilities0 of 10526 of 105
REAL. One model, one way of reading how sure it is, three ways of laying out the question, on both test sets: 111 hard questions and 2,000 decisions. The hollow ring is the layout as I first wrote it. Jev is drawn as a reference line, not as a competitor, because it answered in its own request format. The last panel asks a different question of three systems: not how many answers were right, but how many changed.

Instruments. The harness I started with opens a new connection for every request, which charged the hosted API a fresh handshake each time and made it look about 140 ms slower than it is; every timing here uses a connection that stays open. My first latency loop sent the same input on every call, which flatters anything with a cache: one model read 63 ms that way and 100 ms once I varied the input. Both errors produced the number I was expecting to see, which is why they lasted.

Statistics. Jev and the Qwen probability reading differ on 25 of 111 hard questions, 9 one way and 16 the other, p = 0.23. That is inconclusive, not a tie and not a win. And the hard answer keys were written by two models, 57 by GPT-5.6 Sol and 54 by Claude Opus 5, which I only knew after reading the whole provenance field rather than one row of it.

What the first measurement said, and what the check showed

Prompt layout. The first measurement said the 27B model, read through its own probabilities, got 78.4% of the hard questions right, ahead of the hosted model's 72.1%. In fact, that 78.4% came from one way of listing the options. With the same options in reverse order it got 71.2%. That is 7.2 points from layout alone, as large as the apparent lead.

Prompt layout. The first measurement said reading the model's probabilities won on the hard questions and then lost on the second test set, where it came last. In fact, most of that was the prompt. Yes or no questions had been written as lettered options. Across the three layouts I later tried, the same arm ran from 64.7% to 70.6% on that set, and the layout that did worst on the hard questions did best here.

Prompt layout. The first measurement said Eve's numbers are a lower bound, because the wrapper dropped the criteria text on its yes or no questions. In fact, the wrapper did drop that text, and fixing it did not raise Eve's score. Run again, it got 37 of the 111 hard questions right, against 38 the first time.

Timing. The first measurement said on long single questions the local GPU answered faster than the hosted service. In fact, the benchmark harness opened a new connection for every request, so each hosted call paid for a fresh handshake that the local calls never paid. Measured over a connection that stays open, the hosted model answers one question in a median of 156 ms.

Statistics. The first measurement said the hosted model and the 27B model tied on the hard questions. In fact, they answered 72.1% and 78.4% of the same 111 questions right, and a paired test on the questions where they disagreed could not separate them. Inconclusive is not a tie, and it is not a win either.

Statistics. The first measurement said small trained open models bought the format, not trustworthy confidence. In fact, Eve held 7 of its 73 wrong answers at 90% sure or higher and Laya 2 of 73, while Kev-9b held 22 of 48. Often wrong and confidently wrong are two different failures, and the small models here do not all make the same one.

Every number above is recomputed from the same data file as the other figures.

Show the data
First measurementWhat the check foundWhat the post says now
Prompt layout: the 27B model, read through its own probabilities, got 78.4% of the hard questions right, ahead of the hosted model's 72.1%that 78.4% came from one way of listing the options. With the same options in reverse order it got 71.2%. That is 7.2 points from layout alone, as large as the apparent lead.Layout moved this arm by 7.2 points on the hard set.
Prompt layout: reading the model's probabilities won on the hard questions and then lost on the second test set, where it came lastmost of that was the prompt. Yes or no questions had been written as lettered options. Across the three layouts I later tried, the same arm ran from 64.7% to 70.6% on that set, and the layout that did worst on the hard questions did best here.All three layouts are reported, on both test sets.
Prompt layout: Eve's numbers are a lower bound, because the wrapper dropped the criteria text on its yes or no questionsthe wrapper did drop that text, and fixing it did not raise Eve's score. Run again, it got 37 of the 111 hard questions right, against 38 the first time.Eve's first-run numbers stand, with the rerun reported beside them.
Timing: on long single questions the local GPU answered faster than the hosted servicethe benchmark harness opened a new connection for every request, so each hosted call paid for a fresh handshake that the local calls never paid. Measured over a connection that stays open, the hosted model answers one question in a median of 156 ms.Every timing in this post uses a connection that stays open. Jev reads 156 ms for one question.
Statistics: the hosted model and the 27B model tied on the hard questionsthey answered 72.1% and 78.4% of the same 111 questions right, and a paired test on the questions where they disagreed could not separate them. Inconclusive is not a tie, and it is not a win either.The comparison is reported as inconclusive on hard questions.
Statistics: small trained open models bought the format, not trustworthy confidenceEve held 7 of its 73 wrong answers at 90% sure or higher and Laya 2 of 73, while Kev-9b held 22 of 48. Often wrong and confidently wrong are two different failures, and the small models here do not all make the same one.The small models here fail in at least two different ways, and the post names both.
REAL. 6 first measurements, each with what the check showed. Two more rest on figures that are not in the data file and are left to the text.

Two other AI systems checked the work before publication: Codex, which recomputed every number from the raw files, and Grok, with live web search. They found all of the above. They were wrong too: two of three checkable claims from one of them failed when I checked, a temperature flag that does not exist at that commit and a floor figure from a different definition than mine. A hostile review is a list of leads, not a verdict.

Yes. An AI reviewed an AI’s evaluation of AI models answering questions written by AI. I am not going to defend that. It is why the nine-system table in Part one is scored by people, and why I dropped my plan to label the 111 hard questions myself: a hundred people per item beat one tired author.

The three kinds have one thing in common: every one of these errors made the story simpler than the data did, and not one made it more complicated. I do not think that is chance. It is what motivated reasoning looks like from the inside, where it does not feel like reasoning at all. It feels like being finished.

Of 127 numbers in my methods write-up recomputed independently, 2 were wrong, and both are fixed. The arithmetic held. The interpretations were where the errors lived.

The two review prompts, a numbers audit and a methods review run separately, are generalized and available on request.

Check me. The test items are public (JevBench, ChaosNLI). My scripts, per-item results, raw timings, the pre-registration with its commit timestamp, the three reviews and their prompts are available on request.

Credits and licenses

JevBench, by Florian Standhartinger, MIT. The LocalLLaMA/typed-decisions set, Apache 2.0, whose labels are a teacher model’s rather than a person’s. Kev, by Jared Palmer, Apache 2.0. Decider, by Mark Marosi, Apache 2.0. Eve RLCD, by Anthony Maio, MIT code and Apache 2.0 weights. Laya, by Nandakishor at Convai Innovations, Apache 2.0. Qwen3.8-27B, Apache 2.0. ChaosNLI, by Yixin Nie, Xiang Zhou and Mohit Bansal, CC BY-NC 4.0, the only human labels here. Jev, by TypeSafe, hosted, early access. Thank you.

Found a mistake? Tell me on LinkedIn or X. I will fix it here, dated, with credit to whoever caught it.

If you have human-labelled decisions from real work, not a benchmark, I would like to run this on them.

Let's Build AI That Works

Ready to implement these ideas in your organization?