Your Language Model Is Not-So-Secretly a System 1 Model

Ivan Yee Lee and Taylor Berg-Kirkpatrick · 2026-10-04

TypeSafe’s Jev offers an appealing promise: fast decisions with a probability for every answer option, without generating text. But how much of that capability is unique to Jev? Reading answer probabilities in a single forward pass is already a standard way to evaluate language models.

We tested off-the-shelf open-weight models in that same setting, without additional training. Seven of the 23 models we ranked matched Jev’s overall score across 19,505 decisions. They form a competitive group with different strengths and costs. Costs vary by provider and hardware; some configurations are competitive with Jev or cheaper.

Jev retains advantages in speed through hosted APIs, on some tasks and in the calibration of its delivered probabilities. On our own cards, with no network in the time, one matching model, gemma-4-26B-A4B-it, answered in 0.68× Jev's hosted time. But matching Jev's overall performance on these benchmarks does not require a new kind of model—or even additional training of an existing one.

X-axis:
View:
image/svg+xml1/4×1/2×1×2×4×Cost per decision, relative to Jev's $0.0323 per 1,000 (log scale)0.600.650.700.750.80Overall score (macro-F1, higher is better)no pricelower scores on this lineQwen3.5-9B Score 0.769 [0.756, 0.779]; Jev 0.787 Difference from Jev: -0.018 [-0.029, -0.008] We ran the BF16 version Price: 1.36× Jev's, at the cheapest of 6 OpenRouter listings on 2026-09-29 Darkbloom: FP4, $0.08 per million input tokens Agreement test: untested Log-probabilities: not recordedQwen3-Coder-Next-FP8 Score 0.765 [0.754, 0.774]; Jev 0.787 Difference from Jev: -0.023 [-0.031, -0.014] We ran the FP8 version Price: 2.00× Jev's, at the cheapest of 4 OpenRouter listings on 2026-09-29 Parasail: BF16, $0.12 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: not recordedMistral-Medium-3.5-128B Score 0.747 [0.733, 0.759]; Jev 0.787 Difference from Jev: -0.040 [-0.051, -0.029] We ran the FP8 version Price: 26.06× Jev's, at the cheapest of 3 OpenRouter listings on 2026-09-29 Mistral: precision not stated, $1.50 per million input tokens Agreement test: untested Log-probabilities: not recordedGemini 2.5 Flash-Lite Score 0.735 [0.721, 0.747]; Jev 0.787 Difference from Jev: -0.052 [-0.066, -0.040]; match rule lower bound: -0.068 Price: 2.25× Jev's (Google's list price)MiMo-V2.6-Distill-Qwen-9B Score 0.726 [0.710, 0.739]; Jev 0.787 Difference from Jev: -0.062 [-0.077, -0.048] We ran the BF16 version No price in this view: no OpenRouter listinggemma-4-31B Score 0.722 [0.707, 0.735]; Jev 0.787 Difference from Jev: -0.065 [-0.079, -0.052] We ran the BF16 version No price in this view: no OpenRouter listingQwen3.5-9B-Base Score 0.719 [0.703, 0.732]; Jev 0.787 Difference from Jev: -0.069 [-0.083, -0.056] We ran the BF16 version No price in this view: no OpenRouter listinggpt-oss-120b Score 0.702 [0.688, 0.714]; Jev 0.787 Difference from Jev: -0.085 [-0.099, -0.071] We ran the FP4 version Price: 0.54× Jev's, at the cheapest of 24 OpenRouter listings on 2026-09-29 CoreWeave: FP4, $0.03 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: yesgranite-4.2-8b Score 0.701 [0.686, 0.712]; Jev 0.787 Difference from Jev: -0.087 [-0.100, -0.074] We ran the BF16 version Price: 1.08× Jev's, at the cheapest of 2 OpenRouter listings on 2026-09-29 DeepInfra: BF16, $0.06 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: not recordedLLaDA-1.5 Score 0.698 [0.683, 0.710]; Jev 0.787 Difference from Jev: -0.089 [-0.104, -0.076] We ran the BF16 version No price in this view: no OpenRouter listingLLaDA-8B-Instruct Score 0.694 [0.679, 0.706]; Jev 0.787 Difference from Jev: -0.094 [-0.109, -0.080] We ran the BF16 version No price in this view: no OpenRouter listingNVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 Score 0.591 [0.577, 0.603]; Jev 0.787 Difference from Jev: -0.196 [-0.212, -0.181] We ran the BF16 version Price: 0.68× Jev's, at the cheapest of 5 OpenRouter listings on 2026-09-29 Darkbloom: INT4, $0.039 per million input tokens Agreement test: untested Log-probabilities: not recordedQwen3.5-0.8B-Base Score 0.445 [0.430, 0.457]; Jev 0.787 Difference from Jev: -0.342 [-0.359, -0.327] We ran the BF16 version No price in this view: no OpenRouter listingModernBERT-Large-Instruct Score 0.355 [0.342, 0.368]; Jev 0.787 Difference from Jev: -0.432 [-0.449, -0.414] We ran the BF16 version No price in this view: no OpenRouter listingQwen3.5-0.8B Score 0.341 [0.331, 0.350]; Jev 0.787 Difference from Jev: -0.446 [-0.460, -0.432] We ran the BF16 version No price in this view: no OpenRouter listing26×0.450.360.34Jevopen modelclosed modelno price in this viewmatches Jev within the 0.03 marginerror bar: 95% intervalbest score for the pricewithin 0.03 of Jev's score, at most 1.25× its cost (not a match test)gemma-4-31B-it Score 0.794 [0.781, 0.804]; Jev 0.787 Difference from Jev: +0.006 [-0.001, +0.013]; match rule lower bound: -0.003 We ran the BF16 version Price: 1.46× Jev's, at the cheapest of 15 OpenRouter listings on 2026-09-29 Reka: precision not stated, $0.08 per million input tokens Agreement test: untested Log-probabilities: yesMiMo-V2.6-Pro-MOPD Score 0.792 [0.780, 0.802]; Jev 0.787 Difference from Jev: +0.005 [-0.004, +0.013]; match rule lower bound: -0.005 We ran the FP8 version No price in this view: no OpenRouter listingQwen3.8-Flash-Next-FP8 Score 0.788 [0.775, 0.799]; Jev 0.787 Difference from Jev: +0.001 [-0.009, +0.010]; match rule lower bound: -0.010 We ran the FP8 version Price: 2.56× Jev's, at the only OpenRouter listing on 2026-09-29 Alibaba: precision not stated, $0.15 per million input tokens Agreement test: untested Log-probabilities: not recordedQwen3.8-27B Score 0.787 [0.774, 0.798]; Jev 0.787 Difference from Jev: -0.000 [-0.012, +0.011]; match rule lower bound: -0.013 We ran the BF16 version Price: 0.66× Jev's, at the cheapest of 16 OpenRouter listings on 2026-09-29 Wafer: precision not stated, $0.0307 per million input tokens Agreement test: passed; top answers matched our run's on 561 of 600 questions (93.5% [91.2%, 95.2%]) Log-probabilities: yesgemma-4-26B-A4B-it Score 0.784 [0.771, 0.794]; Jev 0.787 Difference from Jev: -0.004 [-0.012, +0.005]; match rule lower bound: -0.014 We ran the BF16 version Price: 0.77× Jev's, at the cheapest of 14 OpenRouter listings on 2026-09-29 Darkbloom: precision not stated, $0.042 per million input tokens Agreement test: failed; top answers matched our run's on 549 of 600 questions (91.5% [89.0%, 93.5%]) Log-probabilities: yesGPT-6 Luna Score 0.782 [0.769, 0.792]; Jev 0.787 Difference from Jev: -0.006 [-0.016, +0.004]; match rule lower bound: -0.017 Price: 2.06× Jev's (OpenAI's list price)MiMo-V2.6-Flash-RL Score 0.780 [0.768, 0.790]; Jev 0.787 Difference from Jev: -0.007 [-0.016, +0.001]; match rule lower bound: -0.017 We ran the FP8 version Price: 1.35× Jev's, at the cheapest of 5 OpenRouter listings on 2026-09-29 Relace: MXFP4, $0.08 per million input tokens Agreement test: untested Log-probabilities: noQwen3.6-35B-A3B-FP8 Score 0.779 [0.766, 0.790]; Jev 0.787 Difference from Jev: -0.009 [-0.018, +0.000]; match rule lower bound: -0.019 We ran the FP8 version Price: 0.87× Jev's, at the cheapest of 10 OpenRouter listings on 2026-09-29 Darkbloom: FP4, $0.05 per million input tokens Agreement test: failed; top answers matched our run's on 550 of 600 questions (91.7% [89.2%, 93.6%]) Log-probabilities: yesNex-N2.5-mini Score 0.763 [0.751, 0.774]; Jev 0.787 Difference from Jev: -0.024 [-0.032, -0.016] We ran the BF16 version Price: 0.43× Jev's, at the only OpenRouter listing on 2026-09-30 Nex AGI: BF16, $0.025 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: not recordedMistral-Nemo-Instruct-2407-FP8 Score 0.633 [0.614, 0.649]; Jev 0.787 Difference from Jev: -0.154 [-0.172, -0.138] We ran the FP8 version Price: 0.30× Jev's, at the cheapest of 6 OpenRouter listings on 2026-09-30 DekaLLM: FP8, $0.018 per million input tokens Agreement test: untested Log-probabilities: not recordedJev Score 0.787 [0.776, 0.797] Its error bar is the uncertainty of its own score, not the match threshold A match is neither closeness of the scores nor overlap with this 95% interval Cost $0.0323 per 1,000 decisions, TypeSafe's list priceMiMo-V2.6-Flash-RLQwen3.6-35B-A3B-FP8gemma-4-26B-A4B-itQwen3.8-27BJevGPT-6 LunaQwen3.8-Flash-Next-FP8MiMo-V2.6-Pro-MOPDgemma-4-31B-it
image/svg+xml1/4×1/2×1×2×4×Cost per decision, relative to Jev's$0.0323 per 1,000 (log scale)0.6500.6750.7000.7250.7500.7750.800Overall score (macro-F1, higher is better)no pricelower scores on this lineQwen3.5-9B Score 0.769 [0.756, 0.779]; Jev 0.787 Difference from Jev: -0.018 [-0.029, -0.008] We ran the BF16 version Price: 1.36× Jev's, at the cheapest of 6 OpenRouter listings on 2026-09-29 Darkbloom: FP4, $0.08 per million input tokens Agreement test: untested Log-probabilities: not recordedQwen3-Coder-Next-FP8 Score 0.765 [0.754, 0.774]; Jev 0.787 Difference from Jev: -0.023 [-0.031, -0.014] We ran the FP8 version Price: 2.00× Jev's, at the cheapest of 4 OpenRouter listings on 2026-09-29 Parasail: BF16, $0.12 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: not recordedMistral-Medium-3.5-128B Score 0.747 [0.733, 0.759]; Jev 0.787 Difference from Jev: -0.040 [-0.051, -0.029] We ran the FP8 version Price: 26.06× Jev's, at the cheapest of 3 OpenRouter listings on 2026-09-29 Mistral: precision not stated, $1.50 per million input tokens Agreement test: untested Log-probabilities: not recordedGemini 2.5 Flash-Lite Score 0.735 [0.721, 0.747]; Jev 0.787 Difference from Jev: -0.052 [-0.066, -0.040]; match rule lower bound: -0.068 Price: 2.25× Jev's (Google's list price)MiMo-V2.6-Distill-Qwen-9B Score 0.726 [0.710, 0.739]; Jev 0.787 Difference from Jev: -0.062 [-0.077, -0.048] We ran the BF16 version No price in this view: no OpenRouter listinggemma-4-31B Score 0.722 [0.707, 0.735]; Jev 0.787 Difference from Jev: -0.065 [-0.079, -0.052] We ran the BF16 version No price in this view: no OpenRouter listingQwen3.5-9B-Base Score 0.719 [0.703, 0.732]; Jev 0.787 Difference from Jev: -0.069 [-0.083, -0.056] We ran the BF16 version No price in this view: no OpenRouter listinggpt-oss-120b Score 0.702 [0.688, 0.714]; Jev 0.787 Difference from Jev: -0.085 [-0.099, -0.071] We ran the FP4 version Price: 0.54× Jev's, at the cheapest of 24 OpenRouter listings on 2026-09-29 CoreWeave: FP4, $0.03 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: yesgranite-4.2-8b Score 0.701 [0.686, 0.712]; Jev 0.787 Difference from Jev: -0.087 [-0.100, -0.074] We ran the BF16 version Price: 1.08× Jev's, at the cheapest of 2 OpenRouter listings on 2026-09-29 DeepInfra: BF16, $0.06 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: not recordedLLaDA-1.5 Score 0.698 [0.683, 0.710]; Jev 0.787 Difference from Jev: -0.089 [-0.104, -0.076] We ran the BF16 version No price in this view: no OpenRouter listingLLaDA-8B-Instruct Score 0.694 [0.679, 0.706]; Jev 0.787 Difference from Jev: -0.094 [-0.109, -0.080] We ran the BF16 version No price in this view: no OpenRouter listingNVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 Score 0.591 [0.577, 0.603]; Jev 0.787 Difference from Jev: -0.196 [-0.212, -0.181] We ran the BF16 version Price: 0.68× Jev's, at the cheapest of 5 OpenRouter listings on 2026-09-29 Darkbloom: INT4, $0.039 per million input tokens Agreement test: untested Log-probabilities: not recordedQwen3.5-0.8B-Base Score 0.445 [0.430, 0.457]; Jev 0.787 Difference from Jev: -0.342 [-0.359, -0.327] We ran the BF16 version No price in this view: no OpenRouter listingModernBERT-Large-Instruct Score 0.355 [0.342, 0.368]; Jev 0.787 Difference from Jev: -0.432 [-0.449, -0.414] We ran the BF16 version No price in this view: no OpenRouter listingQwen3.5-0.8B Score 0.341 [0.331, 0.350]; Jev 0.787 Difference from Jev: -0.446 [-0.460, -0.432] We ran the BF16 version No price in this view: no OpenRouter listing26×0.590.450.360.340.63Jevopen modelclosed modelno price in this viewmatches Jev within the 0.03 marginerror bar: 95% intervalbest score for the pricewithin 0.03 of Jev's score, at most 1.25× its cost(not a match test)gemma-4-31B-it Score 0.794 [0.781, 0.804]; Jev 0.787 Difference from Jev: +0.006 [-0.001, +0.013]; match rule lower bound: -0.003 We ran the BF16 version Price: 1.46× Jev's, at the cheapest of 15 OpenRouter listings on 2026-09-29 Reka: precision not stated, $0.08 per million input tokens Agreement test: untested Log-probabilities: yesMiMo-V2.6-Pro-MOPD Score 0.792 [0.780, 0.802]; Jev 0.787 Difference from Jev: +0.005 [-0.004, +0.013]; match rule lower bound: -0.005 We ran the FP8 version No price in this view: no OpenRouter listingQwen3.8-Flash-Next-FP8 Score 0.788 [0.775, 0.799]; Jev 0.787 Difference from Jev: +0.001 [-0.009, +0.010]; match rule lower bound: -0.010 We ran the FP8 version Price: 2.56× Jev's, at the only OpenRouter listing on 2026-09-29 Alibaba: precision not stated, $0.15 per million input tokens Agreement test: untested Log-probabilities: not recordedQwen3.8-27B Score 0.787 [0.774, 0.798]; Jev 0.787 Difference from Jev: -0.000 [-0.012, +0.011]; match rule lower bound: -0.013 We ran the BF16 version Price: 0.66× Jev's, at the cheapest of 16 OpenRouter listings on 2026-09-29 Wafer: precision not stated, $0.0307 per million input tokens Agreement test: passed; top answers matched our run's on 561 of 600 questions (93.5% [91.2%, 95.2%]) Log-probabilities: yesgemma-4-26B-A4B-it Score 0.784 [0.771, 0.794]; Jev 0.787 Difference from Jev: -0.004 [-0.012, +0.005]; match rule lower bound: -0.014 We ran the BF16 version Price: 0.77× Jev's, at the cheapest of 14 OpenRouter listings on 2026-09-29 Darkbloom: precision not stated, $0.042 per million input tokens Agreement test: failed; top answers matched our run's on 549 of 600 questions (91.5% [89.0%, 93.5%]) Log-probabilities: yesGPT-6 Luna Score 0.782 [0.769, 0.792]; Jev 0.787 Difference from Jev: -0.006 [-0.016, +0.004]; match rule lower bound: -0.017 Price: 2.06× Jev's (OpenAI's list price)MiMo-V2.6-Flash-RL Score 0.780 [0.768, 0.790]; Jev 0.787 Difference from Jev: -0.007 [-0.016, +0.001]; match rule lower bound: -0.017 We ran the FP8 version Price: 1.35× Jev's, at the cheapest of 5 OpenRouter listings on 2026-09-29 Relace: MXFP4, $0.08 per million input tokens Agreement test: untested Log-probabilities: noQwen3.6-35B-A3B-FP8 Score 0.779 [0.766, 0.790]; Jev 0.787 Difference from Jev: -0.009 [-0.018, +0.000]; match rule lower bound: -0.019 We ran the FP8 version Price: 0.87× Jev's, at the cheapest of 10 OpenRouter listings on 2026-09-29 Darkbloom: FP4, $0.05 per million input tokens Agreement test: failed; top answers matched our run's on 550 of 600 questions (91.7% [89.2%, 93.6%]) Log-probabilities: yesNex-N2.5-mini Score 0.763 [0.751, 0.774]; Jev 0.787 Difference from Jev: -0.024 [-0.032, -0.016] We ran the BF16 version Price: 0.43× Jev's, at the only OpenRouter listing on 2026-09-30 Nex AGI: BF16, $0.025 per million input tokens Agreement test: untested; the host states full precision Log-probabilities: not recordedMistral-Nemo-Instruct-2407-FP8 Score 0.633 [0.614, 0.649]; Jev 0.787 Difference from Jev: -0.154 [-0.172, -0.138] We ran the FP8 version Price: 0.30× Jev's, at the cheapest of 6 OpenRouter listings on 2026-09-30 DekaLLM: FP8, $0.018 per million input tokens Agreement test: untested Log-probabilities: not recordedJev Score 0.787 [0.776, 0.797] Its error bar is the uncertainty of its own score, not the match threshold A match is neither closeness of the scores nor overlap with this 95% interval Cost $0.0323 per 1,000 decisions, TypeSafe's list priceMiMo-V2.6-Flash-RLMiMo-V2.6-Pro-MOPDQwen3.8-Flash-Next-FP8gemma-4-31B-itgemma-4-26B-A4B-itQwen3.6-35B-A3B-FP8JevGPT-6 LunaQwen3.8-27B
Names:Hover over, focus or tap a point for its interval, its difference from Jev and its price; a click, tap or Enter pins it, Escape unpins it.
Every model in Figure 1, best first: score, difference from Jev and cost in each cost view
ModelKindScore [95% interval]Minus Jev [95% interval]OpenRouter cheapest (× Jev)OpenRouter median (× Jev)Rented A6000, estimated from our timing × rental price (× Jev)Rented L40S, estimated from our timing × rental price (× Jev)
same model and precision we ranany listingsame model and precision we ranany listing
gemma-4-31B-itopen0.794 [0.781, 0.804]+0.006 [−0.001, +0.013]2.56×1.46×
gemma-4-31B-it: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Reka, precision not stated, $0.08 per million input tokens, 2026-09-29, the cheapest of the 15 listings of any kind.
Why it is an estimate: Reka does not state its precision; we have not tested it.
2.56×2.56×3.25×4.90×
MiMo-V2.6-Pro-MOPDopen0.792 [0.780, 0.802]+0.005 [−0.004, +0.013]no priceno priceno priceno priceno priceno price
Qwen3.8-Flash-Next-FP8open0.788 [0.775, 0.799]+0.001 [−0.009, +0.010]no price2.56×
Qwen3.8-Flash-Next-FP8: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Alibaba, precision not stated, $0.15 per million input tokens, 2026-09-29, the only listing of any kind.
Why it is an estimate: Alibaba does not state its precision; we have not tested it.
no price2.56×
Qwen3.8-Flash-Next-FP8: the median of all OpenRouter listings
Listing used (OpenRouter): Alibaba, precision not stated, $0.15 per million input tokens, 2026-09-29, the only listing of any kind.
Why it is an estimate: Alibaba does not state its precision; we have not tested it.
no price16.27×
JevJev0.787 [0.776, 0.797]1.00×1.00×1.00×1.00×1.00×1.00×
Qwen3.8-27Bopen0.787 [0.774, 0.798]−0.000 [−0.012, +0.011]2.60×0.66×
Qwen3.8-27B: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Wafer, precision not stated, $0.0307 per million input tokens, 2026-09-29, the cheapest of the 16 listings of any kind.
Why it is an estimate: Wafer does not state its precision; it passed our agreement test.
2.60×3.67×
Qwen3.8-27B: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 16 listings of any kind, midway between the costs at Mancer 2 (FP8, $0.20 per million input tokens, 2026-09-29) and AkashML (FP8, $0.225 per million input tokens, 2026-09-29).
Why it is an estimate: Mancer 2 and AkashML list FP8, not the BF16 version we ran; we have not tested either.
2.08×3.38×
gemma-4-26B-A4B-itopen0.784 [0.771, 0.794]−0.004 [−0.012, +0.005]1.10×0.77×
gemma-4-26B-A4B-it: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, precision not stated, $0.042 per million input tokens, 2026-09-29, the cheapest of the 14 listings of any kind.
Why it is an estimate: Darkbloom does not state its precision; it failed our agreement test.
2.38×1.83×
gemma-4-26B-A4B-it: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 14 listings of any kind, midway between the costs at CoreWeave (BF16, $0.10 per million input tokens, 2026-09-29) and Cloudflare (precision not stated, $0.10 per million input tokens, 2026-09-29).
Why it is an estimate: Cloudflare does not state its precision; we have not tested it.
0.71×1.21×
GPT-6 Lunaclosed0.782 [0.769, 0.792]−0.006 [−0.016, +0.004]2.06×2.06×2.06×2.06×2.06×2.06×
MiMo-V2.6-Flash-RLopen0.780 [0.768, 0.790]−0.007 [−0.016, +0.001]1.96×1.35×
MiMo-V2.6-Flash-RL: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Relace, MXFP4, $0.08 per million input tokens, 2026-09-29, the cheapest of the 5 listings of any kind.
Why it is an estimate: Relace lists MXFP4, not the FP8 version we ran; we have not tested it.
2.31×2.31×no priceno price
Qwen3.6-35B-A3B-FP8open0.779 [0.766, 0.790]−0.009 [−0.018, +0.000]1.72×0.87×
Qwen3.6-35B-A3B-FP8: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the cheapest of the 10 listings of any kind.
Why it is an estimate: Darkbloom lists FP4, not the FP8 version we ran; it failed our agreement test.
2.57×
Qwen3.6-35B-A3B-FP8: the median OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Parasail, FP8, $0.15 per million input tokens, 2026-09-29, the middle of the 7 listings of the same model and precision we ran.
Why it is an estimate: Parasail lists FP8, the precision we ran, but we have not tested it.
2.15×
Qwen3.6-35B-A3B-FP8: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 10 listings of any kind, midway between the costs at Venice (FP8, $0.10 per million input tokens, 2026-09-29) and Parasail (FP8, $0.15 per million input tokens, 2026-09-29).
Why it is an estimate: Venice and Parasail list FP8, the precision we ran, but we have not tested either.
0.34×
Qwen3.6-35B-A3B-FP8: Rented A6000
The A6000 has no FP8 arithmetic; timed with FP8 weights and 16-bit arithmetic.
0.44×
Qwen3.5-9Bopen0.769 [0.756, 0.779]−0.018 [−0.029, −0.008]1.70×1.36×
Qwen3.5-9B: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, FP4, $0.08 per million input tokens, 2026-09-29, the cheapest of the 6 listings of any kind.
Why it is an estimate: Darkbloom lists FP4, not the BF16 version we ran; we have not tested it.
1.70×1.70×
Qwen3.5-9B: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 6 listings of any kind, midway between the costs at SiliconFlow (FP8, $0.10 per million input tokens, 2026-09-29) and Venice (FP8, $0.10 per million input tokens, 2026-09-29).
Why it is an estimate: SiliconFlow and Venice list FP8, not the BF16 version we ran; we have not tested either.
0.46×0.69×
Qwen3-Coder-Next-FP8open0.765 [0.754, 0.774]−0.023 [−0.031, −0.014]3.33×
Qwen3-Coder-Next-FP8: the cheapest OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Novita, FP8, $0.20 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Novita lists FP8, the precision we ran, but we have not tested it.
2.00×
Qwen3-Coder-Next-FP8: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Parasail, BF16, $0.12 per million input tokens, 2026-09-29, the cheapest of the 4 listings of any kind.
Why it is an estimate: Parasail lists BF16, not the FP8 version we ran; it states full precision.
3.33×
Qwen3-Coder-Next-FP8: the median OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Novita, FP8, $0.20 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Novita lists FP8, the precision we ran, but we have not tested it.
3.16×
Qwen3-Coder-Next-FP8: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 4 listings of any kind, midway between the costs at StreamLake (precision not stated, $0.18 per million input tokens, 2026-09-29) and Novita (FP8, $0.20 per million input tokens, 2026-09-29).
Why it is an estimate: StreamLake does not state its precision; we have not tested it. Novita lists FP8, the precision we ran, but we have not tested it.
0.61×
Qwen3-Coder-Next-FP8: Rented A6000
The A6000 has no FP8 arithmetic; timed with FP8 weights and 16-bit arithmetic.
1.03×
Nex-N2.5-miniopen0.763 [0.751, 0.774]−0.024 [−0.032, −0.016]0.43×0.43×0.43×0.43×0.48×0.97×
Mistral-Medium-3.5-128Bopen0.747 [0.733, 0.759]−0.040 [−0.051, −0.029]no price26.06×
Mistral-Medium-3.5-128B: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Mistral, precision not stated, $1.50 per million input tokens, 2026-09-29, the cheapest of the 3 listings of any kind.
Why it is an estimate: Mistral does not state its precision; we have not tested it.
no price26.06×
Mistral-Medium-3.5-128B: the median of all OpenRouter listings
Listing used (OpenRouter): Mistral, precision not stated, $1.50 per million input tokens, 2026-09-29, the middle of the 3 listings of any kind.
Why it is an estimate: Mistral does not state its precision; we have not tested it.
no price94.32×
Gemini 2.5 Flash-Liteclosed0.735 [0.721, 0.747]−0.052 [−0.066, −0.040]2.25×2.25×2.25×2.25×2.25×2.25×
MiMo-V2.6-Distill-Qwen-9Bopen0.726 [0.710, 0.739]−0.062 [−0.077, −0.048]no priceno priceno priceno price0.45×0.70×
gemma-4-31Bopen0.722 [0.707, 0.735]−0.065 [−0.079, −0.052]no priceno priceno priceno price3.21×5.27×
Qwen3.5-9B-Baseopen0.719 [0.703, 0.732]−0.069 [−0.083, −0.056]no priceno priceno priceno price0.54×0.81×
gpt-oss-120bopen0.702 [0.688, 0.714]−0.085 [−0.099, −0.071]0.54×0.54×1.35×2.17×
gpt-oss-120b: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 24 listings of any kind, midway between the costs at Parasail (FP4, $0.10 per million input tokens, 2026-09-29) and SambaNova (precision not stated, $0.14 per million input tokens, 2026-09-29).
Why it is an estimate: SambaNova does not state its precision; we have not tested it.
0.80×1.35×
granite-4.2-8bopen0.701 [0.686, 0.712]−0.087 [−0.100, −0.074]1.08×1.08×1.44×1.44×0.41×0.66×
LLaDA-1.5open0.698 [0.683, 0.710]−0.089 [−0.104, −0.076]no priceno priceno priceno price2.06×3.67×
LLaDA-8B-Instructopen0.694 [0.679, 0.706]−0.094 [−0.109, −0.080]no priceno priceno priceno price2.06×3.69×
Mistral-Nemo-Instruct-2407-FP8open0.633 [0.614, 0.649]−0.154 [−0.172, −0.138]0.30×
Mistral-Nemo-Instruct-2407-FP8: the cheapest OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): DekaLLM, FP8, $0.018 per million input tokens, 2026-09-30, the cheapest of the 4 listings of the same model and precision we ran.
Why it is an estimate: DekaLLM lists FP8, the precision we ran, but we have not tested it.
0.30×
Mistral-Nemo-Instruct-2407-FP8: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): DekaLLM, FP8, $0.018 per million input tokens, 2026-09-30, the cheapest of the 6 listings of any kind.
Why it is an estimate: DekaLLM lists FP8, the precision we ran, but we have not tested it.
0.41×
Mistral-Nemo-Instruct-2407-FP8: the median OpenRouter listing of the same model and precision we ran
Listings used (OpenRouter): the median of the 4 listings of the same model and precision we ran, midway between the costs at DeepInfra (FP8, $0.019 per million input tokens, 2026-09-30) and Parasail (FP8, $0.03 per million input tokens, 2026-09-30).
Why it is an estimate: DeepInfra and Parasail list FP8, the precision we ran, but we have not tested either.
0.55×
Mistral-Nemo-Instruct-2407-FP8: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 6 listings of any kind, midway between the costs at Parasail (FP8, $0.03 per million input tokens, 2026-09-30) and Io Net (FP16, $0.0356 per million input tokens, 2026-09-30).
Why it is an estimate: Parasail lists FP8, the precision we ran, but we have not tested it. Io Net lists FP16, not the FP8 version we ran; it states full precision.
no price0.58×
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16open0.591 [0.577, 0.603]−0.196 [−0.212, −0.181]1.04×0.68×
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, INT4, $0.039 per million input tokens, 2026-09-29, the cheapest of the 5 listings of any kind.
Why it is an estimate: Darkbloom lists INT4, not the BF16 version we ran; we have not tested it.
1.12×1.04×0.51×1.04×
Qwen3.5-0.8B-Baseopen0.445 [0.430, 0.457]−0.342 [−0.359, −0.327]no priceno priceno priceno price0.08×0.14×
ModernBERT-Large-Instructopen0.355 [0.342, 0.368]−0.432 [−0.449, −0.414]no priceno priceno priceno price0.21×0.29×
Qwen3.5-0.8Bopen0.341 [0.331, 0.350]−0.446 [−0.460, −0.432]no priceno priceno priceno price0.07×0.13×

† estimated; click for how we priced it.
Jev and the closed models keep their list prices in every view; hover over “no price” for the reason.

Figure 1. At the cheapest OpenRouter listing, Qwen3.8-27B (0.66×), gemma-4-26B-A4B-it (0.77×) and Qwen3.6-35B-A3B-FP8 (0.87×) cost no more than Jev. Overall score against cost per decision as a multiple of Jev's (log scale); up and to the left is better. Black-ringed open models match Jev (rule in The test); hollow ones have no price; hovering says why. Hosted prices may be for a different or undisclosed precision.

TypeSafe has made a useful idea popular: a language model used as a fast decision function, which it calls a System 1 model, after Daniel Kahneman's name for fast, automatic judgement. It sells one, Jev. The idea fits the many small decisions inside automated workflows: the most probable option is the decision, and the probabilities let a workflow act on confident answers and pass the rest to a person.

That raises an obvious question: can an ordinary language model already do this, if it is read the right way? Comparing the probabilities a language model gives its candidate answers is how GPT-2 and GPT-3 were evaluated on multiple-choice tasks, how PET and Zhao et al. built few-shot classifiers, and how benchmarks such as MMLU are scored.

TypeSafe presents Jev as a new model class, with a new architecture and training for calibrated probabilities. It also says Jev “can’t hallucinate” and is 40 to 200 times faster. Ordinary models can already return probabilities over supplied options without generating text. And TypeSafe’s speed comparison is against models generating an answer token by token, not against a model reading the option probabilities in one forward pass, as we did. We tested how close off-the-shelf models get to Jev without a new architecture or additional training.

What we found

  1. Seven of the 23 open models we ranked match Jev at picking the right option on 19,505 decisions of three kinds: Classification, Agent monitoring and Triage. Jev gets 0.787 by our measure, macro-F1, which counts rare answers as much as common ones.
  2. What a matching model costs depends on where it runs. At the cheapest hosted listings, three of the seven cost less than Jev, and on a rented A6000 two do; at the median hosted listing none does.
  3. Jev does better at Triage (screening prompts and matching records); the open models do better at Agent monitoring (judging what an AI agent is doing).
  4. Through hosted APIs, Jev answers faster: the five models with valid timings took 2.8 to 8.6 times as long per decision. On our own cards, with no network in the time, one matching model was faster than Jev's hosted calls.
  5. As delivered, Jev's probabilities are better calibrated than six of the seven. After temperature scaling, three of the seven are better calibrated than Jev however its zero probabilities are handled; for the other four it depends on that handling.
  6. A closed frontier model, GPT-6 Luna, read from the letter it writes, also matches Jev, at about twice Jev's cost at OpenAI's standard price.
  7. On JevBench, a public benchmark TypeSafe did not make, gemma-4-31B-it beats Jev; the other four matching models we ran there show no measurable difference.
  8. Quantization hurts our decisions less than it hurts standard benchmarks. Four-bit checkpoints keep 98.9% to 100.3% of the overall score in about a third of the memory, with no measurable difference from the benchmarks except for one; every 2-bit and ~1-bit file loses more on the benchmarks than on our decisions.

What a System 1 model does

A decision is one question about one piece of text with a closed set of options, such as which action a support agent takes next.

A language model's output at any point is a probability for every possible next token. End the prompt where the answer begins, with the options labelled, and the probabilities of those labels give the decision: one forward pass, nothing written (Figure 2). In practice an open chat model receives, through its chat template with thinking off, an instruction to reply with only the label, then the text, the question and the labelled options.

Jev receives the same text, question and options through TypeSafe’s API and returns a probability for each option, rounded to 0.01. Both sides get each option by its name alone.

The prompt, as gemma-4-31B-it receives it through its chat template
<bos><|turn>system
You answer one question about the state below. Read the state, then reply with only the label of your answer.<turn|>
<|turn>user
State:
Ford: Monthly Sales Drop, Company Looks To New Vehicles Cruising along the ever-stretching road of decline. Auto giant Ford Motor (nyse: F - news - people ) reported vehicle sales in October that fell 5 from a year ago.

Question: Which option best describes the text?
Options:
A) This example news text is about business news
B) This example news text is about science and technology
C) This example news text is about sports
D) This example news text is about world news
Answer with the label of one option.<turn|>
<|turn>model
<|channel>thought
<channel|>
image/svg+xml 0.0 0.2 0.4 0.6 0.8 1.0 Probability of each option's label as the next token, gemma-4-31B-it A) about business news (correct) B) about science and technology C) about sports D) about world news 0.95 0.02 0.02 0.01
Figure 2. One decision, read in one pass. Left: a Classification question from the benchmark as gemma-4-31B-it receives it. Right: the probability the model gives each option's label as the next token, after the temperature described under calibration. The most probable label, A, is the decision, and it is correct.

The test

The benchmark has three parts:

All three are decisions of the kind TypeSafe markets Jev for. TypeSafe's own headline figures come from workflow tests; Agent monitoring is the part where the open models beat Jev. Each dataset is scored by macro-F1, and the overall score is the mean of the three parts. It weights them equally although they hold 13,053, 5,800 and 652 decisions, which favours Jev, whose strongest part is the smallest.

Every score comes with a 95% bootstrap interval from 10,000 resamples of the test questions, each conversation or agent run kept whole; a difference from Jev uses the same resamples for both models. “Match” means we can rule out the open model being more than 3 points worse than Jev; “beat” means we can rule out its being worse at all. An open model beats Jev when the interval of its difference lies wholly above zero. It matches Jev when the low end of the 97.5% interval of that difference is above −0.03: the test rules out a deficit larger than 0.03. On a single part, a model is better or worse than Jev only when the 95% interval of the difference excludes zero; otherwise there is no measurable difference.

The models we scored on every decision fall into four groups:

Models scored on all 19,505 decisionsNumberIn Figure 1
Ranked: Jev and 23 open models, each with the standard prompt for its kind24yes
Closed models2yes
Quantized checkpoints of ranked models, counted toward their model7no
Prompt variants of ranked models3no
All3626

All seven models that match Jev are read as Figure 2 shows; a few others need a small variation, listed in the methods.

Methods: how the other models are read

Base models, trained only to continue text, get the question as plain text after worked examples. Masked models (one encoder, ModernBERT-Large-Instruct, and two masked diffusion models, LLaDA-8B-Instruct and LLaDA-1.5) are read at a mask token placed where the answer goes. The two closed models' APIs do not return a probability for every option, so we read the letter they write. Every open model read from its probabilities gets one fitted temperature (see calibration).

How Jev was called. We called Jev through TypeSafe's Python SDK (typesafe-sdk 0.7.2): 33,482 calls for the benchmark on 2026-09-28 and 2026-09-29, and 231 for JevBench on 2026-09-30. Every call answered as jev-1.13.0, and every call that records the model it asked for asked for jev-latest.

Seven open models match Jev's overall score

Jev scores 0.787 [0.776, 0.797]. Table 1 lists every open model that matches it, with the quantized checkpoints of it that also match. The best, gemma-4-31B-it, scores 0.794, 0.006 above Jev, a difference too small to measure: its interval, [−0.001, +0.013], includes zero. The low ends that decide a match run from −0.019 to −0.003. Pooled over all decisions instead of averaged over the three parts (point estimates; no intervals), six of the seven are more accurate than Jev (0.740 to 0.753 against 0.735), while pooled macro-F1, dominated by the Classification datasets' many classes, puts Jev ahead of all seven (0.674 against 0.634 to 0.671).

Table 1. Every open model that matches Jev's overall score, with the quantized checkpoints of it that also match. Overall score and difference from Jev, each with its 95% interval; the low end of the 97.5% interval that decides a match; and our exactness check. The two NVFP4 checkpoints that did not pass were each read once; their intervals leave out their run-to-run variation.
ModelVersionOverall scoreMinus JevLow end (97.5%)Exactness check
Jevas sold0.787 [0.776, 0.797]
gemma-4-31B-itas released0.794 [0.781, 0.804]+0.006 [−0.001, +0.013]−0.003pass
MiMo-V2.6-Pro-MOPDas released0.792 [0.780, 0.802]+0.005 [−0.004, +0.013]−0.005did not pass
Qwen3.8-Flash-Nextthe lab's FP8 release0.788 [0.775, 0.799]+0.001 [−0.009, +0.010]−0.010did not pass
Qwen3.8-27Bas released0.787 [0.774, 0.798]−0.000 [−0.012, +0.011]−0.013pass
GGUF file, BF160.787 [0.773, 0.797]−0.001 [−0.012, +0.010]−0.014pass
GGUF file, 4-bit (UD-Q4_K_XL)0.784 [0.770, 0.795]−0.004 [−0.015, +0.008]−0.017pass
the lab's FP8 release0.783 [0.769, 0.794]−0.004 [−0.017, +0.007]−0.019pass
NVIDIA's 4-bit (NVFP4) checkpoint0.779 [0.765, 0.790]−0.008 [−0.022, +0.004]−0.024pass
gemma-4-26B-A4B-itas released0.784 [0.771, 0.794]−0.004 [−0.012, +0.005]−0.014pass
Red Hat's 4-bit (NVFP4) checkpoint0.782 [0.771, 0.792]−0.005 [−0.013, +0.003]−0.014did not pass
Qwen3.6-35B-A3BNVIDIA's 4-bit (NVFP4) checkpoint0.781 [0.769, 0.792]−0.006 [−0.016, +0.003]−0.017did not pass
the lab's FP8 release0.779 [0.766, 0.790]−0.009 [−0.018, +0.000]−0.019pass
MiMo-V2.6-Flash-RLas released0.780 [0.768, 0.790]−0.007 [−0.016, +0.001]−0.017did not pass

Three things qualify the count:

The exactness check's readings

The stored scoring values of the four that passed differ from the check's fresh readings by a median of 0.12 to 0.39 nats, with the same top answer on every question checked. Those of Qwen3.8-27B's NVFP4 checkpoint differ by a median of 0.47 nats, with the same top answer on 11 of 12 questions, and those of Qwen3.8-27B's ~1-bit file (UD-IQ1_S), read one sequence per call, by a median of 0.00 nats.

Jev is better at Triage, the open models at Agent monitoring

The overall match averages over a split (Figure 3). On Agent monitoring, six of the seven beat Jev, by 0.025 to 0.052, and MiMo-V2.6-Flash-RL shows no measurable difference. Most of the lead is in which action an ABCD agent takes next (macro-F1 0.275 for Jev, 0.304 to 0.361 for those six) and whether a web agent looped (0.729 against 0.776 to 0.848).

On Triage, Jev is ahead: six of the seven fall below it, with differences of −0.047 to −0.024, and MiMo-V2.6-Pro-MOPD shows no measurable difference. Two of our scoring choices here matter. In the guardrail data we replaced 25 of the 202 published labels; with the published labels, gemma-4-31B-it's difference narrows from −0.025 to −0.015 and Qwen3.8-27B's from −0.047 to −0.032. For the beer pairs, taking the most probable of the three outcomes instead of our linking rule moves the differences to +0.007 and −0.020. TypeSafe's own example links a pair when the expected level of its three-level score rounds to “same product”; under that rule Jev's beer score falls from 0.926 to 0.840 and its overall score to 0.773; with the open models' probabilities as read, all seven still match and three beat Jev; with their fitted temperatures, six match, Qwen3.6-35B-A3B falling short. A separate search over prompts and quantized checkpoints, each chosen on a dev set, found no model that matches Jev on every part: all 14 of its finalists fall short on Triage, eight only there.

Our two Triage scoring choices

We replaced a published guardrail label only where two AI readers both contradicted it. We count a beer pair as linked when the probability of linking is at least that of leaving the records unlinked.

On Classification, four show no measurable difference from Jev and three are below it (MiMo-V2.6-Pro-MOPD, gemma-4-26B-A4B-it and Qwen3.6-35B-A3B), with differences of −0.015 to −0.007. These datasets are public, so some of their texts may have been in the open models' training data, which would favour them.

image/svg+xml −0.02 0.00 0.02 gemma-4-31B-it MiMo-V2.6-Pro-MOPD Qwen3.8-Flash-Next-FP8 Qwen3.8-27B gemma-4-26B-A4B-it Qwen3.6-35B-A3B-FP8 MiMo-V2.6-Flash-RL Classification 0.000 0.025 0.050 Agent monitoring −0.05 0.00 Triage Open model's score minus Jev's, with its 95% interval (right of zero: the open model is better)
Figure 3. The open models beat Jev on Agent monitoring and fall below it on Triage. Each matching model's part score minus Jev's, with its 95% interval: blue where the open model is better, orange where Jev is, grey where there is no measurable difference.

What it costs depends on where you run it

Jev has one price: TypeSafe's list price of $0.042 per million input tokens, output free, in force during our calls (2026-09-28 to 2026-09-30), which makes $0.0323 per 1,000 decisions here. Figure 1 prices an open model four ways: at the cheapest or the median OpenRouter listing, counting every listing of the model, whatever precision it states; or on a rented RTX A6000 or L40S card, at the median public hourly price on 2026-09-30 ($0.53 and $1.33) over the decisions per hour we measured. Table 2 also counts only the listings of the same model and precision we ran. Closed models keep their list prices.

Which listings count, and our host agreement test

A listing of the same model and precision we ran states full precision or the precision of the file we ran; for a quantized checkpoint, the checkpoint's own precision, or the nearest stated precision above it. We match precision by bit count, so a 4-bit listing can price a 4-bit file in another format: Darkbloom's fp4 listing prices our 4-bit GGUF file of Qwen3.8-27B (UD-Q4_K_XL). Any listing counts every listing of the model, whatever precision it states, if any. The listings are OpenRouter's of 2026-09-29 14:40Z, except for Mistral-Nemo-Instruct-2407-FP8 and Nex-N2.5-mini, which that snapshot lacks; theirs are of 2026-09-30 17:59Z.

Each listing is labelled with the result of our host agreement test, which is separate from the exactness check: on 600 questions sent through OpenRouter, the host's top answers had to agree with our full-precision run's about as often as our own 4-bit checkpoint's did (the top of the 95% interval of the host's agreement had to reach the bottom of the checkpoint's), and its accuracy could not fall below the low end of our full-precision run's interval. We set these two conditions on 2026-09-30, after we had seen Darkbloom's result. The test does not decide which listings count; it is reported beside each one.

Listings count even when the host does not return the log-probabilities that the one-pass reading needs. At the cheapest listing of the same model and precision, three matching models' prices come from such listings: gemma-4-31B-it at Crusoe, Qwen3.8-27B at DeepInfra and MiMo-V2.6-Flash-RL at DeepInfra.

Table 2. Cost per decision of every model in Table 1, as a multiple of Jev's. OpenRouter: the cheapest or median listing of the same model and precision we ran ("same") or of any ("any", as Figure 1 plots), from 2026-09-29 14:40Z (2026-09-28 23:16Z for the text's day-before Qwen3.8-27B price). Rented cards: estimated from our timing and the rental prices of 2026-09-30, card kept busy. "No price": no listing or timing qualifies.
ModelOpenRouter cheapestOpenRouter medianRented A6000, estimated from our timing × rental priceRented L40S, estimated from our timing × rental priceCheapest listing taken: same; any
sameanysameany
Jev1.00×1.00×1.00×1.00×1.00×1.00×TypeSafe's list price
gemma-4-31B-it 2.56×1.46×
gemma-4-31B-it: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Reka, precision not stated, $0.08 per million input tokens, 2026-09-29, the cheapest of the 15 listings of any kind.
Why it is an estimate: Reka does not state its precision; we have not tested it.
2.56×2.56×3.25×4.90×Crusoe, BF16; Reka, precision not stated
MiMo-V2.6-Pro-MOPD no priceno priceno priceno priceno priceno priceno OpenRouter listing
Qwen3.8-Flash-Next, the lab's FP8 release no price2.56×
Qwen3.8-Flash-Next, the lab's FP8 release: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Alibaba, precision not stated, $0.15 per million input tokens, 2026-09-29, the only listing of any kind.
Why it is an estimate: Alibaba does not state its precision; we have not tested it.
no price2.56×
Qwen3.8-Flash-Next, the lab's FP8 release: the median of all OpenRouter listings
Listing used (OpenRouter): Alibaba, precision not stated, $0.15 per million input tokens, 2026-09-29, the only listing of any kind.
Why it is an estimate: Alibaba does not state its precision; we have not tested it.
no price16.27×none; Alibaba, precision not stated
Qwen3.8-27B 2.60×0.66×
Qwen3.8-27B: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Wafer, precision not stated, $0.0307 per million input tokens, 2026-09-29, the cheapest of the 16 listings of any kind.
Why it is an estimate: Wafer does not state its precision; it passed our agreement test.
2.60×3.67×
Qwen3.8-27B: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 16 listings of any kind, midway between the costs at Mancer 2 (FP8, $0.20 per million input tokens, 2026-09-29) and AkashML (FP8, $0.225 per million input tokens, 2026-09-29).
Why it is an estimate: Mancer 2 and AkashML list FP8, not the BF16 version we ran; we have not tested either.
2.08×3.38×DeepInfra, BF16; Wafer, precision not stated
Qwen3.8-27B, GGUF file, BF16 2.60×0.66×
Qwen3.8-27B, GGUF file, BF16: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Wafer, precision not stated, $0.0307 per million input tokens, 2026-09-29, the cheapest of the 16 listings of any kind.
Why it is an estimate: Wafer does not state its precision; it passed our agreement test.
2.60×3.67×
Qwen3.8-27B, GGUF file, BF16: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 16 listings of any kind, midway between the costs at Mancer 2 (FP8, $0.20 per million input tokens, 2026-09-29) and AkashML (FP8, $0.225 per million input tokens, 2026-09-29).
Why it is an estimate: Mancer 2 and AkashML list FP8, not the BF16 version we ran; we have not tested either.
10.61×15.98×DeepInfra, BF16; Wafer, precision not stated
Qwen3.8-27B, GGUF file, 4-bit (UD-Q4_K_XL) 0.92×
Qwen3.8-27B, GGUF file, 4-bit (UD-Q4_K_XL): the cheapest OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Darkbloom lists FP4, not the GGUF file (UD-Q4_K_XL) we ran; it failed our agreement test.
The same bit width in another format: FP4 prices our GGUF file (UD-Q4_K_XL).
0.66×
Qwen3.8-27B, GGUF file, 4-bit (UD-Q4_K_XL): the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Wafer, precision not stated, $0.0307 per million input tokens, 2026-09-29, the cheapest of the 16 listings of any kind.
Why it is an estimate: Wafer does not state its precision; it passed our agreement test.
0.92×
Qwen3.8-27B, GGUF file, 4-bit (UD-Q4_K_XL): the median OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Darkbloom lists FP4, not the GGUF file (UD-Q4_K_XL) we ran; it failed our agreement test.
The same bit width in another format: FP4 prices our GGUF file (UD-Q4_K_XL).
3.67×
Qwen3.8-27B, GGUF file, 4-bit (UD-Q4_K_XL): the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 16 listings of any kind, midway between the costs at Mancer 2 (FP8, $0.20 per million input tokens, 2026-09-29) and AkashML (FP8, $0.225 per million input tokens, 2026-09-29).
Why it is an estimate: Mancer 2 and AkashML list FP8, not the GGUF file (UD-Q4_K_XL) we ran; we have not tested either.
6.21×8.68×Darkbloom, FP4; Wafer, precision not stated
Qwen3.8-27B, the lab's FP8 release 1.59×0.66×
Qwen3.8-27B, the lab's FP8 release: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Wafer, precision not stated, $0.0307 per million input tokens, 2026-09-29, the cheapest of the 16 listings of any kind.
Why it is an estimate: Wafer does not state its precision; it passed our agreement test.
4.14×
Qwen3.8-27B, the lab's FP8 release: the median OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Chutes, FP8, $0.24 per million input tokens, 2026-09-29, the middle of the 7 listings of the same model and precision we ran.
Why it is an estimate: Chutes lists FP8, the precision we ran, but we have not tested it.
3.67×
Qwen3.8-27B, the lab's FP8 release: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 16 listings of any kind, midway between the costs at Mancer 2 (FP8, $0.20 per million input tokens, 2026-09-29) and AkashML (FP8, $0.225 per million input tokens, 2026-09-29).
Why it is an estimate: Mancer 2 and AkashML list FP8, the precision we ran, but we have not tested either.
2.11×
Qwen3.8-27B, the lab's FP8 release: Rented A6000
The A6000 has no FP8 arithmetic; timed with FP8 weights and 16-bit arithmetic.
3.14×Ionstream, FP8; Wafer, precision not stated
Qwen3.8-27B, NVIDIA's 4-bit (NVFP4) checkpoint 0.92×
Qwen3.8-27B, NVIDIA's 4-bit (NVFP4) checkpoint: the cheapest OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Darkbloom lists FP4, the precision we ran, but it failed our agreement test.
0.66×
Qwen3.8-27B, NVIDIA's 4-bit (NVFP4) checkpoint: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Wafer, precision not stated, $0.0307 per million input tokens, 2026-09-29, the cheapest of the 16 listings of any kind.
Why it is an estimate: Wafer does not state its precision; it passed our agreement test.
0.92×
Qwen3.8-27B, NVIDIA's 4-bit (NVFP4) checkpoint: the median OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Darkbloom lists FP4, the precision we ran, but it failed our agreement test.
3.67×
Qwen3.8-27B, NVIDIA's 4-bit (NVFP4) checkpoint: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 16 listings of any kind, midway between the costs at Mancer 2 (FP8, $0.20 per million input tokens, 2026-09-29) and AkashML (FP8, $0.225 per million input tokens, 2026-09-29).
Why it is an estimate: Mancer 2 and AkashML list FP8, not the FP4 version we ran; we have not tested either.
1.82×2.38×Darkbloom, FP4; Wafer, precision not stated
gemma-4-26B-A4B-it 1.10×0.77×
gemma-4-26B-A4B-it: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, precision not stated, $0.042 per million input tokens, 2026-09-29, the cheapest of the 14 listings of any kind.
Why it is an estimate: Darkbloom does not state its precision; it failed our agreement test.
2.38×1.83×
gemma-4-26B-A4B-it: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 14 listings of any kind, midway between the costs at CoreWeave (BF16, $0.10 per million input tokens, 2026-09-29) and Cloudflare (precision not stated, $0.10 per million input tokens, 2026-09-29).
Why it is an estimate: Cloudflare does not state its precision; we have not tested it.
0.71×1.21×DekaLLM, BF16; Darkbloom, precision not stated
gemma-4-26B-A4B-it, Red Hat's 4-bit (NVFP4) checkpoint 1.28×
gemma-4-26B-A4B-it, Red Hat's 4-bit (NVFP4) checkpoint: the cheapest OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): DeepInfra, FP8, $0.07 per million input tokens, 2026-09-29, the cheapest of the 2 listings of the same model and precision we ran.
Why it is an estimate: DeepInfra lists FP8, not the FP4 version we ran; we have not tested it.
0.77×
gemma-4-26B-A4B-it, Red Hat's 4-bit (NVFP4) checkpoint: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, precision not stated, $0.042 per million input tokens, 2026-09-29, the cheapest of the 14 listings of any kind.
Why it is an estimate: Darkbloom does not state its precision; it failed our agreement test.
1.92×
gemma-4-26B-A4B-it, Red Hat's 4-bit (NVFP4) checkpoint: the median OpenRouter listing of the same model and precision we ran
Listings used (OpenRouter): the median of the 2 listings of the same model and precision we ran, midway between the costs at DeepInfra (FP8, $0.07 per million input tokens, 2026-09-29) and SiliconFlow (FP8, $0.14 per million input tokens, 2026-09-29).
Why it is an estimate: DeepInfra and SiliconFlow list FP8, not the FP4 version we ran; we have not tested either.
1.83×
gemma-4-26B-A4B-it, Red Hat's 4-bit (NVFP4) checkpoint: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 14 listings of any kind, midway between the costs at CoreWeave (BF16, $0.10 per million input tokens, 2026-09-29) and Cloudflare (precision not stated, $0.10 per million input tokens, 2026-09-29).
Why it is an estimate: CoreWeave lists BF16, not the FP4 version we ran; it states full precision. Cloudflare does not state its precision; we have not tested it.
0.48×0.62×DeepInfra, FP8; Darkbloom, precision not stated
Qwen3.6-35B-A3B, NVIDIA's 4-bit (NVFP4) checkpoint 0.87×
Qwen3.6-35B-A3B, NVIDIA's 4-bit (NVFP4) checkpoint: the cheapest OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Darkbloom lists FP4, the precision we ran, but it failed our agreement test.
0.87×
Qwen3.6-35B-A3B, NVIDIA's 4-bit (NVFP4) checkpoint: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the cheapest of the 10 listings of any kind.
Why it is an estimate: Darkbloom lists FP4, the precision we ran, but it failed our agreement test.
0.87×
Qwen3.6-35B-A3B, NVIDIA's 4-bit (NVFP4) checkpoint: the median OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the only listing of the same model and precision we ran.
Why it is an estimate: Darkbloom lists FP4, the precision we ran, but it failed our agreement test.
2.15×
Qwen3.6-35B-A3B, NVIDIA's 4-bit (NVFP4) checkpoint: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 10 listings of any kind, midway between the costs at Venice (FP8, $0.10 per million input tokens, 2026-09-29) and Parasail (FP8, $0.15 per million input tokens, 2026-09-29).
Why it is an estimate: Venice and Parasail list FP8, not the FP4 version we ran; we have not tested either.
0.31×0.40×Darkbloom, FP4; Darkbloom, FP4
Qwen3.6-35B-A3B, the lab's FP8 release 1.72×0.87×
Qwen3.6-35B-A3B, the lab's FP8 release: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Darkbloom, FP4, $0.05 per million input tokens, 2026-09-29, the cheapest of the 10 listings of any kind.
Why it is an estimate: Darkbloom lists FP4, not the FP8 version we ran; it failed our agreement test.
2.57×
Qwen3.6-35B-A3B, the lab's FP8 release: the median OpenRouter listing of the same model and precision we ran
Listing used (OpenRouter): Parasail, FP8, $0.15 per million input tokens, 2026-09-29, the middle of the 7 listings of the same model and precision we ran.
Why it is an estimate: Parasail lists FP8, the precision we ran, but we have not tested it.
2.15×
Qwen3.6-35B-A3B, the lab's FP8 release: the median of all OpenRouter listings
Listings used (OpenRouter): the median of the 10 listings of any kind, midway between the costs at Venice (FP8, $0.10 per million input tokens, 2026-09-29) and Parasail (FP8, $0.15 per million input tokens, 2026-09-29).
Why it is an estimate: Venice and Parasail list FP8, the precision we ran, but we have not tested either.
0.34×
Qwen3.6-35B-A3B, the lab's FP8 release: Rented A6000
The A6000 has no FP8 arithmetic; timed with FP8 weights and 16-bit arithmetic.
0.44×AkashML, FP8; Darkbloom, FP4
MiMo-V2.6-Flash-RL 1.96×1.35×
MiMo-V2.6-Flash-RL: the cheapest OpenRouter listing of any kind
Listing used (OpenRouter): Relace, MXFP4, $0.08 per million input tokens, 2026-09-29, the cheapest of the 5 listings of any kind.
Why it is an estimate: Relace lists MXFP4, not the FP8 version we ran; we have not tested it.
2.31×2.31×no priceno priceDeepInfra, FP8; Relace, MXFP4

† estimated; click for how we priced it.

The listing decides the OpenRouter cost. Qwen3.8-27B has 16 listings. The cheapest, Wafer at 0.66× Jev's cost, does not state its precision but passed our agreement test; the day before (Table 2 gives both times), its price made 1.20×, and billing each question's whole prompt, rather than a document asked several questions once, makes 0.80×. The cheapest listing of the same model and precision, DeepInfra's BF16 at $0.15 per million input tokens, makes it 2.60×; the median of all 16 listings makes it 3.67×. At the cheapest listing of the same model and precision, no matching model as ranked costs less than Jev, and three quantized checkpoints of two models do: Qwen3.8-27B's 4-bit GGUF file (UD-Q4_K_XL) and its NVFP4 checkpoint, both at 0.92×, and Qwen3.6-35B-A3B's 4-bit NVFP4 checkpoint (one reading; varies run to run) at 0.87×, all priced at Darkbloom's fp4 listing, from a host that failed our agreement test. At the median listing, every matching model as ranked that is listed on OpenRouter costs more than Jev, whichever listings count.

Renting the card yourself changes the order. On rented A6000 cards, gemma-4-26B-A4B-it costs 0.71× Jev's cost and Qwen3.6-35B-A3B's FP8 release 0.34×; both are mixture-of-experts models, which use only part of their weights for each token. The dense Qwen3.8-27B costs 2.08×. On rented L40S cards, Qwen3.6-35B-A3B's FP8 release stays below Jev (0.44×), and so does gemma-4-26B-A4B-it's 4-bit checkpoint (0.62×; one reading, varies run to run); gemma-4-26B-A4B-it as released costs 1.21× and Qwen3.8-Flash-Next's FP8 release 16.27× on four cards. These costs assume a card kept fully busy, which favours the open models, while the prompts we timed were longer than the benchmark's and read in full, which overstates their cost.

Jev is faster through hosted APIs

TypeSafe's “40x-200x faster” compares Jev with language models that write their answers as text; we timed the models we scored in two ways (Figure 4): through each model's hosted API, as Jev is called, with the network in the time, and on our own L40S cards, one call at a time, with no network. Each case's time is set against Jev's stored time on the same case.

Through a hosted API, every model that passed our checks took longer than Jev. Qwen3.8-27B served by Wafer took 559.9 ms per decision against Jev's 119.8 ms on the same 145 cases. Taken case by case, that is 5.12× [4.87, 5.23] Jev's time. The other models that passed took from 2.82× (Gemini 2.5 Flash-Lite, Google's API) to 8.61× (gemma-4-26B-A4B-it, DekaLLM, BF16) Jev's time (Table 3).

On our own cards the matching models come closer to Jev, but the comparison favours them: their times leave out the network, which Jev's include. With one call in flight, the five matching models we timed there took from 0.68× Jev's hosted time per decision (gemma-4-26B-A4B-it, less than Jev) to 3.86× (Qwen3.8-Flash-Next-FP8), and Qwen3.8-27B 1.63×; their four 4-bit and FP8 checkpoints took 0.55× to 1.85×. We did not time the two MiMo models, the largest of the seven.

Two limits remain. We made no new calls to Jev, so its times are those of its stored calls on the same cases, made from 2026-09-28 to 2026-09-29, days before our passes: the two sides were never timed in the same session, and Jev's own median ranged 92 to 206 ms over three other sessions, a spread the intervals leave out. And Classification's intervals have zero or almost zero width: in the hosted sample, 20 of its 21 datasets contribute one case each, and the bootstrap resamples cases within each dataset, so those datasets never vary; this narrows the overall intervals too.

Hollow points failed a check fixed before the runs: too few valid cases, a pass stopped early, or too little agreement with our scored answers; the note below Table 3 gives each model's cause and counts.

Two scatter plots, one above the other, of overall score against time per decision, on a log scale. Top: models on our own L40S cards, one call at a time, with no network. Bottom: models called through their hosted APIs. Jev, a star, is placed by its stored hosted calls in both.
Figure 4. Through hosted APIs, every model that passed our checks was slower than Jev; on our cards, without the network, some matching setups were faster than Jev's hosted calls. Overall score against time per decision (log scale): our L40S cards above, hosted APIs below. Hollow points failed a check; dotted ones are checks only (Table 3). The band: Jev's median time per call in three other sessions (92–206 ms), not an interval. As SVG; data as CSV and JSON.
Table 3. Time per decision through each model's hosted API, and its ratio to Jev's stored time on the same cases. Medians over the 145 cases (fewer where a case had no valid time), in milliseconds, with 95% intervals; each ratio is taken case by case, overall and on each part. Classification's intervals have zero or almost zero width (see the text). A point that failed a check is drawn hollow in Figure 4, and its numbers are not results.
ModelHostms per decision× Jev× Jev, Classification× Jev, Agent monitoring× Jev, TriageChecks
JevTypeSafe, stored calls120 [116, 123]
Gemini 2.5 Flash-LiteGoogle's API325 [320, 330]2.82 [2.74, 2.98]2.78 [2.78, 2.78]3.13 [2.75, 3.30]2.76 [2.67, 3.00]passed
Qwen3.6-35B-A3B-FP8AkashML, FP8383 [355, 412]3.61 [3.43, 3.99]2.59 [2.59, 2.59]4.11 [3.49, 5.32]3.99 [3.57, 5.84]passed
Qwen3.8-27BWafer560 [551, 570]5.12 [4.87, 5.23]4.17 [4.17, 4.17]6.47 [6.05, 7.83]5.12 [4.77, 5.54]passed
GPT-6 LunaOpenAI's API839 [812, 866]7.28 [6.97, 7.56]6.97 [6.97, 6.97]7.36 [6.56, 8.10]7.47 [6.75, 8.20]passed
gemma-4-26B-A4B-itDekaLLM, BF16945 [867, 989]8.61 [7.99, 9.49]6.40 [6.40, 6.40]8.91 [7.81, 10.48]11.84 [9.92, 14.04]passed
Qwen3.8-27BIonstream, FP8400 [392, 417]3.65 [3.50, 4.34]3.15 [3.15, 3.15]10.62 [4.79, 12.75]4.66 [3.65, 5.86]a second host, as a check
GPT-6 LunaOpenAI's API792 [712, 827]16 requests in flight, as a check
Gemini 2.5 Flash-LiteGoogle's API306 [298, 316]16 requests in flight, as a check
Qwen3.5-9BDeepInfra, BF16650 [630, 679]5.86 [5.56, 6.28]4.29 [4.29, 4.34]6.24 [5.41, 6.61]7.28 [6.02, 8.65]failed: too few valid cases
Nex-N2.5-miniNex AGI, BF16670 [631, 742]5.88 [5.64, 6.62]4.10 [4.10, 4.10]6.41 [5.62, 8.02]7.26 [6.33, 8.96]failed: stopped early
Mistral-Nemo-Instruct-2407-FP8DekaLLM, FP8814 [768, 895]7.84 [6.85, 8.70]5.22 [5.22, 5.22]9.83 [7.64, 11.13]8.64 [7.05, 10.54]failed: agreement too low
gemma-4-31B-itCrusoe, BF16834 [718, 1,211]8.67 [7.70, 10.69]6.05 [6.05, 6.05]8.97 [8.01, 14.60]11.00 [5.78, 17.97]failed: too few valid cases
How we timed the models, and the causes and counts behind the hollow points

Through a hosted API. Every model got the same 145 cases (168 graded decisions), sampled across the three parts and every Classification dataset, in two passes at different times of day on 2026-10-03. The cases went out one at a time, interleaved across the models in a fixed random order, from one client with one kept-open connection per model; open models went through OpenRouter, pinned to one host, and the closed models through their labs' APIs. A case's questions are sent together, and its time runs from the first request to the last answer, including any retries after a rate limit or a server error and the waits before them. A case counts only if every request returned an answer from the pinned host. A model's point needs valid times for 95% of the cases in each pass and agreement with our scored answers on 90% of the graded decisions. Its time per decision is a median over cases, weighted as the overall score weights them, with a 95% bootstrap interval; its ratio to Jev divides each case's time, the median over the passes, by Jev's stored time on the same case, and takes the same weighted median. The band in Figure 4 spans Jev's median time per call in three other sessions, one call at a time, on other cases.

On our own cards. Every model ran a fixed sample of cases (680 graded decisions) three times on L40S cards reserved for it, one case per call with one call in flight, with the cache of earlier prompts emptied before each case. Its point needs agreement with the scored run on 95% of the graded decisions.

No price enters these times, so the dates of the price lists used elsewhere in this post do not apply to them.

The hollow points. Through hosted APIs, Crusoe and DeepInfra rate-limited gemma-4-31B-it and Qwen3.5-9B, leaving too few valid cases: against the 138 valid cases of 145 required in each pass, gemma-4-31B-it had 63 and 101, and Qwen3.5-9B 137 in its second pass, one short. Nex-N2.5-mini's only host kept returning empty answers, so its second pass was stopped early, with 190 of its 370 requests unsent. Through DekaLLM, Mistral-Nemo-Instruct-2407-FP8 agreed too rarely with our scored answers: on 292 of 328 graded decisions, against 296 required. On our cards, gpt-oss-120b and Mistral-Nemo-Instruct-2407-FP8 agreed too rarely with their scored runs despite identical prompts: gpt-oss-120b on 633 and Mistral-Nemo-Instruct-2407-FP8 on 641 of 680, against 646 required. gpt-oss-120b's forward pass gives different log-probabilities for the same input, and its own three passes disagree with each other too; Mistral-Nemo-Instruct-2407-FP8's timing run used a separately compiled copy of the model and one case per call, and the resulting rounding differences reversed decisions that were near ties.

Calibration: Jev leads six of seven as delivered; after temperature scaling, three of seven beat it under every treatment

Calibrated probabilities are what TypeSafe's training method, RLCD, is for, so calibration is the most direct evidence we have about it (Table 4). We report expected calibration error (ECE, 15 bins: when it says 80%, is it right 80% of the time?) and log-loss (minus the log of the probability on the right answer, averaged; near-zero probabilities cost heavily); lower is better for both.

The open models' probabilities need temperature scaling, fitted per model on a calibration split of 30% of the benchmark's questions. A temperature never changes which option ranks first, so Figure 1's scores and matches are the same with or without it. Calibration is measured only on the other 70%: 13,673 questions, leaving out those whose labels an AI wrote.

Table 4. Calibration over all three parts, as delivered and with each model's fitted temperature (T). Expected calibration error (ECE) with its difference from Jev's as delivered, and log-loss with probabilities below 0.5% raised to 0.5%. Lower is better. As delivered: Jev's probabilities as returned, and each open model's at temperature one. Jev's fitted row uses the 0.5% floor with the temperature applied to its probabilities as returned; the note after the text lists the others.
ModelECE as deliveredMinus JevTECE with TMinus JevLog-loss as deliveredLog-loss with T
Jev0.1241.010.1240.9050.903
gemma-4-31B-it0.243+0.1195.850.055−0.0691.2961.072
MiMo-V2.6-Pro-MOPD0.157+0.0322.040.024−0.1001.0100.937
Qwen3.8-Flash-Next-FP80.143+0.0181.770.018−0.1070.9800.921
Qwen3.8-27B0.103−0.0221.570.037−0.0880.9230.910
gemma-4-26B-A4B-it0.252+0.1285.600.043−0.0821.3441.093
Qwen3.6-35B-A3B-FP80.147+0.0231.800.033−0.0911.1161.069
MiMo-V2.6-Flash-RL0.159+0.0341.980.028−0.0961.0470.986

As delivered, Jev's calibration error is 0.124. Only Qwen3.8-27B's is lower, at 0.103 (−0.022 [−0.027, −0.016]); the other six are higher, by 0.018 to 0.128, because their raw probabilities are too confident (fitted temperatures 1.77 to 5.85). Jev's log-loss as delivered, 0.905, is lower than all seven's point estimates.

After temperature scaling, the seven open models' calibration errors are 0.018 to 0.055. Jev's depends on a choice its API forces on us: it rounds every probability to 0.01, with no more precise output, so for 1,013 of the 19,505 decisions it gave the correct option exactly zero, and fitting a temperature by log-loss needs zeros raised to a floor. Depending on the floor and how the temperature is applied (each combination a treatment; the note below lists them), Jev's fitted calibration error runs from 0.040 to 0.124. MiMo-V2.6-Pro-MOPD, Qwen3.8-Flash-Next and MiMo-V2.6-Flash-RL, the three that failed our exactness check, two of them through run-to-run nondeterminism, are better calibrated than Jev under every treatment; for the other four, the result depends on the treatment. Log-loss after fitting depends on the treatment as well.

Jev's floors and treatments

We fit Jev's temperature on the calibration split only, with probabilities below the floor raised to it, and apply it either to the probabilities as returned, where a zero stays zero, or to the floored probabilities rescaled to sum to one, over floors of 0.5%, 0.1% and 0.01%. The lowest fitted calibration error comes with the 0.01% floor, rescaled; the highest with the 0.5% floor, as returned: Table 4's row, whose temperature, 1.0058, changes almost nothing. The full table will be in the preprint.

Two caveats. The temperatures come from other questions of the same benchmarks, which favours the open models in the fitted comparison, not the as-delivered one. And the log-loss raises every model's probabilities below 0.5% to 0.5%, which caps what a confident error costs; with Jev's zeros, that cap helps Jev.

Four-bit quantized checkpoints keep almost all of the overall score

Each of the four 4-bit quantized checkpoints we scored kept 98.9% to 100.3% of its own model's overall score, and Qwen's FP8 release of Qwen3.8-27B kept 99.4% (Figure 5). The parts are less even (Table 5). gemma-4-26B-A4B-it's 4-bit checkpoint scores 0.646 [0.630, 0.659] on Agent monitoring (one reading; varies run to run) against its full model's 0.678 [0.662, 0.690], a loss that its overall 99.8% does not show because it gains on the other two parts. Three of the four 4-bit checkpoints lose measurably on one part, and Qwen3.8-27B's NVFP4 and FP8 checkpoints lose measurably overall. Top answers change on 2.2% to 7.8% of decisions, against 0.9% between two engines running the same BF16 weights; the two highest of these, 7.8% and 6.1%, are the two mixture-of-experts NVFP4 checkpoints, whose shares include their own run-to-run changes.

Table 5. Share of each model's score kept by its quantized checkpoints, per part and overall, with paired 95% intervals on the overall score's 10,000 resamples, and the share of decisions whose top answer changed. Bold: measurably below the full model. Last row: the same BF16 weights on another engine (llama.cpp instead of vLLM), the noise floor. NVFP4 rows of gemma-4-26B-A4B-it and Qwen3.6-35B-A3B: one reading; varies run to run.
ModelCheckpointClassificationAgent monitoringTriageOverallAnswers changed
Qwen3.8-27Bthe lab's FP8 release100.2% [100.0%, 100.5%]100.1% [99.4%, 100.8%]98.2% [96.6%, 99.4%]99.4% [98.8%, 100.0%]2.2%
NVIDIA's 4-bit (NVFP4) checkpoint100.0% [99.5%, 100.4%]99.1% [98.1%, 100.0%]97.9% [95.2%, 100.4%]98.9% [97.9%, 99.9%]4.6%
GGUF file, 4-bit (UD-Q4_K_XL)100.3% [99.9%, 100.8%]99.0% [98.1%, 99.8%]99.3% [97.6%, 100.9%]99.6% [98.8%, 100.2%]3.7%
GGUF file, ~1-bit (UD-IQ1_S)78.9% [77.8%, 80.0%]67.8% [65.3%, 69.9%]80.7% [75.5%, 85.8%]76.3% [74.2%, 78.3%]40.5%
gemma-4-26B-A4B-itRed Hat's 4-bit (NVFP4) checkpoint100.7% [100.1%, 101.2%]95.3% [93.9%, 96.7%]102.5% [100.6%, 104.7%]99.8% [99.0%, 100.7%]7.8%
Qwen3.6-35B-A3BNVIDIA's 4-bit (NVFP4) checkpoint99.3% [98.7%, 99.9%]100.0% [99.0%, 100.9%]101.5% [99.4%, 103.7%]100.3% [99.4%, 101.2%]6.1%
Qwen3.8-27BGGUF file, BF16, on llama.cpp100.1% [99.9%, 100.3%]99.9% [99.7%, 100.1%]99.7% [99.1%, 100.0%]99.9% [99.6%, 100.1%]0.9%

The publishers' own results for the same files are mixed. NVIDIA's cards for its two 4-bit Qwen files report 98.0% to 101.0% kept across their benchmarks, the same range as ours. Red Hat's card for its 4-bit gemma-4-26B-A4B-it shows larger losses on tasks answered by generating text: it keeps 81.0% to 100.0% of the full model's score with thinking and 82.8% to 103.3% without, the lowest on agentic tool use and hard math. These comparisons are suggestive, not controlled: the publishers score other tasks by generating answers, while we read one pass, and for Qwen3.6-35B-A3B our baseline is its FP8 release where NVIDIA's is BF16.

Below four bits, four of the five files lose measurably: the files labelled 2-bit keep 95.4% (Qwen3.8-27B, 2.16 bits per weight measured), 98.1% (Qwen3.6-35B-A3B, 2.48) and 100.4% (gemma-4-26B-A4B-it, 3.14, no measurable loss), and those labelled ~1-bit keep 76.3% (Qwen3.8-27B, 1.84) and 95.3% (Qwen3.6-35B-A3B, 2.12) of their own model's overall score, with the largest losses on Agent monitoring. Qwen3.8-27B's ~1-bit file (UD-IQ1_S) is read one sequence per call, a reading that passes the exactness check; an earlier batched reading of the same file kept 78.0%, but its values depended on which prompts shared a batch, so it failed the exactness check and is not used.

On standard benchmarks, the 8-bit and 4-bit checkpoints are not separated from our decisions: the share of its benchmark score that each keeps does not differ measurably from the share of its decision score, with one exception. gemma-4-26B-A4B-it's 4-bit checkpoint keeps 95.0% [92.9%, 97.3%] of its benchmark score, against 99.8% of its decision score (one reading; varies run to run), with its largest losses on competition math and code; under the rule we registered in advance, a level separates only if two of its three models do, so the 4-bit level does not. Every 2-bit and ~1-bit file loses measurably on the benchmarks, and each loses less on our decisions, from Qwen3.8-27B's ~1-bit file, which keeps 29.7% of its benchmark score and 76.3% of its decision score, to gemma-4-26B-A4B-it's 2-bit file, at 93.7% and 100.4%. No checkpoint loses measurably more on our decisions than on the benchmarks.

The benchmarks form six families, weighted equally: knowledge (MMLU-Pro), reasoning (GPQA Diamond), math (AIME, MATH-500 and GSM8K Platinum), code (LiveCodeBench), instruction following (IFEval and IFBench) and multilingual knowledge (MMMLU). They were read with thinking off for twelve checkpoints, against the baselines of our decisions (BF16; for Qwen3.6-35B-A3B, its FP8 release): the 8-bit (FP8) checkpoints of Qwen3.8-27B and gemma-4-26B-A4B-it, the three 4-bit NVFP4 checkpoints and Qwen3.8-27B's 4-bit GGUF file, the three models' 2-bit files, Qwen3.8-27B's and Qwen3.6-35B-A3B's ~1-bit files, and Qwen3.6-35B-A3B's BF16 release.

Two cautions. A second reading of Qwen3.6-35B-A3B's FP8 release, in a fresh process, keeps 97.5% [95.4%, 99.8%] of the first reading's benchmark score, so on that model rereading alone can produce a loss of that size; the intervals, which resample questions, do not include this variation. The 2-bit result holds on the other two models' files alone. And “not separated” does not mean “the same”: no margin within which two losses count as equal was set in advance.

How the benchmarks were compared, and each checkpoint's result

On each benchmark, a checkpoint keeps its score as a share of its baseline's on the same questions, both read on the same type of card; a family keeps the mean of its benchmarks' shares, and the benchmark share kept is the mean of the families'. The decision share kept is measured as in Figure 5, over our 19,505 decisions. Their difference, decisions minus benchmarks, has an interval from 10,000 paired resamples. A checkpoint loses less on our decisions when that interval lies above zero both for the shares as read and for the shares above chance (10% on MMLU-Pro, 25% on GPQA Diamond and MMMLU, none on the other benchmarks, 33.4% on our decisions), and more when both lie below; a level does so when two of its three models' checkpoints do and none the other way, and the 8-bit and ~1-bit levels, with two models each, need both. Answers cut off or with no answer count as wrong. A second reading of each baseline keeps 99.0% [97.1%, 101.0%] of its benchmark score for Qwen3.8-27B, 99.6% [97.9%, 101.6%] for gemma-4-26B-A4B-it and 97.5% [95.4%, 99.8%] for Qwen3.6-35B-A3B. llama.cpp reads the GGUF files and vLLM the others; llama.cpp's reading of Qwen3.8-27B's BF16 weights keeps 97.0% to 107.3% of vLLM's score across the benchmarks, with no measurable difference on any. NVFP4 rows of gemma-4-26B-A4B-it and Qwen3.6-35B-A3B: one reading; varies run to run. Bold: measurably below the baseline, or a difference that separates the two.

ModelCheckpointDecisions keptBenchmarks keptDifference, points
Qwen3.8-27B8-bit: FP8 (Qwen's release)99.4%100.0% [97.8%, 102.4%]−0.6 [−3.0, +1.7]
4-bit: NVFP4 (NVIDIA)98.9%97.6% [95.1%, 100.3%]+1.4 [−1.5, +4.0]
4-bit: GGUF UD-Q4_K_XL99.6%99.5% [97.1%, 102.2%]+0.0 [−2.7, +2.5]
2-bit: GGUF UD-IQ2_XXS95.4%71.3% [68.6%, 74.2%]+24.1 [+20.8, +27.2]
~1-bit: GGUF UD-IQ1_S76.3%29.7% [27.7%, 31.8%]+46.7 [+43.7, +49.4]
gemma-4-26B-A4B-it8-bit: FP8 dynamic (Red Hat AI; W8A8)99.6%100.0% [98.2%, 102.0%]−0.4 [−2.5, +1.5]
4-bit: NVFP4 (Red Hat AI)99.8%95.0% [92.9%, 97.3%]+4.8 [+2.4, +7.1]
2-bit: GGUF UD-IQ2_XXS100.4%93.7% [91.5%, 96.0%]+6.7 [+4.2, +9.1]
Qwen3.6-35B-A3B16-bit: BF16 (as released)100.2%98.2% [96.0%, 100.5%]+2.0 [−0.3, +4.3]
4-bit: NVFP4 (NVIDIA)100.3%98.4% [96.0%, 101.0%]+1.9 [−0.8, +4.4]
2-bit: GGUF UD-IQ2_XXS98.1%93.3% [90.8%, 95.9%]+4.8 [+2.0, +7.5]
~1-bit: GGUF IQ1_M (bartowski)95.3%68.7% [66.1%, 71.5%]+26.6 [+23.5, +29.5]

A 4-bit checkpoint stores most of its weights in four bits instead of sixteen; the four we scored measure 4.81 to 5.99 bits per weight, about a third of the memory, so they fit on fewer or cheaper cards: on a rented A6000, gemma-4-26B-A4B-it's runs on one card instead of two and costs 0.48× Jev's cost instead of 0.71×. OpenRouter's prices do not pass this on reliably: a 4-bit checkpoint costs 0.35 to 1.17 times its model's price at the cheapest listing of the same model and precision, and exactly its model's price at the cheapest listing of any kind, since that view prices both from the same listings.

image/svg+xml 1 2 3 4 8 16 Bits per weight, measured (log scale) 60% 70% 80% 90% 100% Share of the baseline score kept Qwen3.8-27B: our decisions gemma-4-26B-A4B-it: our decisions Qwen3.6-35B-A3B: our decisions Gemma 3 27B GGUF checkpoints (a sibling, at labelled bits): Unsloth's MMLU the publisher's own tasks for the same 4-bit file
Figure 5. Four-bit quantized checkpoints keep almost all of the overall score. Share of baseline score kept against measured bits per weight. Filled markers joined by lines: our decision scores, against BF16 (for Qwen3.6-35B-A3B, its FP8 release). gemma-4-26B-A4B-it's and Qwen3.6-35B-A3B's NVFP4 points: one reading; varies run to run. Open markers, spread sideways: the publishers' results for the same 4-bit files, every listed task, against BF16. Dotted, at labelled (unmeasured) bits: Unsloth's MMLU for Gemma 3 27B GGUF files, a sibling.
The publishers' full tables (38 rows), read 2026-10-01
SourceTaskSettingFull precisionCompressedKept
RedHatAI/gemma-4-26B-A4B-it-NVFP4IFEval, prompt-level strictwithout thinking89.9687.2497.0%
IFEval, instruction-level strictwithout thinking93.2191.2997.9%
GSM8K Platinumwithout thinking95.4394.4399.0%
MMLU-Prowithout thinking83.4781.6597.8%
MATH-500without thinking84.8087.60103.3%
AIME 2025without thinking80.0066.2582.8%
GPQA Diamondwithout thinking73.2072.3998.9%
LiveCodeBench v6without thinking74.4868.1991.6%
IFEval, prompt-level strictwith thinking94.3994.1599.7%
IFEval, instruction-level strictwith thinking96.1695.8499.7%
GSM8K Platinumwith thinking95.9595.3199.3%
MMLU-Prowith thinking85.1983.2997.8%
MATH-500with thinking85.8785.0799.1%
AIME 2025with thinking88.7581.2591.5%
GPQA Diamondwith thinking80.8177.6196.0%
LiveCodeBench v6with thinking77.9072.7693.4%
BFCLv4 Overallwith thinking67.6262.9693.1%
BFCLv4 Single Turnwith thinking83.8582.0097.8%
BFCLv4 Multi-Turnwith thinking62.1362.13100.0%
BFCLv4 Agenticwith thinking62.1650.3381.0%
nvidia/Qwen3.8-27B-NVFP4GPQA Diamondevery task the card lists88.9288.0199.0%
Terminal-Benchevery task the card lists75.5674.0298.0%
AA-LCRevery task the card lists72.6373.38101.0%
MMMU-Proevery task the card lists75.1474.8699.6%
SciCodeevery task the card lists47.9348.41101.0%
IFBenchevery task the card lists80.0778.9398.6%
nvidia/Qwen3.6-35B-A3B-NVFP4MMLU Proevery task the card lists85.6085.0099.3%
GPQA Diamondevery task the card lists84.9084.8099.9%
τ²-Bench Telecomevery task the card lists95.5094.7099.2%
SciCodeevery task the card lists40.8040.6099.5%
AIME 2025every task the card lists89.2088.8099.6%
AA-LCRevery task the card lists62.0062.00100.0%
IFBenchevery task the card lists62.3062.80100.8%
MMMU PROevery task the card lists74.1074.50100.5%
Unsloth's GGUFs of Gemma 3 27B, a sibling of the models we ranMMLU, UD-Q4_K_XLMMLU, five worked examples71.5071.47100.0%
MMLU, UD-Q3_K_XLMMLU, five worked examples71.5070.8799.1%
MMLU, UD-IQ2_XXSMMLU, five worked examples71.5059.2082.8%
MMLU, UD-IQ1_SMMLU, five worked examples71.5041.8758.6%

A closed frontier model matches Jev, read from the letter it writes

GPT-6 Luna matches Jev, scoring 0.782, −0.006 [−0.017, +0.006] from Jev at the 97.5% level, but costs more: 2.06× Jev's cost at OpenAI's standard price (read 2026-09-30) or 1.03× at its batch price (read 2026-10-01), where answers can take hours. Neither OpenAI's API nor Google's returns a probability for every option, so we read each provider's cheapest current model from the letter it writes. Gemini 2.5 Flash-Lite, with thinking off, does not match: it scores 0.735, −0.052 [−0.068, −0.038] from Jev, at 2.25× Jev's cost at Google's list price (read 2026-10-01).

JevBench

JevBench is a public benchmark for decision models like Jev; it is not TypeSafe's. Its repository ships 231 questions, none of which overlap ours (its judge tier and held-back questions are not included). We scored them by accuracy, outside the overall score. Jev answers 85.7% [81.8%, 89.6%] correctly. gemma-4-31B-it beats it with 91.8%, +6.1 [+2.6, +9.5] percentage points; gemma-4-26B-A4B-it, Qwen3.8-Flash-Next, Qwen3.8-27B and Qwen3.6-35B-A3B show no measurable difference. The two MiMo models were not run on it.

Where Jev is ahead, and what its results require

Jev's results on this benchmark do not require a new kind of model: ordinary open-weight models read in one forward pass reach its overall score and beat it on Agent monitoring, and after temperature scaling, three of the seven are better calibrated under every treatment of Jev's rounded probabilities. This does not tell us what Jev is built on, or show that it contains nothing new.

The choice depends on the work. Screening like Triage, no labelled data, fast answers through a hosted API or nothing to host favour Jev, which also costs less than every matching model listed on OpenRouter, at the median of all OpenRouter listings. Judging agents, labelled examples to fit a temperature on, or a card of your own kept busy favour trying an open model first, served by a host that returns log-probabilities.

What we have not shown

These results come from one benchmark of 25 datasets, built from public data and TypeSafe's examples rather than a company's own workflows. Jev was version jev-1.13.0 throughout; two identical passes of it agreed on 98.3% of top answers, a variation our intervals do not cover. We have not measured Jev's or the open models' response times under concurrent load, whether a cheap listing's price lasts, or why the open models fall below Jev on Triage. TypeSafe’s documentation suggests a short description of each option, and we did not test whether one would change Jev's or the open models' scores.

The interface works with models you already have

The useful idea in TypeSafe's framing is the interface: a decision as one call that returns a probability for every option. Any open-weight model served with its next-token log-probabilities provides it: label the options, end the prompt where the answer begins, read one forward pass, and fit one temperature on labelled examples of your own. What remains is to check its score on your own decisions and its cost where you run it.

How to cite

Ivan Yee Lee and Taylor Berg-Kirkpatrick, “Your Language Model Is Not-So-Secretly a System 1 Model”, blog post, 2026-10-04, https://ivnle.github.io/blog/2026/system-one-encoder/.

@misc{lee2026system1,
  title = {{Your Language Model Is Not-So-Secretly a System 1 Model}},
  author = {Lee, Ivan Yee and Berg-Kirkpatrick, Taylor},
  howpublished = {Blog post},
  year = {2026},
  date = {2026-10-04},
  url = {https://ivnle.github.io/blog/2026/system-one-encoder/}
}

A preprint with full details will follow.