Two weeks ago I tested Jev, TypeSafe's decision model, on 450 Dutch business texts. On the hard cases I could not tell it apart from a general model read through logprobs.
Since then the field moved. OpenAI announced a Decisions API at DevDay on September 29. It is now in public beta. Anthropic shipped Claude Haiku 5.5 on October 7, at a tenth of the old Haiku's input price.
So I ran both on the same 450 texts, with the same questions and the same analysis. I ran Jev again too. The whole round cost $0.42.
The short version: on the hard cases I found no demonstrable difference between Jev, OpenAI's Decisions API and Haiku 5.5. Over all 600 answers Haiku made the fewest mistakes. On speed, price and calibration they split cleanly, and each one wins a different column.
What changed in two weeks
Jev is still what it was: a hosted decision model. You send a text and typed questions, you get a choice, a score or a yes/no back, each with a probability. $0.042 per million input tokens, output free.
OpenAI's Decisions API has the same shape. It runs on gpt-6-luna, is in public beta, and costs $0.10 per million input tokens with output free. It returns a predicate, choice or score. A predicate comes back as one probability. A choice or score comes with a probability distribution and a separate confidence.
Claude Haiku 5.5 is not a decision model. It is a general model that writes text. You ask for JSON and parse it. There is no probability per answer. It thinks before it answers by default, and you can switch that off. $0.10 per million input tokens and $0.50 per million output tokens, for prompts under 100,000 tokens.
The setup
Everything from the first benchmark stayed fixed: the same 450 fictional Dutch texts, the same questions and definitions, the same metric.
- Three tests of 150 texts: helpdesk emails, search queries for a bike shop, and sales call fragments. Each test has 50 easy, 50 medium and 50 hard texts.
- Per text: one or two category questions, one score and one yes/no.
- Primary metric: macro-F1 on the category question, on the hard tier. It weighs every class equally, so a model can't win by overusing one.
- Fixed before the first call: five new arms, five hypotheses and a $2 budget, written down and committed as an addendum to the original preregistration.
The five new arms:
| Arm | What it is |
|---|---|
| Jev, rerun | The same model version as on September 26, to check stability and latency on the same day |
| OpenAI Decisions | gpt-6-luna-decisions through OpenRouter's decisions endpoint, same payload as Jev |
| Haiku 5.5, low | One call per text, all questions, answer as JSON, lowest effort |
| Haiku 5.5, medium | The same at the default effort |
| Haiku 5.5, with confidence | Low effort, plus a self-reported confidence per category answer |
All 2,250 calls succeeded. After the results were in, I added a sixth arm with thinking switched off. It is exploratory and gets its own section below.
The result on the hard tier
Macro-F1 on the category question, 50 hard texts per test (for search: the average of intent and topic cluster).
| Approach | Helpdesk | Search | Sales |
|---|---|---|---|
| Keyword rule | 0.34 | 0.60 | 0.38 |
| Qwen3.7 Flash via logprobs (round 1) | 1.00 | 0.90 | 0.90 |
| Jev, September 26 | 1.00 | 0.92 | 0.92 |
| Jev, October 9 | 1.00 | 0.90 | 0.92 |
| OpenAI Decisions | 0.98 | 0.96 | 0.98 |
| Haiku 5.5, low | 1.00 | 0.96 | 0.98 |
| Haiku 5.5, medium | 1.00 | 0.94 | 0.96 |
Between Jev, OpenAI and Haiku I found no demonstrable difference. OpenAI is ahead of the September 26 Jev run by 0.04 on search and 0.06 on sales. None of the preregistered intervals excludes zero. Haiku and the better of the two decision models land on the same score in all three tests.
One comparison I didn't preregister points the other way. Against the same-day Jev rerun, the search gap is 0.06, and there the interval just clears zero for both OpenAI and Haiku. Jev's two runs are one error apart on those 50 texts. I read it as a lead, not a result.
OpenAI's Decisions API does beat the logprobs approach. Against Qwen3.7 Flash it scores +0.06 on search and +0.08 on sales, both with intervals above zero. Jev did not manage that in round 1.
The helpdesk test sits at the ceiling. All six model rows score 0.98 or 1.00 there. It cannot separate them.
Where they differ

Over all three tiers there are 600 category answers per approach. Here the picture changes.
| Jev | OpenAI Decisions | Haiku 5.5, low | |
|---|---|---|---|
| Wrong answers out of 600 | 41 | 30 | 15 |
| Calibration error (ECE) | 0.014 | 0.054 | no probabilities |
| Median latency per text | 293 ms | 179 ms | 3,105 ms |
| Cost per 1,000 decisions | $0.011 to $0.016 | $0.024 to $0.031 | $0.063 to $0.071 |
| Cost of all 450 texts | $0.020 | $0.041 | $0.101 |
Haiku makes the fewest mistakes. Fifteen wrong out of 600, against 30 for OpenAI and 41 for Jev. Most of the gap sits on the easy and medium tiers, which the hard-tier table doesn't show: 28 errors there for Jev, 10 for Haiku.
Over all 600 answers that gap is larger than noise. A paired test (McNemar) on the answers where two approaches disagree gives p < 0.001 for Haiku against Jev and p = 0.002 for Haiku against OpenAI. OpenAI against Jev: p = 0.09. I didn't preregister this test, so treat it as a strong lead, not a confirmed result.
OpenAI is the fastest. 179 ms per text at the median, against 293 ms for Jev measured on the same day. Haiku takes three seconds, because it thinks first: 176 reasoning tokens on average per call at low effort.
Jev is the cheapest and the best calibrated. When Jev says 90% or more, it is right 98.7% of the time (471 answers). Its calibration error is 0.014, against 0.015 two weeks ago.
Jev is also stable. On October 9 it gave the same answer as on September 26 on 594 of the 600 category questions.
OpenAI is right more often than it thinks
OpenAI's probabilities are useful, but they lean low. Across all category questions it was right 95.0% of the time at an average confidence of 0.897. In the bucket where it said 90% or more, it was right on all 415 answers.
The yes/no question has the opposite problem: there its probabilities run too high. On "should a person handle this email", its ranking is nearly perfect, with an AUC of 0.996. But at the obvious threshold of 0.5 it gets 84% right. Of the 77 emails that did not need a person, 24 still scored 0.5 or higher. A threshold of 0.95, picked afterwards on these same 150 emails, would have made it 97%.
OpenAI's own docs tell you to set thresholds on your own labeled examples. This is why. A probability that ranks well is not automatically a probability you can cut at 0.5.
Haiku's own confidence is the weakest signal
I asked one Haiku arm to add a confidence number to each category answer. It was right 97.5% of the time at an average stated confidence of 0.864. In the bucket where it said 80 to 90%, it was right on 164 of 165.
So the self-reported number points in the right direction but understates by a wide margin. With a calibration error of 0.112, it is the worst of the three. Asking for the number also cost 43% more per text than plain low effort, for no gain in accuracy.
Medium effort didn't help either. It scored the same or slightly lower than low effort on the hard tier, at 10% more cost.
Haiku with thinking switched off
My preregistration said Haiku 5.5's thinking can't be switched off. I took that from early coverage, and it was wrong. Anthropic's docs say thinking: {"type": "disabled"} works at low, medium and high effort. Through OpenRouter that is reasoning: {"enabled": false}.
So I ran one more arm with thinking off, after I had seen the other results. That makes it exploratory. Same texts, prompt and analysis.
| Haiku, low, thinking on | Haiku, thinking off | OpenAI Decisions | |
|---|---|---|---|
| Hard tier, helpdesk / search / sales | 1.00 / 0.96 / 0.98 | 1.00 / 0.93 / 0.94 | 0.98 / 0.96 / 0.98 |
| Wrong answers out of 600 | 15 | 26 | 30 |
| "Should a person handle this", right at 0.5 | 97% | 84% | 84% |
| Output tokens per call | 217 | 47 | output is free |
| Median latency per text | 3.1 s | 1.1 s | 0.18 s |
| Cost per 1,000 decisions | $0.063 to $0.071 | $0.036 to $0.049 | $0.024 to $0.031 |
With thinking off, Haiku is about three times faster and 40% cheaper. It also makes more mistakes: 26 against 15 (p = 0.035, exploratory). The yes/no question suffers most.
The interesting row is the comparison with OpenAI. Thinking off, Haiku lands where OpenAI's Decisions API lands on errors: 26 against 30, p = 0.62. OpenAI is still six times faster and about a third cheaper, and it gives you a probability.
For this kind of task I'd keep Haiku's thinking on at low effort. The thinking is what bought the low error count. Without it, a decision model does the same job faster.
All three hit the same wall
Nine mistakes were shared by all three approaches. All nine are the same thing: search queries labeled "navigational".
- "website fabrikant fietsaccu handleiding" (the manufacturer's website for a battery manual)
- "morrow stadsfietsen reviews pagina" (a brand's reviews page)
- "forum voor stroomfietsers" (a forum for e-bike riders)
By my definition the searcher wants to reach a specific page, so it's navigational. All three models read the first as informational and the second as commercial. So would many people.
This is the same class that broke all three models in round 1. Jev and both newcomers missed the same nine texts, eight of them with the same wrong answer. That says more about my definition than about the models.
What this test does not prove
One model family wrote the texts. Subagents of a Claude model wrote all 450 texts, and the same model checked them. Kimi blind-labeled a sample of 60 as the only outside check. Haiku 5.5 is also a Claude model. Texts can suit a reader from the same family better, so Haiku's low error count may be partly home advantage. This is the most important limit of this round, and I can't rule it out with this data.
Everything ran through OpenRouter, Haiku pinned to Anthropic as the provider. Any translation in OpenRouter's endpoints is part of what I measured.
Fifty texts per tier. A difference of five points does not show up reliably at this size. "No demonstrable difference" means exactly that. It does not mean "equal".
One run per approach, at temperature 0 where the provider allowed it, measured from one machine.
Three Dutch tasks. This is not a general ranking of these models.
Which one I'd pick
The question decides more than the model. That was the lesson of round 1, and it held.
- Jev when cost and calibration matter most: many decisions, and you want a probability you can cut at a fixed threshold. It is the cheapest per decision and its numbers mean what they say.
- OpenAI's Decisions API when latency matters, or when you already run on OpenAI and don't want another vendor. Set your thresholds on your own data, because its probabilities don't line up with a 0.5 cut.
- Haiku 5.5 when every error counts, the volume is modest, and you don't need a probability. With thinking on it is ten times slower and about five times more expensive per decision than Jev, and for this test still a fraction of a cent per text. Switching thinking off makes it faster and cheaper, but then it loses the lead it had on errors.
- A keyword rule where the signal is literal. It is still free and instant.
If you run Claude on a Max or Team plan, the new monthly API credits cover Haiku on the Claude API itself, not through OpenRouter as I ran it here. I wrote up what they cover and when claude -p counts against them.
Frequently Asked Questions
Vincent van Deth
AI Strategy & Architecture
I build production systems with AI — and I've spent the last six months figuring out what it actually takes to run them safely at scale.
My focus is AI Strategy & Architecture: designing multi-agent workflows, building governance infrastructure, and helping organisations move from AI experiments to auditable, production-grade systems. I'm the creator of VNX, an open-source governance layer for multi-agent AI that enforces human approval gates, append-only audit trails, and evidence-based task closure.
Based in the Netherlands. I write about what I build — including the failures.