OpenAI's Decisions API: What It Returns, What It Costs, and Where Jev and Haiku 5.5 Fit

OpenAI's Decisions API doesn't write text. You give it a document and a set of typed questions, and it returns an answer per question with a probability attached.

That is a different tool from a chat model. It is closer to what TypeSafe built with Jev. OpenAI announced it at DevDay on September 29, about two weeks after TypeSafe launched Jev. It is now in public beta.

I ran it on 450 of my own test texts on October 9, next to Jev and Claude Haiku 5.5. This post is the explainer I wanted before I started: what the API returns, what it costs, what it leaves out, and when I would pick something else.

What OpenAI's Decisions API returns

Every question has one of three types. Each answer carries the name you gave the question.

TypeWhat you askWhat you get back
predicateIs this condition true?A probability from 0 to 1
choiceWhich of these options fits?The choice, a probabilities list per option, and a confidence
scoreWhere on this scale does it land?A score, a probabilities list per level, and a confidence

A score is a probability-weighted average over the levels, so it can fall between two levels. The levels are numbered from 0.

There is also a fourth answer type: refusal. OpenAI's own sample code checks for it before reading any numbers. Do the same.

What you don't get is text. There is no reason, no quote from the source, no explanation of why the model chose what it chose.

A request

The endpoint is POST /v1/decisions. The shared text goes into input, the questions into a questions list. This is the shape from OpenAI's Decisions guide, with a helpdesk email as input:

json
{
  "model": "gpt-6-luna",
  "input": "My order arrived broken and I want a refund before Friday.",
  "questions": [
    {
      "type": "choice",
      "name": "category",
      "instructions": "What is the main reason for this email?",
      "choices": [
        {"value": "delivery", "description": "Late, lost or damaged in transit"},
        {"value": "defect", "description": "The product does not work"},
        {"value": "return", "description": "The customer wants to send it back"},
        {"value": "other", "description": "None of the above"}
      ]
    },
    {
      "type": "score",
      "name": "urgency",
      "instructions": "How urgent is this email?",
      "levels": [
        {"label": "low", "description": "No deadline, no damage"},
        {"label": "medium", "description": "Handle today or tomorrow"},
        {"label": "high", "description": "Money or safety at stake right now"}
      ]
    },
    {
      "type": "predicate",
      "name": "needs_human",
      "instructions": "Should a person handle this instead of a template?"
    }
  ]
}

Two things worth copying from OpenAI's own guidance. Add an other option when your choices don't cover every possible input, so the model has somewhere to go. And write levels with criteria you can observe, so two neighbouring levels never describe the same email.

What it costs

OpenAI's guide puts it plainly: "With gpt-6-luna, input costs $0.10 per 1M tokens. You pay only for input tokens." No output charge and no cache charges. Regional processing and long-context multipliers still apply.

In my test that came to $0.024 to $0.031 per thousand decisions. All 450 texts, with three or four questions each, cost $0.041.

For comparison, on the same texts and the same day:

Input per 1M tokensOutputPer 1,000 decisions in my test
Jev (TypeSafe)$0.042free$0.011 to $0.016
OpenAI Decisions$0.10free$0.024 to $0.031
Claude Haiku 5.5, low effort$0.10$0.50$0.063 to $0.071
Claude Haiku 5.5, thinking off$0.10$0.50$0.036 to $0.049

Haiku 5.5 is a general model, so you pay for the tokens it writes. By default it thinks before it answers. At low effort that came to 217 output tokens per call, 176 of them reasoning. With thinking switched off it wrote 47, which cut the cost by about 40%. It also made more mistakes, 26 instead of 15 out of 600.

What it leaves out

A reason. A probability tells you how sure the model is, not why. If a person downstream has to act on the answer, like a client reading a score, they need a sentence from the source. The Decisions API doesn't give you one.

Calibrated numbers by default. OpenAI calls the predicate value "the model's estimate". The docs say nothing about calibration and tell you to "use labeled examples from your application to set thresholds". On my 600 choice questions the numbers leaned low: right 95% of the time while the top probability averaged 0.90. On one yes/no question they leaned high: 24 of the 77 emails that didn't need a person still scored 0.5 or more. A cut at 0.5 got 84% right. A cut at 0.95, picked afterwards on the same emails, got 97%.

A stated input limit. The guide names no context window, no maximum input size and no maximum number of questions per request. The model underneath, gpt-6-luna, has a context window of 1,050,000 tokens. Its model page says prompts over 272K input tokens are billed at twice the input rate for the whole request, and the Decisions guide says long-context multipliers apply. For a long contract that would mean $0.20 per million tokens instead of $0.10. I haven't tested inputs that long, because my test texts were short. Haiku 5.5 also has a million-token window, but its price jumps earlier: above 100K tokens it costs $0.50 per million input.

Hosted images. Images work, but only as inline base64 data URLs. Hosted URLs and file_id inputs are not supported.

A choice of model. gpt-6-luna is the only one. The API is in public beta, and OpenAI says it expects general availability "in the coming weeks". Expect the edges to move.

A speed number you can check. OpenAI says it returns answers "about 10x faster than the Responses API". The docs don't show the measurement. In my test, run through OpenRouter, the median was 179 ms per text with all questions in one call.

On the other side of the ledger, it supports Zero Data Retention and HIPAA use for eligible customers, with data residency in the US and in Europe. For a European business that matters more than a few milliseconds.

Where Jev and Haiku 5.5 fit

I ran all three on the same 450 Dutch texts. The full results are in Jev vs OpenAI's Decisions API vs Claude Haiku 5.5. In short: on the hard texts I found no demonstrable difference in accuracy between them. They split on everything else.

  • Pick OpenAI's Decisions API when latency matters, when you need images, or when you already run on OpenAI and don't want a second vendor. It was the fastest of the three.
  • Pick Jev when you make many decisions and want a probability you can cut at a fixed threshold. It was the cheapest and the best calibrated, and it gave the same answer on 594 of 600 questions two weeks apart.
  • Pick Claude Haiku 5.5 when every error counts and volume is modest. With thinking on it made the fewest mistakes, but it is about ten times slower than Jev and gives no probability per answer. With thinking off it is three times faster, and its error count rises to about OpenAI's level.
  • Keep a keyword rule where the signal is literal. It is still free and instant.

If you build on Claude: Anthropic has no decisions endpoint. The Vercel AI SDK emulates one for Claude, and its docs are clear about the trade-off: Claude's "Boolean answers contain prompted estimates of P(true), and Choice and Score answers omit probability distributions."

When I use a decision model, and when I don't

I use one where the question is closed, the error is cheap and visible, and a number is all the next step needs. Routing an email, tagging a finding, flagging a search query.

I don't use one as the gate itself, or anywhere a person has to understand the answer. A score of 0.73 without a reason is a worse product for a client than a sentence from the source, even when the number is right.

And I start with the question, not the model. In my second round, all three approaches failed on the same nine search queries, all labelled navigational. The first round tripped on the same definition. The definition of the class was the problem, not the model reading it. Fix that first and the choice of API matters less. Choosing where a decision model fits in a real workflow is part of my AI architecture work.

Read also: Jev vs 25 Lines of Python: Typed Decisions Tested on 450 Dutch Texts

Frequently Asked Questions

Vincent van Deth

AI Strategy & Architecture

I build production systems with AI — and I've spent the last six months figuring out what it actually takes to run them safely at scale.

My focus is AI Strategy & Architecture: designing multi-agent workflows, building governance infrastructure, and helping organisations move from AI experiments to auditable, production-grade systems. I'm the creator of VNX, an open-source governance layer for multi-agent AI that enforces human approval gates, append-only audit trails, and evidence-based task closure.

Based in the Netherlands. I write about what I build — including the failures.

Comments

Your email address will not be published. Comments are reviewed before publication.

Loading comments...