October 7, 2026
OpenAI's Decisions API: Why the Next AI Model Your Agents Need Doesn't Write a Single Word
Every AI agent you deploy spends most of its life making small choices: Is this ticket billing or technical? Should this request go to a human? Which tool comes next? For the past two years, enterprises have answered those questions by calling a large, general-purpose chat model and parsing whatever prose came back. That habit is expensive, slow, and awkward to govern.
At DevDay on September 29, 2026, OpenAI signaled it wants to change it. OpenAI’s Developers account announced the Decisions API, “powered by GPT-6 Luna,” which lets developers “define questions and possible answers to classify content, route requests, or choose an agent’s next action.” It is available in limited preview. This post explains what the Decisions API is, why the category matters, what we still do not know, and how to test a decision model before it routes real work.
What OpenAI Actually Announced
The Decisions API narrows a model to a bounded task. Instead of an open prompt, you pass a question, a closed list of allowed answers, and context in the form of text or images. According to reporting on the launch, the model returns a single answer from your defined set, which your code can branch on without parsing free-form text.
The speed claim
OpenAI’s DevDay material claims 150 ms against 1.6 seconds for a standard GPT-6 Luna call, roughly a 10x speedup. For a routing step that runs on every request in a customer-support or agent workflow, that difference is the gap between a snappy product and a laggy one.
What is still missing
The launch is lighter on specifics than the headline suggests. As of September 30, no docs page, price, rate limits, SDK, benchmark or statement on calibration had been published, and one analysis noted that standard API keys still returned a 403 “Decision API is not enabled” error as of October 2. Treat everything beyond the announcement as unconfirmed until OpenAI publishes documentation.
Why a “Decision Model” Category Is Emerging
The Decisions API did not appear in a vacuum. Two weeks earlier, TypeSafe AI launched Jev, a “System One” model that returns typed decisions instead of generated text, and Vercel reported that nearly 13% of paid AI Gateway teams used it within 24 hours. We covered the economics in our look at Jev’s record launch. OpenAI’s move confirms that the largest model vendor now sees the same opportunity: a cheap, fast, constrained model for the high-volume micro-decisions inside agent systems.
Why enterprises should care
Three forces are pushing this pattern forward:
- Cost. If an agent makes dozens of routing calls per task, using a frontier model for each one is wasteful. Jev’s listed price on OpenRouter is $0.042 per million input tokens with no separate output charge, while Luna’s published per-token rates are $0.10 input and $0.50 output per million tokens. This connects directly to the shift we described in the end of per-token pricing.
- Latency. Decisions that block a user-facing flow need to return in a fraction of a second, not seconds.
- Predictability. A model that can only pick from a list is far easier to test, log, and audit than one that can write anything.
That last point matters most for governance. A closed answer set turns an opaque model call into something closer to a typed function, which is exactly what compliance teams ask for when they review agent fleets and their oversight needs.
The Hidden Risk: Confidence You Cannot Trust
Decision APIs often return a confidence score alongside the answer, and teams use that score to decide when to escalate to a human. That design only works if the score is honest. One analysis cautioned that confidence scoring for the Decisions API is not confirmed in official sources, even though press coverage mentions it.
What one third-party test found
A comparison published by a Jev-aligned analysis site tested Luna on the ProofWriter reasoning dataset. It reported that Luna’s accuracy fell from 93-98% on simple problems to about 45-46% on problems requiring five chained inferences, and that when Luna expressed 99%+ confidence, it was correct only 68% of the time. Treat these numbers with care: the author favors a competing product, the test used a general reasoning dataset rather than business routing tasks, and the preview model OpenAI ships may differ from the Luna that was tested.
Still, the lesson holds regardless of vendor. If a model’s stated confidence does not track its actual accuracy, any threshold you set (for example, “auto-approve above 90%”) will quietly push wrong decisions past your reviewers. That is the same failure pattern behind the “Approve” button problem: a control that looks like oversight but is not.
How to Evaluate a Decision Model Before It Routes Real Work
You do not need to wait for OpenAI’s documentation to build a test harness. The same checklist applies to Jev, Luna, or any model that picks from a list.
1. Test on your own decisions
Public benchmarks show how a model handles chained logic; your own tickets, invoices, and requests show how it handles your business. Pull a few hundred historical cases with known correct answers, including the ugly ambiguous ones.
2. Check calibration, not just accuracy
Bucket the model’s answers by stated confidence and compare each bucket to real accuracy. If the 95% bucket is right only 70% of the time, that score cannot drive automation.
3. Check consistency
Run the same inputs repeatedly. The third-party test above reported that Luna changed its answer on about 10% of repeated problems versus 2-3% for Jev. Whatever the true figures for your data, unstable routing is a production incident waiting to happen.
4. Plan the escalation path
Decide in advance what happens when confidence is low or the model disagrees with a rule. Route to a human, log the case, and feed corrections back into your test set.
5. Keep a vendor-neutral abstraction
Because the contract for the Decisions API is not public and the market is moving weekly, wrap decision calls behind your own interface. Swapping models should be a configuration change, not a rewrite. This also protects you from the contract-level surprises we examined in the Cursor cutoff.
What This Means for Your AI Stack
The bigger story is architectural. For years, “AI in the product” meant one big model doing everything. Agent systems are splitting into layers: large models for open-ended reasoning and generation, and small, fast decision models for the branching logic in between. The Decisions API is OpenAI’s endorsement of that split.
For leaders, the practical implications are:
- Audit where your agents spend tokens. Routing and classification calls are the likeliest candidates for cheaper decision models.
- Treat preview features as previews. Limited access, no published pricing, and no documentation are reasons to prototype, not to commit.
- Demand calibration evidence. Ask any vendor how confidence scores were validated, and verify it on your data.
- Keep humans where errors are expensive. Use decision models to speed up low-risk routing, not to approve payments or access.
Conclusion: Small Decisions, Big Consequences
The Decisions API is less about a new model and more about a new job description for AI: not “answer anything” but “choose correctly, quickly, and predictably.” If it ships with clear documentation and honest confidence signals, it could cut cost and latency across agent workflows. If it ships without them, it will join a growing list of features that look like control but are not.
Your next steps this week: inventory the routing decisions in your agent workflows, assemble a labeled test set from real cases, and build a calibration check you can run against any decision model, OpenAI’s or anyone else’s. The teams that can measure a model’s judgment will be able to adopt the next one in days instead of quarters.
Frequently Asked Questions
What is OpenAI’s Decisions API?
It is an API built on GPT-6 Luna that takes a question, a fixed list of possible answers, and text or image context, then returns one answer from the list. It is meant for classifying content, routing requests, and choosing an agent’s next action, and it is currently in limited preview.
Is the Decisions API generally available?
Not yet. Reporting from late September and early October says access is limited to selected API customers, and standard keys were still receiving errors as of October 2. Check OpenAI’s documentation for the current status.
How is it different from calling a normal chat model?
A chat model can generate any text. A decision model is constrained to choose from your predefined options, which makes outputs easier to branch on, test, and audit. OpenAI also claims it responds around 150 ms versus about 1.6 seconds for a standard Luna call.
Can I trust the confidence scores?
Not without testing. Official confirmation of confidence scoring was lacking at launch, and a third-party test reported poor calibration on hard reasoning problems. Always compare stated confidence against measured accuracy on your own data before using it as an automation threshold.
How does it compare with TypeSafe’s Jev?
Jev launched on September 15, 2026, is available on OpenRouter, and has published pricing and schemas. The Decisions API had no public contract at the time of writing. Which performs better on your tasks is something only your own evaluation can answer.
Should we switch our routing to a decision model now?
Prototype now, but do not commit production traffic until pricing, documentation, and your own accuracy and calibration tests are in hand. Wrap the calls behind an abstraction so you can change providers easily.
Sources
- OpenAI Developers on X - OpenAI’s announcement of the Decisions API, powered by GPT-6 Luna, in limited preview
- Firecrawl: OpenAI’s Decisions API vs Jev - Dates, latency claims, Luna pricing, and the unpublished API contract
- eesel AI: OpenAI Decisions API explained - Request shape, preview access status, and confidence-scoring caveat
- AlphaSignal: Decisions API as a constrained GPT-6 Luna router - Coverage of the launch and what remains unpublished
- Dymesty: TypeSafe Jev explained - Jev’s September 15 launch and its decision-model design
- TechBytes: Jev is the fastest-adopted model in AI Gateway history - Vercel AI Gateway adoption figures for Jev
- anth.us: The OpenAI Decisions API Needs a Confidence You Can Trust - Third-party calibration test on ProofWriter (author favors Jev)
Have a project like this in mind?
Tell us what you're building — we'll help you scope it and ship it.
Talk to usKeep reading

October 9, 2026
SAP Buys TechWolf: Why the Grounding Layer, Not the Model, Is Enterprise AI's Next Battleground

October 8, 2026
Oracle Fusion Claw: What a Governed Agent Runtime Inside Your ERP Means for Finance and Operations Teams

October 4, 2026