Andrew TempletonWriting ↗

Decisions API study · October 6, 2026

Fast choices.
Uneven results.

GPT‑6 Luna’s Decisions API asks the model to choose from allowed options and returns probabilities. We tested whether it could take over the structured tasks from our Jev study.

In our 240-case test, Decisions was faster than Luna Responses. It matched the recorded Astra answers less often.

Synthetic tasks only. Luna was measured October 6; Jev results are archived from September.

THE SHORT VERSION

Speed did not guarantee a correct whole answer.

We checked complete answers against recorded answers from the Astra model, our “teacher.” Matching that teacher measures consistency, not real-world correctness.

Complete outputs matching the teacher

Same 240-case mix. Bars are observed matches. White lines contain 95% of the estimated match probability under the assumptions in Methods. These selected cases do not establish production accuracy.

Compare request times and API costs
ConfigurationMedian requestAPI estimate /240 cases

Published-rate estimates, not invoices. September Jev timings and prices are historical, not a simultaneous speed test. All failed outputs stay in the denominator.

Project selection was the weak spot.

Decisions matched 59/60 ticket-triage answers but only 5/60 project-selection answers. Project selection requires choices, totals and feasibility to agree with one another. Our direct port asked for those fields as separate choice questions.

Responses used the same Luna model with high reasoning effort and a strict JSON schema. The API encoding and reasoning settings differ, so this is a comparison of configurations, not an isolated causal effect of the endpoint.

A useful next test is a sequential adapter for dependent choices. OpenAI recommends separate requests when one decision depends on an earlier answer. We did not test that redesign; these results do not establish its quality or cost.

Jev matched 180/240 complete outputs in the archived run. The new Decisions run matched 157/240, including 11 unsuccessful outputs. Use the task breakdown below before judging either system for a specific job.

Explore the experiments

Every model receives the same task contract and cases. Here, a “match” means the entire typed output agrees with the captured Astra teacher. The separate reference check uses generator-intended text labels or executable policy and portfolio rules.

Compare tasks, reference answers and eight-example arms

Complete-output agreement

95% credible interval

Bars show observed proportions; whiskers show a Beta(1,1) prior updated by successes and failures. This working model assumes exchangeable cases. Correlated synthetic families limit population interpretation.

Inspect fees, latency, failures and denominators

Fees are published-rate estimates for these requests, not invoices. Timings are observed request durations. Historical measurements are not contemporaneous speed tests.

ConfigurationMatchesInvalid / failedAPI estimateMedian95th percentile
What exactly changed between Decisions and Responses?

The model ID is gpt-6-luna in both arms. Decisions uses finite choice questions translated from the Jev compiler. Responses uses the original strict JSON schema and high reasoning effort on these typed tasks. Native Decisions has no effort setting in this study. This compares two usable configurations; it does not isolate the causal effect of changing only an endpoint. The direct port asks separate questions for fields that can have joint constraints, especially portfolio feasibility and totals. OpenAI recommends separate requests for decisions that depend on earlier answers. A redesigned sequential adapter could behave differently; it was not tested here.

The example arm receives the same eight archived training demonstrations. Neither arm receives test labels. All refusals, incomplete responses and invalid objects count against the planned denominator. No cloud request is silently retried.

Confidence needs a held-back test.

A gate chooses which outputs to accept. Each new confidence threshold was selected on development or calibration only, saved before new test calls, then left unchanged. The score is the minimum selected-label probability across fields; it is not the probability that the whole object is correct.

Inspect frozen thresholds and explore coverage

Coverage versus disagreement

The full test curve is descriptive. The highlighted point is the frozen calibration choice; no test point was used to choose a policy.

Exploring the slider does not change the reported frozen gate. Zero observed disagreements still leave posterior uncertainty and unknown production risk.

Reuse the earlier input-only routes with Decisions as fallback

These replays preserve the September router decisions and Jev outputs. Rejected cases use the new Luna Decisions output. Fees and durations add the selected stages; they are not observations of a deployed pipeline. Local routing hardware remains unpriced.

Frozen routerSent to JevFinal teacher matchesAPI estimate / 240Median replay

Read the input. Inspect the disagreement.

These are synthetic evaluation cases, not private company requests. A teacher disagreement and a reference error are different events. Expand a case to inspect the complete typed outputs.

Open the synthetic case browser

Methods and evidence

Full study design, uncertainty and limitations

A matched archival extension

Two 48-item readiness callouts; four-task teacher transfer with 48 development and 160 test cases; four-task applicability with 96 calibration and 240 test cases. The investigator had seen historical outcomes. No new untouched holdout or production traffic was collected.

Three meanings of “right”

Readiness compares with original or audited gold. Typed tasks compare with captured Astra outputs. Separate references are authored text semantics or executable policy/portfolio rules. None is an observed completed business outcome. One archived approval-policy ambiguity remains in the denominator.

Bayesian uncertainty

A Beta(1,1) prior and Bernoulli likelihood give Beta(successes+1, failures+1). Whiskers are central 95% credible intervals. Downloaded data include Beta(½,½) sensitivity. Paired configuration differences use a task-stratified family Bayesian bootstrap with 4,000 Dirichlet-weight draws. Both condition on the selected synthetic support.

Economic boundary

Request fee estimates use captured native usage and dated published rates. Decisions costs $0.10 per million input tokens; Responses has separate input, cache and output rates. Local compute, human review, engineering and realized error costs are unmeasured. No dollar savings claim includes those unknowns.

Causal model and identification assumptions

Task and input → reference, model output, error loss
Configuration → model output, request fee, latency
Input → confidence → routing → selected output → loss
Shared ambiguity → teacher error and local error
Hardware, load, network and date → latency
Synthetic selection → observed case distribution

The estimand is the paired change in complete-output agreement when substituting configurations on these inputs. Consistency and no cross-request interference are assumed. Randomized collection order reduces order effects, but reasoning settings and request encodings differ. Historical comparisons confound date and serving conditions. Observational confidence associations do not prove that confidence causes correctness or that deployment will improve business outcomes.

Source lineage and protocol deviations

The original local trial was published in June. Jev teacher transfer and its applicability follow-up were collected in September. This update reuses their contracts, cases, labels and frozen example packs. Jev is the archived 1.13.0 configuration; no new Jev calls were made. We do not describe it as current Jev performance.

New Decisions gates preserve the original selection rules. Responses does not expose the same field probabilities, so it receives no invented confidence gate. The local ladder keeps the original models and prompt; every attempt is retained. Native API refusals count as unsuccessful task outputs. The 30 new unsuccessful cloud outputs were retained, not repaired. The public evidence archive contains synthetic inputs, requests, normalized responses and analysis sources; credential handling and internal operational notes are excluded. The archive includes an offline reproduction check.

The new same-model Responses arm and Bayesian summaries are additions to the original protocols. The teacher-confusion calibration correction is a disclosure of what the original code does, not a new gold-free estimator. There is no fine-tuning or training in this update.

Follow the result back to its evidence.

The two articles share one frozen study and evidence archive. All inputs are synthetic; no private company data is included.

Collection counts and request-cost accounting

Download synthetic evidence (.zip) Machine-readable results · Original-code diagnostic · Numerical sensitivity · Frozen protocol

Sources and original paper
  1. Original paper: When a Free Model Can Replace the Frontier
  2. OpenAI Decisions API: format, probabilities and input-only pricing (retrieved October 6, 2026).
  3. GPT‑6 Luna model and Responses pricing (retrieved October 6, 2026).
  4. Jev native API; archived request usage and contract details are in the downloadable evidence.
  5. Ollama chat transport; local model digests are recorded with the study.