GPT‑6 Luna’s Decisions API asks the model to choose from allowed options and returns probabilities. We tested whether it could take over the structured tasks from our Jev study.
In our 240-case test, Decisions was faster than Luna Responses. It matched the recorded Astra answers less often.
Synthetic tasks only. Luna was measured October 6; Jev results are archived from September.
THE SHORT VERSION
Speed did not guarantee a correct whole answer.
We checked complete answers against recorded answers from the Astra model, our “teacher.” Matching that teacher measures consistency, not real-world correctness.
Complete outputs matching the teacher
Same 240-case mix. Bars are observed matches. White lines contain 95% of the estimated match probability under the assumptions in Methods. These selected cases do not establish production accuracy.
Compare request times and API costs
Configuration
Median request
API estimate /240 cases
Published-rate estimates, not invoices. September Jev timings and prices are historical, not a simultaneous speed test. All failed outputs stay in the denominator.
Project selection was the weak spot.
Decisions matched 59/60 ticket-triage answers but only 5/60 project-selection answers. Project selection requires choices, totals and feasibility to agree with one another. Our direct port asked for those fields as separate choice questions.
Responses used the same Luna model with high reasoning effort and a strict JSON schema. The API encoding and reasoning settings differ, so this is a comparison of configurations, not an isolated causal effect of the endpoint.
A useful next test is a sequential adapter for dependent choices. OpenAI recommends separate requests when one decision depends on an earlier answer. We did not test that redesign; these results do not establish its quality or cost.
Jev matched 180/240 complete outputs in the archived run. The new Decisions run matched 157/240, including 11 unsuccessful outputs. Use the task breakdown below before judging either system for a specific job.
Explore the experiments
Every model receives the same task contract and cases. Here, a “match” means the entire typed output agrees with the captured Astra teacher. The separate reference check uses generator-intended text labels or executable policy and portfolio rules.
Compare tasks, reference answers and eight-example arms
Complete-output agreement
95% credible interval
Bars show observed proportions; whiskers show a Beta(1,1) prior updated by successes and failures. This working model assumes exchangeable cases. Correlated synthetic families limit population interpretation.
Inspect fees, latency, failures and denominators
Fees are published-rate estimates for these requests, not invoices. Timings are observed request durations. Historical measurements are not contemporaneous speed tests.
Configuration
Matches
Invalid / failed
API estimate
Median
95th percentile
What exactly changed between Decisions and Responses?
The model ID is gpt-6-luna in both arms. Decisions uses finite choice questions translated from the Jev compiler. Responses uses the original strict JSON schema and high reasoning effort on these typed tasks. Native Decisions has no effort setting in this study. This compares two usable configurations; it does not isolate the causal effect of changing only an endpoint. The direct port asks separate questions for fields that can have joint constraints, especially portfolio feasibility and totals. OpenAI recommends separate requests for decisions that depend on earlier answers. A redesigned sequential adapter could behave differently; it was not tested here.
The example arm receives the same eight archived training demonstrations. Neither arm receives test labels. All refusals, incomplete responses and invalid objects count against the planned denominator. No cloud request is silently retried.
Confidence needs a held-back test.
A gate chooses which outputs to accept. Each new confidence threshold was selected on development or calibration only, saved before new test calls, then left unchanged. The score is the minimum selected-label probability across fields; it is not the probability that the whole object is correct.
Inspect frozen thresholds and explore coverage
Coverage versus disagreement
The full test curve is descriptive. The highlighted point is the frozen calibration choice; no test point was used to choose a policy.
Exploring the slider does not change the reported frozen gate. Zero observed disagreements still leave posterior uncertainty and unknown production risk.
Reuse the earlier input-only routes with Decisions as fallback
These replays preserve the September router decisions and Jev outputs. Rejected cases use the new Luna Decisions output. Fees and durations add the selected stages; they are not observations of a deployed pipeline. Local routing hardware remains unpriced.
Frozen router
Sent to Jev
Final teacher matches
API estimate / 240
Median replay
Read the input. Inspect the disagreement.
These are synthetic evaluation cases, not private company requests. A teacher disagreement and a reference error are different events. Expand a case to inspect the complete typed outputs.
Open the synthetic case browser
Methods and evidence
Full study design, uncertainty and limitations
A matched archival extension
Two 48-item readiness callouts; four-task teacher transfer with 48 development and 160 test cases; four-task applicability with 96 calibration and 240 test cases. The investigator had seen historical outcomes. No new untouched holdout or production traffic was collected.
Three meanings of “right”
Readiness compares with original or audited gold. Typed tasks compare with captured Astra outputs. Separate references are authored text semantics or executable policy/portfolio rules. None is an observed completed business outcome. One archived approval-policy ambiguity remains in the denominator.
Bayesian uncertainty
A Beta(1,1) prior and Bernoulli likelihood give Beta(successes+1, failures+1). Whiskers are central 95% credible intervals. Downloaded data include Beta(½,½) sensitivity. Paired configuration differences use a task-stratified family Bayesian bootstrap with 4,000 Dirichlet-weight draws. Both condition on the selected synthetic support.
Economic boundary
Request fee estimates use captured native usage and dated published rates. Decisions costs $0.10 per million input tokens; Responses has separate input, cache and output rates. Local compute, human review, engineering and realized error costs are unmeasured. No dollar savings claim includes those unknowns.
Causal model and identification assumptions
Task and input → reference, model output, error loss Configuration → model output, request fee, latency Input → confidence → routing → selected output → loss Shared ambiguity → teacher error and local error Hardware, load, network and date → latency Synthetic selection → observed case distribution
The estimand is the paired change in complete-output agreement when substituting configurations on these inputs. Consistency and no cross-request interference are assumed. Randomized collection order reduces order effects, but reasoning settings and request encodings differ. Historical comparisons confound date and serving conditions. Observational confidence associations do not prove that confidence causes correctness or that deployment will improve business outcomes.
Source lineage and protocol deviations
The original local trial was published in June. Jev teacher transfer and its applicability follow-up were collected in September. This update reuses their contracts, cases, labels and frozen example packs. Jev is the archived 1.13.0 configuration; no new Jev calls were made. We do not describe it as current Jev performance.
New Decisions gates preserve the original selection rules. Responses does not expose the same field probabilities, so it receives no invented confidence gate. The local ladder keeps the original models and prompt; every attempt is retained. Native API refusals count as unsuccessful task outputs. The 30 new unsuccessful cloud outputs were retained, not repaired. The public evidence archive contains synthetic inputs, requests, normalized responses and analysis sources; credential handling and internal operational notes are excluded. The archive includes an offline reproduction check.
The new same-model Responses arm and Bayesian summaries are additions to the original protocols. The teacher-confusion calibration correction is a disclosure of what the original code does, not a new gold-free estimator. There is no fine-tuning or training in this update.
Follow the result back to its evidence.
The two articles share one frozen study and evidence archive. All inputs are synthetic; no private company data is included.