# Local replacement, Jev and GPT-6 Luna Decisions: matched extension

Frozen before experimental requests on October 6, 2026. The separate synthetic
transport smoke is excluded. This is an extension using previously inspected
synthetic datasets, not a new confirmatory holdout. Earlier outcomes and errors
were available to the investigator. No production/company observations enter.

## Preserved experiments

1. **Readiness:** the two original 48-item, four-class callouts; five original
Qwen tiers; three independent requests per item for each new Luna configuration.
Preserve original local prompt, temperature zero, think=false, JSON mode, up to
three attempts for invalid answers, and 2048 output-token limit. Run one fresh
local prediction per model/item, retain every attempt and the original archived
ladder. Additive timing records do not make local compute free. Preserve the
$3 hypothetical misroute loss and strict expected-loss bar below $0.50.
Original hand gold and the archived two-auditor triage consensus are separate.
The new API uses the same instruction and allowed class vocabulary. Gold is
never included in a request. Majority-vote ties use lexical order and are logged.
2. **Teacher transfer:** four original typed contracts, saved Astra targets,
48 development and 160 test inputs, zero examples versus the same eight train-only
examples per task. Collect all new development outputs, freeze each configuration's
gate, then collect its test outputs. New Decisions gates use the lowest minimum
selected-field probability accepting at least four dev inputs with no teacher
disagreements; ties are kept. No qualifying threshold means no acceptance.
3. **Applicability:** the September follow-up's 96 calibration and 240 test
inputs, zero examples, the same four contracts. New Decisions confidence gate
requires at least 12 calibration acceptances and zero teacher disagreements,
then freezes before test calls. Archived input-only routes and Jev outputs remain
unchanged. Compare native Decisions both directly and as fallback on those routes.

## New arms and API adapter

`gpt-6-luna` via native POST /v1/decisions and the same model via POST /v1/responses.
Responses uses high reasoning for the two typed studies, none for the original
simple classification trial, strict JSON Schema, store=false, 4096 output tokens.
Native Decisions uses choice questions mechanically translated from the original
Jev compiler: identical field instruction, reversible JSON-literal choice values,
descriptions and per-question order. No extra examples, repairs or consistency
checks. Decisions has no reasoning-effort parameter in this protocol. Therefore
the estimand compares **configurations**, not an isolated endpoint effect.
All whole-output mismatches, schema/transport failures and incomplete outputs stay
in the planned denominator. One cloud attempt, no silent retries. Randomized order
within phase and bounded concurrency six. Native returned model must match the
requested model. Preserve input, request, response and source hashes.

## Estimands and uncertainty

Primary observed quantities: complete teacher matches / all planned; separate
oracle matches; accepted errors / accepted; accepted / planned; request fee
estimates, median and p95 wall latency. Gold accuracy is available on readiness;
teacher agreement is not business correctness. Decision confidence is a ranking
statistic, never assumed to be the joint correctness probability.

Use Beta(1,1) prior with Bernoulli success likelihood for conditional accuracy,
posterior Beta(success+1,failure+1), 95% credible interval and posterior probability
of meeting the loss bar. Report Beta(0.5,0.5) sensitivity. IID exchangeability is a
working model only; synthetic families and repeated items violate unqualified
population interpretation. Also use task-stratified family Bayesian bootstrap
(Dirichlet(1,...,1) weights, 4,000 draws) for paired configuration differences.
For the original trial, repeated calls are clustered by item, never tripled N.

Recompute the original latent-class constrained EM selection diagnostic and 2,000
paired item bootstrap replicates at the original fixed bar. These are descriptive
stability summaries, not posterior draws or p-values. The teacher confusion matrix
and class frequencies in the historical method are estimated using oracle gold;
gold is absent from model requests but **not** absent from this calibration. This
is an oracle-assisted diagnostic, not an identified gold-free production estimator.
Conditional error independence, teacher calibration transfer and full-rank confusion
are identification assumptions. Report failure/null results and shared-error risk.

## Causal boundary and economics

DAG: task/input -> reference, output and loss; configuration -> output, fee and
latency; input -> confidence -> routing -> selected output -> loss. Hardware,
network, serving load and collection date -> latency; shared ambiguity -> both
teacher and student error. Synthetic selection -> observed task/input distribution.
Paired intervention is substituting the model configuration on the same input.
Within-date paired comparisons identify observed configuration differences under
consistency and no cross-request interference; no randomized production traffic,
reviewer outcomes or realized business loss was observed. Historical timing has
date/hardware confounding. No general causal savings claim.

Loss calculator: expected request fee + user-specified local serving expense +
hypothetical error cost times posterior error probability. Empty unmeasured costs
stay unknown. Do not add field-level losses to whole-case loss. Decisions estimate
uses $0.10/M input tokens only; Responses uses $0.10/M input, $0.01/M cached input,
$0.125/M cache writes, $0.50/M output. These are published-rate estimates, not
billed invoice totals. Cloud reservation cap $20; unknown charges remain unknown.

Public bundle: synthetic cases, prompts, numeric results and reproducibility code.
Exclude keys, account/project identifiers, internal notes and unrelated repository
content. 

Public copy: an operational authorization note is omitted. Experimental methods are unchanged. Frozen full-protocol SHA-256: 783b000cddcf41e73264393381bc6551e4257de77f5724dbc941630c1f295797.
