Andrew TempletonWriting ↗

Local-model study · October 6, 2026

When can a local
model take over?

A small model can handle clear classifications and struggle on ambiguous cases. Test the task you intend to hand over.

Qwen3 4B got 48/48 clear feedback cases right, but 37/48 ambiguous support cases. That is evidence for a task-specific pilot, not a general replacement claim.

Previously inspected synthetic cases, rerun October 6. No fresh holdout or production outcomes.

THE LOCAL TEST

Start with the task.

We reran five local Qwen sizes on two small classification tasks: clear feedback intent and ambiguous support routing. Each task has 48 synthetic cases. Switch tasks to see how the same model changes.

Cloud rows use a three-call majority vote. Repeated calls do not make the 48 cases independent extra observations. Read the separate Decisions API article for the larger typed-output experiments.

Agreement with the reference labels

95% credible interval

Bars show observed accuracy. White lines contain 95% of the estimated accuracy probability under the assumptions in Methods; selected synthetic cases do not establish production accuracy. Support routing uses the archived two-auditor consensus; feedback uses the original labels. Every planned item stays in the accuracy denominator. The gray line is the earlier accuracy target: 83⅓%, equivalent to allowing $0.50 expected loss at $3 per mistake.

A replacement decision needs two prices: serving the model and being wrong. Local hardware and operating costs were not measured. The calculator below lets you supply assumptions; it does not establish production savings.

Try your own cost of mistakes

Price a mistake before choosing a model.

These controls are hypothetical assumptions, not observed business losses. Local hardware, energy and operations were not measured.

ConfigurationExpected error loss / itemProbability of clearing your barTotal cost / 1,000 calls

Expected loss uses posterior mean error. Total adds estimated request or local serving cost. The cloud totals pay for all three requests used by the majority vote. This is a replay of collected predictions, not a measured production policy.

Technical correction: the earlier estimator used gold labels

The earlier article described a correction for a fallible teacher. Its actual code estimates the teacher’s confusion matrix and class frequencies using the evaluation set’s gold labels. Gold is absent from model requests, but supplies the correction’s calibration. This is an oracle-assisted diagnostic; it does not demonstrate readiness without gold calibration.

Corrected selection also assumes teacher and local errors are conditionally independent given the true class, a full-rank teacher confusion matrix, and transferable calibration. Shared ambiguity can break those assumptions. The updated diagnostic below retains the original constrained EM estimator and loss bar.

The 2,000 bootstrap resamples are descriptive stability checks, not independent experiments or posterior draws. Invalid jointly scored cases are reported separately from all-input accuracy above.

Methods and evidence

Full study design, uncertainty and limitations

A matched archival extension

Two 48-item readiness callouts; four-task teacher transfer with 48 development and 160 test cases; four-task applicability with 96 calibration and 240 test cases. The investigator had seen historical outcomes. No new untouched holdout or production traffic was collected.

Three meanings of “right”

Readiness compares with original or audited gold. Typed tasks compare with captured Astra outputs. Separate references are authored text semantics or executable policy/portfolio rules. None is an observed completed business outcome. One archived approval-policy ambiguity remains in the denominator.

Bayesian uncertainty

A Beta(1,1) prior and Bernoulli likelihood give Beta(successes+1, failures+1). Whiskers are central 95% credible intervals. Downloaded data include Beta(½,½) sensitivity. Paired configuration differences use a task-stratified family Bayesian bootstrap with 4,000 Dirichlet-weight draws. Both condition on the selected synthetic support.

Economic boundary

Request fee estimates use captured native usage and dated published rates. Decisions costs $0.10 per million input tokens; Responses has separate input, cache and output rates. Local compute, human review, engineering and realized error costs are unmeasured. No dollar savings claim includes those unknowns.

Causal model and identification assumptions

Task and input → reference, model output, error loss
Configuration → model output, request fee, latency
Input → confidence → routing → selected output → loss
Shared ambiguity → teacher error and local error
Hardware, load, network and date → latency
Synthetic selection → observed case distribution

The estimand is the paired change in complete-output agreement when substituting configurations on these inputs. Consistency and no cross-request interference are assumed. Randomized collection order reduces order effects, but reasoning settings and request encodings differ. Historical comparisons confound date and serving conditions. Observational confidence associations do not prove that confidence causes correctness or that deployment will improve business outcomes.

Source lineage and protocol deviations

The original local trial was published in June. Jev teacher transfer and its applicability follow-up were collected in September. This update reuses their contracts, cases, labels and frozen example packs. Jev is the archived 1.13.0 configuration; no new Jev calls were made. We do not describe it as current Jev performance.

New Decisions gates preserve the original selection rules. Responses does not expose the same field probabilities, so it receives no invented confidence gate. The local ladder keeps the original models and prompt; every attempt is retained. Native API refusals count as unsuccessful task outputs. The 30 new unsuccessful cloud outputs were retained, not repaired. The public evidence archive contains synthetic inputs, requests, normalized responses and analysis sources; credential handling and internal operational notes are excluded. The archive includes an offline reproduction check.

The new same-model Responses arm and Bayesian summaries are additions to the original protocols. The teacher-confusion calibration correction is a disclosure of what the original code does, not a new gold-free estimator. There is no fine-tuning or training in this update.

Follow the result back to its evidence.

The two articles share one frozen study and evidence archive. All inputs are synthetic; no private company data is included.

Collection counts and request-cost accounting

Download synthetic evidence (.zip) Machine-readable results · Original-code diagnostic · Numerical sensitivity · Frozen protocol

Sources and original paper
  1. Original paper: When a Free Model Can Replace the Frontier
  2. OpenAI Decisions API: format, probabilities and input-only pricing (retrieved October 6, 2026).
  3. GPT‑6 Luna model and Responses pricing (retrieved October 6, 2026).
  4. Jev native API; archived request usage and contract details are in the downloadable evidence.
  5. Ollama chat transport; local model digests are recorded with the study.