When Is a Cheaper Model Actually Cheaper?

A cheaper model is cheaper when its lower token bill survives the rest of the bill: repair, waiting, routing, evaluation, and the cost of switching. Compare total expected cost over the same work and time horizon, including the consequences of mistakes.

Spread the one-time switch over the amount of work you actually expect to use it for. A saving on every request can still lose money if the setup takes longer to earn back than the system will last.

An estimated $1.89 per thousand commands

I compared Luna 6 and Sol 6.1, both at high reasoning effort, on 20 authored examples of simple commands. Their typed interpretations matched exactly on 19. These examples were neither everyday traffic nor a separate final test set.

Estimated token feesPer 1,000 comparable commands at the observed workload

Estimate from 20 authored examples

Luna 6 high$0.13
Sol 6.1 high$2.01

Not billing results or general model prices. Bars and savings use unrounded estimates; displayed fees are rounded.

The ratio looks large. The absolute saving is small. Suppose, illustratively, that Sol also checks 10% of Luna’s commands in the background at the same per-command fee. Those checks leave about $1.68 per thousand commands before any additional costs.

At an assumed $100 per hour for repair time, roughly one extra minute of repair across those thousand commands consumes that entire saving. The shadow rate and hourly value are assumptions, not measured operating results. Delay, routing, review, and switching could consume more; fewer errors could create savings.

Agreement does not tell you who is right

In five selected, blinded meaning checks, I judged Luna correct on all five and Sol correct on four. Here is the joint result, using my judgments as the reference.

Both models can be wrongFive selected checks · my fallible reference labels
My judgmentSol correctSol wrong
Luna correct4Both right1Luna only
Luna wrong0Sol only0Both wrong

Selected examples, not production confusion-rate estimates. A zero here does not establish that an error never happens.

The reference can be wrongThe one selected disagreement
“I lost the plot; remind me what is going on.”
Sol proposed

Ask for clarification.

Luna proposed

Read the selected session briefing.

My judgment

The briefing interpretation was right.

Treat Sol as an unquestionable teacher and you penalize Luna for the answer I considered correct. But my judgment is also fallible. These labels record my assessment of meaning; they do not show completed tasks, avoided repairs, or realized savings.

Choose a policy, then price what happens

My selected direction was Luna by default, with sampled Sol shadow comparisons and reconsideration after real use. Sol would see copies without controlling the answer. That was a provisional policy, not a completed rollout or demonstrated efficiency gain.

I would watch actual corrections, repair time, delays, and costly mistakes. A confirmed, costly, hard-to-undo interpretation error would trigger a pause and investigation. For smaller errors, the question is whether additional expected losses and operating costs exceed the fee saving.

Another evaluation is worth buying when its expected improvement to the choice exceeds its cost. The same rule applies beyond models: choose the best available option, name what could beat it, and update when the evidence changes. That is the method in How to Be Freakishly Decisive.

Optional: assumptions, posterior uncertainty, and exact costs

What the posterior describes

With exchangeable trials, a weak Beta(1,1) prior on agreement becomes Beta(20,2) after 19 matches in 20. This estimates agreement, not correctness.

Calibration separates agreements from disagreements. Within agreements, a Beta(1,1) prior on shared correctness becomes Beta(5,1) after four shared-correct labels. Within disagreements, a Dirichlet prior assigns weight 0.5 to each of the four joint correctness categories; the Luna-only label raises that category to 1.5.

Conditional on those labels and assumptions, Luna’s modeled correctness on the authored set has a 95% credible interval of about 49–98%. The calculation does not model my labeling error, remove selection bias, establish everyday accuracy, or measure task completion. Five selected labels leave substantial uncertainty.

The arithmetic

Luna per command
$0.00012619775
Sol per command
$0.00201204
Saving per 1,000
$1.88584225
After illustrative 10% Sol shadow
$1.68463825

All amounts are USD estimates at the observed workload. The last line subtracts 100 Sol calls from the token-fee saving; it assumes comparable shadow calls and excludes other costs.