← Lexicon

Bayesian Alignment Kernel

A typed contract for running an AI loop on two posteriors instead of one score: an operator-preference posterior and a reality-outcome posterior with disjoint evidence channels, a per-evaluator reliability ledger that gates every update, and a divergence statistic between the two that rings when the proxy is being gamed. Makes machine judgment underwritable by forcing every output to carry the evidence channel that produced it and the measured reliability of every grader that scored it.

Bayesian Alignment Kernel - conceptual diagram

Why It Exists

The single number most agent loops optimize - "quality" - silently merges four different objects (what the operator wants, what reality rewards, what the system produces, how much each grader can be trusted), and while they share one number reward hacking is not even statable: there is nothing for the hacked metric to diverge from. Typing the concerns apart turns autonomy from a vibe into a balance sheet: it extends the way a lender extends credit, against measured loss rates instead of demos, with a stated alarm for the day the book is being marked to model instead of market.

Related Terms

Verification Trap - A task that is easy to generate but hard to verify.