rho and sigma: can the measurer be trusted, and was the checker independent
Can the thing that scored this be trusted, in THIS construct domain? Who checked this, what did it cost them to be wrong, and were they correlated with whoever produced it? Two questions about the instrument that produced a reading, not about the claim the reading supports - and, together, the one place on this ladder where a strong composite score can come entirely from the instrument being wrong in a way nothing here can see.
The formal frame
For a balanced binary grader whose κ is measured against a reference standard - gold labels, balanced prevalence, symmetric error, so κ = 1-2p is the channel parameter - the KL contraction coefficient is η = κ². (Inter-rater κ between two conditionally independent graders of equal quality IS η already; do not square it twice.) The per-verdict information rate is κ*log₂((1+κ)/(1-κ)) bits, and the single-verdict discrimination ceiling is the probability (1+κ)/2. At κ=0 both are exactly 0 bits and 1/2: infinite samples buy nothing.
Under conditional independence given truth, the bits you double-count when pooling two checkers equal exactly I(Y₁; Y₂). Observed agreement above a₁a₂ + (1−a₁)(1−a₂) is co-bias and is worth zero bits.
Both formulas describe an instrument, not the evidence it reads. The curve after the rungs below plots rho’s half directly against kappa; the grid after sigma’s rungs plots the co-bias half directly against two illustrative checkers.
The seven rho rungs
rho0, unmeasured
rho1, measured, below floor
rho2, marginal
rho3, in-domain κ ≥ 0.6
rho4, κ ≥ 0.8
rho5, κ ≥ 0.92 + held out
rho6, deterministic by construction
Rung 1 sits ABOVE rung 0, so a grader measured at κ = 0.05 outranks an unmeasured one. That is deliberate - a measured floor is a checkable claim and an unmeasured one is not - but it means measuring badly buys a rung. It also means this ladder cannot represent the κ = −0.23 anti-correlated case at all: that has to be written as ρ = 0, indistinguishable from never having looked. Both are consequences of a scalar. See gap G3.
What a kappa is worth, in bits
rho’s information rate is a pure function of kappa, with no data dependency at all - the curve below is the formula, nothing else. It is also even in kappa: the rate at kappa = -0.23 and the rate at kappa = +0.23 are identical, which is the exact mechanism behind the anti-correlated case the rung warning above just named. A scalar cannot climb out of that; it does not have the sign to climb with.
Pricing figures throughout are synthetic and chosen to make the type of the claim legible; none of them are measurements.
Why rho is orthogonal
hash in the catalog. rho = 1 by construction, and it tells you one bit about one file that transports nowhere (T 0, maxU 0). Honestly a JOINT rho and S witness rather than a pure one, which is why the orthogonality claim is stated pairwise - no axis derivable from the others - rather than as the stronger "high here, low on all five".
The seven sigma rungs
sigma0, self-reported
sigma1, self-eval
sigma2, same-family check
sigma3, independent internal review
sigma4, cross-provider / independent team
sigma5, liability-bearing third party
sigma6, adversarial collaboration, joint prereg
Why sigma is orthogonal
shortseller in the catalog. No method, no identification, unreliable instrument, sigma 5 - and it moves prices, because the author has capital at risk. The cleanest witness in the set.
Where rho and sigma interact
rho asks whether the measurer can be trusted. sigma asks whether the checker was independent of whoever produced the claim. Both questions are about the instrument, and neither is about the evidence the instrument read. That is why they share one route: split them across two pages and the failure mode that matters most - that they fail together - has nowhere to live.
They fail together in a specific and common way. A captured auditor scores high on rho, because its judgments are internally consistent and well calibrated on the cases it has already seen, and scores at or near zero on sigma, because it was never independent of what it was checking. The composite reads as strong verification. That is the co-bias identity above, read as an operational warning rather than a formula: agreement between correlated checkers beyond what chance-corrected independence predicts is exactly the excess that identity prices at zero bits, and a captured auditor is the case that earns the most of it.
Two checkers, each 0.80 accurate and chosen only for legibility. The diagonal cells carry an illustrative 0.03 of excess joint probability beyond the chance-corrected independence prediction; the off-diagonal cells give up exactly that much, so both checkers’ individual accuracies are unchanged. That excess implies a pairwise mutual information of about 0.023 bits between the two checkers - a distinct quantity in a distinct unit, not the same number as the excess. Per sigma’s own identity, those bits are exactly what gets double-counted when the two checkers are pooled, so they are worth zero bits of independent confirmation: the mutual information itself is not zero, it is the size of the apparent agreement that is not independent evidence.
The formal frame above states the reference-standard reading explicitly because the other reading is a live error, not a hypothetical one. eta = kappa² holds only when kappa is measured against gold labels; inter-rater kappa between two conditionally independent graders of equal quality is already eta, and squaring it a second time understates reliability by construction. Settling which reading applies to a given number is not a matter of restating the formula again - it means reading quality-sgd’s own reliability ledger, which nobody has done yet.
The published gap: rho is the wrong shape
rho is the wrong shape. A scalar sees only stochastic degradation. Contrast degeneracy and construct substitution are invisible to it, and the second is the failure mode that motivated the axis.
Index rho by (evaluator, cliff) and estimate it as blind AUC on synthesized matched pairs that toggle exactly that cliff.
The same logic runs the other direction, and it is sigma’s falsifier as much as rho’s gap: if agreement above what chance-corrected independence predicts never turns out to predict a downstream error, the correction the co-bias identity supplies is dead weight - real in the formula, worthless in the decision it exists to inform.
Where rho breaks
Collapses to a single scalar only for stochastic degradation. Two other failure modes exist and this axis is blind to both: contrast degeneracy (zero divergence on the specific contrast, no n fixes it) and construct substitution (κ near 1, sign inverted), which is the κ = −0.23 result above and the reason this axis is not a scalar.
Where sigma breaks
Wrong if the checker bears no realized loss. If nobody has ever been penalized for a bad check, σ is nominal regardless of the checker's title - which is precisely how captured auditors and rating agencies present as high-σ and score near zero.