The Nature of Knowledge: six axes, and what each licenses alone
Two analysts hand you the same elasticity. One read "Elasticity is -1.8." off a regression on last year’s transactions; the other ran a randomized price test and got the same digits. You have to decide whether to reprice, and the number itself won’t tell you which one you’re holding.
You score a claim on the six coordinates below and fold five of them. What comes back is a level, plus the set of axes sitting at it: the ones binding the claim, and the ones to go fix. No single axis licenses a claim by itself.
The six axes
Two of the six describe the claim itself. On U you commit to a shape: a bare point, an interval, a distribution, a set you can’t narrow. On I you ask whether the quantity you’re talking about can be recovered from evidence at all, however carefully you stated it.
The other four are about the edge between the claim and whatever produced it. Between them they cover the check that produced it, the instrument that read it, whether the checker was independent, and how far the evidence’s own region reaches. Scoring an edge doesn’t require settling what the claim is worth, so we fold those four together with I into a single bound at the end.
Each axis, the question it asks, and the record from the evidence catalog that witnesses it. U has none, so its marker is a hollow ring.
Pricing figures throughout are synthetic and chosen to make the type of the claim legible; none of them are measurements.
Five of the six axes each get a witness, and U is the one without. A witness is a catalog record where the axis it witnesses sits at the top or the bottom of that record’s own five scores, at least three rungs clear of the other end. That is a claim about one record’s own spread and nothing more. High here and low on all the rest would be stronger, and it’s false of the reliability witness, which sits at the top of severity too, because determinism maxes both at once. Either direction works, and the identification witness is a low one. U0 in the package’s own ladder is "Elasticity is -1.8." - an error bar deleted, not a small one - and reporting the same number to six decimals leaves it exactly there. U measures the shape of the commitment and not the precision of the display, so it’s a property of how you wrote the claim down.
U: the order of uncertainty
This axis measures commitment, not warrant. Higher levels state forms of uncertainty that lower levels leave out. The two anti-correlate at the extremes: a bare point estimate is the most committed thing you can say, and the least warranted.
- Rung
- U0 Point
- Type signature
- Point<A> = { value: A, provenance: Provenance }
- Cannot say
- Anything about how wrong it might be. A point estimate is a claim with the error bar deleted, not a claim with a small error bar.
- Promoted by
- Any repeated measurement with a stated sampling model.
- Not promoted by
- More decimal places. Precision of display is not precision of estimate.
- Pricing
- "Elasticity is -1.8."
Pick a level for its type signature, what it can’t say, and what does and doesn’t promote a claim to it.
At U5 (Partially identified (set-valued, non-shrinking)) the width doesn’t shrink as n grows. That non-shrinking is the diagnostic: the width belongs to the identification problem rather than to the sample. If your band narrows as data accumulates, it isn’t showing you the identification gap. At U6 (Residual-bounded (named unmodeled terms)) the bound isn’t a probability at all:
Nothing further - but the bound itself is NOT a probability. It is a corroboration claim: a track record of how often this model class has blown through its stated bounds. You cannot bound the residual from inside the model, because a model's residual estimate covers exactly the errors it can represent.
Neither limit gets bought off with a larger sample. That leaves an out-of-model track record as a routinely available promoter at the top of this ladder. It’s why we ship a cap at the last level instead of a probability.
A cap bounds the consequence rather than the distribution, and the two quantify over different things. A distributional bound quantifies over outcomes: every draw from the law lands inside the stated region, a claim you can only make from inside a model class. A cap quantifies over exposure. It names the worst realized loss you’ve agreed to carry and the action that stops the loss there, and says nothing about the shape of the residual. The level-6 pricing line already carries one:
"Competitor repricing is unmodeled and could contribute 4 volume points, +/- $36k on this SKU; our exposure is capped by a 30-day revert clause."
That 30-day revert clause doesn’t predict the elasticity and doesn’t bound the competitor the model never represented. It bounds what a wrong elasticity can cost: thirty days of mispriced margin on one SKU, then the price goes back. A stated cap is valid only if some owner and some system can actually execute the stopping action. If the revert depends on a decision nobody owns, or on a system that can’t reprice inside a month, you never had a cap.
Estimating the error on the error doesn’t escape the regress. The ladder’s U3 entry (Calibrated predictive (error on error)) says why:
A hierarchical model. Dist(Dist(A)) collapses to Dist(A) under the Giry monad multiplication - integrate out the mixing measure - so stacking hyperpriors produces fatter tails and never an escape from the regress. (1) This is monad multiplication, not the tower property; the tower property is the conditional-expectation identity that makes the collapse decision-irrelevant, which is a different fact. (2) The collapse is decision-irrelevant only because Bayes risk is AFFINE in the predictive measure, so a decision maker sees the marginal and nothing else. That is exactly the assumption rung 4 abandons: a credal set has no mixing measure to integrate, so there is no marginal to collapse to, and the regress genuinely stops being a regress.
I: is the quantity recoverable at all
7 levels, running from bare association up to a validated mechanism. What they grade is the estimand, never the report: whether the thing you’re asking about can be pulled out of the evidence you have.
I0, Association
P(Y | X) as observed. Prediction under the status quo policy only.
I1, Adjusted association
P(Y | X, Z). Still zero interventional content; adjustment is not identification.
I2, Ignorability asserted
A causal reading conditional on an untestable assumption you named. Gate on a sensitivity witness.
I3, Partially identified causal
A SET containing the causal effect, under assumptions weak enough to defend.
I4, Point-identified, local
E[Y|do(X)] on a subpopulation you may not be able to name (LATE compliers, RDD at the cutoff, DiD treated units).
I5, Point-identified, target population
E[Y|do(X)] on the population the decision is about.
I6, Mechanism / counterfactual
Rung-3 queries: attribution, mediation, "would this unit have converted anyway". Needs a validated SCM, not a DAG you drew.
U and I don’t substitute for each other, and the reason is definitional rather than a theorem. U types the shape of what a report asserts; I types whether the quantity it’s about can be recovered. Sharpening the shape can’t change which question the report answers, so refining a claim along U moves it up the grid below and never right.
Pearl’s Causal Hierarchy Theorem gives that empirical bite: a rung-k quantity isn’t determined by rung-(k-1) data alone, absent causal assumptions, except on a measure-zero set of structural models. Drop that qualifier and the theorem forbids this ladder’s own I2 through I5 - Ignorability asserted, Partially identified causal, Point-identified, local and Point-identified, target population. Recovering any of those from observational data means importing an assumption the data can’t check. Randomize on the population the decision is about and you need no such import. That imported assumption is itself a claim, scorable on these same six axes.
8 worked claims on the U x I grid (rows run U0 at the bottom to U6 at the top; columns run I0 through I6 left to right). The two leftmost columns are the dead zone, shaded and hatched in rose: I0 (Association) and I1 (Adjusted association). Neither licenses a causal reading, however far up U a claim climbs. The two diamonds are the traps, both inside that dead zone. Hover a dot for the claim and why it sits where it does.
Climbing U can’t rescue a confounded estimand, and the failure is worse than no improvement. The trap "Calibrated elasticity band" above is the worked case. Calibration passes there, because it’s scored against the estimand you happen to have. More data makes it worse: the variance goes to zero and the bias stays. Under confounding, any trust metric that falls as posterior variance falls is anti-correlated with the truth.
S: what a false claim would have had to survive
How hard did the producing transition try to refute this, and did it survive? Severity grades the transition that produced the claim rather than the quantity the claim is about.
Sev(C; test, x) = P(test yields a worse fit for C | C is false). Each rung closes one path by which a false claim could have fit anyway.
Read the conditional as: how likely this test was to come back worse for the claim if the claim were false.
Each level and the route it closes. A level name is shorthand for the specific way a false claim could otherwise have fit anyway.
obs-slope / negcontrol in the catalog. Preregistered, blinded, powered, independently run - and zero identification, because nothing was intervened on. S 5 ("live + powered") and S 6 ("durable") presuppose a deployed randomized intervention, so among empirical-domain kinds S >= 5 forces I >= 5 for every kind except negcontrol, which reaches S 5 at I 2. Outside that domain the presupposition does not apply and maximal S with zero I is directly constructible: refutation and hash both sit at S 6 with I 0.
Severity on its own still doesn’t justify a claim, because a severe test of an unidentified quantity is a severe test of the wrong thing.
Breaks if realized severity is not discounted by search: a test with one-shot severity 0.95 has realized severity near zero if it is the surviving member of an unlogged specification search.
An agent searching specifications runs the search far faster than anyone does by hand, and it can log every branch it took, which turns a judgement call into a log. We haven’t built it (G4).
rho and sigma: can the instrument be trusted, and was the checker independent
Can the thing that scored this be trusted, in THIS construct domain? Who checked this, what did it cost them to be wrong, and were they correlated with whoever produced it?
rho: reliability
For a balanced binary grader whose κ is measured against a reference standard - gold labels, balanced prevalence, symmetric error, so κ = 1-2p is the channel parameter - the KL contraction coefficient is η = κ^2. (Inter-rater κ between two conditionally independent graders of equal quality IS η already; do not square it twice.) The per-verdict information rate is κ*log2((1+κ)/(1-κ)) bits, and the single-verdict discrimination ceiling is the probability (1+κ)/2. At κ=0 both are exactly 0 bits and 1/2: infinite samples buy nothing.
rho0, unmeasured
rho1, measured, below floor
rho2, marginal
rho3, in-domain κ >= 0.6
rho4, κ >= 0.8
rho5, κ >= 0.92 + held out
rho6, deterministic by construction
Rung 1 sits ABOVE rung 0, so a grader measured at κ = 0.05 outranks an unmeasured one. That is deliberate - a measured floor is a checkable claim and an unmeasured one is not - but it means measuring badly buys a rung. It also means this ladder cannot represent the κ = -0.23 anti-correlated case at all: that has to be written as ρ = 0, indistinguishable from never having looked. See gap G3.
That anti-correlated case is a real measurement, and it came out of the transport study below rather than out of anything on this axis. rho’s information rate is a pure function of κ with no data in it at all, so the curve below is the formula and nothing else. It’s also even in κ: the rate at κ = -0.23 and the rate at κ = +0.23 are identical.
hash in the catalog. rho = 1 by construction, and it tells you one bit about one file that transports nowhere (T 0, and a ceiling of U 0 on anything resting on it). It witnesses S as well as rho. The orthogonality claim is therefore the pairwise one: no axis is derivable from the others.
rho collapses to a single scalar only for stochastic degradation. Two other failure modes exist and this axis is blind to both: contrast degeneracy (zero divergence on the specific contrast, no n fixes it) and construct substitution (κ near 1, sign inverted), which is the κ = -0.23 result above and the reason this axis is not a scalar.
sigma: independence
Under conditional independence given truth, the bits you double-count when pooling two checkers equal exactly I(Y1; Y2). Observed agreement above a1a2 + (1-a1)(1-a2) is co-bias and is worth zero bits.
sigma0, self-reported
sigma1, self-eval
sigma2, same-family check
sigma3, independent internal review
sigma4, cross-provider / independent team
sigma5, liability-bearing third party
sigma6, adversarial collaboration, joint prereg
shortseller in the catalog. No method, no identification, unreliable instrument, sigma 5 - and it moves prices, because the author has capital at risk.
Wrong if the checker bears no realized loss. If nobody has ever been penalized for a bad check, σ is nominal regardless of the checker's title - which is precisely how captured auditors and rating agencies present as high-σ and score near zero.
Where they fail together
A captured auditor scores high on rho. Its judgments are internally consistent and well calibrated on the cases it has already seen. It scores near zero on sigma, because it was never independent of what it was checking. Read rho alone and it looks verified. The fold takes the minimum, so a near-zero sigma caps it. Agreement between correlated checkers, beyond what independence predicts, is the excess sigma’s identity prices at zero bits, and a captured auditor earns it in quantity.
Two checkers, 0.80 and 0.80 accurate. The diagonal cells carry an illustrative 0.03 of excess joint probability beyond the independence prediction. The off-diagonal cells give up exactly that much, so both checkers’ individual accuracies are unchanged. That excess implies a pairwise mutual information of about 0.023 bits between the two checkers. That is a distinct quantity in a distinct unit, not the same number as the excess. Per sigma’s own identity those bits are exactly what gets double-counted when the two checkers are pooled, so they buy zero bits of independent confirmation. The mutual information itself is not zero. It’s the size of the apparent agreement that is not independent evidence.
Whether a given κ is already η or still needs squaring is settled by quality-sgd’s own reliability ledger, which we haven’t read yet. Guessing wrong runs one way: η is a contraction coefficient in [0, 1], so squaring a value that was already η returns η^2, at or below η. Double-squaring can understate reliability and never overstate it.
sigma has the same exposure: if agreement beyond what independence predicts never turns out to predict a downstream error, the correction the co-bias identity supplies is dead weight.
Alone, rho and sigma assess the reading, not the underlying quantity.
T: does the evidence cover where the decision lives
Does the region the evidence covers contain the region the decision lives in? Transport grades where the evidence was gathered against where the decision will act. That’s a different question from how hard the evidence was tested, or how reliable the instrument reading it was.
We model transfer as a product: transfer ~ (construct preservation) x (evidence self-containment). On a held-out sample of Cochrane risk-of-bias assessments, the dimension whose supporting evidence sat outside the graded artifact scored κ = -0.23, so the sign inverted rather than degrading toward chance. No dimension reached the human agreement ceiling, and that ceiling is itself inflated. The other two degraded toward it instead, and the two of them ordered the way that product predicts: κ = 0.62 where the construct means the same thing in both domains and the paper states it outright, κ = 0.23 where the construct is adjacent and the evidence has to be inferred.
Transport fails two ways, either enough alone: the claim measures a different construct than your decision needs, or its support never covers where your decision lives. The levels below are keyed to the second failure.
What T asks for is containment, not similarity. A similarity metric is symmetric - it reports how close two things are, and closeness runs the same in both directions. A region can contain another without being contained by it. Evidence whose coverage is a superset of your decision’s region transports, because every configuration the decision can land in was already exercised somewhere inside the evidence. Evidence that merely sits near your decision, however that nearness is measured, justifies no claim about the configurations the sample never visited.
A distance metric cannot represent this failure. It can report that two regions are close on average. What it has no vocabulary for is the actual problem: the decision lives partly where the evidence never looked. Averaging distance over the part that was covered hides exactly the part that was not. The relation is the wrong one only if the underlying regularity is genuinely global. If the same law holds everywhere, any sample transports regardless of where it was drawn, and T collapses to a constant with nothing left to discriminate.
T0, different domain, unstated
T1, different population
T2, same population, stale period
T3, current, but off the decision support
T4, partial support coverage
T5, covers the decision support
T6, measured at deployment fraction 1
One evidence region, held fixed across all three panels, against a decision region that moves: inside it, straddling its edge, or entirely outside it. Only the contained case transports, and the disjoint case clears the sum of the two radii with margin to spare rather than merely sitting mostly outside.
telemetry in the catalog: T 5, I 0. The record matches the population and period the decision lives in, and identifies nothing. It cannot also serve as the I witness, since a record sitting at I 0 witnesses identification only by its absence; obs-slope is the I witness, preregistered and blinded and powered and independently run and still I 0, because nothing was intervened on.
Breaks in reflexive domains, where the sign inverts: more live exposure CONSUMES the regularity rather than confirming it. Pricing is reflexive by construction.
You act on the estimate, and acting on it changes the thing it estimated. Competitors reprice against the move; customers adapt to the new level. Every period the estimate sets a price, the market that generated that number stops existing in the form it was measured in. Your next reading comes off a population already conditioned by your last decision.
Composition: the minimum of five scores
Composition isn’t a seventh axis. It’s the rule we use to fold I, S, rho, sigma, T into a single level, and that fold is their minimum - reported with whichever of them sit at that level, because which axis binds is what tells you what to fix.
We run a gate before any of the five gets weighted. It asks whether the evidence’s construct domain - empirical, formal, synthetic, or any - is one your decision admits. A mismatch doesn’t get a discount. The gate drops the support, and a warrant left with nothing admitted is refused outright. A machine-checked proof about code doesn’t by itself provide evidence about a demand curve, whatever its reliability.
Call an aggregator sound when it never exceeds any of its five inputs. Soundness alone doesn’t single out the minimum: the function that always returns zero is sound too, and it separates nothing. So we ask for both properties at once, sound and greatest among the sound ones. Only the componentwise minimum has both, since a sound aggregator is bounded above by its smallest input and the minimum attains that bound. An average has neither. It can exceed its smallest coordinate, and it scores a claim strong on reliability and weak on severity exactly like those two numbers reversed, though the axis that binds is different.
We exclude U from the fold by construction. It measures how much a claim commits to, not how well the claim survived checking. A claim reported below what its evidence supports is under-sold rather than untrustworthy.
The fold compares level indices, so it’s safe only under a reparameterization that relabels all five ladders the same way. Re-cut one ladder’s cut points and you can move which axis sits at the minimum with no change in the underlying evidence. Every level the composition section names in prose or a caption is printed beside the ladder’s own words for it. The figures print the index alone, because a graph node gets one short line.
8 of the catalog’s 33 kinds have more than one axis tied at the minimum. One of them is llm1, a single unmeasured LLM-as-judge verdict: its meet is level 0, with I (Association) and rho (unmeasured) tied there together. Break that tie by whichever axis sits first in an array and you advise “fix identification” when the instrument reading the claim is equally binding. So the fold returns the full set of axes at the minimum.
llm1’s five scored coordinates. The dashed rose outline marks the axes tied at the meet. The solid violet line marks the arithmetic mean of the same five numbers - the aggregator this section argues against, reporting 0.60 where the floor is 0.
In a conjunctive chain every step has to hold for the conclusion to hold, so the weakest link caps the whole chain however strong the others are. Two of the gate’s checks are specific to chaining. Its four edge coordinates have to equal the fold of the steps it declares. On every edge axis, that fold also has to sit at or below the fold of what its admitted supports license. Attach one more in-domain support that is weaker on some edge axis and the right-hand side drops, so a chain that passed can start failing. Adding evidence can lower what a set licenses, because the bound is a minimum. A chain that clears both checks can still be refused by the identification cap, the U ceiling, the decision’s axis gates or its meet floor, each on its own.
Parallel support pools by the same minimum fold, as a deliberately conservative placeholder: it credits independence nowhere. Three supports that vary in three different ways pool to the componentwise minimum, which sits at or below all three of them, exactly as a chain of three steps would. Take the componentwise maximum instead and a support strong on one axis paired with a support strong on another buys back coverage neither one earned. Shared source data do discount the effective count, without changing the score vector. Supports tracing back to the same underlying panel are close to being one support. A review that counts citations instead of sources overstates its own base, always in the optimistic direction.
Left, a three-step chain: two near-ceiling steps and one weak one fold to S 1 (anecdote), rho 2 (marginal), sigma 1 (self-eval), T 2 (same population, stale period), the weak step’s own vector. Right, three parallel supports fan into one pooled vector of S 4 (held-out), rho 4 (κ >= 0.8), sigma 4 (cross-provider / independent team), T 4 (partial support coverage), at an effective n of 1.8 rather than the raw count of three.
Severity itself doesn’t compose under this rule (G1), and the fix that gap names is an e-process. Optional stopping is the ordinary practice of watching evidence accumulate and deciding, from what you see, whether to keep collecting. It invalidates a fixed-sample p-value outright. That guarantee assumes the sample size was fixed in advance, and a stopping decision made in light of the running result breaks it however innocuous the peek felt. An e-value is a nonnegative statistic whose expected value under the null is at most one. A running product of sequentially valid e-values is a nonnegative supermartingale. Ville’s inequality bounds it at every stopping time at once, so you stay valid under exactly the monitoring that breaks p-values.
Each claim you fold this way is a node in a content-addressed proof graph, hashed to the code and data that produced it.
What we have not solved
Each gap carries a date and a falsifier: what would have to happen for it to close or be shown wrong. Some falsifiers name a test that is nearly free to run; G6’s says outright that none is known. All 10 carry the same date today, because the register was cut in one pass. The date is per gap rather than per register so they can come apart as gaps are re-examined: G4 is close to closed and G6 may never close.
Each of the four edge axes names a condition under which that axis is the wrong tool, and those sit in the axis sections above. Transport’s is where you find that its sign inverts in reflexive domains, pricing among them. U and I carry no equivalent. U states, level by level, what that level can’t say and what won’t promote a claim to it; I states what each level licenses. Both are useful, and neither is a claim about when the axis itself is the wrong construct. Two of the gaps below restate edge-axis conditions: G3 for rho, G4 for severity.
| ID | Gap | The cheapest test that would show it wrong | As of |
|---|---|---|---|
| G1 | Severity does not compose Minimum is right for a conjunctive chain and wrong for redundant parallel support, and no operator handles partially correlated supports without an error-correlation matrix nobody has. | Re-express cliffs as e-processes and combine them. A sequential e-process stays valid under optional stopping, and the mean of several e-values is an e-value under arbitrary dependence, so parallel support can be combined without an error-correlation matrix. Multiplying them cannot: the product needs independence, or a filtration in which each factor is an e-process given the ones before it. | 2026-08-03 |
| G2 | The axis set may be overparameterised, and the correlation estimates are unstable Measured over 33 kinds: the co-moving pair is S~rho at r = 0.627 (partial 0.576), then sigma~T at 0.581 (partial 0.488). Eigenvalues 2.62 / 0.84 / 0.81 / 0.46 / 0.26, so PC1 carries 52.5% and the participation-ratio rank is 2.93 against a nominal 5; on the 28 kinds whose construct domain price admits (before gates and the meet floor), 63.2% at rank 2.26. Three-ish axes do the work of five. The second half of the gap is worse than the first: sweep every one of the 40,920 four-record removals from this catalog and the top co-moving pair comes out S~rho on 29,617 of them, sigma~T on 8,613, and one of four other pairs on the remaining 2,690. Four records out of 33 can change which pair the diagnostic names. These numbers describe the exemplar set. | NOT a corpus correlation - that instrument fails twice, being unstable to catalog composition and unable to separate ladders whose rung definitions share cliffs. Use a discriminant-validity test: synthesize matched pairs holding every cliff on one axis fixed while toggling only the other, and check whether independent labelers move one without the other. Then label a real corpus - our own reliability ledger plus the held-out Cochrane sample - and compute the scree over REALIZED CLAIMS rather than authored kinds. | 2026-08-03 |
| G3 | rho is the wrong shape A scalar sees only stochastic degradation. Contrast degeneracy and construct substitution are invisible to it, and the second is the failure mode that motivated the axis. | Index rho by (evaluator, cliff) and estimate it as blind AUC on synthesized matched pairs that toggle exactly that cliff. | 2026-08-03 |
| G4 | Realised severity is search-discounted and the search is not logged Our severity computation has n_hypotheses = 1 baked in as an unstated auxiliary, so every score it has ever emitted is an upper bound. RoB-2 domain 5 and ROBINS-I have both scored "bias in selection of the reported result" since 2019. The narrower true gap is that they score it by human judgement, one study at a time, and no machine-readable forking-path register exists. | Log the register - for an agent it is a log rather than a judgement, so nearly free - then report between-run SD across ten runs on a frozen panel. | 2026-08-03 |
| G5 | Claim identity has no standard form Nothing decides whether two propositions are the same proposition, and dedup, correlation, and min-cut all inherit the error. | No good answer. Human confirmation at authoring time, and a round-trip rewrite into a standard form measured above 0.95 agreement before trusting any automated dedup. | 2026-08-03 |
| G6 | Blackwell has no completion for credal evidence Once information is a set of measures rather than a measure, "is E more informative than F" has no settled definition. Dominance on every axis, over hand-assigned levels, is a PROXY for Blackwell's order rather than that order, because no garbling kernel (an explicit noisy channel from one experiment to the other) is constructed. The intersection-of-decision-orders claim holds over the full weight simplex, not only the four decisions shipped here. | None known. No cheap test decides whether a Blackwell completion for credal evidence exists. | 2026-08-03 |
| G7 | There is no polarity coordinate, so refutation has nowhere to live All six axes measure how well evidence SUPPORTS a claim; nothing represents evidence that kills one. Express a counterexample's veto force by inflating its support axes and it ranks first under all four decisions, beating an RCT at setting a price. We corrected that record's coordinates. The missing dimension is still missing. Related: min/argmin over rung indices is ordinal-safe but not scale-invariant across axes, so re-cutting one ladder can move the headline. | Declaring polarity means relating evidence to a target proposition, and nothing yet decides when two claims are the same claim (G5). The scale half has a real fix: put all five axes on one numeric scale in bits - rho and sigma already are, S via the log-likelihood ratio a passed cliff supplies, T via the KL between evidence region and decision support - and take the min over bits. That conversion is its own implementation and validation project. | 2026-08-03 |
| G8 | No axis for cross-study inconsistency Six coordinates describe one body of evidence and cannot express "five studies, three positive, two negative" - the most common downgrade reason in GRADE. Pooling is therefore a permanent stub, since averaging coordinates across supports is exactly the buy-back the composition rule forbids. | Either a seventh axis for realized heterogeneity, or the position that this frame types a single body of evidence and meta-analytic structure sits a level above it. The frame currently treats meta-analytic structure as a layer above evidence typing, which is why pooling two supports is still reported as not derived. | 2026-08-03 |
| G9 | Two of the four decisions may not be distinguishable Kendall tau(ship, publish) = 0.934 over the 14 kinds both admit; underwrite and publish agree perfectly on the 7 they share. If naming a decision is what turns the partial order into a complete ranking, decisions producing the same ranking are not different decisions, and "the reordering is the payload" is thinner than four columns suggest. | A ~12-pair Bradley-Terry elicitation over these kinds replaces assertion with measurement and settles whether the near-identical pairs survive contact with a real preference. | 2026-08-03 |
| G10 | The ceiling on U may be a restatement of I r(maxU, I) = 0.566 over the catalog, where maxU is the highest U a claim resting on that record may assert. The case for excluding U from the meet rests on U entering through the ceiling check instead; if the ceiling is mostly the identification rung wearing a different name, that defence is weaker than it reads. | Score maxU and I independently, blind, on the same kinds and check whether labelers move one without the other. Same instrument as G2's falsifier, so it is nearly free to run alongside. | 2026-08-03 |
Where to look next
We score and order 33 kinds of evidence next door, by dominance on every axis (componentwise dominance). It names no decision, and no positive weighting of the 5 scored axes can reverse a pair it decides. 262 of the 528 pairs order that way and 266 do not, so 49.6% is decidable before you pick a loss function. The rest need weights, and the 4 scored decisions there each carry a different set. Read the evidence catalog.
Rosetta Stone
Four circles, four readings of the same object. Each role reads the artifact through its own lens.
Evaluate the evidence behind an opportunity's inputs before you use those inputs to rank it. A single elasticity estimate and a preregistered field experiment can carry the same headline number and completely different admissibility, and you can't tell them apart from the number alone.
A widely repeated metric should record who validated it, on which population, and whether that validation transports to the decision in front of you. Most of them record none of it.
If you're wiring an eval or a verifier into a pipeline, this is the type system underneath 'trustworthy output': six coordinates, a construct-domain gate that runs before any of them get scored, and a composition rule that takes the minimum of five.
The register carries 10 dated gaps, one of them arguing that three-ish axes do the work of five. Read them before relying on any one axis.