Why trust meets and does not average
Composition is not a sixth axis. It is the rule that folds the five scored axes, identification, severity, instrument reliability, adversarial pressure, and transport, into the one number a decision actually uses: their minimum, never their average. The choice is forced, the set of axes tied at that minimum matters as much as the minimum itself, and the rule this page is honest about not yet having is the one that pools redundant parallel support.
Why minimum, and not merely “a” safe rule
Soundness alone does not single out the minimum: the function that always returns zero is sound in the same sense, and it is useless for the reason it is safe. What forces meet to be the minimum is a second requirement stacked on the first: among every sound aggregator, the composite must be the greatest one, the tightest bound that still never over-claims. Minimum is the unique function meeting both halves at once, over the five scored axes only, I, S, rho, sigma, T. Uncertainty is excluded by construction, since it measures how much a claim commits to, not how well it survived checking, and a claim reported below what its evidence supports is under-sold rather than untrustworthy.
One limit rides along with the result. Because the fold compares rung indices, it is safe only under a reparameterization that relabels all five ladders the same way: re-cutting one ladder’s cut points can move which axis is tied at the minimum with no change to the underlying evidence. Every rung this page names is printed beside its human-readable label rather than as a bare number for that reason, the label standing in for the invariance the fold does not have.
The headline is a set, not a scalar
8 of the catalog’s 33 kinds have more than one axis tied at the minimum. llm1, a single unmeasured LLM-as-judge verdict, is one: its meet is rung 0, with I and rho tied there together. Breaking a tie by whichever axis sits first in an array would silently advise “fix identification” when the instrument reading the claim is equally binding, advice about the wrong axis, manufactured by an implementation detail rather than the evidence. That is why the fold’s output is the full argmin set: a scalar headline would launder a tie into a false priority.
llm1’s five scored coordinates. The dashed rose outline marks both axes tied at the meet; the solid violet line marks the arithmetic mean of the same five numbers, the aggregator this page argues against, sitting well above the true floor.
Pricing figures throughout are synthetic and chosen to make the type of the claim legible; none of them are measurements.
Chains fold correctly; pooling does not, yet
A conjunctive chain, every step must hold for the conclusion to hold, is exactly what the minimum was built for: the weakest link caps the whole chain regardless of how strong the others are. Parallel redundant support is a different shape, and the package says so plainly: pooling today falls back to the same minimum fold, a deliberately conservative placeholder that never credits independence it has not verified, because the alternative, a componentwise maximum, would let a support strong on one axis and weak on another buy back exactly the coverage neither support alone earned. What the placeholder does add is a discount on the count for shared provenance: supports that trace back to the same underlying panel are close to being one support, and a review that counts citations instead of provenance overstates its own base, always in the optimistic direction.
Left, a four-step chain: two near-ceiling steps and one weak one fold to S 1, rho 2, sigma 1, T 2, the weak step’s own vector. Right, three parallel supports fan into one pooled vector of S 4, rho 4, sigma 4, T 4, at an effective n of 1.8 rather than the raw count of three.
The published gap
Severity does not compose. Minimum is right for a conjunctive chain and wrong for redundant parallel support, and no operator handles partially correlated supports without an error-correlation matrix nobody has.
Re-express cliffs as e-processes and multiply. E-values are valid under optional stopping and combine under arbitrary dependence.
What the falsifier is actually asking for
e-processes replace the minimum’s chain logic with a betting one. Optional stopping is the ordinary practice of watching accumulating evidence and deciding, from what you see so far, whether to keep collecting or stop, and it invalidates a fixed-sample p-value outright: the p-value’s guarantee assumes the sample size was fixed in advance, and a stopping decision made in light of the running result breaks that guarantee however innocuous the peek felt. An e-value sidesteps this by being a bet rather than a bound: a nonnegative statistic whose expected value under the null hypothesis is at most one, which makes its running product across successive tests a martingale, so its value at whatever moment you actually decide to stop is still bounded the same way a fixed-sample statistic would have been had you never peeked at all. e-values therefore stay valid under exactly the kind of monitoring that breaks p-values, and, unlike p-values, they combine by multiplication under arbitrary dependence between the tests being combined. That second property is the one the minimum lacks and the reason the gap names e-processes rather than a better pooling heuristic: a chain of e-values composes by multiplying, with no assumption about how correlated the underlying supports are. Averaging would buy back precisely the redundancy composition exists to forbid, and any fixed-sample bound on the pooled result already assumes a stopping rule nobody actually followed.
Where the claim and its evidence travel together
Every claim this page folds is itself a node in a content-addressed proof graph, hashed to the code and data that produced it; rebuilding that graph is out of scope for this page.