← The Nature of Knowledge

What this frame cannot do yet

This page is the risk register for the frame described across the other eight. Those pages assert what the six axes measure and how they compose; this one lists what the frame cannot yet do, dated per gap rather than per page, because they do not resolve on the same clock. G4 is close to closed; G6 may never close, and dating the page as a whole would make the closed gap look as open as the far one. The organizing test is falsifiability turned on the frame itself: a model that publishes its own failure conditions is making a claim someone could show wrong, and a model that omits them is making no claim at all, only a description. What this register carries is the cross-cutting gaps. The four edge axes additionally carry an axis-level falsifier, naming the condition under which that one axis breaks, and those render on the axis’s own route rather than here; transport’s is where you find that its sign inverts in reflexive domains, pricing among them. The two node axes carry nothing equivalent: order of uncertainty states, rung by rung, what that rung cannot say and what fails to promote you to it, and identification states only what each rung licenses. Useful, but neither is a claim about when the axis itself is the wrong construct. G2 is the gap whose falsifier would decide which of the six axes to cut. It is read in full below, not summarized here.

11 gaps, each dated on its own clock

IDGapCheapest falsifierAs of
G1

Severity does not compose

Minimum is right for a conjunctive chain and wrong for redundant parallel support, and no operator handles partially correlated supports without an error-correlation matrix nobody has.

Re-express cliffs as e-processes and multiply. E-values are valid under optional stopping and combine under arbitrary dependence.2026-08-03
G2

The axis set may be overparameterised, and the correlation estimates are unstable

Measured over 33 kinds: the co-moving pair is S~rho at r = 0.627 (partial 0.576), then sigma~T at 0.581 (partial 0.488). Eigenvalues 2.62 / 0.84 / 0.81 / 0.46 / 0.26, so PC1 carries 52.5% and the participation-ratio rank is 2.93 against a nominal 5; on the 28 kinds whose construct domain price admits (before gates and the meet floor), 63.2% at rank 2.26. Three-ish axes do the work of five. The second half of the gap is worse than the first: sweep every one of the 40,920 four-record removals from this catalog and the top co-moving pair comes out S~rho on 29,617 of them, sigma~T on 8,613, and one of four other pairs on the remaining 2,690. Four records out of 33 can change which pair the diagnostic names. These numbers describe the exemplar set.

Measured over 33 kinds, 262 of the 528 pairs are comparable and 266 are not, which is 49.6%.

NOT a corpus correlation - that instrument fails twice, being unstable to catalog composition and unable to separate ladders whose rung definitions share cliffs. Use a discriminant-validity test: synthesize matched pairs holding every cliff on one axis fixed while toggling only the other, and check whether independent labelers move one without the other. Then label a real corpus - our own reliability ledger plus the held-out Cochrane sample - and compute the scree over REALIZED CLAIMS rather than authored kinds.2026-08-03
G3

rho is the wrong shape

A scalar sees only stochastic degradation. Contrast degeneracy and construct substitution are invisible to it, and the second is the failure mode that motivated the axis.

Index rho by (evaluator, cliff) and estimate it as blind AUC on synthesized matched pairs that toggle exactly that cliff.2026-08-03
G4

Realised severity is search-discounted and the search is not logged

Our severity computation has n_hypotheses = 1 baked in as an unstated auxiliary, so every score it has ever emitted is an upper bound. Credit where due, since the first draft implied nobody had noticed: RoB-2 domain 5 and ROBINS-I have both scored "bias in selection of the reported result" since 2019. The narrower true gap is that they score it by human judgement, one study at a time, and no machine-readable forking-path register exists.

Log the register - for an agent it is a log rather than a judgement, so nearly free - then report between-run SD across ten runs on a frozen panel.2026-08-03
G5

Claim identity has no normaliser

Nothing decides whether two propositions are the same proposition, and dedup, correlation, and min-cut all inherit the error.

No good answer. Human confirmation at authoring time, and a round-trip normaliser measured above 0.95 agreement before trusting any automated dedup.2026-08-03
G6

Blackwell has no completion for credal evidence

Once information is a set of measures rather than a measure, "is E more informative than F" has no settled definition. Two adjacent honesty notes belong with it: componentwise dominance on hand-assigned rungs is a PROXY for Blackwell's order rather than that order, since no garbling kernel is constructed; and the intersection-of-decision-orders claim holds over the full weight simplex, not over the four decisions shipped here.

None known. Stated on the page rather than papered over.2026-08-03
G7

There is no polarity coordinate, so refutation has nowhere to live

All six axes measure how well evidence SUPPORTS a claim; nothing represents evidence that kills one. The first draft expressed a counterexample's veto force by inflating its support axes, and it ranked first under all four decisions - beating an RCT at setting a price. The coordinates were corrected; the missing dimension was not. Related: min/argmin over rung indices is ordinal-safe but not scale-invariant across axes, so re-cutting one ladder can move the headline.

Polarity is cheap to declare and expensive to mean, because it is a relation to a target proposition and claim identity has no normaliser (G5). The scale half has a real fix: cardinalize all five axes into bits - rho and sigma already are, S via the log-likelihood ratio a passed cliff supplies, T via the KL between evidence region and decision support - and take the min over bits. That is a project, not an edit.2026-08-03
G8

No axis for cross-study inconsistency

Six coordinates describe one body of evidence and cannot express "five studies, three positive, two negative" - the most common downgrade reason in GRADE. Pooling is therefore a permanent stub, since averaging coordinates across supports is exactly the buy-back the composition rule forbids.

Either a seventh axis for realized heterogeneity, or the position that this frame types a single body of evidence and meta-analytic structure sits a level above it. The second is more honest and is the current stance; it is also the reason poolSupports returns derived: false.2026-08-03
G9

Two of the four decisions may not be distinguishable

Kendall tau(ship, publish) = 0.934 over the 14 kinds both admit; underwrite and publish agree perfectly on the 7 they share. If naming a decision is what totalizes the order, decisions producing the same order are not different decisions, and "the reordering is the payload" is thinner than four columns suggest.

A ~12-pair Bradley-Terry elicitation over these kinds replaces assertion with measurement and settles whether the near-identical pairs survive contact with a real preference. This is the same elicitation 4.5 defers, so G9 is its strongest justification.2026-08-03
G10

maxU may be a restatement of I

r(maxU, I) = 0.566 over the catalog. Section 4.4 defends excluding U from the meet on the grounds that U enters through the ceiling check instead; if the ceiling is mostly the identification rung wearing a different name, that defence is weaker than it reads.

Score maxU and I independently, blind, on the same kinds and check whether labelers move one without the other. Same instrument as G2's falsifier, so it is nearly free to run alongside.2026-08-03
G11

A chain derivation binds only part of its claim to its evidence

admit binds a claim to its supports differently for each derivation tag, and for one of the three the binding covers only part of the vector. A single derivation must sit componentwise at or below its one admitted support across every scored axis. A chain and a pool are checked against a fold of edge vectors, which reaches the four EDGE axes only, since identification is not an edge axis and that comparison cannot see it. For a pool the fold is computed FROM the admitted supports, so a pool is bound. For a chain it is computed from the steps the caller declares, and nothing ties those declared steps to the supports, so the gate will admit a claim carrying maximal severity, instrument reliability, adversarial pressure and transport while standing on a real support that licenses far less. Measured over the shipped catalog and decisions, 48 of 132 (kind, decision) pairs admit at maximal edge coordinates; in the worst case the honest derivation is refused outright while the unbound fold admits at score 5.85, so the exploit does not buy a better score than an honest claim - it buys an admission where honesty gets none. Two axes still bind: identification to the minimum its supports license, and order of uncertainty to their ceiling. That the identical hole once existed for identification and was closed with exactly that conjunctive minimum is both the precedent for closing it here and the reason it stays an open question rather than an oversight. The declared steps may each be independently attested, in which case the fold is the warrant and the support list is context.

Extend the conjunctive bound already used for identification to the edge axes: require the fold of the declared steps to sit componentwise at or below the componentwise minimum over the admitted supports own edge vectors, then re-run every case the gate currently admits. If none is refused, the uncapped reading costs nothing to abandon and this closes with one comparison in a shape the module already uses. If a legitimate chain is refused, the declared steps do stand on their own attestation and the cap is too strong - in which case a chain warrant support list is documentation rather than warrant, and the type should say so.2026-08-04