← The Nature of Knowledge

S: what a false claim would have had to survive

How hard did the producing transition try to refute this, and did it survive? Severity types the producing transition, not the quantity it targets - a different question from identification, and the two axes interact in a specific, measured way below.

The formal frame

Sev(C; test, x) = P(test yields a worse fit for C | C is false). Each rung closes one path by which a false claim could have fit anyway.

Each of the seven rungs below closes one specific route by which a false claim could still have produced a good fit. The rung name is a label; the closed route is the argument.

The seven rungs, and the path each closes

S0 unforgedS1 anecdoteS2 pre-registeredS3 controlledS4 held-outS5 live + poweredS6 durableNo test has run yet: a false claim fits because nothing wasever asked to predict something specific enough to fail.One case was checked and it matched. Closes the path wherethe claim fails its own most convenient example; selectionand a missing comparison stay open.Hypothesis and analysis plan fixed before the data arrived.Closes the path where a hypothesis or cutoff gets chosenafter seeing which way the data went.Compared against a control holding the same confounds fixed.Closes the path where an untreated third factor, not theclaim, produced the pattern.Scored on data the claim never touched while it was built ortuned. Closes the path where a claim was quietly fit toflatter its own evaluation sample.Deployed live, at a scale powered to catch the error thatwould matter. Closes the path where a claim survives onlybecause the check was too small to catch it failing.Repeated across independent live runs and held every time.Closes the path where one powered result was itself a luckydraw that would not survive a second attempt.RUNGS RUN S0 (LEAST SEVERE) AT THE BOTTOM TO S6 (MOST SEVERE) AT THE TOP

Pricing figures throughout are synthetic and chosen to make the type of the claim legible; none of them are measurements.

Why severity forces identification here

obs-slope / negcontrol in the catalog. Preregistered, blinded, powered, independently run - and zero identification, because nothing was intervened on. Note the cap: S 5 ("live + powered") and S 6 ("durable") presuppose a deployed randomized intervention, so on the shipped ladders S >= 5 forces I >= 5 for every kind except negcontrol. The witness is therefore S 4 with I 0, not "maximal S with zero I", which is unconstructible here and was overclaimed in the first draft.

The forking-path gap

Realised severity is search-discounted and the search is not logged. Our severity computation has n_hypotheses = 1 baked in as an unstated auxiliary, so every score it has ever emitted is an upper bound. Credit where due, since the first draft implied nobody had noticed: RoB-2 domain 5 and ROBINS-I have both scored "bias in selection of the reported result" since 2019. The narrower true gap is that they score it by human judgement, one study at a time, and no machine-readable forking-path register exists.

Log the register - for an agent it is a log rather than a judgement, so nearly free - then report between-run SD across ten runs on a frozen panel.

An agent that searches specifications runs that search far faster than an analyst ever could by hand, which makes keeping the forking-path register nearly free: for an agent, the register is a log, not a judgement. That is why G4’s falsifier is cheap, and why the register’s absence today is a choice rather than a constraint.

Where severity itself breaks

Breaks if realized severity is not discounted by search: a test with one-shot severity 0.95 has realized severity near zero if it is the surviving member of an unlogged specification search.