S: what a false claim would have had to survive
How hard did the producing transition try to refute this, and did it survive? Severity types the producing transition, not the quantity it targets - a different question from identification, and the two axes interact in a specific, measured way below.
The formal frame
Sev(C; test, x) = P(test yields a worse fit for C | C is false). Each rung closes one path by which a false claim could have fit anyway.
Each of the seven rungs below closes one specific route by which a false claim could still have produced a good fit. The rung name is a label; the closed route is the argument.
The seven rungs, and the path each closes
Pricing figures throughout are synthetic and chosen to make the type of the claim legible; none of them are measurements.
Why severity forces identification here
obs-slope / negcontrol in the catalog. Preregistered, blinded, powered, independently run - and zero identification, because nothing was intervened on. Note the cap: S 5 ("live + powered") and S 6 ("durable") presuppose a deployed randomized intervention, so on the shipped ladders S >= 5 forces I >= 5 for every kind except negcontrol. The witness is therefore S 4 with I 0, not "maximal S with zero I", which is unconstructible here and was overclaimed in the first draft.
The forking-path gap
Realised severity is search-discounted and the search is not logged. Our severity computation has n_hypotheses = 1 baked in as an unstated auxiliary, so every score it has ever emitted is an upper bound. Credit where due, since the first draft implied nobody had noticed: RoB-2 domain 5 and ROBINS-I have both scored "bias in selection of the reported result" since 2019. The narrower true gap is that they score it by human judgement, one study at a time, and no machine-readable forking-path register exists.
Log the register - for an agent it is a log rather than a judgement, so nearly free - then report between-run SD across ten runs on a frozen panel.
An agent that searches specifications runs that search far faster than an analyst ever could by hand, which makes keeping the forking-path register nearly free: for an agent, the register is a log, not a judgement. That is why G4’s falsifier is cheap, and why the register’s absence today is a choice rather than a constraint.
Where severity itself breaks
Breaks if realized severity is not discounted by search: a test with one-shot severity 0.95 has realized severity near zero if it is the surviving member of an unlogged specification search.