Scoring a warning rule the way a warning rule should be scored: objects not grid cells, an explicit opportunity grid for the correct negatives, block-bootstrap confidence intervals, and a verdict that can say no.
Rule scored: issue a WARNING when forecast gust reaches 65 km/h, against 2020–2024 observed gales at London. The WARNING gate is precision ≥ 0.20 with the 95% CI lower bound reaching it. Verdict: SHIP — precision 0.392 [0.278, 0.489], recall 0.872, median lead time 3.0 h, frequency bias 2.06, over 47 events and 97 episodes. Clears every gate.
The same hazard at a lower trigger, scored as a WATCH, where the gate is recall ≥ 0.80. Verdict: SHIP WITH CAVEATS — recall 0.915 [0.823, 0.983] passes, and lead time improves to 4.0 h, but the rule fires 162 episodes for 47 events: frequency bias 3.45 against a wanted range of 0.5–3, so three in four alerts are false. Hard gates cleared, bias flagged rather than hidden.
A probability model fitted on one five-year window and verified out-of-sample on the next, so the verification data was never trained on. Verdict: SHIP WITH CAVEATS — it catches every one of the 47 events (recall 1.000) with a median 11 h of lead time, but at the cost of 419 episodes: frequency bias 8.91. Perfect recall bought with nine times too many alerts is exactly the trade the bias gate exists to surface.
The control case: a synthetic signal whose ground truth is known by construction, so the scoring machinery itself can be checked. Verdict: SHIP — 268 events, recall 1.000, median lead time 13 h, frequency bias 0.70, all gates passed. Not real weather; it exists to prove the scorer, not a rule.