London gale WARNING (gust >= 65 km/h)
WARNING · 2020-01-01 to 2025-01-01 · 47 events · 97 warning episodes
SHIP
Clears every gate for a WARNING.
Gates
PASSsample size47 observed events (need >= 20)
PASSskill (PSS)PSS = 0.830 (need > 0); grid slot 1:00:00
PASSWARNING objective (precision)precision = 0.392 (need >= 0.20)
PASSprecision lower CI bound95% CI 0.392 [0.278, 0.489] (lower bound must reach 0.20)
PASSmedian lead timemedian 3.0 h, IQR 2.0-5.0 h (need >= 1 h)
PASSfrequency biasbias = 2.06 (want 0.5-3); 97 warning episodes for 47 events
Metrics
| metric | value (95% CI) | counted over | definition |
|---|---|---|---|
| recall | 0.872 [0.756, 0.958] | events | a/(a+c) of the events, the share that were warned |
| precision | 0.392 [0.278, 0.489] | episodes | a/(a+b) of the warnings issued, the share that verified |
| far | 0.608 [0.511, 0.722] | episodes | b/(a+b) false alarm RATIO (= 1 - precision) |
| f1 | 0.541 [0.411, 0.638] | mixed | harmonic mean of precision and recall |
| csi | 0.387 [0.273, 0.490] | mixed | a/(a+b+c) critical success index / threat score |
| bias | 2.064 [1.677, 2.774] | episodes/events | (a+b)/(a+c) >1 over-warns, <1 under-warns |
| pofd | 0.042 [0.032, 0.053] | grid | b/(b+d) false alarm RATE (needs correct negatives) |
| hss | 0.040 [0.030, 0.048] | grid | Heidke skill score, 0 = chance, 1 = perfect |
| pss | 0.830 [0.717, 0.917] | grid | Peirce skill score = POD - POFD, base-rate insensitive |
| ets | 0.021 [0.015, 0.025] | grid | equitable threat score (CSI corrected for chance hits) |
Object counts
| counted over | total | right | wrong | drives |
|---|---|---|---|---|
| events | 47 | 41 hits | 6 misses | recall / POD |
| warning episodes | 97 | 38 verified | 59 false alarms | precision / FAR |
| raw alerts | 608 | merged into 97 episodes | nothing directly — merging prevents double counting | |
Multiplicity 1.08 hits per verified episode. At 1.00 the two denominators agree and every object score is unambiguous; well above 1.00 means long episodes are sweeping up several events at once.
Contingency tables
| table | hits | false alarms | misses | correct negatives |
|---|---|---|---|---|
| object, classical mixed | 41 | 59 | 6 | n/a by design |
| grid (slot 1:00:00) | 41 | 1852 | 6 | 41949 |
Lead time
median 3.0 h, IQR 2.0–5.0 h, range 1.0–13.0 h over 41 hits.
Diagnostics
Plain text scorecard
==============================================================================
warncal scorecard -- London gale WARNING (gust >= 65 km/h)
role: WARNING period: 2020-01-01 to 2025-01-01
==============================================================================
VERDICT: SHIP
Clears every gate for a WARNING.
------------------------------------------------------------------------------
GATES
[PASS] sample size
47 observed events (need >= 20)
[PASS] skill (PSS)
PSS = 0.830 (need > 0); grid slot 1:00:00
[PASS] WARNING objective (precision)
precision = 0.392 (need >= 0.20)
[PASS] precision lower CI bound
95% CI 0.392 [0.278, 0.489] (lower bound must reach 0.20)
[PASS] median lead time
median 3.0 h, IQR 2.0-5.0 h (need >= 1 h)
[PASS] frequency bias
bias = 2.06 (want 0.5-3); 97 warning episodes for 47 events
------------------------------------------------------------------------------
OBJECT COUNTS (overlapping alerts merged into episodes)
events 47 = 41 hits + 6 misses -> recall
episodes 97 = 38 verified + 59 false alarms -> precision
raw alerts 608 merged into 97 episodes
multiplicity 1.08 hits per verified episode (1.00 = clean units)
Classical mixed table, for comparison with published studies:
event yes event no
warned yes 41 59
warned no 6 n/a
OPPORTUNITY GRID (slot = 1:00:00) -- the only source of correct negatives
event yes event no
warned yes 41 1852
warned no 6 41949
------------------------------------------------------------------------------
METRICS (95% block-bootstrap CI where available)
metric value [95% CI] counted over
recall 0.872 [0.756, 0.958] events
precision 0.392 [0.278, 0.489] episodes
far 0.608 [0.511, 0.722] episodes
f1 0.541 [0.411, 0.638] mixed
csi 0.387 [0.273, 0.490] mixed
bias 2.064 [1.677, 2.774] episodes/events
pofd 0.042 [0.032, 0.053] grid
hss 0.040 [0.030, 0.048] grid
pss 0.830 [0.717, 0.917] grid
ets 0.021 [0.015, 0.025] grid
base rate 0.00107 grid
------------------------------------------------------------------------------
LEAD TIME (event onset minus episode issue time, hits only)
median 3.0 h IQR 2.0 to 5.0 h range 1.0 to 13.0 h
median 95% CI [2.0, 4.0] h
1.0 h |█▆▅▃▃▂▁▁▂▁ ▃| 13.0 h (n=41)
------------------------------------------------------------------------------
warncal scores warning rules. It is a measurement tool, not a decision
authority: it is not medical, clinical or safety advice, it is not an
emergency service, and a favourable verdict here does not make a rule safe
to deploy. A human who understands the hazard must make that call.
==============================================================================
Not advice. warncal scores warning rules. It is a measurement tool, not a decision authority: it is not medical, clinical or safety advice, it is not an emergency service, and a favourable verdict here does not make a rule safe to deploy. A human who understands the hazard must make that call.