London gale, probabilistic (fitted 2015-2019, verified 2020-2024)
WATCH · 2020-01-01 to 2025-01-01 · 47 events · 419 warning episodes
SHIP WITH CAVEATS
Clears the hard gates; frequency bias is off.
Gates
PASSsample size47 observed events (need >= 20)
PASSskill (PSS)PSS = 0.757 (need > 0); grid slot 1:00:00
PASSWATCH objective (recall)recall = 1.000 (need >= 0.80)
PASSrecall lower CI bound95% CI 1.000 [1.000, 1.000] (lower bound must reach 0.80)
PASSmedian lead timemedian 11.0 h, IQR 6.0-13.0 h (need >= 1 h)
WARNfrequency biasbias = 8.91 (want 0.5-3); 419 warning episodes for 47 events
Metrics
| metric | value (95% CI) | counted over | definition |
|---|---|---|---|
| recall | 1.000 [1.000, 1.000] | events | a/(a+c) of the events, the share that were warned |
| precision | 0.093 [0.064, 0.121] | episodes | a/(a+b) of the warnings issued, the share that verified |
| far | 0.907 [0.879, 0.936] | episodes | b/(a+b) false alarm RATIO (= 1 - precision) |
| f1 | 0.170 [0.120, 0.217] | mixed | harmonic mean of precision and recall |
| csi | 0.110 [0.076, 0.144] | mixed | a/(a+b+c) critical success index / threat score |
| bias | 8.915 [6.727, 13.002] | episodes/events | (a+b)/(a+c) >1 over-warns, <1 under-warns |
| pofd | 0.243 [0.213, 0.273] | grid | b/(b+d) false alarm RATE (needs correct negatives) |
| hss | 0.007 [0.005, 0.009] | grid | Heidke skill score, 0 = chance, 1 = perfect |
| pss | 0.757 [0.727, 0.787] | grid | Peirce skill score = POD - POFD, base-rate insensitive |
| ets | 0.003 [0.002, 0.004] | grid | equitable threat score (CSI corrected for chance hits) |
Object counts
| counted over | total | right | wrong | drives |
|---|---|---|---|---|
| events | 47 | 47 hits | 0 misses | recall / POD |
| warning episodes | 419 | 39 verified | 380 false alarms | precision / FAR |
| raw alerts | 4043 | merged into 419 episodes | nothing directly — merging prevents double counting | |
Multiplicity 1.21 hits per verified episode. At 1.00 the two denominators agree and every object score is unambiguous; well above 1.00 means long episodes are sweeping up several events at once.
Contingency tables
| table | hits | false alarms | misses | correct negatives |
|---|---|---|---|---|
| object, classical mixed | 47 | 380 | 0 | n/a by design |
| grid (slot 1:00:00) | 47 | 10641 | 0 | 33160 |
Lead time
median 11.0 h, IQR 6.0–13.0 h, range 2.0–13.0 h over 47 hits.
Probabilistic
| bin | n | mean forecast | observed freq | 95% CI |
|---|---|---|---|---|
| 0.00–0.05 | 1018 | 0.047 | 0.025 | [0.017, 0.036] |
| 0.05–0.10 | 1121 | 0.066 | 0.060 | [0.047, 0.075] |
| 0.10–0.15 | 960 | 0.131 | 0.118 | [0.099, 0.140] |
| 0.15–0.20 | 943 | 0.173 | 0.195 | [0.171, 0.222] |
| 0.20–0.25 | 0 | – | – | – |
| 0.25–0.30 | 0 | – | – | – |
| 0.30–0.35 | 0 | – | – | – |
| 0.35–0.40 | 0 | – | – | – |
| 0.40–0.45 | 0 | – | – | – |
| 0.45–0.50 | 0 | – | – | – |
| 0.50–0.55 | 0 | – | – | – |
| 0.55–0.60 | 0 | – | – | – |
| 0.60–0.65 | 0 | – | – | – |
| 0.65–0.70 | 0 | – | – | – |
| 0.70–0.75 | 0 | – | – | – |
| 0.75–0.80 | 0 | – | – | – |
| 0.80–0.85 | 0 | – | – | – |
| 0.85–0.90 | 0 | – | – | – |
| 0.90–0.95 | 0 | – | – | – |
| 0.95–1.00 | 0 | – | – | – |
Brier 0.08321 · Brier skill 0.043 · ROC AUC 0.701 · 4042 grid slots, base rate 0.09624
Diagnostics
Plain text scorecard
==============================================================================
warncal scorecard -- London gale, probabilistic (fitted 2015-2019, verified 2020-2024)
role: WATCH period: 2020-01-01 to 2025-01-01
==============================================================================
VERDICT: SHIP WITH CAVEATS
Clears the hard gates; frequency bias is off.
------------------------------------------------------------------------------
GATES
[PASS] sample size
47 observed events (need >= 20)
[PASS] skill (PSS)
PSS = 0.757 (need > 0); grid slot 1:00:00
[PASS] WATCH objective (recall)
recall = 1.000 (need >= 0.80)
[PASS] recall lower CI bound
95% CI 1.000 [1.000, 1.000] (lower bound must reach 0.80)
[PASS] median lead time
median 11.0 h, IQR 6.0-13.0 h (need >= 1 h)
[WARN] frequency bias
bias = 8.91 (want 0.5-3); 419 warning episodes for 47 events
------------------------------------------------------------------------------
OBJECT COUNTS (overlapping alerts merged into episodes)
events 47 = 47 hits + 0 misses -> recall
episodes 419 = 39 verified + 380 false alarms -> precision
raw alerts 4043 merged into 419 episodes
multiplicity 1.21 hits per verified episode (1.00 = clean units)
Classical mixed table, for comparison with published studies:
event yes event no
warned yes 47 380
warned no 0 n/a
OPPORTUNITY GRID (slot = 1:00:00) -- the only source of correct negatives
event yes event no
warned yes 47 10641
warned no 0 33160
------------------------------------------------------------------------------
METRICS (95% block-bootstrap CI where available)
metric value [95% CI] counted over
recall 1.000 [1.000, 1.000] events
precision 0.093 [0.064, 0.121] episodes
far 0.907 [0.879, 0.936] episodes
f1 0.170 [0.120, 0.217] mixed
csi 0.110 [0.076, 0.144] mixed
bias 8.915 [6.727, 13.002] episodes/events
pofd 0.243 [0.213, 0.273] grid
hss 0.007 [0.005, 0.009] grid
pss 0.757 [0.727, 0.787] grid
ets 0.003 [0.002, 0.004] grid
base rate 0.00107 grid
PROBABILISTIC
Brier 0.08321
Brier skill 0.043 (vs sample climatology)
ROC AUC 0.701
decomposition: reliability 0.00029 - resolution 0.00405 + uncertainty 0.08698
reliability (forecast bin -> observed frequency)
bin n mean p obs freq 95% CI
0.00-0.05 1018 0.047 0.025 [0.017, 0.036]
0.05-0.10 1121 0.066 0.060 [0.047, 0.075]
0.10-0.15 960 0.131 0.118 [0.099, 0.140]
0.15-0.20 943 0.173 0.195 [0.171, 0.222]
0.20-0.25 0 - - -
0.25-0.30 0 - - -
0.30-0.35 0 - - -
0.35-0.40 0 - - -
0.40-0.45 0 - - -
0.45-0.50 0 - - -
0.50-0.55 0 - - -
0.55-0.60 0 - - -
0.60-0.65 0 - - -
0.65-0.70 0 - - -
0.70-0.75 0 - - -
0.75-0.80 0 - - -
0.80-0.85 0 - - -
0.85-0.90 0 - - -
0.90-0.95 0 - - -
0.95-1.00 0 - - -
------------------------------------------------------------------------------
LEAD TIME (event onset minus episode issue time, hits only)
median 11.0 h IQR 6.0 to 13.0 h range 2.0 to 13.0 h
median 95% CI [7.0, 12.0] h
2.0 h |▁▃▂▁▂▂▁ ▁▃▂█| 13.0 h (n=47)
------------------------------------------------------------------------------
warncal scores warning rules. It is a measurement tool, not a decision
authority: it is not medical, clinical or safety advice, it is not an
emergency service, and a favourable verdict here does not make a rule safe
to deploy. A human who understands the hazard must make that call.
==============================================================================
Not advice. warncal scores warning rules. It is a measurement tool, not a decision authority: it is not medical, clinical or safety advice, it is not an emergency service, and a favourable verdict here does not make a rule safe to deploy. A human who understands the hazard must make that call.