Static scorecard, regenerated locally on 2026-09-02. warncal is an offline, deterministic backtester; these four pages are its own HTML output, unmodified apart from this note. Nothing here is live or updating.

Weather data: Open-Meteo historical archive at London Heathrow, CC BY 4.0, attribution to Copernicus/ECMWF. Open-Meteo's free tier is explicitly non-commercial; this is a portfolio demonstration.

warncal scores warning rules. It is a measurement tool, not a decision authority: it is not medical, clinical or safety advice, it is not an emergency service, and a favourable verdict here does not make a rule safe to deploy. A human who understands the hazard must make that call.
← all four scorecards

London gale WATCH (gust >= 60 km/h)

WATCH · 2020-01-01 to 2025-01-01 · 47 events · 162 warning episodes

SHIP WITH CAVEATS
Clears the hard gates; frequency bias is off.

Gates

PASSsample size47 observed events (need >= 20)
PASSskill (PSS)PSS = 0.841 (need > 0); grid slot 1:00:00
PASSWATCH objective (recall)recall = 0.915 (need >= 0.80)
PASSrecall lower CI bound95% CI 0.915 [0.823, 0.983] (lower bound must reach 0.80)
PASSmedian lead timemedian 4.0 h, IQR 2.5-9.0 h (need >= 1 h)
WARNfrequency biasbias = 3.45 (want 0.5-3); 162 warning episodes for 47 events

Metrics

metricvalue (95% CI)counted overdefinition
recall0.915 [0.823, 0.983]eventsa/(a+c) of the events, the share that were warned
precision0.247 [0.174, 0.324]episodesa/(a+b) of the warnings issued, the share that verified
far0.753 [0.676, 0.826]episodesb/(a+b) false alarm RATIO (= 1 - precision)
f10.389 [0.291, 0.480]mixedharmonic mean of precision and recall
csi0.254 [0.179, 0.333]mixeda/(a+b+c) critical success index / threat score
bias3.447 [2.660, 4.759]episodes/events(a+b)/(a+c) >1 over-warns, <1 under-warns
pofd0.074 [0.058, 0.089]gridb/(b+d) false alarm RATE (needs correct negatives)
hss0.024 [0.017, 0.030]gridHeidke skill score, 0 = chance, 1 = perfect
pss0.841 [0.749, 0.920]gridPeirce skill score = POD - POFD, base-rate insensitive
ets0.012 [0.009, 0.015]gridequitable threat score (CSI corrected for chance hits)

Object counts

counted overtotalrightwrongdrives
events4743 hits4 missesrecall / POD
warning episodes16240 verified122 false alarmsprecision / FAR
raw alerts1050merged into 162 episodesnothing directly — merging prevents double counting

Multiplicity 1.07 hits per verified episode. At 1.00 the two denominators agree and every object score is unambiguous; well above 1.00 means long episodes are sweeping up several events at once.

Contingency tables

tablehitsfalse alarmsmissescorrect negatives
object, classical mixed431224n/a by design
grid (slot 1:00:00)433237440564

Lead time

median 4.0 h, IQR 2.5–9.0 h, range 1.0–13.0 h over 43 hits.

Diagnostics

verification diagnostics for London gale WATCH (gust >= 60 km/h)

Plain text scorecard

==============================================================================
  warncal scorecard  --  London gale WATCH (gust >= 60 km/h)
  role: WATCH   period: 2020-01-01 to 2025-01-01
==============================================================================

  VERDICT: SHIP WITH CAVEATS
    Clears the hard gates; frequency bias is off.

------------------------------------------------------------------------------
  GATES
    [PASS] sample size
           47 observed events (need >= 20)
    [PASS] skill (PSS)
           PSS = 0.841 (need > 0); grid slot 1:00:00
    [PASS] WATCH objective (recall)
           recall = 0.915 (need >= 0.80)
    [PASS] recall lower CI bound
           95% CI 0.915 [0.823, 0.983] (lower bound must reach 0.80)
    [PASS] median lead time
           median 4.0 h, IQR 2.5-9.0 h (need >= 1 h)
    [WARN] frequency bias
           bias = 3.45 (want 0.5-3); 162 warning episodes for 47 events

------------------------------------------------------------------------------
  OBJECT COUNTS (overlapping alerts merged into episodes)

    events             47   =  43 hits + 4 misses        -> recall
    episodes          162   =  40 verified + 122 false alarms  -> precision
    raw alerts       1050   merged into 162 episodes
    multiplicity     1.07   hits per verified episode (1.00 = clean units)

    Classical mixed table, for comparison with published studies:
                   event yes   event no
    warned yes           43         122
    warned no             4         n/a

  OPPORTUNITY GRID (slot = 1:00:00) -- the only source of correct negatives

                   event yes   event no
    warned yes           43        3237
    warned no             4       40564

------------------------------------------------------------------------------
  METRICS   (95% block-bootstrap CI where available)

    metric       value [95% CI]               counted over
    recall       0.915 [0.823, 0.983]         events
    precision    0.247 [0.174, 0.324]         episodes
    far          0.753 [0.676, 0.826]         episodes
    f1           0.389 [0.291, 0.480]         mixed
    csi          0.254 [0.179, 0.333]         mixed
    bias         3.447 [2.660, 4.759]         episodes/events
    pofd         0.074 [0.058, 0.089]         grid
    hss          0.024 [0.017, 0.030]         grid
    pss          0.841 [0.749, 0.920]         grid
    ets          0.012 [0.009, 0.015]         grid

    base rate    0.00107                      grid

------------------------------------------------------------------------------
  LEAD TIME (event onset minus episode issue time, hits only)
    median 4.0 h   IQR 2.5 to 9.0 h   range 1.0 to 13.0 h
    median 95% CI [3.0, 6.0] h
    1.0 h |▅▄█▂▃▄ ▁▂ ▃▆| 13.0 h   (n=43)

------------------------------------------------------------------------------
  warncal scores warning rules. It is a measurement tool, not a decision
  authority: it is not medical, clinical or safety advice, it is not an
  emergency service, and a favourable verdict here does not make a rule safe
  to deploy. A human who understands the hazard must make that call.
==============================================================================
Not advice. warncal scores warning rules. It is a measurement tool, not a decision authority: it is not medical, clinical or safety advice, it is not an emergency service, and a favourable verdict here does not make a rule safe to deploy. A human who understands the hazard must make that call.