Static scorecard, regenerated locally on 2026-09-02. warncal is an offline, deterministic backtester; these four pages are its own HTML output, unmodified apart from this note. Nothing here is live or updating.

Weather data: Open-Meteo historical archive at London Heathrow, CC BY 4.0, attribution to Copernicus/ECMWF. Open-Meteo's free tier is explicitly non-commercial; this is a portfolio demonstration.

warncal scores warning rules. It is a measurement tool, not a decision authority: it is not medical, clinical or safety advice, it is not an emergency service, and a favourable verdict here does not make a rule safe to deploy. A human who understands the hazard must make that call.
← all four scorecards

London gale WARNING (gust >= 65 km/h)

WARNING · 2020-01-01 to 2025-01-01 · 47 events · 97 warning episodes

SHIP
Clears every gate for a WARNING.

Gates

PASSsample size47 observed events (need >= 20)
PASSskill (PSS)PSS = 0.830 (need > 0); grid slot 1:00:00
PASSWARNING objective (precision)precision = 0.392 (need >= 0.20)
PASSprecision lower CI bound95% CI 0.392 [0.278, 0.489] (lower bound must reach 0.20)
PASSmedian lead timemedian 3.0 h, IQR 2.0-5.0 h (need >= 1 h)
PASSfrequency biasbias = 2.06 (want 0.5-3); 97 warning episodes for 47 events

Metrics

metricvalue (95% CI)counted overdefinition
recall0.872 [0.756, 0.958]eventsa/(a+c) of the events, the share that were warned
precision0.392 [0.278, 0.489]episodesa/(a+b) of the warnings issued, the share that verified
far0.608 [0.511, 0.722]episodesb/(a+b) false alarm RATIO (= 1 - precision)
f10.541 [0.411, 0.638]mixedharmonic mean of precision and recall
csi0.387 [0.273, 0.490]mixeda/(a+b+c) critical success index / threat score
bias2.064 [1.677, 2.774]episodes/events(a+b)/(a+c) >1 over-warns, <1 under-warns
pofd0.042 [0.032, 0.053]gridb/(b+d) false alarm RATE (needs correct negatives)
hss0.040 [0.030, 0.048]gridHeidke skill score, 0 = chance, 1 = perfect
pss0.830 [0.717, 0.917]gridPeirce skill score = POD - POFD, base-rate insensitive
ets0.021 [0.015, 0.025]gridequitable threat score (CSI corrected for chance hits)

Object counts

counted overtotalrightwrongdrives
events4741 hits6 missesrecall / POD
warning episodes9738 verified59 false alarmsprecision / FAR
raw alerts608merged into 97 episodesnothing directly — merging prevents double counting

Multiplicity 1.08 hits per verified episode. At 1.00 the two denominators agree and every object score is unambiguous; well above 1.00 means long episodes are sweeping up several events at once.

Contingency tables

tablehitsfalse alarmsmissescorrect negatives
object, classical mixed41596n/a by design
grid (slot 1:00:00)411852641949

Lead time

median 3.0 h, IQR 2.0–5.0 h, range 1.0–13.0 h over 41 hits.

Diagnostics

verification diagnostics for London gale WARNING (gust >= 65 km/h)

Plain text scorecard

==============================================================================
  warncal scorecard  --  London gale WARNING (gust >= 65 km/h)
  role: WARNING   period: 2020-01-01 to 2025-01-01
==============================================================================

  VERDICT: SHIP
    Clears every gate for a WARNING.

------------------------------------------------------------------------------
  GATES
    [PASS] sample size
           47 observed events (need >= 20)
    [PASS] skill (PSS)
           PSS = 0.830 (need > 0); grid slot 1:00:00
    [PASS] WARNING objective (precision)
           precision = 0.392 (need >= 0.20)
    [PASS] precision lower CI bound
           95% CI 0.392 [0.278, 0.489] (lower bound must reach 0.20)
    [PASS] median lead time
           median 3.0 h, IQR 2.0-5.0 h (need >= 1 h)
    [PASS] frequency bias
           bias = 2.06 (want 0.5-3); 97 warning episodes for 47 events

------------------------------------------------------------------------------
  OBJECT COUNTS (overlapping alerts merged into episodes)

    events             47   =  41 hits + 6 misses        -> recall
    episodes           97   =  38 verified + 59 false alarms  -> precision
    raw alerts        608   merged into 97 episodes
    multiplicity     1.08   hits per verified episode (1.00 = clean units)

    Classical mixed table, for comparison with published studies:
                   event yes   event no
    warned yes           41          59
    warned no             6         n/a

  OPPORTUNITY GRID (slot = 1:00:00) -- the only source of correct negatives

                   event yes   event no
    warned yes           41        1852
    warned no             6       41949

------------------------------------------------------------------------------
  METRICS   (95% block-bootstrap CI where available)

    metric       value [95% CI]               counted over
    recall       0.872 [0.756, 0.958]         events
    precision    0.392 [0.278, 0.489]         episodes
    far          0.608 [0.511, 0.722]         episodes
    f1           0.541 [0.411, 0.638]         mixed
    csi          0.387 [0.273, 0.490]         mixed
    bias         2.064 [1.677, 2.774]         episodes/events
    pofd         0.042 [0.032, 0.053]         grid
    hss          0.040 [0.030, 0.048]         grid
    pss          0.830 [0.717, 0.917]         grid
    ets          0.021 [0.015, 0.025]         grid

    base rate    0.00107                      grid

------------------------------------------------------------------------------
  LEAD TIME (event onset minus episode issue time, hits only)
    median 3.0 h   IQR 2.0 to 5.0 h   range 1.0 to 13.0 h
    median 95% CI [2.0, 4.0] h
    1.0 h |█▆▅▃▃▂▁▁▂▁ ▃| 13.0 h   (n=41)

------------------------------------------------------------------------------
  warncal scores warning rules. It is a measurement tool, not a decision
  authority: it is not medical, clinical or safety advice, it is not an
  emergency service, and a favourable verdict here does not make a rule safe
  to deploy. A human who understands the hazard must make that call.
==============================================================================
Not advice. warncal scores warning rules. It is a measurement tool, not a decision authority: it is not medical, clinical or safety advice, it is not an emergency service, and a favourable verdict here does not make a rule safe to deploy. A human who understands the hazard must make that call.