Results from retrospective tests — the live record grows with every season.
The model was trained exclusively on data from 2012–2025 and then tested day by day against the 2026 Torbole observations.
A Peirce Skill Score of 0 means no discrimination; 1 is perfect discrimination. For each test year, the model was trained without that year’s data.
| Regime | Base rate | Model (Peirce) | Climatology only |
|---|---|---|---|
| Ora | 69 % | +0.54 | +0.22 |
| Peler | 53 % | +0.39 | +0.20 |
For context, Ora fires on roughly 80% of midsummer days, so a blanket “always GO” forecast would also achieve high accuracy. The skill score therefore measures actual discrimination rather than just the base rate.
Tested on archived previous-day forecasts from 2024–2026 (513 days): switching from reanalysis data to the real forecast chain costs almost no forecast skill. The model’s decision thresholds are calibrated on exactly this data.
| Regime | Reanalysis (ERA5) | Previous-day forecast |
|---|---|---|
| Ora | +0.58 | +0.59 |
| Peler | +0.38 | +0.49 |
The evaluation follows the same principles as the Walchensee project (Peirce Skill Score rather than raw accuracy, yearly cross-validation, and calibrated thresholds).