Every performance figure on this site traces back to this page: what we measured, how, on which days, against which baselines — including the full exclusion ledger and the correction we made to our own evaluation when a chart didn't look right. If another vendor quotes you an accuracy percentage, ask them for their version of this page.
Model v2, live since 4 September 2026, is scored walk-forward out-of-sample: for every month in the window the model was trained only on data before that month, on our reference plant (5.94 MW installed inverter AC), daytime hours only (solar elevation > 5°), April to August 2026, with the same exclusion ledger as the issued record. Design choices and tuning were fixed on months before the window; the archived day-ahead weather vintage is bracketed against a 48-hour lead (5.0–5.2%). Its issued-forecast record starts accruing now and will replace these numbers here.
The previous model is scored on issued forecasts: every forecast archived the morning before the operating day, replayed against metered generation — 133 archived days, 113 scored, 11 telemetry-outage and 12 active-curtailment days excluded by mechanical rules from SCADA records. On the 1,596 hours where both exist, it scores 7.24% and v2 scores 5.1%. The full ledger is below.
"Accuracy" as 100-minus-error is not a recognised forecasting metric: it conceals the normalisation basis. We quote nMAE and state the basis — and we will give you the same error under both denominators, because it is the same error.
Mean absolute error divided by installed AC capacity: 0.29 MWh/h ÷ 5.94 MW = 4.9–5.2% for model v2 (0.494 MWh/h ÷ 5.94 MW = 8.32% for the previous model). The field convention, and the headline we use.
The same error divided by mean daytime output instead: ~11% for model v2, ~19% for the previous model. Same forecasts, same days, bigger-looking number. Any vendor quoting one basis can be made to look twice as good — or bad — by switching; we publish both.
Solar elevation > 5°. Counting night hours — when zero is always forecast perfectly — flatters hourly error by roughly 2×.
Scoring daily energy totals instead of intervals hides a median 43% of the real error on our own data (up to 99% on some days) — intraday misses net out. Settlement is per interval, so per-interval error is the number that maps to money.
An 'accuracy %' with no metric, basis, horizon, or day-filter attached is a marketing number. We do not publish one.
Model v2 on our reference plant, the same 122 scored days and the same forecasts, scored four ways. Nothing changes but the convention — and the "headline" moves by a factor of four. Night hours flatter accuracy almost 2×; switching the denominator to mean output doubles it the other way. This is why we state the convention before the number, and why a sub-3% single-plant figure you see elsewhere is almost always an all-hours figure.
An error number means little without a reference. Skill = 1 − (our MAE ÷ baseline MAE), computed on identical intervals, filters and window. Positive is better than the baseline; zero is no better.
One quirk we report rather than hide: smart persistence scores worse than plain persistence on this sunny-summer window — the clear-sky-index transfer adds noise when most days are already near-clear. We beat both.
Two mechanical rules, both derived from SCADA records rather than judgment calls. Telemetry rule: only days where every logger delivered all 96 15-minute bins are scored (11 days excluded) — outages under-record production and contaminate error in either direction. This rule alone made our headline worse, and we published it anyway. Curtailment rule: days with active sub-rated apply_limit commands in the SCADA log are excluded (12 days, mostly curtailed to 0 kW on negative-price days). A curtailed plant deliberately under-produces; scoring the forecast against it counts market decisions as model error. Standing grid-permit caps are normal operation and are not excluded.
What the correction changed: removing the curtailment contamination moved the headline from 10.5% to 8.32% — and, more tellingly, halved the apparent heavy tail (nRMSE 17.4% → 12.0%). It also resolved an anomaly we had been honestly reporting: clear days appeared to forecast worse than cloudy ones. They don't — market curtailments happen on sunny negative-price days and were being misread as clear-sky model failures. We publish the mistake, the trigger, and the fix, because that is what this page is for. And the fix is not just retrospective: all three data-quality guards — night hours, telemetry outages, active curtailments — are now enforced in the production training pipelines too, so the models never learn from contaminated hours again.
Median daily nMAE 8.0%; one day in ten under 2.6%; one in ten over ~15.7%. The residual worst case is a ±20% day when the cloud forecast is wrong — the universal failure mode of PV forecasting, and the honest reason we quote the all-conditions average rather than a good week. These charts are generated from the actual evaluation series, not illustrations.
A new forecasting model went into production on 4 September 2026. It has no issued-forecast record yet, so its number belongs to a different evidence class than the 8.32% above, and we say so: walk-forward out-of-sample. For every test month the model was trained only on data before that month, on the same plant, the same clean hours and the same exclusion ledger the issued figure uses. On the 1,596 portfolio hours where both exist, the previous model scores 7.24% and v2 scores 5.1% — a third of the error removed, on 94 of 122 days.
Before quoting it we ran five checks: design choices and hyper-parameters selected on months before the scoring window only (in-window tuning would have read 4.91%); weather inputs at the archived day-ahead vintage, with the fixed 24-hour lead bracketed against 48 hours and anchored on true 00 UTC runs (bracket 5.0–5.2%); no same-day observations among the features; curtailment and outage hours excluded from the same command-log ledger; one denominator throughout. Issued-forecast scores for v2 will replace this section as they accrue.
On the same plant and month, our site-trained forecasts showed roughly 39% lower error than a leading commercial forecast provider's generic feed (7.41% vs 12.13% portfolio nMAE, hourly, daytime). Site training is the edge — and we made it practical.
The caveats, stated by us before you ask: their product in this comparison is a regenerated historic dataset their own documentation says to use with care for direct comparisons; and our site configuration handicapped their feed, which is why we also computed a capacity-corrected number — and win on that too. We will not claim "we beat provider X" unqualified. We will show you this comparison in full.
Multi-source numerical weather prediction, median-combined per field, corrected by per-plant gradient-boosted models retrained nightly on each site's own history. Today's product is day-ahead; the intraday nowcasting path is in active development and will be measured here before it is sold. On-site irradiance and temperature sensors are ingested and stored today and enter the feature set next.
We measured the empirical coverage of our own served P10–P90 bands and found 0.53 against a nominal 0.80 — materially too narrow. We diagnosed it (conformal widening computed at training but never applied at serving), deployed rolling conformal calibration on 2026-08-18, and validated the fix by replaying five months of issued forecasts: coverage 0.85 at ~70% wider bands. Honest caveat: rolling calibration targets average coverage — a sudden regime shift can still bust a band before the window adapts.
Intraday-horizon accuracy — it lands together with the nowcasting path now in development. CRPS and pinball loss on the probability bands. Fleet availability, MTTR, and performance ratio, extracted from the edge databases we already run. Imbalance cost in euros under a controlled before/after protocol. Curtailment-day revenue impact. Every one of these will appear on this page first, measured — and none of them will appear anywhere on this site as a number before it has been. That is the deal this page makes with you.
Scoring scripts and per-interval series are maintained internally and reproducible on request under NDA. Figures on this page regenerate from the evaluation series, not from a designer's keyboard. See the forecasting product →
We'll measure your plant with exactly this protocol in a pilot — issued forecasts, both bases, published baselines — and you decide with the numbers in hand.