Every performance figure on this site traces back to this page: what we measured, how, on which days, against which baselines — including the full exclusion ledger and the correction we made to our own evaluation when a chart didn't look right. If another vendor quotes you an accuracy percentage, ask them for their version of this page.
Every forecast we serve is archived before delivery, the morning before the operating day. We score those archived files against metered generation — a production replay of what the plant actually experienced. The test period was the live future when the forecast was made, so no tuning, no hindsight weather, and no leakage can touch it.
The evaluation: our reference plant (5.94 MW installed inverter AC), 2026-04-01 to 2026-08-17, daytime hours only (solar elevation > 5°), scored at the hourly native resolution and the 15-minute settlement resolution. Of 133 archived forecast days, 113 are scored — 11 telemetry-outage days and 12 active-curtailment days are excluded by mechanical rules from SCADA records, not judgment calls. The full ledger is below.
"Accuracy" as 100-minus-error is not a recognised forecasting metric: it conceals the normalisation basis. We quote nMAE and state the basis — and we will give you the same error under both denominators, because it is the same error.
Mean absolute error divided by installed AC capacity: 0.494 MWh/h ÷ 5.94 MW = 8.32%. The field convention, and the headline we use.
The same error divided by mean daytime output instead. Same forecasts, same days, bigger-looking number. Any vendor quoting one basis can be made to look twice as good — or bad — by switching; we publish both.
Solar elevation > 5°. Counting night hours — when zero is always forecast perfectly — flatters hourly error by roughly 2×.
Scoring daily energy totals instead of intervals hides a median 43% of the real error on our own data (up to 99% on some days) — intraday misses net out. Settlement is per interval, so per-interval error is the number that maps to money.
An 'accuracy %' with no metric, basis, horizon, or day-filter attached is a marketing number. We do not publish one.
From our second plant's 18-month backtest (labelled as a backtest): the same model produces four different "headline" numbers depending purely on scoring conventions. Night hours and coarse resolution flatter accuracy ~2.5×. This is why we state conventions first. One caveat we volunteer: the historical meter record behind this backtest may itself contain trader curtailments we cannot identify from meter data alone — treat these numbers as conservative.
An error number means little without a reference. Skill = 1 − (our MAE ÷ baseline MAE), computed on identical intervals, filters and window. Positive is better than the baseline; zero is no better.
One quirk we report rather than hide: smart persistence scores worse than plain persistence on this sunny-summer window — the clear-sky-index transfer adds noise when most days are already near-clear. We beat both.
Two mechanical rules, both derived from SCADA records rather than judgment calls. Telemetry rule: only days where every logger delivered all 96 15-minute bins are scored (11 days excluded) — outages under-record production and contaminate error in either direction. This rule alone made our headline worse, and we published it anyway. Curtailment rule: days with active sub-rated apply_limit commands in the SCADA log are excluded (12 days, mostly curtailed to 0 kW on negative-price days). A curtailed plant deliberately under-produces; scoring the forecast against it counts market decisions as model error. Standing grid-permit caps are normal operation and are not excluded.
What the correction changed: removing the curtailment contamination moved the headline from 10.5% to 8.32% — and, more tellingly, halved the apparent heavy tail (nRMSE 17.4% → 12.0%). It also resolved an anomaly we had been honestly reporting: clear days appeared to forecast worse than cloudy ones. They don't — market curtailments happen on sunny negative-price days and were being misread as clear-sky model failures. We publish the mistake, the trigger, and the fix, because that is what this page is for. And the fix is not just retrospective: all three data-quality guards — night hours, telemetry outages, active curtailments — are now enforced in the production training pipelines too, so the models never learn from contaminated hours again.
Median daily nMAE 8.0%; one day in ten under 2.6%; one in ten over ~15.7%. The residual worst case is a ±20% day when the cloud forecast is wrong — the universal failure mode of PV forecasting, and the honest reason we quote the all-conditions average rather than a good week. These charts are generated from the actual evaluation series, not illustrations.
On the same plant and month, our site-trained forecasts showed roughly 39% lower error than a leading commercial forecast provider's generic feed (7.41% vs 12.13% portfolio nMAE, hourly, daytime). Site training is the edge — and we made it practical.
The caveats, stated by us before you ask: their product in this comparison is a regenerated historic dataset their own documentation says to use with care for direct comparisons; and our site configuration handicapped their feed, which is why we also computed a capacity-corrected number — and win on that too. We will not claim "we beat provider X" unqualified. We will show you this comparison in full.
Multi-source numerical weather prediction, median-combined per field, corrected by per-plant gradient-boosted models retrained nightly on each site's own history. Today's product is day-ahead; the intraday nowcasting path is in active development and will be measured here before it is sold. On-site irradiance and temperature sensors are ingested and stored today and enter the feature set next.
We measured the empirical coverage of our own served P10–P90 bands and found 0.53 against a nominal 0.80 — materially too narrow. We diagnosed it (conformal widening computed at training but never applied at serving), deployed rolling conformal calibration on 2026-08-18, and validated the fix by replaying five months of issued forecasts: coverage 0.85 at ~70% wider bands. Honest caveat: rolling calibration targets average coverage — a sudden regime shift can still bust a band before the window adapts.
Intraday-horizon accuracy — it lands together with the nowcasting path now in development. CRPS and pinball loss on the probability bands. Fleet availability, MTTR, and performance ratio, extracted from the edge databases we already run. Imbalance cost in euros under a controlled before/after protocol. Curtailment-day revenue impact. Every one of these will appear on this page first, measured — and none of them will appear anywhere on this site as a number before it has been. That is the deal this page makes with you.
Scoring scripts and per-interval series are maintained internally and reproducible on request under NDA. Figures on this page regenerate from the evaluation series, not from a designer's keyboard. See the forecasting product →
We'll measure your plant with exactly this protocol in a pilot — issued forecasts, both bases, published baselines — and you decide with the numbers in hand.