METHODOLOGY / EVALUATION

Most vendors publish a number. We publish the method.

Every performance figure on this site traces back to this page: what we measured, how, on which days, against which baselines — including the full exclusion ledger and the correction we made to our own evaluation when a chart didn't look right. If another vendor quotes you an accuracy percentage, ask them for their version of this page.

8.32%
Day-ahead nMAE, hourly
113 days
Issued forecasts scored
+0.43
Skill vs smart persistence
21.5%
Our worst day — published
EVIDENCE CLASS

Issued forecasts. Not a backtest.

Every forecast we serve is archived before delivery, the morning before the operating day. We score those archived files against metered generation — a production replay of what the plant actually experienced. The test period was the live future when the forecast was made, so no tuning, no hindsight weather, and no leakage can touch it.

The evaluation: our reference plant (5.94 MW installed inverter AC), 2026-04-01 to 2026-08-17, daytime hours only (solar elevation > 5°), scored at the hourly native resolution and the 15-minute settlement resolution. Of 133 archived forecast days, 113 are scored — 11 telemetry-outage days and 12 active-curtailment days are excluded by mechanical rules from SCADA records, not judgment calls. The full ledger is below.

nMAE, day-ahead, hourly8.32%capacity basis · all conditions
nMAE, 15-min settlement9.30%the resolution money settles at
nRMSE, hourly12.00%ratio to nMAE ≈ 1.44 — errors are mostly small
Clear days / partly cloudy8.59% / 8.68%all-conditions average quoted, not a good week
R² (out-of-sample)0.801secondary metric · hourly
Mean absolute error0.494 MWh/hn = 1,079 scored hours
Per-plant spread6.5% – 12.6%newest tracker plant highest — least training history
DEFINITIONS

The metric, defined — with both denominators.

"Accuracy" as 100-minus-error is not a recognised forecasting metric: it conceals the normalisation basis. We quote nMAE and state the basis — and we will give you the same error under both denominators, because it is the same error.

=

nMAE, capacity basis — 8.32%

Mean absolute error divided by installed AC capacity: 0.494 MWh/h ÷ 5.94 MW = 8.32%. The field convention, and the headline we use.

=

nMAE, mean-production basis — ~19%

The same error divided by mean daytime output instead. Same forecasts, same days, bigger-looking number. Any vendor quoting one basis can be made to look twice as good — or bad — by switching; we publish both.

Daytime only

Solar elevation > 5°. Counting night hours — when zero is always forecast perfectly — flatters hourly error by roughly 2×.

15

Interval-level, not daily totals

Scoring daily energy totals instead of intervals hides a median 43% of the real error on our own data (up to 99% on some days) — intraday misses net out. Settlement is per interval, so per-interval error is the number that maps to money.

!

Never 100 − error

An 'accuracy %' with no metric, basis, horizon, or day-filter attached is a marketing number. We do not publish one.

THE CONVENTIONS BRIDGE — SAME MODEL, FOUR NUMBERS

From our second plant's 18-month backtest (labelled as a backtest): the same model produces four different "headline" numbers depending purely on scoring conventions. Night hours and coarse resolution flatter accuracy ~2.5×. This is why we state conventions first. One caveat we volunteer: the historical meter record behind this backtest may itself contain trader curtailments we cannot identify from meter data alone — treat these numbers as conservative.

4.95%
all hours, hourly
night zeros included
~8.9%
daytime, hourly
night hours removed
11.06%
true day-ahead weather
no hindsight forecasts
12.64%
15-min settlement
the number that maps to money
SKILL SCORES

Beating baselines you can define.

An error number means little without a reference. Skill = 1 − (our MAE ÷ baseline MAE), computed on identical intervals, filters and window. Positive is better than the baseline; zero is no better.

One quirk we report rather than hide: smart persistence scores worse than plain persistence on this sunny-summer window — the clear-sky-index transfer adds noise when most days are already near-clear. We beat both.

vs persistence+0.331yesterday's actual at the same hour · baseline nMAE 12.44%
vs smart persistence+0.432yesterday's clear-sky index × today's clear-sky curve · 14.66%
vs clear-sky × PR+0.325cloudless curve scaled by fitted performance ratio · 12.34%
THE EXCLUSION LEDGER

133 days archived. 113 scored. Every exclusion published.

Two mechanical rules, both derived from SCADA records rather than judgment calls. Telemetry rule: only days where every logger delivered all 96 15-minute bins are scored (11 days excluded) — outages under-record production and contaminate error in either direction. This rule alone made our headline worse, and we published it anyway. Curtailment rule: days with active sub-rated apply_limit commands in the SCADA log are excluded (12 days, mostly curtailed to 0 kW on negative-price days). A curtailed plant deliberately under-produces; scoring the forecast against it counts market decisions as model error. Standing grid-permit caps are normal operation and are not excluded.

Curtailment day — 2026-04-26 · excluded from scoring — issued forecast vs metered actual, hourly MW Curtailment day — 2026-04-26 · excluded from scoring curtailed to 0 kW at midday by SCADA apply_limit on negative prices forecast 40.8 MWh vs actual 11.5 MWh 0 1 2 3 4 5 6 06:00 08:00 10:00 12:00 14:00 16:00 18:00 LOCAL TIME MW forecast (issued D-1) actual (metered) forecast actual
2026-04-26, preserved from the pre-correction evaluation: production tracks the forecast until mid-morning, then the SCADA log shows the plant ordered to 0 kW for the negative-price midday block, then output returns. An earlier revision of our own evaluation scored days like this as forecast error and called the worst of them a 43% 'missed front'. A reader challenging that chart is what triggered the correction.

What the correction changed: removing the curtailment contamination moved the headline from 10.5% to 8.32% — and, more tellingly, halved the apparent heavy tail (nRMSE 17.4% → 12.0%). It also resolved an anomaly we had been honestly reporting: clear days appeared to forecast worse than cloudy ones. They don't — market curtailments happen on sunny negative-price days and were being misread as clear-sky model failures. We publish the mistake, the trigger, and the fix, because that is what this page is for. And the fix is not just retrospective: all three data-quality guards — night hours, telemetry outages, active curtailments — are now enforced in the production training pipelines too, so the models never learn from contaminated hours again.

THE DISTRIBUTION

A typical day runs around 8%. One day in ten is near-perfect.

Median daily nMAE 8.0%; one day in ten under 2.6%; one in ten over ~15.7%. The residual worst case is a ±20% day when the cloud forecast is wrong — the universal failure mode of PV forecasting, and the honest reason we quote the all-conditions average rather than a good week. These charts are generated from the actual evaluation series, not illustrations.

Distribution of daily nMAE across 83 scored days Daily nMAE distribution — 83 scored days daily mean hourly error ÷ 5.94 MW AC · outage- and curtailment-free days, ≥10 scored daylight hours 5 9 14 18 0% 5% 10% 15% 20% 25% DAILY nMAE (% OF CAPACITY) DAYS p10 2.6% median 8.0% p90 15.7% worst 21.5%
Daily nMAE across all 83 fully-scored days (≥10 daylight hours, outage- and curtailment-free). Real evaluation output — the same series the 8.32% headline comes from.
Worst genuine day — 2026-06-08 · nMAE 21.5% — issued forecast vs metered actual, hourly MW Worst genuine day — 2026-06-08 · nMAE 21.5% an under-forecast morning, then an afternoon the weather models missed forecast 33.0 MWh vs actual 38.9 MWh 0 1 2 3 4 5 6 06:00 08:00 10:00 12:00 14:00 16:00 18:00 LOCAL TIME MW forecast (issued D-1) actual (metered) forecast actual
2026-06-08 — the worst genuine forecast day on record: daily nMAE 21.5%. The plant out-produced an under-forecast morning, then an afternoon system arrived that the weather models missed; actual 38.9 MWh vs forecast 33.0 MWh, worst single hour missed by 2.24 MWh. This is what a genuinely missed day costs — and why the error tail, not the average, is what imbalance exposure cares about.
Our best day — 2026-07-31 · daily nMAE 1.17% — issued forecast vs metered actual, hourly MW Our best day — 2026-07-31 · daily nMAE 1.17% a clear summer day — clear days are essentially solved forecast 41.8 MWh vs actual 42.2 MWh 0 1 2 3 4 5 6 06:00 08:00 10:00 12:00 14:00 16:00 18:00 LOCAL TIME MW forecast (issued D-1) actual (metered) forecast actual
2026-07-31 — the best day on record: daily nMAE 1.17%, forecast 41.8 MWh vs actual 42.2 MWh. Clear days are essentially solved; the three best days on record are all ≤1.2%.
BENCHMARK

We argue against ourselves — and still win.

On the same plant and month, our site-trained forecasts showed roughly 39% lower error than a leading commercial forecast provider's generic feed (7.41% vs 12.13% portfolio nMAE, hourly, daytime). Site training is the edge — and we made it practical.

The caveats, stated by us before you ask: their product in this comparison is a regenerated historic dataset their own documentation says to use with care for direct comparisons; and our site configuration handicapped their feed, which is why we also computed a capacity-corrected number — and win on that too. We will not claim "we beat provider X" unqualified. We will show you this comparison in full.

DYNVOLT, site-trained7.41%same plant, same month, hourly daytime nMAE
Commercial generic feed12.13%leading provider, unnamed by policy
Their feed, capacity-corrected8.9%independent re-check with the handicap removed — still behind
Relative error reduction~39%site training is the edge
W

Inputs, described honestly

Multi-source numerical weather prediction, median-combined per field, corrected by per-plant gradient-boosted models retrained nightly on each site's own history. Today's product is day-ahead; the intraday nowcasting path is in active development and will be measured here before it is sold. On-site irradiance and temperature sensors are ingested and stored today and enter the feature set next.

P

Probabilistic bands — measured, fixed, re-validated

We measured the empirical coverage of our own served P10–P90 bands and found 0.53 against a nominal 0.80 — materially too narrow. We diagnosed it (conformal widening computed at training but never applied at serving), deployed rolling conformal calibration on 2026-08-18, and validated the fix by replaying five months of issued forecasts: coverage 0.85 at ~70% wider bands. Honest caveat: rolling calibration targets average coverage — a sudden regime shift can still bust a band before the window adapts.

What we are measuring next

Intraday-horizon accuracy — it lands together with the nowcasting path now in development. CRPS and pinball loss on the probability bands. Fleet availability, MTTR, and performance ratio, extracted from the edge databases we already run. Imbalance cost in euros under a controlled before/after protocol. Curtailment-day revenue impact. Every one of these will appear on this page first, measured — and none of them will appear anywhere on this site as a number before it has been. That is the deal this page makes with you.

Scoring scripts and per-interval series are maintained internally and reproducible on request under NDA. Figures on this page regenerate from the evaluation series, not from a designer's keyboard. See the forecasting product →

Ask any other vendor for this page.

We'll measure your plant with exactly this protocol in a pilot — issued forecasts, both bases, published baselines — and you decide with the numbers in hand.

Start a pilotTrust & security →
FAQ

Methodology — frequently asked

Can't find what you're looking for? Talk to our team →

Why do you publish your forecast error?
Because a number without a method is marketing, and this industry has too much of that. Every figure on this site traces to the same evaluation: issued forecasts, daytime hours, capacity-normalised, clean telemetry days. If a vendor quotes you an accuracy percentage, ask for their equivalent of this page.
Why nMAE and not accuracy or MAPE?
"Accuracy" as 100-minus-error conceals the normalisation basis, and MAPE is undefined and explosive near zero output — solar spends every morning and evening near zero. nMAE with a stated basis is the field convention, and we quote both bases: 8.32% of installed capacity, ~19% of mean production. Same error, different denominator.
Are these backtest numbers?
No. Every headline figure is scored on forecasts we actually issued in production, archived before delivery and replayed against metered generation — no hindsight weather, no tuning on the test period. Where we cite a backtest (our second plant's 18-month history), we label it as one.
What is next on the measurement roadmap?
Intraday-horizon accuracy — it arrives together with the nowcasting path now in development. Probabilistic scores like CRPS on the bands. Fleet availability and MTTR, already being extracted from the edge databases. And imbalance cost in euros under a controlled protocol. Each will be published here first, measured, before it appears anywhere else on this site.