The full basis of preparation for the public benchmark — what is measured, how, from which sources, what it does and does not claim, and how any figure can be reproduced.
Version 1.0 · basis of preparation for the open day-ahead capture mesh and FCR reserve · ← back to the board
This is the methodology and limitations statement for a public, non-monetised benchmark. It describes an internally-reproducible measurement produced by Ikenga. It is not an independent assurance report, an audit opinion, investment advice, or a forward performance promise. Where the benchmark reports a figure, this page tells you exactly how it was derived so a third party can check it.
CausalGrids Lab measures how much of the revenue each market actually made available a forecast-led battery strategy captured — every market, every day, on regulated data. The open benchmark covers the day-ahead energy-arbitrage mesh (43 price-zones across 28 countries) and the FCR reserve (8 pan-EU cooperation zones). The intraday value stack (auctions and continuous order book) and the aFRR / mFRR / MSD balancing legs, together with any market without a free regulated feed, are delivered under the Commercial tier and are not scored on the public page.
All public figures are computed for one common, notional yardstick: a 1 MW / 4-hour battery, one cycle per day, operated as a price-taker. This is a deliberately conservative, comparable standard chosen so every market is measured on the same basis and the result is currency-invariant — it is not a claim about any specific asset. Real assets often cycle more, stack ancillaries, and face duration, degradation, availability and route-to-market constraints the standardised asset does not model. The Commercial tier replaces the yardstick with the owner's real asset parameters.
Capture is the headline number: the revenue the forecast-led strategy achieved as a fraction of the most revenue obtainable with perfect foresight of prices.
100% is the unreachable maximum; the gap to 100% is revenue lost to forecast error. Capture is reported per market, against that market's own ceiling and persistence baseline — never blended or summed across streams. Because it is a ratio of two revenue figures in the same currency, it is unit-free and currency-invariant.
The ceiling is the revenue an operator of the standardised asset could have earned with perfect knowledge of the prices that ultimately cleared — the optimal one-cycle-per-day dispatch computed ex-post on the realised price curve. It is a theoretical maximum, unreachable in practice, and sets the 100% reference. Because the standardised task (one cycle against a daily price shape) is intentionally simple, absolute capture is high in most markets; the informative signal is the edge over persistence (Section 5), not the raw level.
The strategy commits a physical dispatch schedule for the delivery day from an out-of-sample price forecast made before the market cleared, then that committed schedule is settled on the prices that actually cleared. Achieved is what the committed schedule banks on realised prices — a modelled dispatch settled on real prices, net of standardised transaction-cost and imbalance assumptions. It is a measured counterfactual, not realised cash.
The forecast only ever sees data from before the moment it commits. There is no in-sample fitting to the day being scored and no look-ahead. Each day is scored once, at the point of commitment, and the record is fixed before the outcome exists. A walk-forward backtest removes hindsight bias; it is still not a forward promise.
The floor the forecast must beat is persistence: a strategy that commits the battery assuming today's prices repeat yesterday's curve — the standard “no-forecast” reference in the electricity-price-forecasting literature (Lago et al.). It is a genuinely competitive baseline and can go negative when blind timing destroys value. Persistence compresses in the same low-spread markets the forecast does, so the difference between the two — the forecast's edge over persistence — is the like-for-like read of skill, more informative than the absolute capture level. Persistence is a deliberately transparent floor, not the strongest available commercial forecast; head-to-head comparison against a commercial forecast is a Commercial-tier exercise.
The headline capture is money-weighted over a 180-day walk-forward window that rolls forward one day at a time. Money-weighting means high-value days carry proportionally more weight than thin days, so the aggregate reflects where the revenue actually was. Each zone's capture is computed within that zone in its own local currency and is therefore unaffected by exchange rates. The cross-zone aggregate weights each zone by the euro-equivalent revenue the market made available, converted at daily FX reference rates; FX enters the cross-zone weighting only, never a zone's own capture figure.
All inputs are primary settlement data published by the market operator that cleared each market. No modelled, forecast or third-party estimated prices enter the reported figures.
| Stream | Source(s) | Tier |
|---|---|---|
| Day-ahead — EU price-zones | ENTSO-E Transparency Platform (day-ahead prices, load & RES forecasts) | Open |
| Day-ahead — Great Britain | Elexon BMRS (APX / N2EX volume-weighted MID) + demand & wind vintages | Open |
| Day-ahead — North America | CAISO OASIS (DAM LMP), PJM Data Miner 2 (DA hub LMP) | Open |
| FCR reserve | regelleistung.net (four German TSOs), €/MW per 4-hour block, pan-EU cooperation | Open |
| Intraday auctions | GME MI (Italy), OMIE (Spain / Portugal) | Commercial |
| Intraday continuous order book | EPEX M7, Nord Pool | Commercial |
| Balancing energy (aFRR / mFRR / RR) | ENTSO-E 17.1.F (PICASSO / MARI / TERRE); Elia, RTE, APG, TenneT national portals | Commercial |
| Ancillary dispatch (MSD) | GME MSD (Italy — Terna dispatch market) | Commercial |
Different streams carry different histories; each is reported against its own window, and the board shows the sample explicitly rather than blending them into one figure.
| Stream | Window / sample |
|---|---|
| Day-ahead mesh | 180-day money-weighted walk-forward window, rolled one day at a time |
| FCR reserve | Rolling verified daily window (currently a summer run-rate on the standardised asset — a run-rate, not yet a published annual) |
| Balancing / ancillary (Commercial) | Early, limited windows disclosed per row; several are net-negative and reported as-is |
Two robustness statistics are shown on the board, computed on the out-of-sample daily series of the strategy's revenue edge over persistence (pooled across the validated mesh), and disclosed here so they are not read as marketing:
Deflated Sharpe Ratio (DSR) — the probability (0–1) that the edge over persistence is genuinely positive once the raw Sharpe is corrected for multiple testing (the number of candidate configurations evaluated in the model tournament), short samples, and non-normal returns (after López de Prado). It is displayed as “> 0.99” whenever it would round to 1.00, because a bare “1.00” overstates certainty; the underlying value and the trial count are held in the version-pinned ledger.
Probability of Backtest Overfitting (PBO) — the chance that the configuration that looked best in-sample underperforms out-of-sample, estimated by combinatorially-symmetric cross-validation. Below 0.5 is low; lower means the edge is structure rather than curve-fitting.
These are strategy-level robustness diagnostics, not a promise of future return.
Each daily figure is deterministically reproducible from version-pinned inputs: the committed ex-ante record is content-hashed, stamped with the versioned policy that produced it, and independently re-derived from freshly-pulled public prices to assert byte-identical reproduction. Anyone with the committed record and the public data can obtain the identical number. This is an internal reproducibility control — not independent assurance under a recognised standard (e.g. ISAE 3000). The distinction is deliberate and material.
Independence. The public benchmark is not monetised and no fee is contingent on the results it reports; independence here means Ikenga is not paid by the optimisers or route-to-market providers being measured. Conflict. The inherent conflict — the same party builds the strategy and scores it — is disclosed openly and mitigated by the reproducibility control above and by the invitation to third parties to replicate any zone. Completeness. All in-scope results are reported without exception, including periods of underperformance; markets are not added or removed to shape the aggregate, and losing periods (e.g. weak-capture Nordic zones, net-negative balancing windows, the FCR zones where persistence beat the forecast) are published as-is.
The methodology is version-controlled under a dated change log. Figures are not restated other than with disclosure of the change and its effective date. Material methodology changes are versioned on this page.
The open benchmark is this public page — a standardised asset on the free regulated markets. The Commercial tier applies the same discipline to an owner's real asset and extends to the licensed feeds (continuous intraday order book) and markets with no free feed. Under the Commercial tier, limited- and reasonable-assurance grades (ISRE 2410, ISAE 3000 / 3410) denote engagements performed, and any opinion signed, by an independent accredited practitioner. Valuation and performance are reconciled on a GIPS-consistent basis; GIPS® compliance, where presented, rests on independent verification and is not self-certified. The free board itself is not an assurance engagement and carries no such opinion.
← back to the CausalGrids Lab board