CausalGrids Lab · BESS revenue-assurance benchmark

Methodology, Limitations & Audit

The full basis of preparation for the public benchmark — what is measured, how, from which sources, what it does and does not claim, and how any figure can be reproduced.

Version 1.0 · basis of preparation for the open day-ahead capture mesh and FCR reserve · ← back to the board

What this document is — and is not

This is the methodology and limitations statement for a public, non-monetised benchmark. It describes an internally-reproducible measurement produced by Ikenga. It is not an independent assurance report, an audit opinion, investment advice, or a forward performance promise. Where the benchmark reports a figure, this page tells you exactly how it was derived so a third party can check it.

1Purpose & scope

CausalGrids Lab measures how much of the revenue each market actually made available a forecast-led battery strategy captured — every market, every day, on regulated data. The open benchmark covers the day-ahead energy-arbitrage mesh (43 price-zones across 28 countries) and the FCR reserve (8 pan-EU cooperation zones). The intraday value stack (auctions and continuous order book) and the aFRR / mFRR / MSD balancing legs, together with any market without a free regulated feed, are delivered under the Commercial tier and are not scored on the public page.

2The standardised asset

All public figures are computed for one common, notional yardstick: a 1 MW / 4-hour battery, one cycle per day, operated as a price-taker. This is a deliberately conservative, comparable standard chosen so every market is measured on the same basis and the result is currency-invariant — it is not a claim about any specific asset. Real assets often cycle more, stack ancillaries, and face duration, degradation, availability and route-to-market constraints the standardised asset does not model. The Commercial tier replaces the yardstick with the owner's real asset parameters.

3The capture metric

Capture is the headline number: the revenue the forecast-led strategy achieved as a fraction of the most revenue obtainable with perfect foresight of prices.

capture = achieved_revenue ÷ perfect_foresight_ceiling

100% is the unreachable maximum; the gap to 100% is revenue lost to forecast error. Capture is reported per market, against that market's own ceiling and persistence baseline — never blended or summed across streams. Because it is a ratio of two revenue figures in the same currency, it is unit-free and currency-invariant.

Perfect-foresight ceiling (the denominator)

The ceiling is the revenue an operator of the standardised asset could have earned with perfect knowledge of the prices that ultimately cleared — the optimal one-cycle-per-day dispatch computed ex-post on the realised price curve. It is a theoretical maximum, unreachable in practice, and sets the 100% reference. Because the standardised task (one cycle against a daily price shape) is intentionally simple, absolute capture is high in most markets; the informative signal is the edge over persistence (Section 5), not the raw level.

Achieved revenue (the numerator)

The strategy commits a physical dispatch schedule for the delivery day from an out-of-sample price forecast made before the market cleared, then that committed schedule is settled on the prices that actually cleared. Achieved is what the committed schedule banks on realised prices — a modelled dispatch settled on real prices, net of standardised transaction-cost and imbalance assumptions. It is a measured counterfactual, not realised cash.

4Walk-forward discipline

The forecast only ever sees data from before the moment it commits. There is no in-sample fitting to the day being scored and no look-ahead. Each day is scored once, at the point of commitment, and the record is fixed before the outcome exists. A walk-forward backtest removes hindsight bias; it is still not a forward promise.

5Persistence baseline

The floor the forecast must beat is persistence: a strategy that commits the battery assuming today's prices repeat yesterday's curve — the standard “no-forecast” reference in the electricity-price-forecasting literature (Lago et al.). It is a genuinely competitive baseline and can go negative when blind timing destroys value. Persistence compresses in the same low-spread markets the forecast does, so the difference between the two — the forecast's edge over persistence — is the like-for-like read of skill, more informative than the absolute capture level. Persistence is a deliberately transparent floor, not the strongest available commercial forecast; head-to-head comparison against a commercial forecast is a Commercial-tier exercise.

6Aggregation, window & currency

The headline capture is money-weighted over a 180-day walk-forward window that rolls forward one day at a time. Money-weighting means high-value days carry proportionally more weight than thin days, so the aggregate reflects where the revenue actually was. Each zone's capture is computed within that zone in its own local currency and is therefore unaffected by exchange rates. The cross-zone aggregate weights each zone by the euro-equivalent revenue the market made available, converted at daily FX reference rates; FX enters the cross-zone weighting only, never a zone's own capture figure.

7Source data & provenance

All inputs are primary settlement data published by the market operator that cleared each market. No modelled, forecast or third-party estimated prices enter the reported figures.

StreamSource(s)Tier
Day-ahead — EU price-zonesENTSO-E Transparency Platform (day-ahead prices, load & RES forecasts)Open
Day-ahead — Great BritainElexon BMRS (APX / N2EX volume-weighted MID) + demand & wind vintagesOpen
Day-ahead — North AmericaCAISO OASIS (DAM LMP), PJM Data Miner 2 (DA hub LMP)Open
FCR reserveregelleistung.net (four German TSOs), €/MW per 4-hour block, pan-EU cooperationOpen
Intraday auctionsGME MI (Italy), OMIE (Spain / Portugal)Commercial
Intraday continuous order bookEPEX M7, Nord PoolCommercial
Balancing energy (aFRR / mFRR / RR)ENTSO-E 17.1.F (PICASSO / MARI / TERRE); Elia, RTE, APG, TenneT national portalsCommercial
Ancillary dispatch (MSD)GME MSD (Italy — Terna dispatch market)Commercial

8Measurement windows & samples

Different streams carry different histories; each is reported against its own window, and the board shows the sample explicitly rather than blending them into one figure.

StreamWindow / sample
Day-ahead mesh180-day money-weighted walk-forward window, rolled one day at a time
FCR reserveRolling verified daily window (currently a summer run-rate on the standardised asset — a run-rate, not yet a published annual)
Balancing / ancillary (Commercial)Early, limited windows disclosed per row; several are net-negative and reported as-is

9Statistical significance — DSR & PBO

Two robustness statistics are shown on the board, computed on the out-of-sample daily series of the strategy's revenue edge over persistence (pooled across the validated mesh), and disclosed here so they are not read as marketing:

Deflated Sharpe Ratio (DSR) — the probability (0–1) that the edge over persistence is genuinely positive once the raw Sharpe is corrected for multiple testing (the number of candidate configurations evaluated in the model tournament), short samples, and non-normal returns (after López de Prado). It is displayed as “> 0.99” whenever it would round to 1.00, because a bare “1.00” overstates certainty; the underlying value and the trial count are held in the version-pinned ledger.

Probability of Backtest Overfitting (PBO) — the chance that the configuration that looked best in-sample underperforms out-of-sample, estimated by combinatorially-symmetric cross-validation. Below 0.5 is low; lower means the edge is structure rather than curve-fitting.

These are strategy-level robustness diagnostics, not a promise of future return.

10Reproducibility & controls

Each daily figure is deterministically reproducible from version-pinned inputs: the committed ex-ante record is content-hashed, stamped with the versioned policy that produced it, and independently re-derived from freshly-pulled public prices to assert byte-identical reproduction. Anyone with the committed record and the public data can obtain the identical number. This is an internal reproducibility control — not independent assurance under a recognised standard (e.g. ISAE 3000). The distinction is deliberate and material.

11Independence, conflicts & completeness

Independence. The public benchmark is not monetised and no fee is contingent on the results it reports; independence here means Ikenga is not paid by the optimisers or route-to-market providers being measured. Conflict. The inherent conflict — the same party builds the strategy and scores it — is disclosed openly and mitigated by the reproducibility control above and by the invitation to third parties to replicate any zone. Completeness. All in-scope results are reported without exception, including periods of underperformance; markets are not added or removed to shape the aggregate, and losing periods (e.g. weak-capture Nordic zones, net-negative balancing windows, the FCR zones where persistence beat the forecast) are published as-is.

12Change governance

The methodology is version-controlled under a dated change log. Figures are not restated other than with disclosure of the change and its effective date. Material methodology changes are versioned on this page.

13Limitations

  • Paper, not cash. The book is a shadow book. Figures are a modelled dispatch settled on realised prices — a measured counterfactual, not realised trading revenue.
  • Price-taker, perfect fill. The standardised asset is assumed to transact at the cleared price with no market impact, no bid rejection and no partial fills. A 1 MW unit is too small to move the clearing price, so this holds at the margin. Note that BESS in aggregate is increasingly a price-maker — the fleet-wide compression of spreads (cannibalisation) is already embedded in the realised prices scored against; what is not modelled is the market impact of a large individual book, which would move prices and face execution risk, particularly in thin venues (continuous intraday, balancing).
  • Notional asset. One cycle per day on a 1 MW / 4 h battery ignores multi-cycling, real duration and power, round-trip degradation, availability, outages and warranty throughput limits — all of which reduce realisable revenue and are addressed only in the Commercial tier.
  • Standardised frictions. Transaction costs and imbalance exposure are applied on standardised assumptions, not an individual asset's real contracts.
  • Baseline choice. Persistence is the academic-standard floor, not the best available commercial forecast; beating persistence is not the same as beating every competing forecaster.
  • Sample length. Some streams (balancing, ancillary) rest on short windows; short samples are noisy and are disclosed as such.
  • Not a forward promise. A walk-forward backtest removes hindsight but does not guarantee future performance.

14Tiers & assurance standards

The open benchmark is this public page — a standardised asset on the free regulated markets. The Commercial tier applies the same discipline to an owner's real asset and extends to the licensed feeds (continuous intraday order book) and markets with no free feed. Under the Commercial tier, limited- and reasonable-assurance grades (ISRE 2410, ISAE 3000 / 3410) denote engagements performed, and any opinion signed, by an independent accredited practitioner. Valuation and performance are reconciled on a GIPS-consistent basis; GIPS® compliance, where presented, rests on independent verification and is not self-certified. The free board itself is not an assurance engagement and carries no such opinion.

← back to the CausalGrids Lab board

CausalGrids Lab · Ikenga / Altinium Invest · part of the SSI Index — an independent, non-profit foundation project · Methodology, Limitations & Audit v1.0 · public benchmark not monetised · internally reproducible, not independent assurance · not realised cash, not investment advice · “Walk forward, don't look ahead.”