Most energy-data feeds stop at "here's the number." We test whether a signal actually predicts anything, out-of-sample, before it ever reaches the API - and we publish the honest misses alongside the real ones. This page is the full version of what the homepage summarizes in one paragraph.
A backtest that's statistically sound but leaks future information is worthless. Every factor is built on a pipeline that enforces this before any statistics are computed:
| Step | What it does |
|---|---|
| Source publishes | A source reports a value days after the "as of" date it actually describes - e.g. EIA's weekly storage figure. |
| Reporting lag respected | A real, per-source min_lag_days safety margin - a factor is never paired against a target move that happened before the observation was actually public. |
| Nothing overwritten | Every raw fact and factor observation is append-only, insert-if-not-exists - a later revision never silently replaces history. |
| Walk-forward tested | One held-out fold per calendar year (or quarter, for shorter-history sources) - never fit and tested on the same window. |
| Scored, not just shipped | A five-dimension quality scorecard decides the tier - see below. |
The live API exposes this directly: every factor observation
carries a known_at field distinct from the period it
describes, and GET /v1/factors/latest only ever
returns a value that was genuinely knowable as of the requested
date - not the nominal period date.
Every candidate factor goes through the same five checks before it earns any tier at all:
A factor's tier comes from five scored dimensions, not a single correlation number:
| Dimension | What it measures |
|---|---|
| Sign consistency | Fraction of walk-forward folds that agree on direction. |
| Signal strength | Average magnitude of the rank correlation across horizons. |
| Crisis robustness | Whether the sign holds during real crisis-labeled regimes specifically, not just on average. |
| Horizon consistency | Whether every tested horizon agrees on direction. |
| OOS stability | Whether the held-out test window is at least as strong as the training window - a sign flip scores zero. |
Verdicts use fixed, published thresholds, not a threshold chosen after seeing a specific factor's own result: core factor requires sign consistency ≥85%, horizon consistency exactly 100%, OOS stability ≥50%, and crisis robustness ≥50% (or no crisis regime in the sample at all); regime-dependent component requires sign consistency ≥60% and horizon consistency ≥75%; anything weaker is reported as insufficient evidence - and published as such, not quietly dropped.
A flawless standalone scorecard is not on its own evidence a factor is independent information. Before any factor reaches the public API, we run a partial-correlation redundancy check against the factors we already serve - does it still add predictive power once a related, already-validated signal is controlled for. This is the same check that caught a factor scoring a clean "core factor" on every standalone dimension, then found it retained only ~6% of its raw predictive power once controlled for a related signal already in the library - a real result, published, not hidden because the standalone number looked strong.
Backtested effect sizes are historical, gross of transaction costs, slippage, and capacity - not a forward return guarantee. A factor combined with others does not automatically beat the single strongest factor alone on a risk-adjusted basis - we tested this directly, twice, under both a naive equal-weight and a risk-normalized construction, and report the honest result rather than a composite "Intelligence Score" the evidence doesn't support.
Every real number behind every verdict - full backtest tables, regime splits, and walk-forward fold counts - lives in FACTOR_CATALOGUE.md. To see this framework applied to a specific factor, start with ERCOT spark spread (a clean 100% pass) or storage response to percentile (a real, honest near-miss that's still our strongest cross-market signal) - or browse the full Factor Library.
Every claim on this page traces to a real, dated backtest result - not marketing copy. Full derivation: FACTOR_CATALOGUE.md.
Free early access to the API: Get early access →