← Ordinex Energy Intelligence

Factor Validation Methodology

How every published factor earns its tier - not just whether it "worked" once.

Most energy-data feeds stop at "here's the number." We test whether a signal actually predicts anything, out-of-sample, before it ever reaches the API - and we publish the honest misses alongside the real ones. This page is the full version of what the homepage summarizes in one paragraph.

1. Point-in-time correctness comes first

A backtest that's statistically sound but leaks future information is worthless. Every factor is built on a pipeline that enforces this before any statistics are computed:

StepWhat it does
Source publishesA source reports a value days after the "as of" date it actually describes - e.g. EIA's weekly storage figure.
Reporting lag respectedA real, per-source min_lag_days safety margin - a factor is never paired against a target move that happened before the observation was actually public.
Nothing overwrittenEvery raw fact and factor observation is append-only, insert-if-not-exists - a later revision never silently replaces history.
Walk-forward testedOne held-out fold per calendar year (or quarter, for shorter-history sources) - never fit and tested on the same window.
Scored, not just shippedA five-dimension quality scorecard decides the tier - see below.

The live API exposes this directly: every factor observation carries a known_at field distinct from the period it describes, and GET /v1/factors/latest only ever returns a value that was genuinely knowable as of the requested date - not the nominal period date.

2. The five-step statistical framework

Every candidate factor goes through the same five checks before it earns any tier at all:

  1. Predictive content - decile bucket analysis on forward returns, plus the long-short spread and rank correlation across the full sample, not just the 10 bucket means.
  2. Monotonicity - does the effect move in one direction across buckets, or is it a fluke concentrated in one bucket.
  3. Regime stability - named calendar-window splits, including real crisis periods (2020 COVID, 2022 Ukraine energy shock, Winter Storm Uri where applicable) - does the sign hold up in stress, not just on average.
  4. Mandatory walk-forward validation - one held-out fold per calendar year (or quarter) - never fit and tested on the same window. A factor with only a handful of real folds is marked with an explicit confidence tier (LOW/MODERATE/FULL) rather than treated as equally proven.
  5. Economic significance - is the statistically real effect large enough to matter, measured against the target's own realized volatility - not just "significant," but worth trading.

3. The Factor Quality Score

A factor's tier comes from five scored dimensions, not a single correlation number:

DimensionWhat it measures
Sign consistencyFraction of walk-forward folds that agree on direction.
Signal strengthAverage magnitude of the rank correlation across horizons.
Crisis robustnessWhether the sign holds during real crisis-labeled regimes specifically, not just on average.
Horizon consistencyWhether every tested horizon agrees on direction.
OOS stabilityWhether the held-out test window is at least as strong as the training window - a sign flip scores zero.

Verdicts use fixed, published thresholds, not a threshold chosen after seeing a specific factor's own result: core factor requires sign consistency ≥85%, horizon consistency exactly 100%, OOS stability ≥50%, and crisis robustness ≥50% (or no crisis regime in the sample at all); regime-dependent component requires sign consistency ≥60% and horizon consistency ≥75%; anything weaker is reported as insufficient evidence - and published as such, not quietly dropped.

The scorecard is designed to catch exactly the failure mode a raw-effect-size screen would miss. One storage region (Mountain) had the single largest raw economic effect of any region we tested (+0.74x) - the number that would look best in a naive screen. Two independent tests flagged it as untrustworthy anyway: the worst out-of-sample stability of any region tested (3%), and a redundancy check showing its apparent edge evaporates to a coin flip once the national aggregate is controlled for. It's excluded from the live API despite the headline number. Full story: Why we didn't ship our best backtest result.

4. Redundancy testing before promotion

A flawless standalone scorecard is not on its own evidence a factor is independent information. Before any factor reaches the public API, we run a partial-correlation redundancy check against the factors we already serve - does it still add predictive power once a related, already-validated signal is controlled for. This is the same check that caught a factor scoring a clean "core factor" on every standalone dimension, then found it retained only ~6% of its raw predictive power once controlled for a related signal already in the library - a real result, published, not hidden because the standalone number looked strong.

5. What this doesn't claim

Backtested effect sizes are historical, gross of transaction costs, slippage, and capacity - not a forward return guarantee. A factor combined with others does not automatically beat the single strongest factor alone on a risk-adjusted basis - we tested this directly, twice, under both a naive equal-weight and a risk-normalized construction, and report the honest result rather than a composite "Intelligence Score" the evidence doesn't support.

See it applied

Every real number behind every verdict - full backtest tables, regime splits, and walk-forward fold counts - lives in FACTOR_CATALOGUE.md. To see this framework applied to a specific factor, start with ERCOT spark spread (a clean 100% pass) or storage response to percentile (a real, honest near-miss that's still our strongest cross-market signal) - or browse the full Factor Library.

Every claim on this page traces to a real, dated backtest result - not marketing copy. Full derivation: FACTOR_CATALOGUE.md.

Free early access to the API: Get early access →