← Blog

Why we didn't ship our best backtest result

2026-08-28 · Ordinex US Energy Data

We validate every candidate trading signal in our energy-market data API before it goes near a customer - forward-return bucket analysis, regime splits, walk-forward out-of-sample testing, and economic significance measured against realized volatility, not just whether a backtest looks profitable. One candidate had the largest effect size in our entire factor library. We didn't ship it. Here's why.

The setup

One factor family we tested: natural gas storage percentile - how full a storage region is relative to its own 5-year seasonal average - against forward Henry Hub returns. We tested it across all 8 real EIA-tracked storage regions: a national aggregate (Lower 48) plus 7 sub-regions (Salt, Nonsalt, East, Midwest, South Central, Mountain, Pacific).

The headline number

One region stood out immediately: Mountain. Economic effect +0.74x realized volatility at a ~1-quarter horizon - the single largest effect size of any factor anywhere in our library, bigger than the national aggregate, bigger than our crude-oil spread factor, bigger than anything else we'd tested. On a naive effect-size screen, Mountain wins outright and ships first.

What actually happened

We didn't ship it. Two independent tests, built to answer different questions, both flagged the same region before it ever reached that point.

First - out-of-sample stability. A held-out train/test split showed Mountain's own relationship barely surviving into the test window:

RegionOut-of-sample stability
Salt100%
South Central92%
Lower 48 (national)91%
Nonsalt86%
Pacific74%
Midwest52%
East51%
Mountain3%

Every other region we tested scored between 51% and 100%. Mountain scored 3%. Train and test were telling almost opposite stories.

Second, independent test - redundancy. We ran a partial-correlation test asking a different question: does Mountain's storage level tell you anything about future gas prices that the national aggregate doesn't already tell you, once you control for it? Mountain's raw correlation with forward returns was a real +0.25. Controlling for the national aggregate, it collapsed to +0.065 - and its own walk-forward hit rate against that controlled signal was 50%. A coin flip.

Two tests, built for entirely different questions - one about temporal stability, one about redundancy with an existing signal - independently flagged the same region as untrustworthy. That's not a coincidence, and it's not something a single backtest run would have caught.

The boring number was the real one

The smaller headline number - the national aggregate, at +0.47x - turned out to be the strongest, most cross-family-robust factor in our entire library. We confirmed it four separate times across unrelated tests: this same redundancy check, two independent combined-signal backtests (naive equal-weight and volatility-targeted), and by elimination in a later, entirely different cross-market investigation.

Mountain would have looked great in a pitch. It would have been wrong.

Why this is worth talking about

"Select on effect size, skip the out-of-sample check, and never ask whether a regional signal is actually independent of the aggregate it's drawn from" is a quiet, common failure mode in factor research. We built our validation pipeline specifically to catch it before it reaches a customer - and this is the clearest example it's produced so far.

Full breakdown of every factor we've tested - including the ones that survived, and the exact numbers behind this one - is in our public factor catalogue.

We're opening early access to a small number of quant developers, energy/gas traders, and researchers who want to stress-test this further. Get early access →