We validate every candidate trading signal in our energy-market data API before it goes near a customer - forward-return bucket analysis, regime splits, walk-forward out-of-sample testing, and economic significance measured against realized volatility, not just whether a backtest looks profitable. One candidate had the largest effect size in our entire factor library. We didn't ship it. Here's why.
One factor family we tested: natural gas storage percentile - how full a storage region is relative to its own 5-year seasonal average - against forward Henry Hub returns. We tested it across all 8 real EIA-tracked storage regions: a national aggregate (Lower 48) plus 7 sub-regions (Salt, Nonsalt, East, Midwest, South Central, Mountain, Pacific).
One region stood out immediately: Mountain. Economic effect +0.74x realized volatility at a ~1-quarter horizon - the single largest effect size of any factor anywhere in our library, bigger than the national aggregate, bigger than our crude-oil spread factor, bigger than anything else we'd tested. On a naive effect-size screen, Mountain wins outright and ships first.
We didn't ship it. Two independent tests, built to answer different questions, both flagged the same region before it ever reached that point.
First - out-of-sample stability. A held-out train/test split showed Mountain's own relationship barely surviving into the test window:
| Region | Out-of-sample stability |
|---|---|
| Salt | 100% |
| South Central | 92% |
| Lower 48 (national) | 91% |
| Nonsalt | 86% |
| Pacific | 74% |
| Midwest | 52% |
| East | 51% |
| Mountain | 3% |
Every other region we tested scored between 51% and 100%. Mountain scored 3%. Train and test were telling almost opposite stories.
Second, independent test - redundancy. We ran a partial-correlation test asking a different question: does Mountain's storage level tell you anything about future gas prices that the national aggregate doesn't already tell you, once you control for it? Mountain's raw correlation with forward returns was a real +0.25. Controlling for the national aggregate, it collapsed to +0.065 - and its own walk-forward hit rate against that controlled signal was 50%. A coin flip.
The smaller headline number - the national aggregate, at +0.47x - turned out to be the strongest, most cross-family-robust factor in our entire library. We confirmed it four separate times across unrelated tests: this same redundancy check, two independent combined-signal backtests (naive equal-weight and volatility-targeted), and by elimination in a later, entirely different cross-market investigation.
Mountain would have looked great in a pitch. It would have been wrong.
"Select on effect size, skip the out-of-sample check, and never ask whether a regional signal is actually independent of the aggregate it's drawn from" is a quiet, common failure mode in factor research. We built our validation pipeline specifically to catch it before it reaches a customer - and this is the clearest example it's produced so far.
Full breakdown of every factor we've tested - including the ones that survived, and the exact numbers behind this one - is in our public factor catalogue.
We're opening early access to a small number of quant developers, energy/gas traders, and researchers who want to stress-test this further. Get early access →