In our last post, we explained why we excluded our library's largest raw effect size - a natural gas storage signal that looked great until it failed two independent robustness tests. This post is about the opposite problem: a factor that passed every test we threw at it, with the largest economic effect we've measured - and we still didn't publish it for months.
In August, we validated ercot_spark_spread - ERCOT
wholesale power prices against Henry Hub gas costs - through our
standard five-step pipeline: predictive content, monotonicity, regime
stability, mandatory walk-forward out-of-sample testing, and economic
significance.
It cleared every gate we use, including sign consistency, horizon consistency, out-of-sample stability, and crisis robustness - 100% on each. Its economic effect ranged from -0.60x to -0.89x realized volatility at 20 trading days, depending on the settlement point. That's the largest economic effect in our factor library, ahead of our previous best, WTI/ Brent crude spread mean reversion (-0.82x).
By the headline numbers, this looked like a clear ship-it result. We didn't ship it.
Our validation pipeline measures five dimensions. It also tracks something just as important: how much independent evidence sits behind the score.
Our other core factors - crude spread, storage - are backed by 12 to 27 real annual walk-forward folds, built from years of price history. ERCOT was different. When we first ran the validation, the real settlement-price history available to us was about six months old. That produced real results, but only a handful of independent out-of-sample folds behind them.
A 100% hit rate across two folds means "both times we checked, it worked." That's useful evidence. It is not the same claim as a 100% hit rate across twelve years of independent checks - even though the printed number looks identical either way.
So we added an evidence-depth tier to every scorecard: LOW, MODERATE,
FULL. The factor was classified core factor (PROVISIONAL)
rather than plain core factor - not because the evidence
was bad, but because there wasn't enough of it yet.
Instead of changing the methodology or lowering the threshold, we let the dataset grow. ERCOT prices kept refreshing automatically, and a manual archive backfill extended real coverage back to January 2023. Every few weeks we re-ran the same validation pipeline - same methodology, same thresholds, just more data behind it.
We also added two checks aimed specifically at false confidence. First, a pooled-sample test across all 15 ERCOT settlement points, to confirm the apparent agreement between nodes wasn't just the same correlation counted fifteen times. Second, a bucket-count sensitivity sweep from a 2-way split through a 50-way split, to confirm the result wasn't an artifact of choosing exactly 10 buckets. Both held.
By late August, the walk-forward had grown to 12 real quarterly folds - the same evidence depth as our other core factors. We ran the pipeline one more time. The verdict didn't change. The confidence level did: LOW → MODERATE → FULL.
There's another reason we waited. The early, thin-evidence version of this factor reported an economic effect of -1.39x to -1.81x - by far the largest number in our library. With the deeper dataset, the estimate fell to -0.60x to -0.89x. Still the largest effect in the library, but roughly half the original size.
The shorter sample had, by chance, contained a disproportionate share of ERCOT's biggest price swings. More data gave us a less spectacular number - and a more trustworthy one.
These two posts describe opposite failure modes. Mountain storage: an attractive number that didn't survive scrutiny, excluded despite looking like our best result. ERCOT spark spread: an attractive number that did survive scrutiny - repeatedly, across five separate robustness checks - but wasn't published until the evidence behind it was as deep as everything else in the library.
That's what the validation process is for. Not to find the biggest number. To determine which numbers deserve to be trusted.
ercot_spark_spread is live now, across all 15 ERCOT hubs
and load zones, in the public API. It took longer to trust the factor
than it did to write the code.
Full breakdown of this factor, its real validation tiers, and the exact numbers behind it: ERCOT spark spread in our factor library.
We're opening early access to a small number of quant developers, energy/gas traders, and researchers who want to stress-test this further. Get early access →