More Than Half of Published Trading-Strategy "Alphas" Are False Discoveries
query explainer from the socialgood index, 2026-09-14.
query-explainer · 2026-09-14 · 1 numbers checked against the index
More than 50% of published factor alphas are false discoveries. That is the finding, not a rough guess: Campbell Harvey, Yan Liu, and Heqing Zhu, in their 2016 paper, showed that when you re-run the standard significance test used across the finance literature with a corrected bar — a t-statistic hurdle of t > 3.0 rather than the conventional threshold — a majority of the "market-beating" factors published in peer-reviewed journals fail to clear it.
The seller's claim and the incumbent answer look identical: a peer-reviewed paper tested a strategy and it beat the market. That is treated as the gold standard of proof. The audited reality, per Harvey, Liu, and Zhu, is that this standard was never adjusted for how many strategies were tested to find the one that got published. Multiple hypothesis testing without a corrected t-statistic hurdle means researchers (and vendors citing them) can run dozens or hundreds of factor variations, report the one that clears a weak significance bar by chance, and call it a discovery. Harvey, Liu, and Zhu's t > 3.0 hurdle is the correction — and applying it flips more than half of the published factor alphas from "significant" to "false discovery."
The mechanism for retail traders is the same one, dressed differently: any claim of a "tested" or "backtested" edge inherits this same multiple-testing problem unless the person making the claim discloses how many variations were tried before this one was reported. A single backtest result, without that disclosure, carries the same false-discovery risk Harvey, Liu, and Zhu quantified in the academic literature — a risk that exceeds 50% by their measurement.
The operational rule: never accept a single backtest or a single published alpha at face value. Require the t-statistic and compare it against the t > 3.0 hurdle, and treat any strategy report that omits how many variations were tested as unverified.