01 / RESEARCH FOUNDATION
Start with an economic mechanism, not a data pattern.
A systematic signal is a repeatable rule that converts observable information into a portfolio preference. The rule may use valuation, quality, trend, revisions, positioning or another measurable characteristic. But statistical association alone does not explain why the pattern should persist after discovery, publication and implementation.
A serious research memo begins with the mechanism: who is constrained, which information is processed slowly, what risk might be compensated and why the signal should survive competition. It then states the conditions under which that explanation would be wrong. This order matters. If the story is written after the backtest, it can become a flexible justification for noise.
A factor is not established because it worked once. It earns attention when the idea, data, test design and implementation evidence agree.
Classic asset-pricing research provides useful reference points for common risk factors, but the published factor library has expanded enormously. Work by Campbell Harvey, Yan Liu and Heqing Zhu shows why multiple testing raises the evidentiary hurdle: the more ideas researchers try, the easier it becomes to find apparently significant results by chance.
02 / THE RESEARCH PIPELINE
Make every transformation reproducible.
Write the hypothesis
Define the economic channel, investable universe, rebalance rule, expected horizon and falsification test before looking at the final result.
Control the data clock
Use point-in-time constituents and publication dates. Adjust for restatements, delistings and corporate actions without allowing future knowledge into the past.
Define the signal
Specify formulas, lags, winsorization, neutralization and missing-value treatment. Small discretionary choices can create large historical differences.
Choose a benchmark
Compare against a simple alternative that represents what could have been done without the new signal, after matching risk where practical.
Reserve unseen evidence
Separate development, validation and final evaluation periods. Avoid repeatedly consulting the holdout until it becomes part of development.
Translate to positions
Apply liquidity, turnover, sector, factor, concentration and exposure limits before interpreting the result as an investable portfolio.
Versioning is part of research integrity. The code, source data snapshot, parameter set and result should be linked. A reader should be able to distinguish the originally tested specification from later improvements.
03 / BACKTEST OVERFITTING
Treat every research choice as a hidden trial.
Backtest overfitting occurs when a process selects the best-looking result from many alternatives and then evaluates that same result as if it were the only test. Trials include not just explicit models, but also changes in universe, start date, rebalance frequency, feature definitions, portfolio construction and exclusion rules.
| Warning sign | Why it matters | Better practice |
|---|---|---|
| Many nearby specifications fail | The chosen result may depend on a narrow parameter accident. | Map the response surface and prefer broad, stable regions. |
| Performance is concentrated in one period | The result may be a regime exposure rather than a persistent signal. | Use subperiod, cross-market and leave-one-period-out tests. |
| The holdout is checked repeatedly | It gradually becomes training data through human feedback. | Maintain a research log and a genuinely untouched final test. |
| Reported metrics omit failed trials | Conventional significance measures can overstate confidence. | Record the search process and adjust interpretation for selection. |
Research by David Bailey and co-authors formalizes how repeated selection can produce an optimized backtest that performs poorly out of sample. No single correction eliminates this risk. The practical defense is cumulative: fewer degrees of freedom, transparent trial counts, stable neighborhoods, independent review and evidence that survives realistic implementation.
04 / IMPLEMENTATION REALITY
A forecast becomes valuable only after the portfolio pays for it.
Gross simulated return is not the same as a tradable outcome. Turnover creates commissions, spreads, market impact, delay and opportunity cost. These frictions vary by security, order size, volatility, venue and market state. A fixed cost assumption can therefore make a fragile strategy appear more scalable than it is.
- Liquidity: size positions relative to realistic traded volume and model the time needed to enter or exit.
- Capacity: examine how expected cost and signal decay change as capital increases.
- Risk overlap: decompose exposure to market, sector, size, value, momentum, volatility and crowded holdings.
- Execution timing: align the assumed trade price with when the source data was actually observable and orders could reasonably be placed.
- Financing and borrow: include short availability, borrow cost, margin and funding conditions when relevant.
The SEC’s staff report on algorithmic trading describes how automation can improve market efficiency while also creating operational and market risks. For a research team, the implication is straightforward: model logic, portfolio logic and execution controls should be tested as one chain.
05 / LIVE MONITORING
Monitor the thesis before monitoring the score.
Live deterioration does not automatically prove that a strategy is broken; normal uncertainty can produce difficult periods. Equally, a recent profit does not prove that the underlying mechanism remains sound. Monitoring should separate four questions: Is the data healthy? Is the signal behaving as designed? Are exposures within mandate? Is the economic explanation still credible?
- Compare live feature distributions with the research sample and investigate sudden coverage or scale changes.
- Track forecast calibration, turnover, realized costs, exposure drift and contribution by market environment.
- Predefine review thresholds and the actions attached to them: investigate, reduce, pause or retire.
- Record overrides and exceptions so discretionary intervention can be evaluated rather than forgotten.
- Retest only with a controlled change process; do not repair a drawdown by searching until history looks attractive again.
A model should be retired when its economic basis, data integrity or controlled live evidence no longer supports use—not merely because a recent chart is uncomfortable.
06 / SOURCES AND FURTHER READING
Primary research behind this guide.
This article is an original synthesis of public research and official market-structure material.
- 01Fama Research Readings↗
University of Chicago collection linking foundational work on common risk factors and empirical asset pricing.
- 02…and the Cross-Section of Expected Returns↗
NBER paper by Harvey, Liu and Zhu on the multiple-testing problem created by a large factor search.
- 03The Probability of Backtest Overfitting↗
Research by Bailey and co-authors on selection bias in investment simulations.
- 04SEC Staff Report on Algorithmic Trading↗
An official review of algorithmic trading benefits, risks and market-structure considerations.