01 / RESEARCH FOUNDATION

Start with an economic mechanism, not a data pattern.

A systematic signal is a repeatable rule that converts observable information into a portfolio preference. The rule may use valuation, quality, trend, revisions, positioning or another measurable characteristic. But statistical association alone does not explain why the pattern should persist after discovery, publication and implementation.

A serious research memo begins with the mechanism: who is constrained, which information is processed slowly, what risk might be compensated and why the signal should survive competition. It then states the conditions under which that explanation would be wrong. This order matters. If the story is written after the backtest, it can become a flexible justification for noise.

Research standard

A factor is not established because it worked once. It earns attention when the idea, data, test design and implementation evidence agree.

Classic asset-pricing research provides useful reference points for common risk factors, but the published factor library has expanded enormously. Work by Campbell Harvey, Yan Liu and Heqing Zhu shows why multiple testing raises the evidentiary hurdle: the more ideas researchers try, the easier it becomes to find apparently significant results by chance.

02 / THE RESEARCH PIPELINE

Make every transformation reproducible.

01

Write the hypothesis

Define the economic channel, investable universe, rebalance rule, expected horizon and falsification test before looking at the final result.

02

Control the data clock

Use point-in-time constituents and publication dates. Adjust for restatements, delistings and corporate actions without allowing future knowledge into the past.

03

Define the signal

Specify formulas, lags, winsorization, neutralization and missing-value treatment. Small discretionary choices can create large historical differences.

04

Choose a benchmark

Compare against a simple alternative that represents what could have been done without the new signal, after matching risk where practical.

05

Reserve unseen evidence

Separate development, validation and final evaluation periods. Avoid repeatedly consulting the holdout until it becomes part of development.

06

Translate to positions

Apply liquidity, turnover, sector, factor, concentration and exposure limits before interpreting the result as an investable portfolio.

Versioning is part of research integrity. The code, source data snapshot, parameter set and result should be linked. A reader should be able to distinguish the originally tested specification from later improvements.

03 / BACKTEST OVERFITTING

Treat every research choice as a hidden trial.

Backtest overfitting occurs when a process selects the best-looking result from many alternatives and then evaluates that same result as if it were the only test. Trials include not just explicit models, but also changes in universe, start date, rebalance frequency, feature definitions, portfolio construction and exclusion rules.

Warning signWhy it mattersBetter practice
Many nearby specifications failThe chosen result may depend on a narrow parameter accident.Map the response surface and prefer broad, stable regions.
Performance is concentrated in one periodThe result may be a regime exposure rather than a persistent signal.Use subperiod, cross-market and leave-one-period-out tests.
The holdout is checked repeatedlyIt gradually becomes training data through human feedback.Maintain a research log and a genuinely untouched final test.
Reported metrics omit failed trialsConventional significance measures can overstate confidence.Record the search process and adjust interpretation for selection.

Research by David Bailey and co-authors formalizes how repeated selection can produce an optimized backtest that performs poorly out of sample. No single correction eliminates this risk. The practical defense is cumulative: fewer degrees of freedom, transparent trial counts, stable neighborhoods, independent review and evidence that survives realistic implementation.

04 / IMPLEMENTATION REALITY

A forecast becomes valuable only after the portfolio pays for it.

Gross simulated return is not the same as a tradable outcome. Turnover creates commissions, spreads, market impact, delay and opportunity cost. These frictions vary by security, order size, volatility, venue and market state. A fixed cost assumption can therefore make a fragile strategy appear more scalable than it is.

  • Liquidity: size positions relative to realistic traded volume and model the time needed to enter or exit.
  • Capacity: examine how expected cost and signal decay change as capital increases.
  • Risk overlap: decompose exposure to market, sector, size, value, momentum, volatility and crowded holdings.
  • Execution timing: align the assumed trade price with when the source data was actually observable and orders could reasonably be placed.
  • Financing and borrow: include short availability, borrow cost, margin and funding conditions when relevant.

The SEC’s staff report on algorithmic trading describes how automation can improve market efficiency while also creating operational and market risks. For a research team, the implication is straightforward: model logic, portfolio logic and execution controls should be tested as one chain.

05 / LIVE MONITORING

Monitor the thesis before monitoring the score.

Live deterioration does not automatically prove that a strategy is broken; normal uncertainty can produce difficult periods. Equally, a recent profit does not prove that the underlying mechanism remains sound. Monitoring should separate four questions: Is the data healthy? Is the signal behaving as designed? Are exposures within mandate? Is the economic explanation still credible?

  • Compare live feature distributions with the research sample and investigate sudden coverage or scale changes.
  • Track forecast calibration, turnover, realized costs, exposure drift and contribution by market environment.
  • Predefine review thresholds and the actions attached to them: investigate, reduce, pause or retire.
  • Record overrides and exceptions so discretionary intervention can be evaluated rather than forgotten.
  • Retest only with a controlled change process; do not repair a drawdown by searching until history looks attractive again.
Decision rule

A model should be retired when its economic basis, data integrity or controlled live evidence no longer supports use—not merely because a recent chart is uncomfortable.

06 / SOURCES AND FURTHER READING

Primary research behind this guide.

This article is an original synthesis of public research and official market-structure material.

  • 01
    Fama Research Readings

    University of Chicago collection linking foundational work on common risk factors and empirical asset pricing.

  • 02
    …and the Cross-Section of Expected Returns

    NBER paper by Harvey, Liu and Zhu on the multiple-testing problem created by a large factor search.

  • 03
    The Probability of Backtest Overfitting

    Research by Bailey and co-authors on selection bias in investment simulations.

  • 04
    SEC Staff Report on Algorithmic Trading

    An official review of algorithmic trading benefits, risks and market-structure considerations.

Continue learning