Just The Markets

Walkthrough · Step by step · Stock Trading

How to Run a Walk-Forward Test on a Trading Strategy

A walk forward test tunes a strategy on one stretch of history and scores it on the next, over and over, so every result you keep comes from data the settings never saw.

AI-assisted, reviewed by the Just The Markets human editor: John James → 4 min read Published

Short answer

Split your history into windows, tune the strategy's settings on one window, then run those exact settings on the following period without changes. Roll the windows forward and repeat, stitch together only the test periods, and compare that record with the tuned results. Include costs and slippage from the first test.

  1. 1

    Split the history into windows

    Choose a tuning length and a test length before you run anything, such as a four-year tuning window and a one-year test.

  2. 2

    Tune on the first window

    Optimize the strategy's settings on the tuning window only, with trading costs and slippage already included.

  3. 3

    Test on the next period unchanged

    Run those exact settings on the period that follows. No adjustments, however tempting the result.

  4. 4

    Roll forward and repeat

    Move both windows forward by one test length, retune on the new window and test on the next unseen period, until the data runs out.

  5. 5

    Combine only the test periods

    Stitch the out-of-sample results together into one record. The tuning results are left out entirely.

  6. 6

    Compare with the tuned results

    Set the stitched test record against the performance in the tuning windows. A large drop means the settings were fitted to noise.

A backtest tuned on ten years of data and then scored on the same ten years has measured one thing, how well the settings fit the past. It says nothing about next year. Walk-forward testing asks what the rules would have done on data they had never seen, and it asks repeatedly across the history, one unseen stretch at a time, so a single lucky year cannot carry the verdict.

How do you split the history?

Settle the lengths before you look at a single result: a tuning window, where the settings are chosen, and a test window, where they are used. The choice depends on how often the strategy trades. Each tuning window needs enough trades to pick settings on, and each test window needs enough to judge.

With ten years of daily data and a strategy that trades a few dozen times a year, four years to tune and one year to test is a reasonable start.

How does one pass work?

Tune on years 1 to 4. Find the moving-average length, the stop distance, whatever the strategy’s settings are, that performed best in those years. Write them down.

Then run those exact settings on year 5. Don’t touch them. If year 5 looks bad and a small tweak would fix it, the tweak is the thing the test is designed to catch.

Costs go in from this first pass. A hypothetical strategy trading 50 times a year with $40 of commission and slippage per round trip gives up $2,000 a year before it earns anything, and a tuning run without costs will happily choose settings that trade too often to survive them.

How do you roll it forward?

Slide both windows along by one test length and do it again.

Each pass may choose different settings. That’s fine. It is how the strategy would really be run, retuned on recent data at regular intervals. What you are testing is the whole process, the rules plus the way you choose their settings, and that is the thing you would actually be trading with real money.

This version rolls, dropping the oldest year each time. An anchored version keeps year 1 as the start and lets the tuning window grow. Both are valid. Pick one before you start.

What do you do with the results?

Stitch the test years together into one record. Leave every tuning result out. That record is the one to judge: its return, its drawdown, its longest losing streak.

Then compare it with the tuning windows. Suppose, hypothetically, the tuning windows averaged 20% a year and the stitched test years averaged 8%. The test kept 8 / 20 = 40% of the tuned performance. Some drop is normal, since tuned results always flatter. A test record that falls to nothing, or turns negative, says the settings were fitted to noise in each window and the strategy has no edge you can find this way.

Look at the settings each pass chose, as well. If the best moving-average length drifts gently, say from 45 to 50 days across the passes, the strategy has a stable region worth trusting. If it jumps from 20 to 90 and back, each window’s winner was a fluke of that window. That instability is a finding in its own right.

Watch the spread of test years too. One great year and five flat ones is a different strategy from six modest ones, even with the same average, and the losing streak calculator will tell you how deep a run of losses the stitched record implies for your position size.

Where does it stop protecting you?

It can’t show you a market regime missing from all ten years. It can’t catch a data error that runs through the whole history, such as a list of stocks that leaves out companies that were delisted. And six test years is still a small sample of years.

The case for trusting only out-of-sample numbers is argued in out-of-sample testing is the only part that counts. For building rules a test can actually run, the course Rules First, Money Later starts from scratch, and its lesson on backtesting biases and costs covers the errors a walk-forward test won’t catch on its own.

People also ask

How long should the tuning and test windows be?

Long enough that each tuning window holds a fair number of trades and a mix of market conditions, and each test window holds enough trades to mean something. A strategy that trades a few times a year needs years per window; one that trades daily can use months. There is no standard ratio, so decide before you look at results.

What is the difference between anchored and rolling walk-forward testing?

In a rolling test every tuning window has the same length and slides forward, dropping the oldest data. In an anchored test the start date stays fixed and each tuning window grows to include everything up to the next test period. Rolling adapts faster to change; anchored uses more history for each set of settings.