Just The Markets

Viewpoint · The case · Swing Trading

Out-of-Sample Testing Is the Only Part of a Backtest That Counts

Out of sample testing means judging a trading method on data it never saw while you tuned it. Everything else in a backtest is the method describing its own past.

AI-assisted, reviewed by the Just The Markets human editor: John Todora → 4 min read Published

A gloved hand holding a glass beaker marked in milliliters in a laboratory
Photo by RephiLe water on Unsplash

The position

Judge a trading method only on data held back from tuning; the in-sample result mostly measures how well you fitted the past.

Average trade times expected trades. That’s the projection. On a tuned method it’s usually the most flattering number you will ever see. The part of a backtest worth trusting is the stretch of history the rules never saw while you were building them, and on a hypothetical swing method, the gap between the two can look like the difference between a business and a hobby.

The tuned years fit the tuned years

Every rule in a trading method gets chosen somehow. You try a 20-day moving average, then a 50-day, then a 30-day, and keep the one that made the most money on the history in front of you. Each choice fits that history a little better. Some of the fit is a real pattern. Some is noise that happened to line up. Say the 30-day setting won because of a handful of trades in one violent month. That month won’t come back in the same shape, and the 30-day average has no particular reason to beat the 50-day next time.

There’s no way to tell those apart by looking at the tuned results. Both show up as profit. One test separates them. Run the finished rules on data that played no part in choosing them.

A hypothetical method on ten years of data

Here’s the setup, with illustrative numbers. You have ten years of daily data for a swing method that holds positions for days to weeks, and you tune every rule on the first seven years while the last three stay locked away, unopened, until the rules are finished and written down.

In the tuned years, the method takes 140 trades averaging +0.6R, where R is the amount risked per trade. That is a strong result. The held-back years produce 50 trades, so here’s the projection against what they delivered.

The method still made money. It made a sixth of what the tuned years promised. At +0.1R a trade, slightly worse fills could erase it. If you had sized positions, set an account target or quit a job on the +0.6R figure, the plan would have been built on a number that described the past and nothing else.

These are made-up figures. The shape is familiar, though. Expect a large drop out of sample by default.

Every extra knob makes it worse

Each parameter you adjust is another chance to fit noise. A method with a trend filter, an entry trigger, a stop rule, a profit target, a volume condition and a day-of-week rule, each tuned separately, can be made to look spectacular on any history. The in-sample line climbs with every tweak. The out-of-sample result, meanwhile, tends to drift the other way, because the extra precision was spent explaining accidents that won’t repeat.

Two habits help. Keep the rule count low and the parameter values round, so a 50-day average stays 50 and never becomes 47 just because 47 scored a little better on the tuned years, and when a small change to one value swings the result a lot, read that as a sign the result is fragile and the method is leaning on a lucky setting.

One look spoils the held-back data

The held-back years only work as a test once. Run the finished rules on them, see a weak result, go back and adjust a rule, and run them again: you’ve now tuned on those years too. They’ve become in-sample data, and you no longer have an honest test at all.

So the order is strict. Finish the rules. Write them down. Then open the held-back data, run it once, and accept the number. If it’s poor, the right response is a new idea tested on fresh data, and the temptation to rescue the old idea is exactly what the split exists to stop.

A walk-forward test stretches the same principle across the whole history: tune on one window, test on the next, roll forward, and stitch the test windows together. It gives more out-of-sample trades from the same data.

The objection: markets change, so old data misleads anyway

A fair counter is that any backtest, in-sample or out, describes markets that no longer exist. Volatility regimes shift, and so do the players. So why treat the held-back years as special?

Because they answer a narrower question, and they answer it honestly. The held-back result can’t tell you the method will work next year. It can tell you whether the method’s edge survived contact with data it wasn’t fitted to, which is the minimum any method has to pass. A method that fails out of sample has no evidence behind it. One that passes has a little. Charge realistic commissions and slippage on every trade in both halves, or the held-back figure flatters too.

The out-of-sample expectancy is also the right number to size from. Plugging +0.1R into your plans means a thinner edge and deeper drawdowns before it shows, which is the sober version; see sizing for the losing streak you haven’t had and run the figures through the losing streak calculator.

Where the held-back number stops meaning much

The argument needs enough trades. Fifty trades is thin, and a method that produced only 12 trades in its held-back years gives a figure one lucky or unlucky run can decide. Two big winners out of 12 can make a poor method look good, and two outsized losers can sink a sound one, so a small out-of-sample count calls for a longer history, a walk-forward test, or simply more humility about the result.

Judge the method on the held-back data, size from the held-back expectancy, and treat the tuned result as a description of how well you fitted the past. The rules-first course and the swing trading hub cover building the rules before any money goes in.

People also ask

How much data should you hold back for out-of-sample testing?

There is no fixed rule. A common split holds back somewhere between a fifth and a third of the history, and the held-back stretch needs enough trades to mean something. A method that trades rarely may need a longer history overall, or a walk-forward test that reuses the data in rolling windows.

What is the difference between in-sample and out-of-sample results?

In-sample results come from the data used to choose and tune the rules, so they show how well the rules fit that particular history. Out-of-sample results come from data the rules never touched during tuning, which makes them the nearest thing a backtest has to live trading.

Why do backtests look better than live trading?

Every parameter chosen because it improved the historical result also fits some of that history's random noise. Live trading has different noise, so part of the tuned edge disappears. Costs and slippage that the test understated take away more.