Backtesting a Trading Strategy: Steps and Limits
Backtesting a trading strategy means applying your entry, exit and risk rules to historical bars so you can inspect the resulting trades, drawdowns and statistics. It is a decision-support exercise, not a forecast: markets change, costs bite, and past sequences do not repeat on demand. Trading involves risk of loss.
What backtesting a trading strategy actually does
You take a written rule set (signals, filters, position size, stops, targets) and replay it bar by bar on a chosen symbol and timeframe. The output is a list of hypothetical fills, equity curve, win rate, average win versus average loss, maximum drawdown and similar numbers. Trading strategy backtesting does not prove the next month will look the same. It only shows how those rules behaved on that sample under the assumptions you coded.
Use it to reject obviously broken ideas, to size risk, and to see whether a rule set survives realistic costs. Do not use it as permission to scale up.
Choose the sample before you touch the rules
Split data. Typical split for a liquid futures or FX pair: 60 percent in-sample for design, 20 percent validation, 20 percent untouched hold-out. For crypto, prefer more recent years because microstructure and sessions shifted. For stocks, include both bull and bear regimes; a 2017 - 2021 only sample is not a test.
Bar type matters. Regular session versus 24-hour, tick versus 1-minute versus daily: each changes fill timing. If you trade the New York open, your backtest must include that session’s first prints, not a smoothed daily close.
Minimum length that is still useful
- Intraday scalps: at least 6 - 12 months of 1-minute or 5-minute bars, plus a second symbol for a sanity check.
- Swing on daily: 8 - 15 years if the market has that history; fewer years if the instrument is new.
- Fewer than 100 closed trades in the whole test is usually noise, not a result.
Code the rules so they cannot cheat
Look-ahead bias is the usual killer. A signal that uses the close of bar t must not fire until bar t+1 opens, unless you explicitly model intra-bar fills. Non-repainting Pine Script v6 studies help here because they do not rewrite past plots when new bars arrive. If an indicator redraws, your historical trades are fiction.
Fill model: market orders at next open, or limit at a stated offset. Never assume you always get the exact high or low of a wick. Slippage: 0.5 - 2 ticks on liquid futures, 0.5 - 3 pips on major FX, wider on thin crypto pairs. Commission: your broker’s round-turn, not zero. For US stocks, include the spread plus a few cents per share; for crypto, include taker fees (often 0.04 - 0.10 percent per side on major venues).
- Write the rules in plain English first: if X and Y, enter long, stop at Z, target at W, max risk R percent of equity.
- Translate to code or to a TradingView strategy() script with process_orders_on_close or explicit barstate checks.
- Add costs and a conservative fill delay.
- Run in-sample. If the equity curve is a straight line up, you overfit. Stop and simplify.
- Freeze parameters. Run validation, then hold-out. If performance collapses, the idea failed the test.
Metrics that actually constrain you
Win rate alone is useless. A 70 percent win rate with 1:0.4 reward-to-risk can still lose money after costs. Expectancy (average R per trade) and maximum drawdown in R or in percent of equity matter more. Profit factor above 1.3 after costs on a decent sample is a starting filter, not a trophy.
| Metric | What to look at | Typical trap |
|---|---|---|
| Expectancy | Mean win minus mean loss, in R or currency | Ignoring costs turns a small edge into a loss |
| Max drawdown | Peak-to-trough on the equity curve | One lucky streak hides a 40 percent hole |
| Trade count | Closed trades in each split | Fewer than 30 in hold-out is anecdote |
| Exposure | Percent of bars in a position | Always-in systems look great until a gap |
| Outlier share | Percent of P&L from the best 5 percent of trades | One news spike is not a strategy |
Walk-forward: re-optimise on a rolling window (for example 12 months), then trade the next 3 months with frozen parameters. If the walk-forward equity is still acceptable after costs, you have slightly more evidence than a single split. You still do not have a guarantee.
Overfitting and other ways to lie to yourself
Too many parameters, too many filters, and curve-fitting to one instrument are the standard failures. If you add a session filter, a volume spike, a FVG, a BOS and a VWAP band until the in-sample looks perfect, you have a description of the past, not a rule. Cap the number of free knobs. Prefer structure and risk rules that you can explain in one paragraph.
Regime change: a mean-reversion rule that worked in a range will fail in a trend. Test across both. If you only backtest the last bull run in crypto, you have not tested the strategy.
Position sizing in the test should match how you will trade: fixed fractional (for example 0.5 - 1 percent of equity at risk per idea) rather than fixed lot. Otherwise the equity curve is not comparable to live risk.
Manual versus coded tests
Manual chart markup is fine for a first pass on market structure (BOS/CHOCH, liquidity, FVG). It is slow and biased: you skip the ugly trades. Coded tests on TradingView or a local engine are repeatable. Use both: visual check that the code matches the idea, then the full sample with costs.
Precision indicators sold as one-time purchases can supply non-repainting structure, VWAP, sessions or volume overlays you then wrap in a strategy() script. They remain analysis tools. They do not place your orders or promise a return.
After the numbers: what you still cannot know
Live execution, partial fills, halted sessions, news spikes and your own behaviour are outside the backtest. Paper trade the same rules for a few dozen live bars before any real size. Keep a journal of slippage versus the model. If live fills are systematically worse, cut size or stop.
Trading involves risk. A clean backtest is a filter, not an all-clear. Size so that a repeat of the worst historical drawdown, plus a buffer, does not take you out of the game.
Frequently asked questions
What is backtesting a trading strategy in one sentence?
It is replaying written entry, exit and risk rules on historical bars to produce a list of hypothetical trades and statistics, under explicit cost and fill assumptions.
How many trades do I need before I trust the result?
Aim for well over 100 closed trades across in-sample plus hold-out. Below about 30 in the untouched sample you are looking at noise. More trades still do not make the future match the past.
Should I include commissions and slippage?
Yes. Zero-cost tests overstate expectancy. Use your broker’s round-turn plus a conservative tick or pip slippage. Crypto taker fees of a few basis points per side are enough to erase a thin edge.
Is a high win rate a good sign?
Not by itself. Combine win rate with average win versus average loss (expectancy) and with maximum drawdown. A high win rate with tiny winners and occasional large losers often fails after costs.
Can I backtest on the free TradingView plan?
Yes for many strategy() scripts and bar replays, with the usual bar and history limits of that plan. Heavier history and more concurrent charts sit on paid plans. The test quality still depends on your fill model, not the subscription.
Does a strong backtest mean I should trade live size?
No. Use it to reject bad ideas and to set risk. Paper the same rules, compare live slippage, and size so a worse-than-historical drawdown is survivable. Past sequences are not a promise of profit.