Backtest Looks Great. That Doesn’t Mean the Strategy Is Good

Wait 5 sec.

Backtest Looks Great. That Doesn’t Mean the Strategy Is GoodBitcoin / US DollarCOINBASE:BTCUSDChristopherDownie A lot of traders see a clean equity curve and immediately think they found something. The profits are going up. The drawdown looks manageable. The win rate is decent. So, the strategy must be good, right? Not really. A profitable backtest is only the start of the conversation. It tells you what happened inside one specific set of conditions, with one set of assumptions, over one period of market history. It doesn’t tell you how easily the whole thing can fall apart. And honestly, thats the part traders should probably care about more. Profit Is Usually the First Thing We Look At It makes sense. You test a strategy because you want to know if it can make money. Nobody is spending hours building a system just to admire the entry logic. But total profit can hide a lot. Maybe most of the money came from one strong year. Maybe the strategy performed really well on one asset, then did basically nothing everywhere else. Maybe a handful of oversized winners carried hundreds of average or losing trades. The final number still looks good though. Thats the problem. A backtest can be profitable without being dependable. Those are not the same thing. First, What Is the Strategy Actually Taking Advantage Of? Before looking at indicators, settings or optimisation results, try explaining the strategy in plain language. Why should it work? Not how does it enter. Not which moving average it uses. Why should the behaviour keep showing up in the market? A trend strategy might work because traders are slow to react to new information, causing price movements to continue longer than expected. A mean-reversion strategy might work because short-term panic pushes price too far away from its normal range. A session strategy could depend on liquidity entering the market at certain times. That is the actual idea. The indicators are just how you decided to measure it. This matters because indicators can be adjusted until almost anything looks good historically. Especially if your testing enough combinations. Try completing this: > This strategy works because market participants repeatedly... If you cant finish that sentence in a way that makes sense, theres a chance the strategy is mostly finding patterns that happened by accident. Not always. But its worth questioning. --- Stop Trying to Make the Backtest Look Better Most strategy development naturally moves in one direction. Make the results better. Change the stop loss. - Add another confirmation. - Remove bad trading hours. - Adjust the lookback. - Add a volatility filter. Then maybe another filter to fix the period where the first filter stopped working. Eventually the backtest looks amazing. But the strategy can also become a detailed explanation of what already happened. The better approach is sometimes doing the opposite. Try to make the backtest worse. - Increase the spread. - Add more slippage than you normally expect. - Delay the entry. - Move the parameters slightly. - Change the starting date. - Remove the best year. - Remove the best performing asset. Then see what is left. A strong strategy will usually weaken under pressure. Thats normal. The important part is whether it weakens gradually or just completely collapses. If changing a parameter from 20 to 21 destroys everything, that isnt a great sign. If slightly higher trading costs turn the system from profitable to deeply negative, the edge probably wasnt very large to begin with. A robust strategy doesn’t need to be perfect. It just shouldnt be hanging on by one tiny thread. The Best Parameter Is Often the One You Should Question Most Optimisation tools naturally draw your attention toward the strongest result, whether that is the highest profit, the best return-to-drawdown ratio, the highest Sharpe ratio or the most attractive win rate. Once those figures appear on the screen, it is easy to assume that the top-performing version is also the most reliable one. The issue is that when you test enough combinations, something will usually stand out simply because the sample is so large. That does not necessarily mean you found a real edge. It may only mean you found the luckiest result among thousands of possibilities. A useful way to think about this is to imagine 10,000 people flipping coins. Someone in that group is likely to produce an unusually high number of heads, but that does not mean they developed a repeatable skill. Their result looks impressive because the number of attempts was large enough for an outlier to appear. Parameters can behave in much the same way. A single setting may produce an exceptional backtest, but that result becomes far more convincing when the values around it also perform reasonably well. If a lookback of 30 produces strong results, check what happens at 27, 32 or 35. The performance does not need to be identical across each variation, but the general behaviour should remain consistent. What you are looking for is a stable region rather than one perfect point. If several nearby settings all produce acceptable returns with similar drawdown characteristics, the strategy is more likely responding to a real market tendency. If one exact value performs well while everything around it falls apart, the result may be more dependent on noise than genuine structure. Find Out Where the Profit Actually Came From An equity curve can look smooth and convincing while still hiding a very uneven distribution of returns. That is why the overall result should never be viewed as one complete block. Breaking the strategy down by year, month, session, trade direction, volatility level, asset and market condition can reveal what was actually driving the performance and whether that performance came from several independent sources or one narrow period. You may discover that the strategy was profitable only on long trades while the short side consistently lost money. It may perform well during periods of high volatility but slowly give those gains back when conditions are quieter. In some cases, one unusually strong year or a small number of large trades may account for most of the total return. None of those findings automatically make the strategy invalid, but they change how the system should be understood and managed. A strategy that earns small, steady returns across many different periods is very different from one that depends on a few large opportunities each year. Both can work, but they will create very different experiences in live trading. The first may require patience during frequent but manageable fluctuations, while the second may go through long periods of little activity before producing a meaningful gain. Understanding that difference affects position sizing, risk expectations and how long you may need to wait before deciding whether the strategy is still behaving normally. Out-of-Sample Data Is Not Always as Unseen as It Looks Out-of-sample testing is often treated as the final proof that a strategy is valid. If the system works on data that was not used during development, the result feels more trustworthy because it appears to show that the rules can generalise beyond the original sample. That idea only holds, however, if the out-of-sample period remains genuinely separate from the development process. Consider what happens when a strategy performs poorly on the out-of-sample section. You adjust the rules, test again, make another change and continue until the result improves. Even though that data was not included in the original optimisation, it has still influenced the final design. Over time, the strategy becomes indirectly fitted to the out-of-sample period as well, which removes much of the value that the test was supposed to provide. This is why it helps to keep a final untouched section of data that is not repeatedly checked while the strategy is being built. When you eventually test against it, you also have to be willing to accept a disappointing result without immediately changing the system. That is the uncomfortable part, but it is also what makes the test meaningful. Unseen data is only useful if it has the power to reject the strategy. Otherwise, it becomes another dataset you gradually fit the system around until the numbers look good enough. A Strategy Can Be Right and Still Be Untradable Some strategies look profitable before costs and become almost worthless once real execution is included. This is especially common with scalping systems, high-frequency strategies and setups with a very small average profit per trade. The direction of the trade may be correct and price may move exactly as expected, but spread, commission and slippage can consume most or all of the edge. Execution should not be treated as something that is added after the strategy is complete. It is part of the strategy itself. If the average trade makes $4 before costs and takes $3.50 to execute, there is almost no margin for error. A slightly wider spread, a slower fill or a missed order can remove the entire advantage. On paper, the system may still look technically correct, but in practice it may not be capable of producing a meaningful return. Backtests often assume that trades are filled at the intended price, yet live markets do not offer that guarantee. Orders can be delayed, partially filled or missed altogether, and the final result may differ materially from the theoretical version. A strategy should therefore be tested using realistic and slightly conservative cost assumptions. If it only works when execution is nearly perfect, it probably does not have enough room to survive actual market conditions. Not Every Live Failure Means the Strategy Is Broken When live results begin falling behind the backtest, the first reaction is often to assume that the edge has disappeared. Sometimes that is exactly what happened, but not every gap between live and historical performance points to a problem with the underlying strategy. Wider spreads, delayed fills, missed signals, data-feed differences or incorrect session settings can all produce weaker results without changing the original market behaviour the strategy was designed to capture. Even relatively small operational issues can have a significant effect. A strategy may be using the wrong session time because of daylight saving changes, or it may be entering slightly later because of platform latency. The broker may also be providing a different price feed from the one used during testing. These are execution problems rather than strategy problems, and they should be investigated separately. The only reliable way to tell the difference is to compare what the strategy was expected to do with what actually happened. Track the intended entry price, the actual fill, any missed trades, slippage and changes in trading costs. If the signals are still behaving as expected but execution is poor, changing the trading logic may make things worse because you would be fixing the wrong part of the system. If both the theoretical signals and the live results begin weakening in the same way, however, the problem may be more closely related to the original edge. Decide What Failure Means Before You Are in a Drawdown Every strategy will eventually go through a difficult period. The hard part is deciding whether that period is a normal drawdown or evidence that the strategy is genuinely breaking down. That decision becomes much harder once money is already being lost because emotion begins to influence how the results are interpreted. Some traders stop a good strategy too early because the drawdown feels uncomfortable, while others become attached to the system and continue trading it long after the evidence has changed. Both reactions usually come from making decisions under pressure. A better approach is to define failure conditions before the strategy goes live, when the rules can be set more objectively. Those conditions might include a drawdown moving far beyond the expected range, a large drop in average profit per trade, a meaningful change in trade frequency or a persistent breakdown in the market environments where the strategy previously performed well. You may also decide to reduce risk if live execution consistently differs from the backtest assumptions or if the original market behaviour behind the system no longer appears to be present. One bad week is rarely enough to make that decision, and sometimes one bad month is not enough either. At the same time, blind faith is not a process. The important thing is to decide in advance when you will continue, reduce exposure, pause the strategy or retire it completely. Those choices are much easier to make while you are calm than while watching a losing account and trying to decide whether the next trade will fix everything. A Robust Strategy Should Survive Imperfect Assumptions No backtest will match live trading perfectly. Your estimate of spread may be wrong, slippage may be worse than expected and future market conditions will not look exactly like the past. The best historical parameter may not remain the best one, and there may also be small errors in the data, differences between brokers or changes in the way orders are executed. A strategy does not need to survive unlimited damage, but it should be able to tolerate normal levels of imperfection. That is one of the clearest ways to think about robustness. A robust strategy is not one that performs perfectly in every environment. It is one that continues to function when some of the assumptions behind it are slightly wrong. This does not mean the strategy should ignore changing conditions or continue trading through every possible problem. It means the edge should not disappear because of one small adjustment to spread, one slightly delayed entry or one minor parameter change. If ordinary real-world variation is enough to destroy the result, the original backtest was probably more fragile than it appeared. The Goal Is Not to Prove That Failure Is Impossible No amount of testing can prove that a strategy will work forever. Markets change, edges weaken, trading costs increase and other participants begin using similar ideas. A system that performed well for ten years can still struggle in the next one, especially if the behaviour it relied on becomes less common or market structure changes. Robust testing does not make a strategy permanent. What it does is help you understand what the strategy depends on, how it is likely to struggle and which warning signs deserve attention. That understanding is often more valuable than the headline return because it gives you a clearer idea of how the system should be managed once real money is involved. A strong backtest may be enough to convince you to deploy a strategy, but understanding why it works, where it is vulnerable and how quickly it deteriorates is what gives you a chance to manage it properly. The more useful question is not simply how much the strategy made. It is how much pressure the strategy could absorb before the edge disappeared. That answer usually tells you far more than the profit figure on its own.