Light mode is here. · Every tab, chart and panel now switches from the tab bar. Still free to run your first validation, no card required. Try it now →

Same trades. Different order. Very different worst case.


There is a number in every cBot backtest that developers treat as a fact, and it is not one.

Maximum drawdown. You read it, you size your positions around it, you decide whether the strategy is survivable. Mine said 7.1 percent across a four-cBot portfolio, 444 trades, ten years of history, Sharpe 3.11. Comfortable.

Then I resampled the same trades fifty thousand times. At the 95th percentile the drawdown was 10.8 percent. At the 99th, 13.4.

Nothing about the strategies changed. The trades were the same trades. The only thing that changed was the order they arrived in.

Your backtest is a sample of one

Your cBot traded once, through one sequence of market conditions, and produced one equity curve. Every statistic you read afterwards describes that single path.

Run the same logic through a market that delivered its losing trades in a slightly different order, and the profit is identical while the drawdown is not. Profit is a sum, and sums do not care about order. Drawdown is a path property. It cares enormously.

This is why the drawdown figure in a backtest is the least reliable number on the page and the one most people anchor to hardest. It is not wrong. It is one observation from a distribution you never saw.

The question worth asking is not “what was the drawdown.” It is “how bad could the drawdown reasonably have been, given the same edge and different luck.”

Three ways to reshuffle, and why they answer different questions

Once you decide to resample, there are three sensible approaches.

Permutation reshuffles your exact set of trades. Every trade is used once. The final balance is identical on every run, because you are multiplying the same numbers in a different order. Only the path changes, and therefore the drawdown. This isolates one thing: sequence risk. Was the smooth ride real, or was it helpful ordering?

Bootstrap draws a fresh sample with replacement. Some trades appear twice, some never appear at all. The final balance varies run to run, which is the point. It answers a different question: how much rested on the particular trades you happened to get? Its working assumption is that trades are independent, which is where it becomes interesting.

Block bootstrap also draws with replacement, but lifts runs of consecutive trades whole rather than picking one at a time. Losses that bunched together in your backtest can still bunch together in the simulation. If your strategy has volatility clustering, and most do, this is the honest one.

The comparison between bootstrap and block is the single most useful diagnostic here, because the only difference between them is whether runs were kept together. Whatever gap opens between the two results is what clustering cost you.

One detail worth insisting on: the block length should be read from your trades, not chosen by you. A parameter you pick is a parameter you can fit. When a strategy genuinely shows no clustering, an honest implementation reports a run length near one trade and says plainly that block bootstrap described the same thing as bootstrap, rather than implying a difference that is not there.

Correlation is the portfolio question nobody runs

Running four cBots feels like diversification. Whether it is depends entirely on whether they lose on the same days.

In the portfolio above, the pairwise correlations on daily profit and loss sat between negative 0.07 and positive 0.05. That is close to independent, and one pair was slightly negative, which acts as a partial hedge. That portfolio was doing what its owner believed it was doing.

The reason to measure it anyway is that the failure case looks identical from the outside. Four cBots on the same instrument, in the same sessions, built from the same logic family, produce four equity curves that each look fine and one account that behaves like a single leveraged strategy. Every drawdown arrives on the same day.

You cannot see this by validating strategies one at a time. It only appears when their daily results sit side by side.

The parameters had a shelf life too

The same portfolio suggested a re-optimization window somewhere around five years, arrived at three separate ways: rolling variance stability, asset class and timeframe benchmarks, and memory in the return series. The three estimates ranged from 3.3 to 7.5 years before being reconciled.

That spread is the finding, not a flaw in the measurement. A strategy whose three estimates cluster tightly has a stable shelf life. One where they scatter does not, and re-optimizing on a fixed calendar habit was always going to be arbitrary for it.

Most developers pick this interval by feel. Quarterly, because quarterly sounds sensible. It is worth at least knowing what the strategy’s own behaviour suggests before overriding it.

What this changes in practice

Size against the tail, not the backtest. If your cBot reported 7.1 percent and the 99th percentile said 13.4, the second number is the one your account has to survive. The first was a sample of one.

Run bootstrap against block before trusting a smooth curve. The gap between them tells you how much the result depended on losses staying politely spread out.

Check correlation before adding the fourth cBot, not after the account explains it to you.

Treat trade count as a hard constraint. Four hundred trades over ten years is a thin sample for several of these questions, and no amount of resampling manufactures information that was not in the data to begin with.

What none of this does

It does not predict anything. Every figure here describes trades that already happened, or rearrangements of them. A 99th percentile drawdown is a statement about a historical distribution, not a floor.

It does not detect look-ahead bias in your code, only the statistical fingerprint if one reached the results. It cannot fix bad tick data, and a confident validation of a flawed backtest is worse than none at all, because it feels like progress.

And a strategy that passes is not a strategy that will work. It is a strategy that survived a set of checks built to catch specific failure modes. That is a genuinely weaker claim than it sounds, and anyone telling you otherwise is selling something.

Run it on a cBot you already trust

That is the useful test. Not one you suspect. One you were about to put on a live account.

Edge Matrix reads cTrader strategy tester reports directly, alongside MT4, MT5 and MQL5 signal exports. The first validation is free, no card, and it is the full analysis on one strategy rather than a limited preview. There is also a browser-based Monte Carlo tool that runs entirely on your own machine and needs no account at all.

The worst outcome is finding out your drawdown estimate was optimistic before your account demonstrates the same thing.


Edge Matrix is a statistical analysis tool. It evaluates historical backtest data using quantitative methods. It does not predict future performance, does not provide investment advice, and does not recommend whether to deploy any strategy. All trading involves risk.

Tags: , , ,

Risk Disclosure

Edge Matrix is a statistical analysis tool. It evaluates historical backtest data using quantitative methods but does not predict future performance or provide investment advice. Edge Matrix does not recommend whether to deploy, modify, or discontinue any trading strategy. All trading involves substantial risk, including the risk of loss. Past performance, whether analyzed or validated, is not indicative of future results. Users are solely responsible for their trading and investment decisions.

Trading foreign exchange carries a high level of risk that may not be suitable for all investors. Past performance is not indicative of future results. The high degree of leverage can work against you as well as for you.