BLACK·GOLD MARKET Join on Telegram

Evidence first · Before the money

How to Test Your Trading Strategy Without Fooling Yourself

Test twenty worthless strategies and, 88 percent of the time, one of them clears a 60 percent win rate on luck alone. Here is how to tell an edge from a coincidence before you fund it.

Black Gold Market, Raphael, XAU/USD trader
Black Gold Market
Protect. Master. Grow.
PILLAR 01

Protect

An untested strategy is an unpriced risk. Testing costs nothing but patience, and it is the cheapest capital protection there is.

PILLAR 02

Master

Judge the evidence, not the equity curve. The sample size, the assumptions, and the number of attempts, stated before you believe it.

PILLAR 03

Grow

Growth follows a real edge. Nothing compounds on a strategy that only ever worked in the sample it was built from.

How to test your trading strategy, separating a real edge from a lucky sample

The Backtest That Lies to You

Almost everyone who asks me how to test your trading strategy has already done the test. They have a rule set, they have scrolled back through the chart, and they have a result they are pleased with. What they want from me is confirmation.

I am not going to give it, and not because I doubt the work. I doubt the procedure. Scrolling back through a chart and counting wins feels like evidence, and under the right conditions it produces a number that means almost nothing at all. The problem is not that people test badly. It is that the ordinary way of testing has a specific mathematical flaw, and once you see it you cannot unsee it.

So let us do this properly. Not with encouragement, with arithmetic.

How to Test Your Trading Strategy When Luck Is Also Applying for the Job

Start with the flaw, because everything else follows from it.

Suppose you build a strategy with no edge whatsoever. A coin flip. It wins exactly 50 percent of the time and every trade is independent. You test it over 50 trades and ask whether it cleared a 60 percent win rate.

A single coin flip strategy clears that bar 10.1 percent of the time. Roughly one attempt in ten. That is already higher than most people expect, but on its own it is manageable.

Now be honest about what you actually did. You did not test one variant. You tried a different stop distance. You moved the moving average from 20 to 50. You added a session filter, then removed it. You tested it on gold, then on the index. Each of those is a separate attempt, and the question is no longer "did my strategy beat the bar" but "did any of my attempts beat the bar".

That is a completely different question, and the arithmetic is brutal.

Test enough worthless strategies and one will look brilliantBar chart on how to test your trading strategy, showing the chance that at least one strategy with no real edge clears a 60 percent win rate by luck alone. Testing 1 variant gives 10.1 percent, 5 variants 41.4 percent, 10 variants 65.6 percent, 20 variants 88.2 percent, and 50 variants 99.5 percent.Test enough worthless strategies and one will look brilliantHow often pure luck produces a 60% win rate, by number of variants triedChance at least one clears 60% wins by luck alone1 variant tested10.1%5 variants41.4%10 variants65.6%20 variants88.2%50 variants99.5%Every variant here has NO edge at all: a true 50% win rate, each trade independent.Each variant judged on 50 trades. Computed in Python from the binomial distribution.The strategies did not improve. Only the number of attempts changed.EDUCATIONAL ILLUSTRATION · NO PRICES, NO SIGNALS
How to test your trading strategy honestly means counting your attempts: with 20 variants tried, luck alone clears a 60 percent win rate 88.2 percent of the time.

Twenty variants, none of which has any edge at all, and 88.2 percent of the time at least one of them looks like a winner. Fifty variants and it is 99.5 percent, a near certainty. The strategies never improved. The only thing that changed was how many times you asked.

You did not find an edge. You found the luckiest member of a group you never counted.

I computed these in Python from the binomial distribution, assuming a true 50 percent win rate and independent trades, and you can recalculate them with different thresholds in a few lines. The assumptions are deliberately generous to the trader: real trading has costs, which makes the picture worse rather than better.

This is not a fringe concern. It has a name in the literature, and four mathematicians put it plainly in the Notices of the American Mathematical Society: it takes a relatively small number of trials to identify an investment strategy with a spuriously high backtested performance. Bailey, Borwein, López de Prado and Zhu were writing about professional fund research, where the same failure destroys real money at scale. The retail version is the same mistake with a smaller budget.

The Second Trap: A Real Edge That Looks Broken

The first trap makes nothing look like something. The second does the reverse, and it is the reason people abandon strategies that were working.

Take a strategy with a genuine edge: it truly wins 55 percent of the time at one to one. That is a good, realistic, unspectacular edge. Now ask how often a sample of it will look like a losing strategy, showing under 50 percent wins.

I simulated 200,000 samples at each length:

  • After 20 trades, it looks like a loser 24.8 percent of the time.
  • After 50 trades, 19.8 percent.
  • After 100 trades, 13.5 percent.
  • After 200 trades, 6.8 percent.
  • After 500 trades, 1.1 percent.

Read the first line again. One time in four, a genuinely profitable strategy will show you a losing record after twenty trades. If your rule is "give it twenty trades and see", you will throw away good work regularly, and you will never know you did it.

Put the two traps together and you get the cycle I watch traders repeat for years. Test many variants, adopt the one that looks best, run it live for twenty trades, watch it disappoint, discard it, and start testing again. Both ends of that loop are driven by sample size, and neither has anything to do with skill.

What Actually Makes a Test Worth Believing

Four things, and none of them are complicated. They are just tedious, which is why they get skipped.

Count your attempts and write the number down. Before you start, decide how many variants you will try, and record every one you test including the ones you abandon. The count is the single most important number in the whole exercise, and it is the one nobody keeps. A result from three attempts and the identical result from fifty attempts are not the same evidence.

Hold data back before you look at it. Split your history in two, build on the first part, and do not touch the second until the rules are final and written. If the strategy only works on the half it was built from, you have described the past rather than found an edge. This one check catches more self deception than any other, and it costs nothing but discipline.

Judge the size of the sample, not just the result. The table above is your guide. Twenty trades tells you almost nothing in either direction. A hundred starts to be informative. If you cannot generate a few hundred trades in a test, you must hold your conclusion loosely and say so out loud.

Include the costs from the first trade. Every backtest that ignores spread and swap is testing a strategy that does not exist. Costs fall hardest on high frequency rules, which are exactly the rules that produce large sample sizes, so the correction is biggest precisely where you were most confident.

Forward Testing Is the Only Test With Teeth

Everything above is still history. History is cooperative: it sits still, it does not requote you, and you already know how it ends even when you are trying not to.

Forward testing removes that. You write the rules down, then you apply them to bars that have not printed yet, and you record the result whether or not you like it. On a demo account it costs nothing except the thing traders are least willing to spend, which is time.

Two rules make forward testing honest. Write the rules before the test, in full, so there is nothing left to interpret in the moment. And log every trade the rules produced, including the ones you decided to skip, because the trades you skip are where your real strategy hides. A record of what your rules said versus what you did is the most useful document you will ever keep, and it is the foundation of consistency.

When the forward test disagrees with the backtest, believe the forward test. It is the only one that was not built with the answer already visible.

What This Changes About How You Work

The practical shift is small and it changes everything. You stop asking "does this strategy win" and start asking "how much evidence do I actually have, and how many times did I ask".

That reframe kills the search for the perfect setup, because you can now see that the search itself manufactures false positives. It also makes you patient with a strategy in a rough patch, because you know what a real edge looks like over a short sample. And it protects capital, which is the point of all of it: an untested rule set sized properly can survive being wrong, while a well tested one sized carelessly still cannot. Testing sits alongside position sizing, never instead of it.

None of this makes trading certain. It makes you honest about how uncertain it is, which turns out to be the part that keeps accounts alive.

Frequently Asked Questions

How many trades do I need to test a trading strategy?

More than most people use. At a true 55 percent win rate, a sample of 20 trades still shows a losing record 24.8 percent of the time, and 50 trades shows one 19.8 percent of the time. By 200 trades that falls to 6.8 percent. There is no magic threshold, but treat anything under 100 trades as a hint rather than a conclusion.

Why does my backtest work but live trading does not?

The most common reason is the number of variants you tried before settling on that one. Testing 20 no-edge variants produces at least one that clears a 60 percent win rate 88.2 percent of the time. The backtest was not measuring your strategy, it was measuring your persistence at searching. Costs left out of the test and rules quietly adjusted mid-test are the other two usual causes.

What is the difference between backtesting and forward testing?

A backtest applies rules to history you can already see, which makes it fast but easy to fool yourself with. A forward test applies written rules to bars that have not printed yet, so nothing can be adjusted after the fact. Forward testing is slower and considerably more trustworthy, and when the two disagree the forward test is the one to believe.

Is a 60 percent win rate a good result in a test?

Not on its own, and not without knowing your sample size and how many variants you tried. A coin flip clears 60 percent over 50 trades one time in ten. Win rate also says nothing about the size of wins against losses, so a lower win rate with larger winners can be the stronger strategy.

Can I test a strategy on a demo account?

Yes, and for forward testing it is the sensible place to do it. Demo will not reproduce how you behave when the money is real, so treat it as a test of the rules rather than a test of you. Testing the rules first is still worth doing, because it stops you paying real money to learn that the rules never worked.

Where did these numbers come from?

I calculated all of them in Python: the variant figures from the binomial distribution assuming a true 50 percent win rate and independent trades, and the sample-size figures from 200,000 simulations at a true 55 percent win rate. The assumptions are stated in the text so you can change them. The supporting paper by Bailey, Borwein, López de Prado and Zhu is linked above.

A Word on Risk, and How to Use This

Everything above is general education about evidence and probability in leveraged markets. Nothing here is financial advice, and no entry, stop or target discussed should be treated as a signal. Trading gold and other leveraged products carries a high risk of losing money quickly.

Cut to the bone: how to test your trading strategy is really a question about counting. Count your attempts, count your sample, and hold data back before you look at it. Do that and you will discard fewer good strategies and adopt far fewer bad ones.

If you want the risk-first companion to this way of working, it is the Black Gold Market Blueprint, a plain walk-through of protecting an account before trying to grow one. It is free, there is nothing to join, and there is no promised return anywhere in it.

Grab the Blueprint here, and for the foundation underneath all of it, start with how to protect your capital when gold gets volatile.

Protect. Master. Grow.

Raphael, Black Gold Market

About the Author

Raphael, founder of Black Gold Market

Raphael runs a XAU/USD channel built on one idea: the level, the context, and the risk, stated before the trade rather than after it. More about how the channel works. It is free to follow, with an optional Kit; he does not sell certainty and does not publish profit claims.

Risk disclaimer: This article is for educational purposes only and is not financial advice, an offer, or a recommendation to buy or sell any instrument. Trading gold, CFDs and leveraged products carries a high risk of rapid loss. No entry, stop or target discussed should be treated as a signal. External figures are linked to their source, and calculated figures are shown with their assumptions so you can check them yourself.

Start here, it's free

Count the attempts. Then decide what you believe.

Follow along on Telegram for daily gold analysis and the macro context behind it, session liquidity, real rates, the dollar, central-bank demand, and the forces that move gold, so you understand the backdrop instead of guessing. Free to follow, with an optional Kit. No hype, no promises, and nobody will ever ask you for a fee to release your own money.

Join Black Gold Market on Telegram No guaranteed profit. No pressure to copy anything blindly.

More from the journal