u/Zestyclose-Eagle1809

▲ 2 r/Forex

Nobody in the prop firm industry will tell you when to stop. Here's the rule.

Nobody who sells resets will ever tell you to walk away. Most of the "don't give up, learn and retry" advice is basically pointing you at another fee. Sometimes a reset is the best move. Often it's throwing good money after bad, and there's a smooth way to tell which one you're in.

The one thing that decides it: did anything change between attempts, and was your failure variance or structure.

Reset makes sense when the failure was bad timing on a normal drawdown. Take your worst backtested day at challenge risk. If it fits under the daily limit and you happened to hit a rough loss that surpassed your backtesting metric, that's variance (only if you have enough samples backing that number up tho). Reset, run it again, the math is on your side. Same if you really fixed the thing that failed you, cut your size, changed the pair, whatever, because now it's a different attempt, not the same one repriced.

Walk away when nothing changed and you're paying to re-roll the same dice. This is the trap. If you failed the same challenge three times at like the same point, with basically the same approach each time, the fourth attempt is not a fresh chance. You already have the data. That approach doesn't pass this evaluation, and the fee is buying you a feeling, not a funded.

One key question. Between the last attempt and the next one, what specifically is different. If the answer is "I'll be more disciplined" or "I'll be more careful," that's not a change, that's a hope, and hope has a bad win rate against a rulebook. A real change is a smaller size, a different market, a longer sample proving the edge, a firm whose rules fit how you trade, or simply confirming the failure was just variance (which goes han din hand with a longer sample).... If you can't name a concrete one, you're not resetting a strategy, you're funding the same failure again.

And as reality check.... Sometimes the repeated failure is telling you the edge is not there yet, not that the firm is unfair or you were unlucky again. Four resets at 100 to 150 each is 500 to 600 spent to learn something a proper backtest would've told you for free. At that point the cheapest move is to stop paying the challenge to test your strategy and go test it properly first.

So before the next reset, name the one thing that's different. If you can point to a real, concrete change, reset with the odds behind you. If you can't, keep the money and go fix the thing, because the resets never going to.

reddit.com
u/Zestyclose-Eagle1809 — 2 days ago
▲ 1 r/Forex

Stop paying challenge fees until your worst backtested streak fits under the max limit loss

Everyone treats the challenge fee like the cost of trying... but in reality it's the cost of skipping the checks that would've told you not to buy yet. I've watched people reset 3/4 times, paying each round, when ten minutes of work up front would've said "not ready" or "size down first." Here's the list you can run before spending a cent on a challenge.

The first, very obvious.. Your worst backtested day has to fit under the daily loss limit. Take your single worst day in backtest, at the risk you'll use in the challenge. If it's already bigger than the firm's daily limit, you don't have a bad luck problem, you have a sizing problem, and no number of resets fixes that problem... Cut your risk per trade until your worst day fits under the limit with space to move.

Your risk per trade is the real lever, not your strategy. Same system, run at 1% then at 0.5%, and the pass rate often jumps easy. The strategy didn't change, the variance did.

Your edge has to survive real costs. Take your real spread, triple it, add a tick of slippage, and rerun. A true edge decreases a bit and keeps going. A fake one just disappears. If your backtest only works at zero cost, the challenge was never the thing standing between you and funded, your cost were doing all the drama.

You need enough trades to know the edge is real. A backtest of 60 or 80 trades is a coin that landed heads a few times. If that's all you have, you're paying the fee to test if you have an edge at all, which is the most expensive way possible to find out. Get the sample up before you get the account.

The profit target has to be reachable at your safe size. Once you've cut risk to fit the daily limit, check you can still hit the target in the number of trades you take. If safe sizing means you'd need 400 trades to reach the target and you take 3 a day, the math doesn't make any sense. That's something to solve now, not after you've paid.

Bonus: pay for a challenge you can afford to lose for 3 times in a row with the capital you decide to allocate to propfirms... going all in is the easiest way to kill hope on trader. Last one, more technical but very useful, run a Monte Carlo and make sure the risk per trade can surive the 5th percentile in terms of drawdown depth and length. If it doesn't, is just a matter of time that a bad streak comes and breaches the daily/max limits.

Run all these and one of two things happens. Either you buy the challenge already knowing your real odds, or you save the fee and fix the thing that would've failed you. Both beat finding out at reset number 5 and literally burning cash.

reddit.com
u/Zestyclose-Eagle1809 — 10 days ago

Your strategy does not have an edge. It has an edge in one regime, and your backtest hid it by averaging.

Your expectancy is an average across market regimes. If your backtest window was heavy on one regime, your edge is mostly that regime showing up a lot. Split it and you often find one regime carrying the whole average while another loses. That makes your live results a bet on the future regime mix, not on your strategy.

Pretty self explanatory already. Keep reading if you want to see the idea developed.

What does it mean for an edge to be regime dependent?

It means your strategy makes money in one type of market and gives it back in another, and a single average number combines the two together into something that looks stable.

I had a system with a clean 1.5 Sharpe that died the week I traded it live. It was not overfit and the sample was fine. It had a genuine edge, in exactly one regime, and my backtest window happened to be full of that regime. The average hid the bet completely.

Most strategies are like this. Trend systems print in trends and bleed in ranges. Mean reversion does the opposite. Your backtest reports one blended expectancy across all of it, and that blend is only meaningful if the future looks like the past. It usually doesn't.

Why does a blended backtest number hide a regime bet?

Because an average has no memory of what produced it. Watch what one number is hiding.

Say your strategy took 300 trades. In trending conditions it earned +0.30R per trade. In ranging conditions it lost 0.10R per trade. Your backtest window was trend heavy, 200 trending trades to 100 ranging.

Regime Trades in backtest Expectancy per trade
Trending 200 +0.30R
Ranging 100 -0.10R
Blended, what you see 300 +0.17R

That +0.17R looks like a solid edge. It isn't a property of your strategy. It is a property of your strategy plus a market that trended two thirds of the time. One regime is carrying the entire average, and the other is a net loser you cannot see behind the blend.

A blended expectancy is only an edge if the future regime mix matches your backtest. That is a bet, not a strategy.

Why does this show up live as the strategy suddenly not working?

Because the regime mix reverts, and your edge moves with it. The market does not owe you the same balance of conditions your backtest catched.

Here is the same strategy, unchanged, as the future regime mix drifts away from that trend heavy backtest.

Similarity to backtest Your real expected edge
67%, same as the backtest +0.17R
50% +0.10R
40% +0.06R
30% +0.02R
25% break even
20% 0.02R loss

Nothing about the rules changed. The moment trending days fall below a quarter of the time, the same strategy that backtested at +0.17R is a losing system. This is one of the most common reasons a real edge dies in live trading, and it looks exactly like the strategy breaking when it is actually the weather changing.

Practical step: How do you test if your own edge is regime dependent?

Split your own trades and look. You do not need a fancy classifier, you need a simple, consistent proxy applied at entry.

Tag every trade in your backtest by the regime at the moment you entered. A basic split is fine: trending versus ranging using something like ADX above or below 25, or price above or below a long moving average, plus a volatility bucket from ATR percentile. Then compute expectancy separately in each bucket.

If your edge is positive in every bucket, you may have a genuinely robust strategy. If one bucket is strongly positive and another is flat or negative, you don't have a universal edge, you have a regime bet wearing an average. Also check the mix itself. If one regime dominated your test window, means your period must be longer than what it is right now until ideally you have the same samples for both regimes.

Why is filtering to the good regime a trap?

Because the moment you slice your results and keep only the regime that worked, you added a parameter and selected on it. That is overfitting with an extra step.

If you discovered the good regime by looking at the results, you ran another trial, and your real edge needs to survive that. Validate the filtered version out of sample, not on the same data that suggested the filter. Run it through a Deflated Sharpe that counts the regime choice as one of your trials. And remember regime is lagging. You only know the regime after it has partly happened, and transitions, the moments the filter is most wrong, are exactly when the biggest losses cluster. A filter that is perfect after the fact can still bleed in live price action.

So is a regime dependent edge worth trading?

Yes, often more than a supposed universal one, but only if you trade it honestly. A regime specific edge that you understand beats a blended number you don't.

Three rules make it work. Size for the regime, smaller or flat when conditions don't favor you rather than forcing trades into the losing bucket. Accept slower times as part of the strategy, because sitting out the wrong regime is the edge, not a failure to trade. And never quote your blended backtest number as if it were stable, because it is a snapshot of one regime mix. Price the strategy on the regime you can expect, not on the one your history happened to catch.

What this doesn't mean

Not every edge is a regime bet. Some strategies are genuinely positive across conditions, and those are the ones worth the most, precisely because they don't depend on the weather. The test is the split, not the assumption.

And regime dependence isn't a flaw to be ashamed of. A well understood, single regime edge, validated honestly and traded only in its conditions, is often more robust than a strategy that claims to work everywhere. The danger isn't the regime dependence. It is not knowing it is there, because the average never told you.

reddit.com
u/Zestyclose-Eagle1809 — 17 days ago
▲ 4 r/Forex

Your live results do not match your backtest. Here are the 6 causes, ranked by how often they are actually the one.

6 reasons your live trading does not match your backtest, ranked, and how to tell which one is yours

TLDR: The gap between a great backtest and a bad live account has six usual causes. Every article lists them and none rank them or tell you which is yours. They split into two families: your backtest was fiction, or your edge was real and got taken away.

What is the fastest way to tell which cause is yours?

Start with one question... did the gap appear immediately, or did it show up after a while?

If your live results were off from the very first trades, the problem is baked into the backtest or your costs. The edge was never as big as the number said. If the strategy worked for weeks and then decayed, the edge was real once and something changed, either the market or you.

That sorts the six causes into two families. Family one, your backtest was fiction, overfitting, too small a sample, and look ahead bias. These never had an edge to lose. Family two, your edge got taken away: costs, regime change, and execution mess. These had an edge and decreased over time until it disappeared.

Here is the table.

Symptom you see Most likely cause
Off from the first trades, constant drag per trade Costs weren't properly tested
Backtest looked almost too perfect, live is total collapse Look ahead bias
Great in the test window, dies on any fresh data Overfitting
The good backtest was under 150 trades Sample too small
Worked for weeks, then slowly stopped working Regime change
Your live trades don't match the trades the rules would take You did not follow the rules

Number 1, most likely: your costs weren't properly tested

This is the most common and the most underrated. Your backtest applied one fixed, tight spread to every trade, including the ones where spread triples. Live, you pay the real number, on winners and losers alike.

The gap is there immediately and feels like a constant drag on every trade. Test this: pull 30 real trades, add spread and commission, convert to R by dividing cost in pips by your stop in pips, and subtract it from your backtested expectancy. Fix it: judge the strategy on the net number, and if it dies, widen the stop, cut frequency, or trade a cheaper pair. Tight stop, high frequency systems are the ones this kills.

Basically if deviation breaches 10% on average with live trades... the way you currently do it isn't gonna make you money in the long run.

Number 2: your backtest used data it could not have had

Look ahead bias means your logic peeked at information that didn't exist yet at the moment of the trade.

The backtest is suspiciously flawless, a very high win rate and a smooth curve, and live is a total, immediate collapse that no cost model could explain. Test: audit your signal timing and confirm every decision uses only closed, past data. Fix: correct the timing and rerun. This is rare in simple manual systems and common in coded ones, and it is the most catastrophic because the entire backtest was a fantasy.

Number 3: you optimized until it looked good

You ran a parameter sweep and kept the best looking combination. That winner was partly skill and partly the luckiest result out of everything you tried, and luck does not repeat live.

The strategy is beautiful in the test window and falls apart the instant it touches any data it wasn't tuned on. The test: change each parameter slightly and watch the result. If a small tweak collapses it, you fit noise. Even better, run a Deflated Sharpe on it, which corrects your Sharpe for how many combinations you tested. Fix: there is no fix for an overfit strategy, only prevention. Fewer parameters, out of sample validation, and honesty about how many versions you really tried.

Number 4: your sample was too small to mean anything

A backtest on 60 trades isn't evidence, it's a coin landing heads a few times in a row. The result sits inside the range of pure luck.

The impressive backtest covered a short window or a small number of trades, and the great stretch was really one good month doing the heavy lifting. Test: count the trades. Under about 150 and the confidence interval on your expectancy is too wide to act on. Fix: test across far more trades and multiple market conditions before you believe any number, and never size up on a strategy proven by a lucky quarter.

Number 5: the market regime changed

Sometimes the edge was real and the market simply moved on. A trend system stops working when the market goes to range. A volatility system starves when volatility dies.

This is the one that worked live for weeks or months, then decayed, and the decay lines up with a shift in volatility or trend. Test: split your backtest by regime and check whether the strategy ever survived the current one historically. If it only ever worked in conditions that are now gone, it is not broken, it is out of season. Fix: trade it only in the regime it fits, or accept that it will have negative periods in your account until its conditions return.

Number 6: you didn't trade your rules

The backtest followed the rules perfectly, without fear, on every signal. You didn't. You skipped the setup after two losses, entered late, moved a stop, or closed a winner early because you didn't want to give it back.

Your live trade log doesn't match the trades the strategy would have taken over the same period. Test: put your actual entries and exits next to the mechanical signals for the last month and count the mismatches. Fix: this is an execution problem, not a strategy problem, so the answer is automation or a hard rule that removes the discretion, not a new system.

So which fixes actually matter?

Work the two families in order. First rule out fiction, because there is no point optimizing execution on a strategy that never had an edge. Check the sample size, run the Deflated Sharpe, and audit for look ahead bias. If it survives all three, the edge is probably real.

Only then work the fees. Recompute costs in R and subtract them, confirm the current regime is one the strategy has actually survived before, and compare your live trades to the rules to catch execution drift. Most blown accounts aren't one cause, they're a real but thin edge that costs and a bad regime pushed under water together.

What this doesn't mean

The ranking is my judgment for retail forex, not a law. If you run a heavily coded system, look ahead bias climbs the list. If you trade with heavy discretion, execution drift does. Reorder it for your own situation.

And it is rarely a single cause. Usually two or three appear together: a slightly overfit edge, real costs, and a regime turn arriving at once. The value of the list isn't picking one winner, it is checking all six instead of blaming the market and rebuilding a strategy that was actually fine.

I wish I had this post when I started 9 years ago.

reddit.com
u/Zestyclose-Eagle1809 — 22 days ago
▲ 9 r/Forex

How much does spread really cost you? Convert it to R and most retail systems are already dead

Why the same strategy is profitable with a 50 pip stop and dead with a 5 pip stop

TLDR: Costs are not a rounding error, they're a fixed tax on every trade, and the tax gets bigger the tighter your stop. Convert spread and commission into R and it becomes visible. The same 55% win rate system nets +0.1R with a 50 pip stop and loses money with a 5 pip stop. Nothing changed except the stop.

Now let's dive deeper for the ones that want to see this developed.

Why does a real edge still lose money live?

I traded a system for a quarter that backtested at plus 0.10R per trade. It finished flat. The edge was real and it showed up in the results exactly as expected. I handed every bit of it to my broker and couldn't see it happening, because I was measuring cost in dollars, where it looks like pocket change.

One pip on a standard lot is 10 dollars. That number feels harmless next to a 5 figure account. So you glance at it, decide it is noise, and move on.

The problem is that dollars are the wrong unit. Your edge is measured in R, your risk is measured in R, and your cost is the only part of the equation you keep measuring in something else. Put it in the same unit and the picture changes completely.

How do you convert spread and commission into R?

Simple. Add your spread and your commission together to get the all in round trip cost. On EURUSD that is roughly 1 pip, whether you pay it as a raw spread plus about 7 dollars per lot, or as a wider spread with no commission. Then divide by your stop.

Trade a 10 pip stop and your cost is 1 divided by 10, so 0.10R per trade. Every position you open starts 0.10R in the hole before price does anything at all.

Even simpler, if after all commissions, instead of +3,5R you won +3,23R, you just calculate 3,23/3,5=0,92 ; 1-0,92 = 8 ; 8% deviation in that trade.

Now compare that to your edge. A 55% win rate at 1:1 rr produces a gross expectancy of 0.10R per trade. Your cost is 0.10R. You are working for free.

Cost in R equals round trip cost in pips divided by your stop in pips. That makes an invisible tax visible.

Why does your stop distance decide your real cost?

Because the cost in pips barely moves, but R is defined by your stop, so a tighter stop makes the same 1 pip a much bigger fraction of everything you risk.

Here is the same strategy, a 55% win rate at 1:1rr, worth 0.10R per trade gross. Only the stop distance changes.

Stop Cost in R Net edge per trade Share of your edge gone
5 pips 0.20R 0.10R loss 200%
10 pips 0.10R break even 100%
20 pips 0.05R 0.05R gain 50%
50 pips 0.02R 0.08R gain 20%
100 pips 0.01R 0.09R gain 10%

Same entries, same exits, same win rate, same broker. At a 100 pip stop the strategy keeps 90% of its edge. At a 5 pip stop it is a losing system but the strategy never changed, only the commission.

What win rate do you need just to break even?

This is the number I wish someone had shown me first. Your costs raise the bar your system has to clear before it earns anything.

At 1:1 on EURUSD with that 1 pip cost, here is the win rate required to arrive at exactly zero.

Stop Win rate needed to break even
5 pips 60%
10 pips 55%
20 pips 52.5%
50 pips 51%
100 pips 50.5%

A scalper on a 5 pip stop needs to win 60% of trades at 1:1 just to finish flat. Every point above 60 is profit, and everything below is a slow bleed. Move to a pair with a 3 pip cost like GBPJPY on that same 10 pip stop and the requirement jumps to 65%.

A 5 pip stop needs a 60% win rate to break even. Most people building scalping systems have no idea that is the bar.

Why is your real cost worse than the spread you see?

First. You pay the spread at the exact moment your order fills, and spreads widen during news, at the rollover hour, and in thin liquidity. Those are also the moments price is moving fast enough to hit your stop. So the trades that stop you out are the ones where you paid the widest spread. Your realized cost is worse than the advertised number.

Second, your backtest almost certainly understates this. Most platforms apply one fixed typical spread across the whole history, which quietly assumes calm conditions on every trade including the violent ones. Then you go live, pay the real number, and blame the strategy.

What does the same cost do to you over a year?

Style Stop Trades per year Cost per trade Gross R you must out earn
Scalper 5 pips 2,500 0.20R 500R
Day trader 10 pips 1,000 0.10R 100R
Intraday 20 pips 500 0.05R 25R
Swing 50 pips 150 0.02R 3R

Same broker, same pair, same 1 pip. The swing trader pays 3R a year and never notices. The scalper has to generate 500R of gross edge annually just to arrive at zero. That is the same tax, and it is why frequency and stop distance decide whether costs are trivia or the whole game.

So what do you actually do about it?

Measure it before you trade the system, not after a bad quarter.

Pull 30 of your real fills and compute what you actually paid, spread plus commission, rather than trusting the advertised number. Convert it to R using your typical stop. Subtract that from your backtested expectancy and judge the strategy on what is left, because the gross number was never the number you get.

Then work out your break even win rate and compare it honestly to your actual one. If the gap is thin, widen your stop, trade a cheaper pair, or cut frequency, because those three levers move cost in R far more than switching brokers ever will. And if your gross edge is under 0.10R per trade, tight stops are not available to you at any broker.

Most obvious solution is to also look for a platform that can actually support your system's edge. Multiple times is not a system's problem, is a platform's problem.

What this does not mean

Scalping is not impossible, and this is not a case against tight stops. Plenty of people trade 5 pip stops profitably. They just do it with a much larger gross edge than 0.10R, because they know the bar is 60% and they built to clear it.

It is also not a broker bashing post. The spread is the price of access and everyone pays it. The mistake is not that costs exist, it is that we measure them in dollars where they look like nothing, instead of in R where they sit right next to the edge they're eating.

Bottom line

Convert your costs to R and most of the mystery about live results not matching backtests disappears. Cost in R is your round trip cost in pips divided by your stop, so tighter stops multiply the same tax. A 55% win rate system keeps 90% of its edge on a 100 pip stop and loses money on a 5 pip stop. Compute your break even win rate before you risk anything, subtract real costs from your backtest, and remember you pay the widened spread precisely when it hurts.

reddit.com
u/Zestyclose-Eagle1809 — 28 days ago
▲ 5 r/Forex

I simulated 200,000 tilt sessions. One a month cuts a +96% year down to +49%. Here is the revenge trade math nobody shows you.

TLDR: Every post about revenge trading tells you to breathe and take a walk. None of them show you the number. So I ran it. You grind up about 0.4% a day. One tilt session averages a 4% loss and a bad one takes 30%. It also turns a profit 26% of the time, which is exactly why you cannot stop. Tilting once a month cuts a +96% year down to +49%.

What does a month of disciplined edge actually look like?

Slow. Boring. Barely visible day to day. A disciplined trader with a real edge, risking 1% per trade at a 55% win rate, makes about 0.4% on an average day. Over a 20 day month that compounds to roughly 8%. Over a year, close to doubling the account.

That is what a real edge feels like from the inside. It is not exciting. There is no single day where you feel like a genius. You just show up, take your setups, and let a small statistical advantage grind out a good result across hundreds of trades.... thats it.

I am setting this baseline first on purpose, because the whole point of the revenge trade math is the mismatch. You build this edge one slow day at a time. You can hand it back in one fast hour.

A real edge makes you about 0.4% on an average day. That is the entire thing you are risking every time you tilt.

What does one tilt session actually cost?

Far more than the loss that triggered it. A tilt session isn't one oversized trade. It is a cluster of them. You size up, you overtrade, and your wr drops because anger is not a strategy.

I modeled a realistic one: 5 times your normal size, 8 trades in a session instead of your usual few, and a win rate that falls from 55% to 45% because you are chasing, not selecting. Then I ran it 200,000 times.

One tilt session, 5% risk, 8 trades, degraded win rate Result
Average outcome 4% loss
Median outcome break even
Bad case, 1 in 20 30% loss
Chance it turns a profit 26%
Chance it erases a full disciplined month 48%

Read the last row. Almost half the time, one tilt session wipes out an entire month of the slow grind above. In a single sitting.

How many green days does it take to undo one tilt session?

This is the number that actually scared me straight. Your edge makes 0.4% a day. An average tilt session loses 4%. So one average tilt session costs you 10 clean, disciplined green days. Two weeks of doing everything right, erased by one hour of doing everything wrong.

And that is the average. The bad case, the 1 in 20 session, loses 30%. At 0.4% a day, that is 75 green days to climb back. Nearly four months of flawless trading to repair one afternoon.

The asymmetry is the whole story. You earn in small steady increments and you lose in rare violent chunks. The downside of a tilt session is 25 to 75 times the size of a good day.

If the math is this bad, why people keep doing it?

Because it works often enough to feel like skill, and that is the trap. Look back at the table. A tilt session turns a profit 26% of the time, and the median outcome is break even. So more than a quarter of the time, you tilt, you get away with it, and your brain files it as a win.

That is textbook intermittent reinforcement, the same wiring that makes slot machines addictive. If tilting lost money every single time, you would stop after two tries. Instead it pays out just often enough to keep you pulling the lever, while the average and the tail quietly bleed your account.

A tilt session profits 26% of the time. That occasional win is not luck saving you. It is the hook that keeps you doing the one thing that will end your account.

What does tilting once a month do to your whole year?

It roughly halves it. I ran a full year both ways. Pure disciplined trading returns about 96%. The same trader who tilts just once a month, one session, 12 times a year, ends at about 49%.

One bad hour every four weeks costs 47% decrease over the year. You didn't lose your edge. You still had it every disciplined day. You just handed half of it back 12 times, in 12 short bursts you could have prevented.

So what actually stops it?

Not willpower, because willpower is exactly what a losing streak strips from you. The math says the only reliable defense is structural, a hard cap that removes your ability to keep trading before the tilt starts.

Set a daily loss limit sized to your edge, not your ego. If a good day makes 0.4%, a daily stop around 2%, roughly five clean losing trades, is already generous. Hit it and your done, not as a guideline you renegotiate at 2pm, but as a rule that closes the platform. The traders who survive aren't calmer. They built a wall between the angry version of themselves and the account, so the tilt session never gets to happen.

What this doesn't mean

This is not a lecture about being weak. The urge to win it back is wired into everyone, and feeling it doesn't make you a bad trader. The point is the arithmetic, not the shame.

And one tilt session does not define you. What the numbers show is that the pattern compounds. A single slip is recoverable in ten green days. A monthly habit is the difference between doubling your account and treading water. The fix is not to feel less. It is to make the expensive action impossible before the feeling arrives.

Bottom line

Revenge trading is not just a psychology problem, it is an arithmetic one. You build your edge at 0.4% a day and a tilt session hands back 4% on average and 30% in the tail, which is 10 to 75 green days per episode. It fools you because it profits 26% of the time, and it halves your year if it becomes monthly. Willpower won't hold on the day it matters, so cap the downside structurally and make the tilt session impossible to start.

reddit.com
u/Zestyclose-Eagle1809 — 1 month ago
▲ 0 r/Forex

Why the 1 percent per trade rule is lying to you in forex.

I simulated 5,000 traders following the 1% rule in forex. With correlated pairs the safe 1% becomes a 64% drawdown.

TLDR: The 1% rule promises your worst loss is small and survivable. It quietly assumes your trades are independent. In forex they almost never are, because most pairs share a currency leg, so four positions at 1% each behave like one bet of 4%. I simulated 5,000 disciplined traders and correlation roughly tripled the bad case drawdown, from 26% to 64%, with the same rule and the same edge.

Breaking everything in detail now.

What does the 1% rule actually promise?

It promises that no single trade can seriously hurt you. Risk 1% per trade and one loss costs 1%, ten losses in a row cost about 10%, and you would need 100 straight losses to blow up. It is the first rule every forex trader learns, and for good reason, because it kills the fastest way to die, betting too big on one idea.

The promise rests on a hidden assumption almost nobody states out loud. It assumes your trades are independent, that each 1% bet is its own separate roll of the dice. Under that assumption the math is beautiful and the rule is close to bulletproof.

The problem is that in forex, the assumption is usually false, and when it breaks, the whole promise breaks with it.

Why does the 1% rule break when your trades are correlated?

Because most currency pairs are not separate bets. They share a leg. If you are long EURUSD, short USDJPY, and long GBPUSD, you are short the dollar three times wearing three different jerseys. When the dollar rips, all three move against you at once.

So your four 1% positions aren't 4 independent bets. They're closer to one big bet of 4% that stops out all at once. The 1% rule sized each trade as if it stood alone, but the market treats them as one position.

I ran the numbers. 5000 disciplined traders, 4 pairs a day, 1% risk each, a real edge of 55% wins at even money. The only thing I changed was the correlation between the pairs.

Correlation between pairs Bad case drawdown (95th percentile) Days per year all 4 lose together
0, truly independent 26% 10
0.5, moderate 46% 40
0.85, typical dollar pairs 64% 73

Same rule, same edge, same 1% per trade. The trader who imagines independent bets expects a bad case around 26%. The trader actually holding correlated dollar pairs gets 64%, and the worst 1% of them hit 84%. Correlation didn't change the rule. It changed what the rule was hiding.

Four correlated pairs at 1% each isn't four small bets. It is one 4% bet that stops out together.

Why do the losing days cluster?

Look at the last column. Independent pairs produce a day where all four lose about 10 times a year. Correlated dollar pairs produce that same all red day 73 times a year.

That is the engine of the deeper drawdown. When your positions are independent, a bad pair is usually offset by a decent one, so your equity curve is smooth. When they move together, the good days are bigger but the bad days are total, and strings of total red days stack into drawdowns the 1% rule swore you would never see.

Why does a market gap break the promise too?

Even a single position can blow past 1%. The rule assumes your stop fills at your stop. Over a weekend, or on a rate decision, price gaps. Your 1% trade with a 20 pip stop can fill 60 pips lower and cost you 3%, not 1.

So the 1% in the 1% rule is a fair weather number. It holds on a normal Tuesday and fails on exactly the days that matter, when volatility spikes and everything moves at once. Plan your real size for the gap, not for the calm.

Is 1% even the right number?

Not necessarily, because it ignores your edge entirely. 1% is a survival floor, not an optimal size. Against a genuinely strong edge, 1% leaves money on the table. Against a weak or negative edge, 1% doesn't save you, it just makes the bleeding slower and more comfortable. A losing system risked at 1% is still a losing system.

The rule answers how not to die on a single trade. It doesn't answer whether you should be sizing up, sizing down, or not trading the idea at all. That answer comes from your edge, not from a round number everyone repeats.

So what should you actually do?

Size the idea, not the trade. If 4 positions all depend on the dollar, treat them as one bet and cap the whole cluster near your real per trade limit, not 1% each. Check the correlation of your open pairs before you pile in. If they all share a leg, you are not diversified, you are concentrated.

Then size for the gap, not the calm, by assuming your worst fill is worse than your stop on volatile days. And set your risk from your edge, using a smaller fraction when the edge is thin. The 1% rule is a fine starting point for one independent trade. It falls apart the moment you hold several that are secretly the same trade.

What this doesn't mean

The 1% rule isn't useless, and this isn't a case for risking more. For a single, independent position it is close to perfect, and betting too big remains the fastest way to blow up an account.

The point is narrower. The rule protects you per trade, but risk isn't per trade in forex, it is per idea, and one idea often lives across several correlated pairs. Apply 1% to each of them and you have quietly built a position several times larger than you think. Fix the correlation blind spot and the rule works again.

To wrap this up

The 1% rule promises small, survivable losses, and it delivers only if your trades are independent. In forex they rarely are, because pairs share currency legs, so four positions at 1% each behave like one 4% bet. Across 5,000 simulated traders, correlation pushed the bad case drawdown from 26% to 64% with no change to the rule or the edge. Size the idea, not the individual trade, watch your pair correlations, and plan for the gap. The rule isn't wrong. It is just measuring the wrong unit of risk.

This is for forex traders using percentage based position sizing across multiple pairs. The correlation blind spot applies to any market where instruments share a common driver.

reddit.com
u/Zestyclose-Eagle1809 — 2 months ago

The Deflated Sharpe Ratio explained simply: why testing 100 strategies can turn a 1.5 Sharpe into a coin flip

The Deflated Sharpe Ratio explained simply: why testing 100 strategies can turn a 1.5 Sharpe into a coin flip

TLDR: The Deflated Sharpe Ratio takes your backtest Sharpe and corrects it for how many strategy variations you tested before keeping the best one. The more combinations you try, the more luck inflates your winner, so the DSR lowers your real odds. It turns a Sharpe number into a probability: the chance your edge is real and not luck.

What is the Deflated Sharpe Ratio?

The Deflated Sharpe Ratio is a corrected Sharpe ratio that accounts for how many times you tried before you found a good result. The first time I ran a parameter sweep, I found a backtest with a 2.5 Sharpe and felt like a genius. Then I ran the DSR on it and learned my genius strategy had close to a coin flip chance of being real.

A normal Sharpe ratio measures return per unit of risk and quietly assumes you tested one strategy. The DSR drops that assumption. It asks the question that actually matters: out of all the variations you tried, is your best one genuinely good, or is it just the luckiest result from a lot of attempts?

Instead of handing you a ratio, it hands you a probability. A DSR of 0.95 means a 95 percent chance your edge is real.

Why does a normal Sharpe ratio lie when you test many strategies?

Because the moment you test many variations and keep the best, the best one is partly luck, and the plain Sharpe has no idea that happened.

Say you test one strategy and it posts a 1.5 Sharpe. That is one honest result. Now say you test 200 versions, tweaking the moving average, the stop, the session, and you keep the single best. That winner did not just beat the market. It beat 199 siblings, and some of that winning is skill while some is just the luckiest draw in a big sample.

The plain Sharpe reports the winner as if it were your only attempt. It can't see the 199 you threw away. So it overstates how good the strategy is, and the more you tested, the bigger the lie.

Why do more combinations lower your Deflated Sharpe?

This is the part that confuses people, so here is the simplest version I know.

Picture a coin flipping contest. One person flips ten coins and gets eight heads. Mildly impressive, maybe a bit of luck. Now put a thousand people in the contest, all flipping ten coins each. Someone will flip ten heads out of ten. Are they a coin flipping genius? Of course not. With a thousand tries, a perfect run was basically guaranteed by luck alone.

Your parameter combinations are the contestants. Every variation you test is another person flipping coins. Test one combination and a strong Sharpe means something. Test a thousand and the best Sharpe in the pile is exactly what pure luck produces, because you gave luck a thousand chances to look like skill.

The DSR knows how many contestants you entered. It works out the Sharpe you would expect from the luckiest of your N tries, then checks whether your result beats that luck benchmark. More combinations push the benchmark higher, so your result has to clear a taller bar to prove it is real.

Here is the same 1.5 Sharpe run through the DSR, changing only how many combinations I tested before keeping the best:

Combinations tested DSR, the chance your edge is real
1 99%
5 86%
20 60%
50 43%
100 32%
500 14%
1000 9%

Same strategy, same Sharpe, every single row. The only thing that changed is how many times I tried before keeping the winner.

Test one combination and a 1.5 Sharpe is almost certainly real. Test 1000 and the same 1.5 Sharpe is a coin flip.

What number does the DSR actually give you?

A probability, not a ratio, and that is the whole point. The output sits between 0 and 1, and it reads as the chance your true Sharpe is above zero once luck is stripped out.

My cutoff is 0.95. If the DSR says there is less than a 95% chance the edge is real, I do not trade it, no matter how pretty the equity curve looks. A 1.5 Sharpe that comes back at 43% is not a strategy. It is a maybe, and a maybe doesn't get capital.

What is the best protocol if your first config is already profitable?

Treat it as the most valuable result you will ever get, and protect it by not searching further. A profitable first try is clean precisely because you did not go hunting, so there is no selection luck to correct, and that makes it the strongest kind of result in trading.

Write down that this was config number 1 before you touch anything, because the moment you start tweaking you lose the right to claim it. Run the DSR at a trial count of 1 anyway, since even one result can be noise from a short sample, and confirm it clears your bar. Check the sample is big enough too, because a profitable first try on 60 trades is luck, not an edge.

Then be honest about your stopping rule. If this config had failed, how many more would you have tried? If the answer is a lot, your real trial count is that planned hunt, not 1, so rerun the DSR at that number. If it survives all of this, validate it on out of sample data rather than more combinations, then size small and trade it. Do not optimize a clean winner into an overfit one.

A profitable first config is clean precisely because you did not search. Protect it by not searching now.

When should you actually use the DSR?

Any time you picked a winner out of many tries. That covers more situations than people admit.

Parameter optimization is the obvious one. So is any machine learning model, which secretly tests huge numbers of combinations during training. So is the everyday habit of running ten backtests and trading the best looking one. If you searched, selected, or optimized in any way, you ran multiple trials, and the DSR is how you correct for it before risking money.

What do you need to compute it?

Four inputs: your reported Sharpe, the number of variations you tested, and the skew and kurtosis of your returns, plus the length of your sample. Open source code exists, so you do not need to derive anything.

The hardest input is honesty about the trial count. People remember the 5 backtests they saved and forget the 80 they deleted. If you lowball the number of combinations, the DSR flatters you, and you are back to fooling yourself. Count every variation you actually tried, including the dead ones.

What the DSR does not fix

It is a correction, not a crystal ball. The DSR tells you whether your result survived the luck of testing many combinations. It says nothing about whether your edge keeps working after costs, after slippage, or after the market regime changes. A high DSR on a strategy that dies the moment fees touch it is still a dead strategy.

It also depends entirely on an honest trial count. Feed it a lie about how many combinations you tested and the output is meaningless. Garbage in, garbage out.

Bottom line

The Deflated Sharpe Ratio corrects your Sharpe for how many strategy variations you tried before keeping the best. More combinations means more chances for luck to fake skill, which raises the bar your result has to clear and lowers the probability your edge is real. The same 1.5 Sharpe can read as 99% real after one test and 9% real after a thousand. Use it before trusting any optimized backtest, and be brutally honest about how many combinations you actually ran.

This is for traders and researchers who optimize, search, or select among backtests. The correction applies to any strategy, any market, and any model trained on price data.

Updated June 2026

reddit.com
u/Zestyclose-Eagle1809 — 2 months ago
▲ 31 r/Forex

I simulated 10,000 FTMO challenges with a profitable strategy. The pass rate went from 93% to 9% on one variable.

TLDR: I built a genuinely profitable strategy, perfect discipline, no revenge trades, and ran it through the FTMO two step rules 10,000 times. The funded rate was 93 percent at 0.5 percent risk per trade and 9 percent at 2 percent risk per trade. Same edge both times. The only thing that changed was position size. The rules do not punish your strategy. They punish how big your bet is.

Why did I bother simulating this?

Everyone repeats the same line: most people fail prop challenges because they have no discipline. Maybe. But that stat never answers the question I actually had. If you take a real edge, trade it perfectly, never tilt, never revenge trade, and just run it into the FTMO rules thousands of times, what passes?

So I built it. A strategy with a clear positive edge, a fixed rule set, and a computer that never gets emotional. 10k challenge attempts. No psychology, no mistakes, just the math of a real edge meeting a narrow rule box.

The answer surprised me, and it changed how I size every funded account I run.

What strategy and rules did I test?

I kept the strategy simple, so nobody could say the edge was the problem.

The strategy has 50% WR, makes 1.5R when it wins, loses 1R when it loses. That is a positive expectancy of 0.25R per trade and a profit factor of 1.5, a genuinely good system that any trader would be happy to own. 5 trades per day.

The rules are the standard FTMO two step. Phase 1 needs +10%. Phase 2 needs +5%. You cannot lose more than 5% in a single day. You can't draw the account down more than 10% from the start, a static floor that does not move. No time limit. Getting funded means passing both phases.

Then I ran the whole thing 10,000 times at four different position sizes.

What was the real pass rate?

Position size, not the edge, decided almost everything. Here is the clean run, identical strategy each time:

Risk per trade Pass Phase 1 Get funded
0.5 percent 99.9 percent 99.8 percent
1 percent 79 percent 68 percent
2 percent 43 percent 20 percent
3 percent 31 percent 10 percent

Read that again. The strategy never changed. Same win rate, same edge, same trades. At half a percent risk it gets funded almost every time. At 3% it fails 9 times out of 10. The only variable was how much was bet on each trade.

Same profitable strategy. Funded 99% of the time at 0.5% risk, 10% of the time at 3% risk.

Why does position size decide everything?

Because the challenge does not score your edge. It scores your path, and bet size is the volume knob on your path.

Your expectancy decides where the equity ends up after thousands of trades. The rules only care about the trip. Double your risk per trade and you double the size of every swing on the way, which means a normal losing streak that used to be 4% now is 8% and trips the daily limit or the drawdown floor. You did not make the strategy worse. You made its variance bigger, and the rule box has no tolerance for variance.

This is why a profitable trader can fail repeatedly and conclude their system is broken. The system is fine. The size is feeding it into a box it cannot fit through.

What happens when you add real world shocks?

The clean run assumes every trade behaves. Real markets gap, slip, and spike on news. So I added a 4 percent chance of a shock loss of 2.5R to every trade and ran it again. The strategy is still profitable, just more realistic.

Risk per trade Pass Phase 1 Get funded
0.5 percent 96 percent 93 percent
1 percent 61 percent 43 percent
2 percent 29 percent 9 percent
3 percent 21 percent 5 percent

The shape is identical, the numbers just drop. At 1 percent risk the funded rate falls from 68 to 43 percent once you allow for the occasional ugly trade. The lesson holds harder, not softer: the smaller you size, the more room you leave for the bad streak that always eventually comes.

What does this mean if you are about to buy a challenge?

Size first, strategy second. Before you pay a fee, the most important number is not your win rate, it is your risk per trade against the daily limit and the floor.

Work it backwards. Take the daily loss limit, look at the worst losing run your strategy produces, and size so that run cannot put you on the floor. For most edges that lands well under 1 percent per trade, far smaller than what feels normal on a personal account. The traders who pass are rarely the ones with the best strategies. They are the ones who sized for the rules instead of for their ego.

And the gap between my 93% and the 10% industry pass rate everyone quotes? That gap is oversizing, broken edges that were never real, and the indiscipline the common stat blames. Position size is the part you control today.

What this is not

This is a model, not a promise. It assumes a real, persistent edge, which most strategies do not actually have. It assumes you follow the rules perfectly, which humans do not. Real trading adds correlated losing streaks, regime shifts, and emotional sizing that no clean simulation captures.

So treat these numbers as the ceiling, the best case for a profitable, disciplined trader. Your real odds sit below them. That is the point. If even the idealized version of a good strategy gets funded only 9% of the time at 2% risk, the size you choose matters more than almost anything else you do.

Bottom line

A profitable strategy isn't enough to pass a prop challenge. In 10,000 simulated FTMO attempts, the same positive edge got funded anywhere from 93% to 5% of the time depending only on risk per trade. The challenge scores your path, and position size is the loudest input to that path. Before you pay, size backwards from the daily limit and the drawdown floor, not from what feels normal. You can simulate your own strategy against the exact rules the same way, in an afternoon, for free, before risking the fee.

This is for traders evaluating funded account challenges. The simulation approach works for any firm, any rule set, and any strategy with a defined trade distribution.

Updated June 2026

reddit.com
u/Zestyclose-Eagle1809 — 2 months ago

Drawdown depth vs duration: why time underwater breaks more traders than the size of the loss

Drawdown depth vs duration: why time underwater breaks more traders than the size of the loss

TLDR: Max drawdown tells you how deep your worst losing streak was. It says nothing about how long you sat in it. Duration, the time spent underwater before recovering, is what actually makes people quit, ties up capital, and fails funded accounts. Two systems with the same max drawdown can be completely different to live with.

Why does drawdown duration matter more than depth?

The drawdown that almost made me quit when I started trading wasn't the deepest one. It was a mid teens loss on a system I trusted, and it just would not come back. Months of flat, grinding, lower highs that never broke even.

Max drawdown is a single number. It marks the one worst point your equity ever hit. It is useful, but it is one snapshot of one bad moment. It tells you nothing about whether you spent a week down there or a year.

Duration is the part you actually have to survive. A loss you recover from quickly is easy game. A loss that lingers for a year is the thing that makes you abandon a working system at the worst possible time.

Max drawdown records one bad moment. Duration records the whole experience of living through it.

What is the difference between drawdown depth, duration, and time underwater?

Depth is the max drawdown. The largest percentage drop from a peak to the following trough. How far down you went.

Duration is how long it took to get back. The time from the peak, through the trough, all the way back to a new high. A 10 percent drawdown that recovers in a month and a 10 percent drawdown that takes two years to recover have identical depth and nothing else in common.

Time underwater is the share of the whole period your equity sat below a previous peak. If your account spent 7 of the last 10 months below an earlier high, you were underwater 70% of the time, even if today you are at a record.

Why does a long shallow drawdown break more traders than a short deep one?

Because people do not quit at the bottom. They quit during the grind.

A sharp 25% drop is frightening, but it is fast, and the recovery usually comes while you still remember why you took the trades. A shallow drawdown that drags on for a year is a slow erosion of belief. Every flat month adds doubt. You start skipping signals. You shrink your size right before the recovery. You blow up the system by hand long before the math would have failed you.

There is a capital cost too. Money stuck recovering an old loss is money not compounding. A system that spends most of its time clawing back to even has a worse real return than its headline numbers suggest, because the equity was dead weight for long stretches.

Two strategies with the same max drawdown that are not the same risk

Strategy A Strategy B
Max drawdown depth 20 percent 20 percent
Longest time to recover 3 months 19 months
Time underwater 25% of the period 68% of the period
Longest losing streak 8 trades 31 trades

If you only compared max drawdown, these two look like equal risk. They are not close. Strategy B will make you question your life choices for a year and a half while it grinds back to even. Strategy A takes its hit and moves on. The number that separates them is duration, and it is invisible on a max drawdown line.

Why is duration the silent killer for funded accounts?

Because a funded account does not just need you to be profitable eventually. It needs you to be profitable now, and to stay inside the rules the whole time you are underwater.

A long drawdown keeps you exposed to the daily loss limit and the max loss floor for far more trading days, so your cumulative chance of tripping a rule climbs the longer you stay down. You also don't get paid while you are underwater. Profit splits come from new highs, so a system that spends 68 percent of its time below a prior peak is a system that pays you almost nothing even when its long run edge is real.

How do you measure drawdown duration and time underwater?

You read it off your equity curve, not your summary stats. Three numbers do most of the work.

Longest drawdown duration: the maximum number of days, or trades, between a peak and the recovery back to that peak. Time underwater: the percentage of all bars where equity sat below a prior high. Longest losing streak: the most consecutive losing trades, which is the day to day version of the same problem.

For one number that blends depth and duration, use the Ulcer Index, built by Peter Martin in 1987. It is the root mean square of your drawdowns from prior peaks, which means deep and long drawdowns both push it up, and a fast recovery scores well. Two systems with the same max drawdown but different recovery times get very different Ulcer Index values, which is exactly the gap max drawdown hides.

Should you stop caring about drawdown depth?

No. Depth still matters, and ignoring it is its own mistake.

Depth is your ruin risk. A drawdown deep enough to hit a margin call or a hard account floor ends the game before duration ever gets a vote. You cannot recover from a loss that takes you to zero, no matter how patient you are. So depth sets the hard limit you cannot cross.

The point is not to swap one number for the other. It is to stop treating max drawdown as the whole risk picture when it is one corner of it. Track depth for survival, track duration for whether you will still be trading the system when the recovery finally shows up. You need both, and almost everyone only watches one.

Bottom line

Max drawdown is one bad moment. Duration is the part you have to live through. Depth tells you whether a loss can wipe you out, duration tells you whether you will quit before it recovers, and time underwater tells you how much of your life the system spends below water. Track all three, not just the headline drop. For funded accounts especially, the slow grinding drawdown does more damage than the sharp one, because it keeps you exposed to the rules and pays you nothing while you wait.

This is for systematic and discretionary traders evaluating their own equity curves and backtests. The depth, duration, and time underwater split applies to any market and any timeframe.

Updated June 2026

reddit.com
u/Zestyclose-Eagle1809 — 2 months ago
▲ 14 r/Forex

Your strategy can be profitable and still fail every funded account. Here is the math prop firms do not explain.

TLDR: A prop challenge is not a profitability test. It is a test of whether your equity path fits inside a box: reach the profit target, never touch the daily loss floor, never touch the max drawdown floor. A genuinely profitable strategy can have a pass probability near a coin flip, and you can simulate yours against a firm's exact rules before paying the fee.

Why does a profitable strategy fail a prop firm challenge?

Because profitability is about the average, and the challenge is about the path. Your expectancy tells you where the equity curve ends up after thousands of trades. The challenge does not care where you end up. It cares whether your curve stays inside three lines on the way there.

Picture a box. The top edge is the profit target you must reach, usually 10 percent. The bottom edge is the max drawdown floor you cannot touch, usually 10 percent down. A second floor, higher up, is the daily loss limit you cannot touch on any single day, usually 5 percent. Passing means your equity path touches the top before it touches any floor.

A profitable strategy guarantees the curve drifts upward eventually. It guarantees nothing about whether the path stays inside the box.

In simpler words... Profitability is where your curve ends up. The challenge only scores where your curve goes on the way there.

What are the three rules actually testing?

Each rule tests a different property of your strategy, and they are not the same property.

The profit target tests your drift. Can your expectancy times your trade frequency cover the target distance before you run out of room? Firms have dropped the old time limit, so this is no longer a race against a clock. The cost of a small edge is subtler: you need more trading days to arrive, and every extra day you trade is another roll against the two floors below.

The daily loss limit tests your variance inside a single day. On a 5 percent daily limit, how many normal losing trades in one session put you on the floor? Risk 2 percent per trade and three losers in a day, which is routine, ends the attempt. That's the first reason of why controlling your risk on challenges matters depending on your risk tolerance.

The max drawdown tests your variance across the whole run. The deepest dip your equity takes, in any ordering of your trades, has to stay off the floor.

Does a static or trailing drawdown change the odds?

It changes them a lot, and it is the first thing to check before you pay. There are two models and they punish different things.

A static floor is fixed at account opening and never moves. FTMO's two step and most forex firms use this. Every dollar of profit you make is a real cushion that grows the gap between your equity and the floor. This is the trader """"friendly"""" model.

A trailing floor climbs up behind your high water mark. TopStep, FTMO's one step, and some others use this. Here profit raises the floor under your own feet. You start a 100,000 account with a 10 percent trailing floor at 90,000, trade well to 107,000, and your floor has climbed to 97,000. Now a normal pullback your strategy survives a hundred times in backtest drops you to 96,800 and the attempt is dead. You were up 7 percent and you failed.

So basically, on a trailing floor, every dollar of profit raises the floor you can fall onto. Your winners build the trap.

Two strategies that earn the same and pass at completely different rates

Identical annual return. Very different pass probability. The difference is the shape of the path, not the size of the edge.

Strategy A (smooth) Strategy B (spiky)
Annual return 43% 43%
Risk per trade 1% 1%
Worst single day down 2.1 percent down 5.2 percent
Pass probability (simulated) around 78 percent around 22 percent

Strategy B is not worse at making money. It makes exactly as much. It fails because its worst day blows through a 5 percent daily limit and its bigger swings sit closer to the max floor. The challenge is not ranking these two by profit. It is ranking them by path variance, and on that ranking B loses.

How do you compute your real pass probability before paying?

You simulate it. Take your strategy's actual trade distribution and run the challenge thousands of times under the exact rules of the firm you are considering.

The method: pull your win rate, average win, average loss, and trades per day from your backtest. Generate a random sequence of trades from that distribution. Walk it forward day by day, applying the real rules: stop the run if the daily loss limit is hit, stop it if the max drawdown floor is hit, mark it passed if the profit target is reached.

In a typical example, a profitable strategy risking 2 percent per trade against a 5 percent daily limit and a 10 percent floor lands near a 40 percent pass probability. A trailing floor pushes it lower. Drop the risk to 0.5 percent and the same strategy can climb past 70 percent, because you shrank the path variance without touching the edge, it will just take more time.

What this does not mean

This is not a claim that discipline is irrelevant. Oversizing, revenge trading, and trading through news end more attempts than anything else. If you break the daily limit by piling into a loss, that is on you, not the math.

The point is narrower and it survives perfect discipline. Two traders can both follow every rule, risk what they planned, and skip the news, and still face very different odds purely because of the variance their strategies carry. The structural failure mode does not show up in your backtest return, and no amount of composure removes it. You size it down or you compute around it.

Bottom line

A prop challenge scores your path, not your average. The profit target tests your drift, the daily limit tests your variance inside a day, the max drawdown tests your variance across the run, and a trailing floor turns your own profits into a trap. Rules vary by firm and product, so check the model before you pay, then simulate your strategy against those exact rules and read the pass probability. A profitable system near a coin flip is common, and the fix is almost always lower risk per trade, not a better edge.

This is for systematic and discretionary traders evaluating funded account challenges. The simulation method applies to any firm and any rule set, forex, futures, or crypto.

Updated June 2026

reddit.com
u/Zestyclose-Eagle1809 — 3 months ago

5 gates I run before risking capital on a strategy, ranked by how cheaply each one rejects

How I decide if a strategy is live-ready: the 4 gate process that killed 37 of my last 41 strategies

TLDR: I run every strategy through 4 gates in cost order, cheapest rejection first. Economic hypothesis, sample size floor, three statistical tests and cost and regime stress test. Of 41 strategies I logged last year, only 4 reached live capital. The process is built to kill, not to bless a system.

Why run a fixed process instead of judging each strategy on its merits?

Discretion is where overfitting hides. If you evaluate each strategy by looking at it, you will find a reason to trade the ones you are attached to.... A fixed sequence removes it. Every candidate faces the same gates in the same order, and the order is deliberate: cheapest disqualifier first, so most strategies die before I spend hours on walk-forward.

Last year I logged 41 strategies through this. Gate 1 killed 12. Gate 2 killed 9. Gate 3 killed 14. Gate 4 killed 2. 4 survived to live capital.

Here is the catch... The 37 that died had a median in sample Sharpe of 2.1. The 4 that survived had a median of 1.4. The strategies that looked best on paper were the ones this exact process rejected.

Gate 1: Is there an economic reason the edge should exist?

This gate has no code, and that is the point. Before any test, I write one paragraph naming why the inefficiency exists, who is on the other side of the trade, and why they keep losing. A liquidity premium... A structural hedger who trades regardless of price... A behavioral bias with real flow behind it.

If I can't name the loser, I do not test the strategy. A pattern with no mechanism is a pattern you found by looking, which means you will find one.

This is the cheapest gate and it rejects the most garbage. 12 of my 41 never cleared it. They were useless systems dressed up as ideas. Most people skip this gate because it is the only one you cannot automate, which is exactly why it filters what the automated gates can't.

Gate 2: Does the backtest have enough data to mean anything?

This gate asks one thing. Do you have enough data to tell skill from luck?

First, enough trades. A backtest with 80 trades swings too much to trust. I want at least 400 before I believe any metric.

Second, a long enough time period. When you test many versions of a strategy and keep the best one, that winner looks good partly by chance, the way the luckiest player in a coin flipping contest looks skilled. Ruling that out takes years of data, and how many years depends on how many versions you tried. After testing around a hundred versions, a 1.0 Sharpe needs about six years of data before you can trust it. A flashy 2.0 Sharpe needs only about two. A high score on a short backtest is the most dangerous thing in a research log, because luck fakes it easily.

9 strategies died here. A 1.3 Sharpe on 14 months is not an edge. It is a sample too small to tell.

It's not that low Sharpe needs more testing because it's weaker. It needs more testing because it's closer to the level random chance can imitate.

Gate 3: Does it survive the three overfitting tests?

This is the expensive gate, so it runs third, only on survivors. 3 tests, each catching a different failure, run cheapest first.

Deflated Sharpe Ratio first, because it is one calculation. A normal Sharpe assumes you only tried one strategy. But if you tested 80 versions and kept the best, that winner is partly lucky, and the plain Sharpe has no idea you ran 80 attempts. The Deflated Sharpe fixes that. It takes your reported Sharpe, accounts for how many versions you tried, and adjusts for fat tails and lopsided returns, then hands back a single probability: the chance your edge is real rather than the luckiest of your tries. Same enemy as Gate 2, caught at a different step. My cutoff is 95%. If the best of your 80 versions scores only 60%, that still leaves a 40% chance the edge is imaginary, so it never reaches my account.

Monte Carlo is next. Resample the trade sequence 10,000 times and read the 95th percentile of max drawdown. Your backtest shows one drawdown, but that is just the order your trades happened to land in. If the drawdown in your 95th percentile drawdown is more than your account can take, it doesn't pass

Walk-forward last, because it is a full reoptimization. 5y build, 1y test, rolled forward. Profit factor holds above 1.3 on at least 7 of 10 out of sample windows or it dies. Fourteen strategies died at this gate, most on the deflated Sharpe before walk forward ever ran.

14 systems died here.

Gate 4: Does the edge survive real costs and a regime split?

A clean backtest assumes free, instant fills. I model realistic costs first, then stress them: triple my real commissions, add a tick of slippage, re-run.

The metric I watch is not profit factor, it is expectancy and Sortino retention. The edge has to keep at least 70% of its expected value after the stressed costs. A real edge degrades gracefully. A fake one inverts the moment friction touches it.

Then I split the history by volatility regime and require profit factor above 1 in both the calm and the stressed halves.

2 strategies died here.

Bottom line

4 gates, cheapest rejection first.

The process is designed to reject, and last year it rejected 37 of 41. The survivors looked worse on paper than most of the strategies it killed, which is exactly why I trust them.

This is for systematic traders deciding whether a backtest deserves live capital. It applies to single strategies, parameter sweeps, and machine learning models trained on price data.

Updated June 2026

reddit.com
u/Zestyclose-Eagle1809 — 3 months ago

3 tests that catch an overfit backtest, in the order you should run them

3 tests that catch an overfit backtest, in the order you should run them

TLDR: Run three tests before trading any backtest, cheapest first. Deflated Sharpe Ratio (one calculation, catches multiple-testing luck), Monte Carlo on your trade sequence (minutes, catches lucky ordering), then walk-forward analysis (a full re-optimization, catches curve-fit parameters). Each has a kill-threshold below. Fail the cheap test and you stop before wasting hours on the expensive one.

Why do most profitable backtests fail live?

The more parameter combinations you test, the higher the chance your best result is luck, not edge. Bailey, Borwein, López de Prado, and Zhu (2014) proved that past a small number of trials, expected out-of-sample return turns negative, not zero. The strategy loses money live.

Harvey, Liu, and Zhu (2016) documented 316 published factors claimed to predict returns. After multiple-testing correction, most fail. If you test 50 strategies and keep the best, you are running the same machine that manufactures false positives.

The fix is not testing less. It is running three specific tests on the survivor, in cost order, and killing it the moment it fails a threshold.

Which test should you run first, and why does order matter?

Run the cheapest disqualifier first. Most strategies die on test one, so there is no reason to spend hours on walk-forward before a one-line check.

The cost order is clear. The Deflated Sharpe Ratio is a single formula evaluation. Monte Carlo is a few minutes of resampling. Walk-forward analysis is a full re-optimization across many windows, often hours of compute.

Order Test Effort Catches Kill-threshold
1 Deflated Sharpe One calculation Multiple-testing luck P(true Sharpe > 0) under 95%
2 Monte Carlo Minutes Lucky trade ordering 95th-pct drawdown breaks your risk limit
3 Walk-forward Hours Curve-fit parameters Profit factor under 1.3 on most windows

Most competing guides explain these tests. None tell you to run the cheap one first.

What is the Deflated Sharpe Ratio and why run it first?

The Deflated Sharpe Ratio is one calculation and it kills more strategies than any other single test. Run it first.

A normal Sharpe assumes one trial. If you tested 100 parameter sets and reported the best, the maximum Sharpe across those trials is inflated, and the inflation grows with the number of trials. Bailey and López de Prado (2014) built the correction. It adjusts for number of independent trials, skewness and kurtosis of returns, and sample length.

Worked example: a 2.5 Sharpe selected from 100 trials on five years of fat-tailed daily returns drops to a 90% probability that its true Sharpe is even positive. A 1.5 Sharpe from 50 trials drops to 43%, worse than a coin flip. Even a gaudy Sharpe can fail this bar.

Open-source Python exists, and the function is short enough to paste in the comments. If P(true Sharpe > 0) is under 95%, stop here.

A 1.5 Sharpe chosen from 50 trials lands at 43% odds the edge is real. Worse than a coin flip.

How does Monte Carlo simulation expose a lucky backtest?

Monte Carlo resamples your trades to show the outcomes you did not happen to get. Run it second, after the Deflated Sharpe passes.

Your equity curve is one ordering of your trades out of thousands. Take the list of trade returns, resample with replacement 10,000 times, and plot the distribution of final equity and maximum drawdown.

The kill-threshold is drawdown against your risk limit. If your backtest shows 18% max drawdown but the 95th percentile of the simulation is 35%, your live account will likely see 35% within a few years. Position sizing built on the backtest number forces you out at the worst moment.

Second check: shuffle trade order. If profit factor falls below 1 when reordered, the sequence carried the result, not the edge. That is regime dependence.

What does walk-forward analysis prove that the other two cannot?

Walk-forward analysis is the expensive final gate. It tests whether your parameters survive on data they were never fit to. Run it last, only on strategies that passed the first two.

A single train/test split is one data point. Walk-forward optimizes on a rolling window, tests on the next, then rolls forward across the whole history. You get out-of-sample results across many regimes instead of one.

Concrete setup: five-year build window, one-year test window, roll forward in one-year steps. The kill-threshold is consistency. If profit factor drops below 1.3 on most out-of-sample windows, the parameters were curve-fit. A strategy profitable in 9 of 10 windows is a signal. One that works in 4 of 10 is noise you tuned.

What do these three tests catch that a standard backtest hides?

Failure mode What a standard backtest shows What the test reveals
Selection from many trials Sharpe of 1.5 DSR true-Sharpe probability near 43%
Lucky trade sequence Smooth equity curve Monte Carlo 95th-percentile drawdown
Curve-fit parameters One clean hold-out Walk-forward collapse across windows

Each isolates a different failure. A strategy can pass two and fail the third. The order saves you time: most die on the Deflated Sharpe before you ever build the walk-forward.

Bottom line

Three tests, cheapest first. Deflated Sharpe for multiple-testing luck, Monte Carlo for trade-ordering luck, walk-forward for curve-fit parameters. Each has a concrete kill-threshold, and failing the cheap one means you stop before the expensive one. None of this guarantees live profit. Slippage, fees, and regime change still kill survivors. What it does is remove the strategies that were never real, which is most of them.

This is for systematic and discretionary traders who backtest before committing capital. The tests apply to single strategies, parameter sweeps, and machine-learning models on price data alike.

Updated May 2026

reddit.com
u/Zestyclose-Eagle1809 — 3 months ago
▲ 0 r/Forex

Two years of losing prop firm challenges taught me one thing: a good backtest means almost nothing

I've been trading systematic forex for +8 years. For about two years my routine was the same:

Build a strategy. See a clean equity curve. Take a prop firm challenge or push it live. Watch it fall apart inside 100 trades... Receipe for disaster.

Every time I blamed something different. Slippage. News. My discipline. Bad luck. None of that was the real reason. The real reason was that I had no honest way to tell whether a strategy actually had edge, or whether I had just found a shape in the noise.

Here is what I wish someone had told me two years and several thousand euros earlier.

A profit factor of 2.0 on 80 trades is not edge

It is a sample. The number of trades behind a result matters more than the result itself. A strategy with a 1.2 profit factor across 600 trades is more trustworthy than a 2.5 profit factor across 60. The math is boring but it is the math. Below roughly 100 trades, almost any conclusion you draw is statistically fragile.

Max drawdown is the wrong fear

Most people look at max drawdown depth and think they have measured risk. Depth is the easy part. The metric that actually ends prop challenges and live accounts is drawdown duration. How many trades does the strategy spend underwater before recovering? A 15% drawdown over 40 trades is recoverable. A 15% drawdown that takes 300 trades to recover from will break your patience before it breaks your account.

A high Sharpe across one good year is noise

If a strategy has a 2.0 Sharpe ratio but 80% of the profit came from a 3-month window, you do not have an edge, you have a regime that happened to favor you. Edge is what survives across months and conditions. The test I use now: split the equity curve into monthly buckets and look at how many of them are positive. If it is less than two-thirds, the headline number is hiding something.

Curve fit is invisible from the inside

The hardest part of all of this is that a curve-fit strategy looks identical to a real one until it doesn't. The only honest defense is to run the strategy on data it has never seen, and to ask whether the performance degrades gracefully or falls off a cliff. Graceful degradation is edge. Cliff is fit.

What I do now?

Before I ever risk capital on a strategy, I score it on four things: how big the edge is, how consistent it is month to month, what the downside looks like in real terms, and whether it is psychologically tradable. If any of those four are weak, I don't trade it. I rebuild it or I throw it out.

I built a tool with my cofounders for this exact workflow because we couldn't find one that did it the way we wanted. All the info is in my bio, but honestly, the framework above is what matters. The tool just automates it.

Happy to answer questions on any of these in the comments.

reddit.com
u/Zestyclose-Eagle1809 — 3 months ago