

9 pre-registered falsification studies later, here's what actually survived (spoiler: almost nothing did, but one pattern keeps showing up everywhere)
Been doing this for a while now — pre-register the rule, lock the cost model and decision threshold before touching data, correct for multiple comparisons whenever I'm testing more than one variant/asset/parameter at once. Figured it's time to actually lay out what's come out of it instead of posting each one separately.
Short version: eight of nine falsified outright. The ninth is sitting in a forward test right now, waiting on real data instead of anything I could've peeked at.
Started with the boring stuff — Friday close to Monday open fade on BTC futures, four different variants of it (ATR filters, weekend gap sizing, that kind of thing). Dead on arrival once you correct for testing four things at once. Same story with a standard vol-breakout setup, and with SMT divergence across gold, silver, BTC, SOL — ran four separate pre-registered versions of that one too, all four failed.
The SMC/ICT signal family was probably the most thorough falsification of the bunch — sweep, absorption, cascade, fair value gaps, opening window bias, tested on actual tick data instead of bars. Gross price moves were something like 7-10x smaller than round-trip cost across almost every setup. There was one cell that looked interesting (a cascade variant, BUY plus rising open interest at the 1-minute mark) but it was one cell out of over 30 tested, and it only showed up in one regime, so I'm not calling that a finding, just noting it exists.
Funding rate carry was the one where the "why" mattered more than the topline number. BTC and ETH just can't clear the cost floor, funding income never gets close. SOL was different though — in-sample the funding actually beat the cost floor, 0.55% vs a 0.48% floor. Should've worked. Didn't, because basis drift between the two legs ate the difference and it ended up net negative anyway. Two totally different ways to die.
Then there's the session-timing stuff, which took up most of the last few weeks. Gold vs dollar index divergence in the first two hours of London or NY — pooled test wasn't confirmed, came in at p=0.13 against a 0.10 bar. Splitting it open showed London was basically noise (looked great at 5 minutes, completely gone by 15), NY held up better across timeframes but that wasn't the headline test so I'm not claiming it as one.
Followed that up properly though. Locked NY-only as the actual headline this time, and made the confirmation period strictly forward — meaning everything I'd already looked at, including data from the first study, got reclassified as discovery-only and thrown out for confirmation purposes. Before any forward data even existed I ran the same exact rule on silver instead of gold, just to sanity check the mechanism. Came back p=0.0005 in the opposite direction. Not exactly a great sign for the gold version, but at least I found that out before wasting months waiting on it.
Then tried to formalize the "6 ICT session profiles" thing everyone posts screenshots of — Asia/London/NY blocks, turned each profile into an actual numeric rule instead of vibes, tested on BTC/SOL/XAU/XAG across three different threshold settings, 72 cells total once you multiply it out. Nothing survives Bonferroni. Lowest p-value in the entire sweep was 0.0025, which sounds good until you remember it's one cell out of 72.
Last one reused the same data to test something narrower — the Range/Manipulation/Expansion three-candle idea, two specific timing setups. This one's almost funny in how rare the setup actually is: even with the loosest thresholds I tried, 90%+ of days matched neither scenario. Under strict thresholds it's north of 99%. So beyond just failing the significance test, there's barely enough sample to ever say anything meaningful about it in the first place.
The thing that keeps nagging me across three of these now, using three completely different signal constructions: silver does not behave like gold. Opposite sign in the forward test, worst performer across all 72 cells in the profile study, negative in 7 of 8 cells in the reversal-day one. I've stopped assuming XAU and XAG are interchangeable for anything session-timing related, because apparently they're not.
All the repos (pre-registration docs, data collection scripts, the analysis code, raw results) are linked from my profile
Portfolio of 8 pre-registered falsification studies on retail/ICT crypto & forex signals — summary of what survived (almost nothing) and the one cross-study pattern that keeps repeating
Methodology across all 8: lock signal definition, cost model, and decision rule before touching data. Chronological or discovery/confirmation split for OOS. Bonferroni correction whenever multiple variants/assets/parameters get tested in parallel. Full code + data for each on GitHub, linked from profile.
Friday-Monday CME futures fade (base + 4 pre-registered variants: ATR filter, weekend gap, MA200 regime, extreme close position) — all falsified after Bonferroni, threshold p<0.025.
SMT divergence across XAU/XAG, XAU vs DXY-proxy, XAU vs synthetic DXY, BTC/SOL — four pre-registered tests, all falsified.
Standard vol-breakout entry logic, same discipline — falsified after correction, nothing new to add there.
Full SMC/ICT signal family (sweep, absorption, cascade, FVG, opening-window bias) on real tick data — gross moves ~7-10x smaller than round-trip cost across the board. One cascade cell looked suggestive (BUY+OI-up at +1min, t=2.76), flagged as regime-locked, one cell out of 30+ tested, not treating it as a claim.
Delta-neutral funding carry, BTC/ETH/SOL — falsified, but via two different mechanisms. BTC/ETH die on the cost floor outright (funding never clears ~0.48% round-trip). SOL in-sample funding actually beat the cost floor (0.55% vs 0.48%), but basis drift between the two legs ate the edge anyway, net -0.16%.
XAU/DXY divergence, first 2h of London/NY sessions — headline pooled test not confirmed (p=0.1265). London turned out to be resolution-dependent noise (significant on 5min, dead by 15min). NY held up across timeframes but wasn't the pre-registered headline, so it's reported as a lead, not a finding.
Direct follow-up on that NY lead — locked NY-only as headline this time, confirmation restricted strictly to data collected after the registration date (everything prior reclassified as discovery-only, no reuse of already-inspected data). Ran a cross-asset check before any forward data existed: same rule on silver came back p=0.0005, opposite direction from gold.
The "6 ICT session profiles" thing (Asia/London/NY blocks) — turned it into actual numeric rules, tested BTC/SOL/XAU/XAG across 3 threshold bundles, 72 cells total. Nothing clears the Bonferroni bar. Lowest p-value in the whole sweep is 0.0025, one cell, doesn't mean anything at that search width. No-match rate alone is telling: even the loosest thresholds leave 35-40% of days matching none of the six profiles.
Across two completely unrelated studies now, XAG just doesn't move like XAU on session-timing/liquidity signals. Opposite sign in the forward-test, worst performer across all 72 cells in the profile study. Different signal constructions, same conclusion both times. Not treating gold and silver as interchangeable for this kind of thing anymore.
Repos linked from profile.
Pre-registered XAU/DXY session-divergence: headline null result, but a NY-vs-London split worth documenting — plus a cross-asset silver check that flags a warning sign before the follow-up forward test even begins
Sixth in a series of pre-registered falsification studies (prior work in profile/repo). This one tests whether XAU/DXY divergence during the first 2 hours of London or NY sessions is tradeable net of costs.
Headline result: not confirmed. Pooled London+NY, 15min, k=1.5 — p=0.1265 against the pre-registered p<0.10 threshold. Per the locked decision rule, that's a null result on the actual registered claim.
What the full 36-cell sweep shows (diagnostic, not confirmatory): London and NY behave completely differently across timeframes. London is "significant" on 5min only, then flat/negative on 15min and 30min — that's the signature of microstructure noise, not a real session effect. NY holds up on 5min and 15min, weaker but still directionally consistent on 30min. That inconsistency between the two sessions is what pooling them hid.
Before chasing the NY pattern into a new pre-registration, ran a cross-asset plausibility check on the existing data: same exact rule, applied to silver (XAG) instead of gold. If the mechanism were a general NY-liquidity effect on precious metals vs the dollar, silver should show the same direction, maybe weaker. Instead: p=0.0005, 1,779 trades, mean return -0.15% — strong effect, opposite sign from gold.
That's now documented as a known warning sign before the actual forward test starts, in the new repo's README, not discovered and buried after a positive result came in. New pre-registration locks NY-only as the headline hypothesis, treats all existing historical data (including what was previously "confirmation" data in the parent study) as discovery-only, and only counts forward data collected after today as real confirmation. Deadline + minimum trade count enforced in code so it can't be checked early and reported as clean.
Both the parent study (headline null + full diagnostics) and the forward test (pre-registered today, pending) are up on GitHub — links are in my profile since this sub doesn't allow linking directly.
Curious if anyone's seen a mechanistic reason gold and silver would diverge in sign on the same session-timing signal — safe-haven vs industrial-commodity flow difference is my best guess but haven't dug into it properly.
Pre-registered falsification of delta-neutral funding rate carry (BTC/ETH/SOL) — structural edge fails, but via two distinct mechanisms (cost floor vs basis risk)
Crossposting from r/algotradingcrypto with mod approval — fifth in a series of pre-registered falsification studies on crypto trading signals (prior work: friday-monday-fade, volatility-breakout, turtle-soup-smt, retail-crypto-alpha — links in profile/repo).
This one tests a structural mechanism instead of a technical pattern: delta-neutral carry (long spot / short perp, or the inverse) triggered when funding rate exceeds a trailing percentile/z-score threshold. Full methodology, cost model, and decision rule locked in PREREGISTRATION.md before any data was touched.
Result: falsified across all three assets and all pre-registered variants. OOS mean net return per trade is negative in every single symbol x variant combination — no exceptions.
What's more interesting than the topline number: the failure mode is different per asset.
BTC/ETH: classic cost-floor problem. Round-trip trading costs are ~0.48% per trade, funding income never gets close, even in-sample.
SOL in-sample: funding income actually cleared the cost floor (0.55% vs 0.48%). Died anyway — basis drift between the two legs (price divergence at exit) ate the whole edge (-0.23%), net -0.16%.
Rebalancing mechanism (pre-registered 1% drift trigger) fired on 0-2.7% of trades — essentially inactive as specified, so it's not what explains the failure.
Full pre-registration, data collection scripts, analysis code, and component breakdown: https://github.com/Mykola-Quant/funding-rate-carry-falsification
Genuinely curious if anyone here has looked at basis-risk mitigation (tighter hedge ratio maintenance, dynamic rebalancing) on structurally similar trades — that seems like the actual open question this leaves, not funding-timing sophistication.