I'm a student learning quant. I built a strategy router as a project, it kind of worked surprisingly well with what it had, here's what I learned!
I'm a student at NYU studying to get into quant. Last year I started a project that I thought would be small, building a router that picks the best strategy for a given market in DeFi and sizes it with Fractional Kelly Criterion. Like an automatic mini allocator. I thought it would be a good way to actually learn how position sizing and edge estimation work instead of just reading about them.
I had maybe 5 strategies I wrote running on ETH paper data and the router would pick whichever had the best recent risk-adjusted return and size it with fractional Kelly.
The first thing I learned: most strategies do not have edge.
Out of maybe 30 strategies I tested initially, 1 or 2 had any real edge after costs (or so I thought :), they got absolutely destroyed after realistic trading fees and friction). The rest were noise, though diversified. Crypto round-trips are like 6-7 bps per side depending on the pair, and if the edge is 10 bps per trade, I'm losing money.
I thought AI could help me with this and tried to improve the existing strategies with it. It gave worse results, it overcomplicates strategies a lot. What I surprisingly found is that dumb and small code works much better than overfit models in the real world. And the tiny "dumb" strategies with on-chain data proved to be much much better than the rest, some even profitable on 2 years of trading data!
I added a cost-adjusted validation stage and regime decomposition. Seeing where a strategy bleeds (chop vs trend vs crisis) helped explain why backtests fail live.
The second thing: the router was actually decent.
Once I had enough strategies, I built a perfect-foresight benchmark (an oracle that picks the best strategy for each window, kinda like God or Congress :) to see if my allocator was doing anything.
The router captures about 86% of the perfect-foresight ceiling (95% lower bound ≈ 39%) and this is on the same volume of trading, around 7-12 trades per day on both the oracle and my router on a very diversified roster of bots. I also found that best available edge scales with roster size at r=0.986 against extreme-value theory (the √(2·ln N) scaling).
This mechanism could technically make money, but the allocator wasn't the bottleneck, the roster quality was.
The network effect (this is the part I'm most excited about):
I wanted to know does adding more strategies actually help or am I just diluting? I subsampled my 93-bot roster down to smaller sizes (10, 20, 35, 50, 70, 93 bots) and re-ran the entire pipeline from scratch on each one.
The best bot's true edge climbs monotonically as you add more: −6 bps at 10 bots → +3 bps at 93 bots. When I fit that against extreme-value theory (the √(2·ln N) scaling that predicts the maximum of N random draws), the correlation is 0.986… almost a perfect match. More strategies = higher ceiling, and it follows theory almost exactly.
But here's the catch: the router only captures that rising ceiling if you use an absolute quality bar, not a relative percentile. If you filter "top 30% of whatever roster exists," the router's edge stays flat no matter how many bots you add because the percentile just re-centers on whatever population is there. If you use a fixed quality threshold instead, the router's edge climbs with the roster. Extrapolating (with caveats, this is beyond the range I actually tested): ~+10 bps net edge at 1,000 bots, ~+16 bps at 10,000.
That's the quantitative argument for why I want creators :) Every good strategy added raises the ceiling for everyone.
How it expanded into a creator studio:
The router dynamically updates its own parameters as the roster changes but it does this via offline re-tuning on a cadence, not via real-time ML yet. The reason is that at 93 bots and ~8 trades/day, you can't detect effects smaller than ~47 bps with any statistical power. A real-time ML model would just be fitting noise. It also has self-capacity awareness so it doesn't frontrun itself.
Once the router worked, I began noting down everything scientifically and made a bunch of changes to my initial project. I added real-time on-chain data feeds with historical data as well. These became obvious next steps:
- 6 active domains: ETH, BTC, SOL direction + scalp (6 more registered but dormant, yield, tail hedge, liquidation arb, memcoins (This one might be insanely hard to get right tbh), etc, waiting for strategies)
- 5-stage validation pipeline: static check, in-sample, out-of-sample, walk-forward, cost-adjusted
- Creator IDE: write Python strategies with custom stop-loss, take-profit, and trailing stops directly in the browser or via API/MCP
- Arena & Leaderboard: strategies that pass validation compete on live paper data for capital allocation
- Non-custodial: API keys stay encrypted in user vaults
Where it is now:
- 80+ default strategies running on live paper data (real prices, paper execution, Not great bots:)
- 3 are currently net profitable (best: ETH Squeeze Breakout, +184bps, 73% win rate). The rest are negative.
What I'm looking for: I'm posting here because this subreddit has people with real domain expertise, and I'd love your feedback:
- Does the router / Kelly allocation approach make sense, or is there an obvious flaw I haven't seen?
- Is capturing ~86% of a foresight ceiling considered typical or decent for this setup (on a volume of 7-12 trades per day, on my quite diversified roster of bots)?
- What features would you actually need in a Python strategy sandbox to make it worth testing your own models?
I'm a student and not charging for anything. Oh, and importantly nobody can see any strategy code, it runs in a confidential VM! Happy to share more details if anyone's curious.
TL;DR:
I'm a student at NYU. Built a router that allocates capital across Python trading strategies using Fractional Kelly for DeFi. Tested 80+ strategies on live paper data, most have no edge after costs (shocking!! I know). The router captures ~86% of a perfect-foresight ceiling at matched volume. Found a network effect: best-available edge scales with roster size at r=0.986 vs extreme-value theory... so more strategies = higher ceiling for everyone. It snowballed into a full creator platform (5-stage validation, arena, non-custodial, confidential VMs for strategy privacy and more). Would love feedback from people who actually know what they're doing. Does the approach make sense, is 86% of ceiling decent, what would you need in a Python strategy sandbox?