u/ryanturbine

My Kalshi Momentum Strategy for 15 minute BTC also backtested well for the 15 minute ETH market

My Kalshi Momentum Strategy for 15 minute BTC also backtested well for the 15 minute ETH market

I had already seen this momentum strategy backtest well on Kalshi's 15-minute BTC markets a week ago (report here), so I wanted to know whether the same idea would carry over to ETH. In this historical run, it did: all 100 completed ETH variants were profitable.

The strategy was like this: During the final 5 minutes, the strategy only enters when the contract is priced from 0.45 to 0.55 and three Coinbase signals agree: ETH's 5-minute change, ETH's 1-minute velocity, and BTC's 1-minute velocity. All three positive means buy YES; all three negative means buy NO. Each entry is 10 contracts, the maximum position is 30, and the bot exits if unrealized P&L falls to -$4.50 or fewer than 5 seconds remain. The only difference with this ETH strategy was the coinbase signals for ETH instead of BTC (and of course the market being traded was the 15 minute ETH market instead of BTC).

https://preview.redd.it/smatmzpkc4kh1.png?width=1080&format=png&auto=webp&s=2ed3d6006b50499ca2c5d259fbfe25b37c183569

I ran 100 variants over the same 30-day historical period, and 100/100 finished profitable. The best returned 477.93% ROI and +$143.38 P&L over 105 trades, with a 66.7% win rate, 0.60 Sharpe, and -$21.54 max drawdown. The weakest still returned 191.60% ROI and +$57.48 P&L, but it needed 250 trades and came with a 62.7% win rate, 0.15 Sharpe, and -$92.32 max drawdown. So profitability was widespread in this sample, but the risk varied a lot by configuration.

https://preview.redd.it/uofg91omc4kh1.png?width=1080&format=png&auto=webp&s=d533e07f3a774e5fc4f3a2c037a3570a0696e619

The parameter sensitivity test crossed price floors from 0.05 to 0.45 with price ceilings from 0.55 to 0.95. All 100/100 cells succeeded, with net P&L ranging from +$57.48 to +$143.38. The winner used a 0.41 floor and 0.59 ceiling. Its neighbors were only 3.8% worse on average, so the 3D surface looks more like a local plateau than a lone spike. The heatmap and marginal curves tell the more useful story: raising the floor helped a little, while widening the ceiling hurt much more. Deflated Sharpe was 0.835 versus an expected maximum Sharpe of 0.390, but it stayed below the 0.95 significance threshold used by the report.

https://preview.redd.it/3j34g6kwc4kh1.png?width=1080&format=png&auto=webp&s=757f3c6a1191a9aa40ec14987610cafb0b5ae0b2

I also ran a permutation test that shuffled the timing of the edge feed and repeated the full sweep. The real best net P&L was +$143.38 and beat 99.4% of 153 shuffled re-sweeps, with an upper-tail p-value of 0.0065. That sounds strong, but the test hit its time limit and used only the 153 completed permutations, so its status is degraded. It also left market prices untouched. This tests whether the ETH and BTC signal timing mattered, not whether the 0.45 to 0.55 entry band itself was valid.

My read is that the upper price bound did real work in this historical sample. A ceiling near 0.59 kept drawdown much lower, and nearby settings held up reasonably well. Still, this is one 30-day sample and the statistical checks were mixed, so I would not call the edge proven. The ETH version held up well enough to justify testing it on unseen data next, which is exactly what I wanted to learn from this experiment.

Full Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 1 day ago

Testing 100 price floor/ceiling ranges on my Kalshi BTC 15-minute momentum strategy

I tested a momentum strategy on Kalshi's 15-minute BTC markets. It trades only when three signals agree: BTC's 5-minute change, BTC's 1-minute velocity, and ETH's 1-minute velocity. When all three are positive it buys YES, and when all three are negative it buys NO. Entries are limited to the final 5 minutes while the contract price is between 45 and 55 cents. Each trade is sized at 10 contracts, with a 30-contract position cap, a -$4.50 stop loss, and an exit just before settlement.

https://preview.redd.it/dg540psn5yih1.png?width=1080&format=png&auto=webp&s=7834aed966dd7177bb44e85b0d49360297b9e522

The research ran 100 backtests over the same 30-day historical period, and all 100 finished profitable. The best-ranked version returned 324.10% ROI and +$97.23 P&L, with a 66.7% win rate across 57 trades, a 1.03 Sharpe, and -$6.96 max drawdown. Even the weakest version stayed green at 131.27% ROI and +$39.38 P&L across 71 trades. That matters because the result did not depend on finding one profitable configuration among a pile of losing ones.

https://preview.redd.it/qb3ts2kq5yih1.png?width=1080&format=png&auto=webp&s=1e85975b19ac31397ef3c5ec9b50efe7907b3a4d

A parameter sensitivity test reruns the same strategy while changing nearby settings. The point is to see whether performance holds across a range or collapses as soon as one number moves. We swept the risk price floor from 0.05 to 0.45 and the risk price ceiling from 0.55 to 0.95, using 10 values for each and producing a 100-cell grid. The 45-to-55-cent entry rule and the momentum signals stayed fixed. All 100 cells completed, with Net PnL ranging from +$39.38 to +$97.23. The top-ranked cell used a 0.05 floor and 0.59 ceiling, but +$97.23 repeated at every tested floor when the ceiling was 0.59. On the chart, that creates a flat ridge across the floor axis and a sharper peak along the ceiling axis. That shape says the floor had very little effect, while the ceiling mattered much more. P&L peaked at a 0.59 ceiling and generally declined as the ceiling moved higher. The Deflated Sharpe was 0.99 versus an expected maximum of 0.33 across the 100 trials.

https://preview.redd.it/c9jx85kv5yih1.png?width=1080&format=png&auto=webp&s=dead37832f57ed7c73fe977ddb92a637c9e6ef64

The permutation test asked a different question: could random timing in the edge feed produce a result this good? I scrambled the edge-feed timing and reran the full parameter sweep. The real winner's Net PnL was +$97.23 and beat 99.9% of the 976 completed reruns, giving an upper-tail p-value of 0.001. Randomized timing rarely matched the real result in this historical sample, which supports the idea that the timing of the momentum inputs carried useful information. The test hit its time limit, so the p-value uses a reduced sample of 976 reruns. It also left market prices untouched, which means it does not validate the strategy's price-based conditions.

My read is that the strategy's main strength is consistency. Every tested configuration made money, the sensitivity sweep shows which parameter drove the variation, and the permutation result suggests the edge-feed timing was not easily reproduced by chance. I'll be paper trading this next to see how the results hold up.

Full Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 8 days ago

My Kalshi 15 minute BTC market strategy’s filtering was too strict, so I tried fixing it by running 100 backtests with looser filters

I had a strategy for Kalshi’s 15 minute BTC market making trades filtered by BTC’s VWAP, EMA, and 5-minute momentum price shifts. The problem with it was it was rejecting too many decent entries so I wanted to refine it by loosening entry bands, lowering momentum thresholds, and simplifying edge conditions. The process I followed was a research run that I ran 100 backtests across different price floors and ceilings while keeping the loop timing, position sizing, momentum thresholds, and other strategy rules fixed.

91 of the 100 variants finished profitable.

The best variant used a 0.45 price floor and a 0.82 price ceiling. It made $31.24, returned 0.62% ROI, had a 63.4% win rate, and placed 177 trades. The weakest completed variant lost $9.17 and returned -0.18% ROI.

I also checked the price bounds with a 10 by 10 sensitivity grid: 10 floor values from 0.05 to 0.45 and 10 ceiling values from 0.55 to 0.95. Net P&L ranged from -$9.17 to +$31.24, with the winner at a 0.45 floor and 0.82 ceiling.

https://preview.redd.it/o08df1er65ih1.png?width=1080&format=png&auto=webp&s=a0336a507845e78cf51963a4de374c4b596ab180

The heatmap and 3D surface show a fairly broad positive area near the upper-right corner, not one isolated winning cell. Adjacent cells were 18.48% worse than the winner. The marginal charts also show that mean and best P&L generally rose with the price floor. Since 0.45 was the highest floor tested, I would treat it as a boundary result, not a settled optimum. The ceiling behaved differently: performance was weak at 0.55, improved after that, and the best P&L flattened once the ceiling reached roughly 0.82.

The statistical caveat is the main reason I don't trust the headline result yet. Deflated Sharpe was 0.6319, while the expected maximum Sharpe from the 100 tested cells was 0.1585. The report flags the deflated result as below its significance threshold, so the winner is still hard to separate from selection noise.

This 100 variant run showed me that the strategy’s price filters were probably too restrictive: higher price floors and wider ceilings consistently produced better results. But it didn’t necessarily prove the edge is reliable, since the best result sat at the edge of the tested range and wasn’t statistically significant.

Full Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 12 days ago

This Kalshi 15-minute BTC strategy looked promising until I tested 100 parameter combinations.

I tested a BTC momentum idea on Kalshi's 15-minute markets. The setup only entered during the final 120 seconds, when the expensive side was at least $0.90 and BTC's 5-minute Coinbase move agreed with its position above or below the 1-hour VWAP.

The idea sounded plausible: if the market was already heavily tilted and the spot signal agreed, maybe the move would hold through expiry. I wanted to see whether that result survived nearby parameter settings.

All 100 variants used the same entry logic. The only changes were the price floor, tested from 0.05 to 0.45, and the price ceiling, tested from 0.55 to 0.95. Every cell completed successfully.

The best cell used a 0.05 floor and 0.95 ceiling:

- Net P&L: +$40.03

- ROI: 400.30%

- Win rate: 67.7%

- Trades: 62

- Max drawdown: -$2.19

Across the full grid, 36 variants finished profitable. The other 64 placed no trades, so there were no losing variants.

Even though no variants came back negative, the ones that weren't profitable never placed any trades in the simulation so I wouldn't consider that a good sign.

Net P&L ranged from $0.00 to +$40.03, with most of the useful performance concentrated below roughly a 0.10 floor and at the maximum 0.95 ceiling.

The sensitivity plots also made me less comfortable with the winner. It sat at the extreme corner of the grid. The surface was more of a tilted ridge than a single spike, but most of the grid was flat at zero or close to it.

The marginals tell the same story. Raising the floor from 0.05 to 0.09 cut average Net P&L from $29.02 to $16.07. Above that, the average dropped to about $1.41. The ceiling barely mattered until the final 0.95 setting, where both average and maximum Net P&L jumped.

Overall, the sweep suggests the apparent edge is narrow and parameter-sensitive, with the best result at the grid boundary and a low Deflated Sharpe, so it should not be considered robust.

The real winner's Net P&L was +$40.03 and beat 99.8% of the 415 completed reshuffles, with an upper-tail p-value of 0.002.

The research run also reran the full search against time-scrambled versions of the BTC edge feed. That permutation run was degraded because only 415 reshuffles finished before the time limit, so the p-value uses the reduced sample. More importantly, only the edge-feed timing was scrambled. Kalshi market prices were not permuted, which means this test does not validate the price-based entry conditions or the fill assumptions.

My read is that the BTC feed may contain some timing information in this historical sample, but the 400% result is not a live-trading expectation. The winner depends on boundary settings and simulated fills that deserve much more scrutiny.

Full Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 15 days ago

My Kalshi LAX daily high market maker backtested great but the parameter sensitivity analysis suggests it could be overfitting

I wanted to see whether wide spreads in Kalshi's LAX daily high temperature markets could support a passive market maker without needing a weather forecast as an edge source.

The bot posted a 25-contract YES bid and a 25-contract NO bid at the same time. I kept that rule set fixed and ran 100 combinations of quote offset (-2 cents to +2 cents), minimum spread (5 cents to 9 cents), and refresh interval (30 to 120 seconds).

I ran 100 backtests and 93 of the 100 variants made money in the simulation.

https://preview.redd.it/aqv5vp8qzyfh1.png?width=1080&format=png&auto=webp&s=746e29464f8f3c1337dcdb271bce92fa8746ad4c

This looked like a great strategy just looking at these results, but our research flow runs a parameter sensitivity test to sweep the minimum spread and quote offset and that results suggested there could be some overfitting at play.

https://preview.redd.it/jaibzgsu1zfh1.png?width=1080&format=png&auto=webp&s=d19c46d2b267da72b7433fe2717dada06c3957b1

Evidence for overfitting:

  • The winner is an isolated peak; neighboring cells lose an average of 77.6% of its value.
  • It has only 19 resolved trades.
  • The best minimum spread, 5 cents, is at the grid boundary, so the apparent optimum is not confirmed.
  • The deflated Sharpe is 0.81, below the report’s 0.95 confidence threshold after accounting for multiple trials.
  • You selected the winner from 100 variants tested on the same historical window.

The counterpoint is that 93/100 full variants and 23/25 sensitivity cells were profitable. That suggests the general thesis may have something behind it, even if the +$57.08 winner is mostly parameter luck.

I still plan on paper trading this on our site with the winning parameters and that should hopefully give me the next steps for tweaking this market maker further.

Full Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 23 days ago

I tried building and rigorously testing a two sided market maker on Kalshi's daily BTC markets

Market making sounds almost directionless: leave passive orders on both sides and collect the spread. Prediction markets make the clean version easy to see since a YES share and a NO share pay $1 between them at settlement. If both buy orders fill for less than $1 total, the pair has a locked in gross profit before fees.

The problem is that the two fills rarely arrive together. One side may trade because somebody knows more, reacts faster, or simply has a stronger view. Until the other side fills, the maker owns a directional position. Queue priority, stale quotes, fees, and cancellation speed can turn an apparent spread into a bad trade.

The bot I tested is a deliberately simple take on that setup for Kalshi's daily BTC markets. Once a minute, while flat, it checks that the spread is at least $0.02 and that more than one minute remains before expiry. If both conditions pass, it posts two passive orders: 25 YES contracts at the YES best bid and 25 NO contracts at the NO best bid. Both are post only, so the bot provides liquidity instead of crossing the book.

If the spread narrows below $0.02 or the market gets within one minute of expiry, the bot cancels its quotes. After a fill, it stops refreshing them. It then holds the inventory through settlement unless unrealized P&L falls below -$25, which triggers an exit. Position size is capped at 50 contracts, with no more than three portfolio positions open at once.

That last part is where most of the risk lives. A matched YES/NO pair can capture the gap between the combined entry price and $1. A one-sided fill is just a bet on the outcome, even if it started life as a market-making quote. The strategy is therefore betting that wide spreads pay enough to cover the occasions when only the wrong side fills.

I was mainly interested in the price bounds. Quoting very cheap contracts may offer more room, but those markets can also be thin for a reason. Raising the floor avoids the cheapest outcomes, while lowering the ceiling keeps the bot away from expensive contracts. I swept both to see whether the strategy worked across a reasonable area or depended on one narrow corner.

Here are the results from the 100 backtest research run:

Of the 100 variants tested, 61 were profitable, with the best producing $42.76 in net P&L.

The test covered a 100 cell grid. Price floors ran from 0.05 to 0.45 and price ceilings from 0.55 to 0.95. All 100 cells completed, and 61 were profitable.

The best cell used a 0.05 floor and 0.55 ceiling. It returned +$42.76, or 85.52% ROI, with a 42.9% win rate across 16 trades. The weakest cell returned -$27.17, or -54.34% ROI.

The heatmap is less flattering than the headline. Positive results cluster around the two lowest floor settings. Mean net P&L was +$30.87 at both 0.05 and 0.09, +$12.53 at 0.14, then -$11.46 at 0.18. The ceiling response was much less consistent. The lowest ceiling produced the best cell, but changing the ceiling did not create the same clear pattern.

The winner sits in a corner rather than on a broad plateau. Some of the sharp drop offs in the parameter 3D sweep surface suggest some possible overfitting so I'm not taking the profitable backtests directly for their word. That being said, this did not model liquidity rewards, so in live trading that could counter balance that reality.

For that reason I'm currently paper trading this on our site before deploying it so I will follow up with the results of that if anyone is interest.

Full Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 29 days ago

100/100 backtests of a BTC momentum strategy came back positive

One of our users ran 100 backtests and a parameter sensitivity analysis on their momentum strategy and the results came back extremely promising.

One of the first brutal truths you learn when you start getting into automated trading is that believing a single backtest can determine the future profitability of a strategy is misguided. Many who come to learn this become pessimistic about wether or not its even worth backtesting. But this pessimism is mostly rooted in ignorance since backtesting is just a tool the same way knowing how to code in python is (neither can guarantee anything alone but used correctly they become essential in successful automated trading).

This user's strategy works like this: it trades the Kalshi 15 minute BTC market and buys YES when BTC's 1-minute EMA 12 is above its 1-minute SMA 20 and spot is above the 5-minute SMA 50. The bearish rule does the reverse and buys NO.

Rather than trusting one strong backtest with a single set of parameters, the user put the strategy through our research flow. It tested 100 variants using different combinations of two inputs: the price floor and price ceiling. The parameter sweep showed how performance changed across the full grid, and a permutation test checked whether the result could reasonably be explained by chance. This made it possible to see whether the strategy was genuinely robust or whether its performance depended on one lucky combination of settings.

The sweep covered 100 combinations of two execution bounds: price floor from $0.05 to $0.45 and price ceiling from $0.55 to $0.95. Every cell completed, and all 100 were profitable in the simulation.

The best cell used a $0.05 floor and $0.59 ceiling:

- Net P&L: +$446.84

- Reported ROI: 8,936.8%

- Sharpe: 1.21

- Win rate: 57.7%

- Trades: 1,796

- Max drawdown: -$26.37

The weakest cell still made +$273.66, with 5,473.2% reported ROI, a 0.91 Sharpe, 58.3% win rate, 1,438 trades, and -$26.35 max drawdown. Average net P&L across the grid was +$389.70. The P&L range was +$273.66 to +$446.84.

https://preview.redd.it/gs2mn864k0dh1.png?width=1080&format=png&auto=webp&s=188753538e10f566bcfa2faad3ff76dc84b351b4

The shape of the parameter sweeps 3D surface came back very positive. The winner's neighborhood degraded by only about 1%, and the surface stayed fairly flat until the floor moved above roughly $0.35 which suggests more of a plateau than a single lucky spike. In addition to a robust looking surface, the Deflated Sharpe was 0.999989 versus an expected maximum Sharpe of 0.207778 under the report's 100 trial assumption.

One interesting finding however was that the ceiling barely mattered. Mean net P&L stayed between about $383.88 and $401.89 across the ceiling values, while raising the floor steadily hurt the result. That suggests the floor did most of the work and the ceiling may be redundant.

https://preview.redd.it/41y45hy1l0dh1.png?width=1080&format=png&auto=webp&s=37aab59e3319eb10f1e5ed3486495464651e5777

The research run also checks the result against time-scrambled versions of the edge feed. The real +$446.84 winner beat 99.6% of 227 completed re-sweeps, with an upper-tail p-value of 0.004386. That is encouraging, but this was a degraded run because only 227 permutations finished before the time limit. More importantly, the test scrambled edge-feed timing only. It did not scramble market prices, so it does not validate the strategy's price-based conditions.

Overall, this research run looks really good and definitely adds a degree of confidence for the strategy over just a single backtest. That being said, even this level of rigor doesn't guarantee real life performance. Only time will tell if this analysis predicted a winning strategy, or if its gaps in accuracy and coverage were still too wide to call the strategy complete.

Full Research Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 1 month ago

One profitable backtest is not proof. I ran 100, then tried to break the winners.

It is easy to stop testing as soon as a backtest turns green.

I could have done that with this simple mean reversion strategy on Kalshi's 15-minute BTC markets. One version made $57.21 across 488 trades and reported 572.1% ROI. If that were the only result I looked at, I might have called the strategy profitable and expected similar performance going forward.

So I tried to break it instead.

First, I ran 100 backtests. I varied the lower entry level from 20 to 40 cents, the upper level from 50 to 70 cents, and the minimum entry price across 5, 10, 15, and 20 cents.

https://preview.redd.it/dlaiov94cbch1.png?width=1080&format=png&auto=webp&s=b7894a6d35494843afb3b1849c42c3da99b39816

Only 25 of the 100 variants made money. Average P&L was -$307.76, and the worst run lost $1,497.43. The profitable backtest existed, but it was not representative of the strategy family.

Next, I plotted the 25 entry-band combinations. Net P&L ranged from -$1,492.92 to +$57.21. The winner sat in one narrow corner rather than a broad profitable area, and performance fell quickly as the parameters moved away from it. Deflated Sharpe was effectively 0 versus an expected maximum of 0.97 for the 25 combinations. Deflated Sharpe adjusts for the fact that testing many settings makes it easier to find one good-looking result by chance. A value near 0 means the winning Sharpe did not clear that multiple-testing bar, so there is very little statistical support for treating it as skill rather than parameter selection.

https://preview.redd.it/hczngimecbch1.png?width=1080&format=png&auto=webp&s=11ea0035adb3c82d37b0499f5008a28dcdc6b055

Then I reshuffled the historical market data 1,000 times and retested all 25 combinations on each reshuffled version. The real winner beat 74.0% of those runs. Its upper-tail p-value was 0.260, so reshuffled data produced a result this good or better about 26% of the time. Each reshuffled run was allowed to choose its own best setting, just as the real sweep did. Beating 74% may sound decent, but it is not statistically unusual. A conventional 5% significance threshold would require the real winner to beat roughly 95% of the reshuffled runs.

https://preview.redd.it/bs2tu2ticbch1.png?width=1080&format=png&auto=webp&s=9a7893083cf424e3e4b00772bc9be081493ad21b

That is the level of rigor I want after seeing a profitable backtest. A single result tells me that one configuration worked on one historical path. It does not tell me that nearby settings work, or that the result is unusual enough to separate from luck.

If I had stopped at the positive backtest, I would have gone in expecting far more than the evidence supported. The 100 runs, parameter sweep, and permutation test all pointed to the same problem: this strategy was fragile, and the headline result made it look more dependable than it was.

Full research report:
BTC Kalshi Mean Reversion

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 1 month ago

Trying to find the right entry band for 15 minute Kalshi mean reversion

I've been trying to perfect a mean reversion strategy on the Kalshi 15 minute markets so I recently ran a test to find the optimal entry band: the price range where I’m willing to buy after an early move, with Entry Low setting how deep the dip needs to be and Entry High setting the upper limit where the trade still looks attractive. I ran two 100 variant backtests on the BTC and ETH 15 minute markets, and did a parameter sweep test for both.

For each market, I tested Entry Low values of 0.20, 0.25, 0.30, 0.35, and 0.40, crossed against Entry High values of 0.50, 0.55, 0.60, 0.65, and 0.70. All 25 combinations ran for BTC and all 25 ran for ETH. Interestingly what was found to be the optimal entry band in both BTC and ETH were similar.

BTC's best backtested band was Entry Low = 0.30 and Entry High = 0.55. That setup produced +$57.38 Net PnL in the entry-band search. Across the tested BTC bands, Net PnL ranged from -$8.65 to +$57.38.

ETH's best backtested band was Entry Low = 0.20 and Entry High = 0.55. That setup produced +$37.14 Net PnL. Across the tested ETH bands, Net PnL ranged from -$21.12 to +$37.14.

So the upper side matched exactly with both BTC and ETH liking Entry High 0.55, but the lower side was close but not identical. BTC's best Entry Low was 0.10 higher than ETH's, meaning BTC did better with a less aggressive dip entry, while ETH needed to wait for a deeper move.

The shape of the results was quite different too. BTC looked more forgiving with 88/100 BTC variants finishing profitable, and the better results formed a broader area around Entry Low 0.30 and Entry High 0.50 to 0.55. ETH was more fragile with 45/100 ETH variants finishing profitable. The best area was tighter, centered around Entry Low 0.20 and Entry High 0.55, and performance fell off as Entry Low moved higher.

Even though the parameter sweep didn't score great for overall strategy robustness, I feel like my goal of finding a good entry band range was achieved and validated by the fact that it was very close on two different markets. I'll probably run the same test on the other 15 minute markets as well to see how it holds up (let me know if you'd want a follow up post about).

Full research reports if you're interested:
BTC 15 Minute Mean Reversion
ETH 15 Minute Mean Reversion

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 1 month ago

A Kalshi BTC bot hit a +54.7% live snapshot by only trading the last 4 minutes

I pulled one anonymized TurbineFi user run because the setup produced a big short-term move within its first day.
Public strategy

It trades Kalshi's 15-minute BTC markets. The bot waits until there are less than 4 minutes left, then buys the side that is already winning

  • buy YES if YES is 85c or higher
  • buy NO if YES is 15c or lower
  • use 50 contracts
  • sell everything with 30 seconds left

The thesis is very simple: near the end of the market, follow the side already priced as likely to settle in the money instead of trying to catch a reversal.

The historical backtest showed +229.96% ROI with a 92.8% win rate over the test window. In live telemetry, the bot hit a +54.7% snapshot after about 14 hours. However, a later snapshot for the same run was only +6.31%, and the the run even dipped slightly negative on one occurrences. The backtest actually clearly shows a large drawdown and not so great sharpe so this doesn't come as a big surprise to me.

What are some tweaks you would make to improve the drawdown on this strategy?

u/ryanturbine — 2 months ago

A Kalshi BTC VWAP bot did $100 in 1 day after adding guardrails

I pulled an anonymized run from one of our TurbineFi users that shows how regime dependent these 15-minute crypto markets can be. Same VWAP idea, but the live behavior changed a lot once the user started messing with sizing and exits.

The first version traded Kalshi KXBTC15M using Coinbase BTC-USD as the signal. If BTC was above its 1h VWAP and the 5m move was positive, it bought YES. If BTC was below VWAP and the 5m move was negative, it bought NO.

The base rules were 1 contract per signal, max position 11, spread filter at 0.03, entry band from $0.25 to $0.75, and max loss at $100. No take-profit. No “only enter when flat” check.

That was the issue. It could keep firing whenever VWAP agreed. On these short crypto markets, that can go from “nice signal” to churn pretty quickly.

So the user added restrictions:
https://www.turbinefi.com/backtest/coinbase-vwap-momentum-db0239e9238d

They did increase size from 1 contract to 2 contracts per signal, with max position 12 instead of 11. But the more important change was making the bot pickier. Spread had to be 0.02 or tighter. It could only enter when there was no open position. It took profit at +$1 unrealized PnL. Max loss went from $100 to $20.

Funny part: the base version had the better headline backtest PnL.

Base version:

  • +$1,980.71 PnL
  • 55.7% win rate
  • 0.54 Sharpe

Updated version:

  • +$470.51 PnL
  • 71.1% win rate
  • 0.59 Sharpe

So the update gave up a lot of raw historical PnL, but got a much higher win rate and better guardrails.

The first Coinbase VWAP live run netted +$20.46. The later live run showed +$99.61 after 1 day, with 50 fills and 250 trades/fills observed.

Curious how others would tune this first: VWAP threshold, spread filter, take-profit, or sizing?

u/ryanturbine — 2 months ago

Has anyone had success automating sports prediction markets?

Most of what I see people build on TurbineFi is around crypto or weather markets, where the extra signals are pretty obvious i.e. Coinbase/market data for crypto, NWS/METAR-style data for weather, etc.

I’m curious about sports markets though. Has anyone here had real success with automated sports strategies on Kalshi or Polymarket?

What types of sports markets seem most bot-friendly; game winners, spreads/totals, player props, tournament outcomes, in-play markets, something else?

And do external signals like sportsbook odds movement, injury/lineup feeds, or public betting data actually create an edge after fees, slippage, and liquidity or does the edge mostly get arbitraged away too fast?

reddit.com
u/ryanturbine — 2 months ago

What do you look for in a backtesting engine

For people who build trading strategies on Kalshi or Polymarket, what are some nonnegotiables and things that stand out to you in a great backtesting engine.

I’m currently trying to make our own backtesting engine on TurbineFi the best it can be and I’m trying to figure out what parts are most worth improving next.

Right now we use L2 orderbook candle data, simulate fills against depth, handle partial fills, model platform specific fees, and track slippage from signal price to fill price.

Where do you think backtesting products usually fall short? Fill modeling? Queue position? Latency? Settlement handling? Overfitting? Better trade logs?

Would love to know what things you would expect before you could trust a backtest for your own strategies.

reddit.com
u/ryanturbine — 2 months ago

100 variants of one Kalshi weather thesis all finished green in my backtest

I ran a full 100-variant backtest on TurbineFi on a Kalshi Chicago high-temperature strategy using NWS data.

The market family was KXHIGHCHI, the weather feed was NWS station KMDW, and the idea was to fade expensive YES prices when the official weather data did not support a hot outcome. In plain English; if the market was still pricing a meaningful chance of 85°F or higher, but the current temperature and NWS forecast high were both capped below that level, the bot bought NO.

The base entry rules were:

  • Broad NO entry: buy NO when YES was above 0.40, current_temp_f was below 85°F, forecast_high_f was at or below 85°F, and the NWS observation was fresh.
  • Stronger NO add: add when YES was above 0.50, current_temp_f was below 83°F, forecast_high_f was at or below 84°F, and the NWS observation was fresh.
  • NWS freshness mattered. The base rule capped observation age at 3600 seconds so the bot was not trading on stale weather readings.

The exits rules wee:

  • Take profit when YES repriced below 0.30.
  • Exit if the current temperature reached 85°F or higher, since that invalidated the NO thesis.
  • Hard stop if unrealized P&L hit -$10.
  • Flatten near expiry, within 30 minutes of market close.

Across the 100 runs, I varied price bounds, max position, position size, loop timing, observation-age limits, and some exit triggers.

https://preview.redd.it/u98g0zr4bm9h1.png?width=1080&format=png&auto=webp&s=f8cc59601083d9bfe4c879297f9c0b7c8f0f3713

  • Completed variants: 100/100
  • Profitable variants: 100/100
  • Total simulated trades: 999
  • Best ROI: 54.6%
  • Weakest completed ROI: 3.44%
  • Average ROI: 23.3341%

The top performers split into two types. The best ROI runs were selective where variants 042 and 044 returned 54.6% ROI on only 8 trades. The higher P&L winners traded more often where variants 032 to 035 took 17 to 19 trades leading to an overall higher P&L. Both groups kept a 100% win rate in the simulation.

The bottom performers were still green, but they showed where the signal got weaker. Variants 038, 039, and 040 returned 3.44% ROI, took 28 trades, and had a 50% win rate. The weak variants usually let NWS observations get stale or entered closer to the 85°F danger zone. Variant 017 shows the other failure mode. It kept a 100% win rate but only reached 21.12% ROI because it scaled down too much.

Overall the strongest variants enforced fresh NWS observations, kept the broad entry at YES > 0.40 with current temperature below 85°F and forecast high at or below 85°F, and used the stronger add only when YES was above 0.50 with current temperature below 83°F and forecast high at or below 84°F.

Position sizing seemed to scale with confidence. The stronger entry added exposure without creating many losses, while stale data and threshold erosion were the main ways the weaker variants gave back edge. The strategy looked best when it waited for a clear gap between market pricing and the NWS read, then exited quickly if YES repriced lower or the temperature actually reached the danger zone.

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 2 months ago

I ran 100 Kalshi BTC variants fading the ETH-leads-BTC hunch. The best ROI was -46.08%.

I tested a simple Kalshi BTC idea: maybe ETH impulses are not useful as a BTC continuation signal. What if ETH leading BTC is a trap?

The setup used the KXBTC15M market. When Coinbase ETH had a clean 5-minute and 15-minute impulse but Coinbase BTC was lagging or failing to confirm, the bot would fade the expected BTC catch-up. ETH up and BTC not confirming meant buying NO. ETH down and BTC not confirming meant buying YES.

I kept that rule family fixed and swept the risk bounds, price filters, max position, and loop timing across 100 variants.

These were the results:

https://preview.redd.it/lpm54ji1q29h1.png?width=1080&format=png&auto=webp&s=803a66315ba96a606ea8e26c02e13139e914fff9

  • Completed variants: 100
  • Failed variants: 0
  • Profitable variants: 0/100
  • Total trades: 7,000
  • Average P&L: -$11.10
  • Average ROI: -96.816%
  • Best ROI: -46.08%
  • Best P&L: -$11.52
  • Best win rate: 3.0%
  • Best run trades: 71
  • Weakest completed ROI: -188.40%
  • Weakest completed P&L: -$9.42

https://preview.redd.it/rtaef757q29h1.png?width=1080&format=png&auto=webp&s=784c9c82507d933af34ea36094ad3fbce48af7bd

As you can see, this did not just fail on one bad setting. The entire sweep was red. The other part that stood out to me was the win rate. The best run only won 3.0% of trades across 71 trades. That suggests the entry condition was catching a lot of bad spots, not just paying too much on execution.

These results obviously aren't proof that "ETH leads BTC" is guaranteed to work either; they just show that this contrarian version did not hold up in this backtest. Even so, I was quite surprised by how poorly this contrarian strategy performed.

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

reddit.com
u/ryanturbine — 2 months ago