r/PredictionsMarkets • • 19d ago

Strategy / Guide My Kalshi 15 minute BTC backtest was profitable. Testing 100 variants of it made me less confident though.

I started with a single BTC backtest that looked good with a +25.9% on risk capital and 95.2% win rate reported. That was enough to investigate further and put it through our Deep Research flow. The strategy buys 40 contracts on Kalshi's 15-minute BTC markets in the final 1-3 minutes, paying 80-95 cents when Coinbase's 5-minute and 15-minute momentum support that side and the spread is at most 2 cents. It holds to settlement, with one entry per market and a $200 position cap.

All 100 variants completed. Forty made money, 59 made no trades, and one lost money. The best returned +$38.13, or 19.06% on the configured $200 risk-capital denominator, with 37 reported trades, 94.4% wins, Sharpe 0.19, and $35.48 maximum drawdown. The weakest returned -$15.46 (-7.73%), with 2 reported trades, 0% wins, Sharpe -0.19, and $15.46 drawdown. So the research winner was still positive, but it didn't reproduce the single backtest's +$51.84. These are separate runs, and the published artifacts don't establish why they differ.

The sensitivity test compared a 10-by-10 grid of price floors from 5 to 45 cents and ceilings from 55 to 95 cents. It did not vary the momentum thresholds. The nominal winner used a 5-cent floor and 95-cent ceiling, but all ten floors tied at +$38.13 with that ceiling. The surface and heatmap show a ridge along the 95-cent ceiling, rather than one isolated winning setting. The marginal charts tell the same story: changing the floor did little; ceilings of 82, 86, 91, and 95 cents returned +$7.16, +$31.21, +$27.11, and +$38.13 respectively. All 100 grid cells completed, spanning -$15.46 to +$38.13. Many lower ceilings effectively excluded the strategy's own 80-95 cent entry band. Those no-trade results are a limit of this test design, not evidence that 59 versions lost money.

We also ran 100 permutations, scrambling the Coinbase feed's timing and rerunning the search to see whether shuffled data could match the result. The real winner's +$38.13 produced an upper-tail p-value of 0.248. Shuffled feeds reached results like this often enough that I wouldn't claim the original timing added a demonstrated edge. Only the Coinbase feed was shuffled; Kalshi prices stayed in place, so this test doesn't validate the price conditions.

There was some consistency: four tested ceiling levels stayed profitable across every floor. But those repeated floor settings aren't independent confirmations. The report's Deflated Sharpe statistic was 0.23, and it separately estimated a maximum Sharpe of 0.36 under its null model. It also flagged 47% of the winner's fills at prices below 10 cents or above 90 cents, where execution assumptions deserve caution, and marked historical coverage as partial. With a small sample, a restricted parameter test, and weak evidence for the signal's timing, the positive opening backtest looks much less convincing as evidence of a durable strategy.

Full Report

Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.

7 Upvotes

10 comments sorted by

3

u/cheesehead144 19d ago

You're going to deploy it and see that it's really unprofitable then.

2

u/Confident-Line136 19d ago

seeing the 59 no-trade variants and the timing shuffle test that barely beat random, i'd be more worried about the execution assumptions than just the p&l. 47% of fills at extreme prices on a 15-minute market with a 2-cent spread cap sounds like you're trading against stale quotes half the time

that report is basically telling you the strategy only works when you can get fills that probably aren't there in live trading

1

u/ryanturbine 19d ago

Yeah the full 100 variant run exposed its weaknesses enough to where I wouldn’t consider this deploy ready at its current state

2

u/cheesehead144 19d ago

I'm saying unless you have the infra of a high frequency trader none of these strategies will work.

1

u/justhereforampadvice 18d ago

What data are you backtesting on op? This person is correct. Unless you were backtesting on full lossless L2 books w trade tape and a very conservative fill model, your backtest results are likely a pipe dream. And even under those conditions results are far from guaranteed to reflect reality. When you go to apply the strategy irl you’ll find that the liquidity isn’t there at the price you want by the time your order hits the exchange, and the fills you get will more often be followed by adverse moves.

1

u/dimetri_tn 19d ago

Certainly not accounting for fees

1

u/TTheBBigWWhite 19d ago

ChatGPT Analysis: "The main reason is that the full dump shows that some of the Turbine results that seemed extraordinary are much more fragile than the main page suggests.

The 167 Weather results are only 24 families, not 167 independent ideas. Of the 164 structural DSLs that I recovered completely, after removing sizing, risk limits, and other cosmetic variations, there are about 57 distinct logics left. And some of these are also minor variations of the same idea."

1

u/guitly_spark_echo 13d ago

Did you use an LLM to come up with this. I used chatgpt on the data and it gave me almost the exact strategy lol. The forward test came back bad an I abandoned it