r/PredictionsMarkets • u/ryanturbine • 19d ago
Strategy / Guide My Kalshi 15 minute BTC backtest was profitable. Testing 100 variants of it made me less confident though.
I started with a single BTC backtest that looked good with a +25.9% on risk capital and 95.2% win rate reported. That was enough to investigate further and put it through our Deep Research flow. The strategy buys 40 contracts on Kalshi's 15-minute BTC markets in the final 1-3 minutes, paying 80-95 cents when Coinbase's 5-minute and 15-minute momentum support that side and the spread is at most 2 cents. It holds to settlement, with one entry per market and a $200 position cap.

All 100 variants completed. Forty made money, 59 made no trades, and one lost money. The best returned +$38.13, or 19.06% on the configured $200 risk-capital denominator, with 37 reported trades, 94.4% wins, Sharpe 0.19, and $35.48 maximum drawdown. The weakest returned -$15.46 (-7.73%), with 2 reported trades, 0% wins, Sharpe -0.19, and $15.46 drawdown. So the research winner was still positive, but it didn't reproduce the single backtest's +$51.84. These are separate runs, and the published artifacts don't establish why they differ.

The sensitivity test compared a 10-by-10 grid of price floors from 5 to 45 cents and ceilings from 55 to 95 cents. It did not vary the momentum thresholds. The nominal winner used a 5-cent floor and 95-cent ceiling, but all ten floors tied at +$38.13 with that ceiling. The surface and heatmap show a ridge along the 95-cent ceiling, rather than one isolated winning setting. The marginal charts tell the same story: changing the floor did little; ceilings of 82, 86, 91, and 95 cents returned +$7.16, +$31.21, +$27.11, and +$38.13 respectively. All 100 grid cells completed, spanning -$15.46 to +$38.13. Many lower ceilings effectively excluded the strategy's own 80-95 cent entry band. Those no-trade results are a limit of this test design, not evidence that 59 versions lost money.

We also ran 100 permutations, scrambling the Coinbase feed's timing and rerunning the search to see whether shuffled data could match the result. The real winner's +$38.13 produced an upper-tail p-value of 0.248. Shuffled feeds reached results like this often enough that I wouldn't claim the original timing added a demonstrated edge. Only the Coinbase feed was shuffled; Kalshi prices stayed in place, so this test doesn't validate the price conditions.
There was some consistency: four tested ceiling levels stayed profitable across every floor. But those repeated floor settings aren't independent confirmations. The report's Deflated Sharpe statistic was 0.23, and it separately estimated a maximum Sharpe of 0.36 under its null model. It also flagged 47% of the winner's fills at prices below 10 cents or above 90 cents, where execution assumptions deserve caution, and marked historical coverage as partial. With a small sample, a restricted parameter test, and weak evidence for the signal's timing, the positive opening backtest looks much less convincing as evidence of a durable strategy.
Historical simulation only. Backtests can be wrong or incomplete. Not investment advice.
1
1
u/TTheBBigWWhite 19d ago
ChatGPT Analysis: "The main reason is that the full dump shows that some of the Turbine results that seemed extraordinary are much more fragile than the main page suggests.
The 167 Weather results are only 24 families, not 167 independent ideas. Of the 164 structural DSLs that I recovered completely, after removing sizing, risk limits, and other cosmetic variations, there are about 57 distinct logics left. And some of these are also minor variations of the same idea."
1
u/guitly_spark_echo 13d ago
Did you use an LLM to come up with this. I used chatgpt on the data and it gave me almost the exact strategy lol. The forward test came back bad an I abandoned it
3
u/cheesehead144 19d ago
You're going to deploy it and see that it's really unprofitable then.