Price Threshold Re-verify
On BTC 15-minute markets, buying YES when price dips to 0.50 and selling at 0.70 is a mean-reversion edge. Re-verify the previously winning threshold pair and search nearby entry/exit thresholds to confirm whether 0.50/0.70 is the robust optimum or a nearby band performs comparably.
Historical research only. Not investment advice.
Top strategy variants
Bottom strategy variants
Kalshi BTC 15-Minute Mean-Reversion Threshold Sweep: Research Report
Short Disclaimer
Historical simulation only. This report does not imply future profits. All results are in-sample and subject to overfitting.
Intro / Thesis
The original idea was simple: on Kalshi's BTC 15-minute markets (KXBTC15M), buy YES when price dips to 0.50 and sell at 0.70. That's a mean-reversion assumption — short-dated crypto binary prices overshoot on dips, then snap back.
We ran the full 10×10 threshold sweep to test two things. First, does the previously winning 0.50/0.70 pair still hold up? Second, is 0.50/0.70 the robust optimum, or is there a nearby band performing comparably?
The short answer: 0.50/0.70 is not the optimum. A nearby pair performed better on raw PnL. But the deeper finding is that none of this looks robust once you account for selection effects.
Variant and Strategy Explanation
Each variant is a threshold pair: a price floor for entry and a price ceiling for exit. The base DSL does mean-reversion with three rules:
- Buy dip: when YES price ≤ floor, buy 1 contract
- Sell target: when YES price ≥ ceiling, sell all
- Flatten close: when time to expiry ≤ 1 minute, sell all
The sweep held the strategy structure fixed and varied two parameters:
risk.price_floor: 0.05, 0.09, 0.14, 0.18, 0.23, 0.27, 0.32, 0.36, 0.41, 0.45risk.price_ceiling: 0.55, 0.59, 0.64, 0.68, 0.73, 0.77, 0.82, 0.86, 0.91, 0.95
That's 100 cells. All 100 completed. The loop ran on a 10-second interval, max position 20 contracts, price bounds 0.05 to 0.95.
Each successful variant is saved as a runnable Turbine strategy with its own strategy_id and slug.
Top Results
The top variant by net PnL isn't the original 0.50/0.70 thesis pair. It's floor 0.32 / ceiling 0.55:
| Rank | Label | Total PnL | ROI | Win Rate | Trades | Max DD | Sharpe |
|---|---|---|---|---|---|---|---|
| 1 | floor 0.32 / ceil 0.55 | $10.27 | 51.4% | 54.6% | 5,128 | -$217 | -0.50 |
| 2 | floor 0.27 / ceil 0.55 | $5.03 | 25.2% | 52.8% | 6,177 | -$272 | -0.57 |
| 3 | floor 0.36 / ceil 0.55 | $3.10 | 15.5% | 55.9% | 4,484 | -$193 | -0.50 |
| 4 | floor 0.32 / ceil 0.59 | $3.09 | 15.5% | 55.7% | 5,406 | -$237 | -0.55 |
The top handful cluster tightly around ceiling 0.55 — the lowest exit in the sweep — and floors in the 0.27 to 0.36 range. That is a long way from the original 0.50 entry. The market facts drove this: on BTC 15-minute binaries, buying YES at 0.50 wasn't where the edge lived in this backtest window. Lower entries paired with a tighter exit target produced the best raw outcomes.
But there's a catch. Several catches, actually.
Fees eat the winner. The gross PnL of the top variant is almost entirely consumed by fees — 102% of winner gross PnL goes to fee drag. This edge is, net of costs, essentially zero.
Sharpe is negative everywhere at the top. The best variant has a Sharpe of -0.50. Every top-8 variant has negative Sharpe. That's not a mean-reversion edge — that's a strategy with high variance and negative risk-adjusted returns, even where nominal PnL is positive.
Deflated Sharpe is near zero. The Monte Carlo permutation test produced an expected max Sharpe of 0.216 for a sweep of this size if the strategies were skill-less. The deflated Sharpe for the winner is 0.018. That is an order of magnitude below the luck ceiling. The top results are indistinguishable from the best outcome you'd get picking 100 random threshold pairs over the same data.
Neighborhood degradation is high. Move slightly away from the winner coordinates and performance drops off substantially — 64% degradation. The peak is narrow, not a plateau. That's a classic overfitting signature.
Bottom Results
The worst performers are concentrated at two edges: very low floors (0.05–0.09) and high ceilings (0.82–0.95):
| Rank | Label | Total PnL | ROI | Win Rate | Trades | Max DD |
|---|---|---|---|---|---|---|
| 100 | floor 0.05 / ceil 0.55 | -$155.07 | -775% | 54.4% | 14,031 | -$467 |
| 99 | floor 0.45 / ceil 0.95 | -$51.15 | -256% | 59.6% | 3,652 | -$161 |
| 98 | floor 0.45 / ceil 0.91 | -$49.61 | -248% | 59.7% | 3,628 | -$159 |
| 97 | floor 0.45 / ceil 0.82 | -$48.91 | -245% | 60.1% | 3,538 | -$169 |
A striking feature of the bottom performers: many have the highest win rates in the entire sweep. Floor 0.45 / ceiling 0.95 wins 59.6% of trades. Floor 0.45 / ceiling 0.77 wins 60.4%. They lose money because when they lose, they lose big — holding long as the 15-minute contract expires worthless wipes out many small wins. Backtest win rate is close to useless here.
The very low floor variants also trade enormous volume: 14,031 trades for floor 0.05. Chasing every dip from 0.05 accumulates fee drag and adverse selection on contracts that are cheap for a reason.
The marginal means confirm this. Average PnL across all ceilings for each floor is negative by the time you reach floor 0.41, and it's deeply negative for floors 0.05 and 0.09. The ceiling marginals show the best average outcomes at the lowest ceiling (0.55) and deterioration from there. The original 0.50/0.70 pair sits in a region where the sweep shows no edge.
Conclusion
The mean-reversion thesis — buy cheap BTC 15M YES contracts, sell on any bounce — did not survive a proper robustness check.
The previously winning 0.50/0.70 pair is not the optimum in this sweep. The best raw net PnL comes from floor 0.32 / ceiling 0.55. But calling that a better strategy would be misleading. The deflated Sharpe of the winner is 0.02, against an expected luck-max Sharpe of 0.22. The permutation test indicates the top result is fully consistent with selection noise. Fees consume the entire winner's gross PnL. All top variants have negative Sharpe.
This is exactly what a null result looks like. No threshold pair in this sweep is strong, validated, or promising. The parameter surface is negative on average, and the positive cells are narrow, cost-sensitive, and indistinguishable from random.
If there's an edge in KXBTC15M mean-reversion, this sweep did not find it.
Long Disclaimer
This report is historical simulation research produced for internal use by Turbine's research function. It is not investment advice, not a trading recommendation, and not a solicitation to buy or sell any security, contract, or financial instrument.
All performance figures are in-sample backtest results. They do not account for execution slippage, market impact, fill uncertainty, exchange outages, or other real-world frictions not modeled in the simulation environment. The fee drag warning in the robustness statistics indicates that simulated fees consume more than 100% of gross PnL for the top variant — actual net results may differ even from the negative figures shown here.
The permutation test and deflated Sharpe metric are statistical tools for detecting overfitting in parameter sweeps. They do not prove or disprove the existence of any underlying market inefficiency. A deflated Sharpe below the expected maximum for a sweep of equivalent size indicates that the top result is not distinguishable from the luckiest draw among skill-less strategies. This report explicitly states that the results are consistent with selection noise and overfitting.
Each variant listed is saved as a runnable Turbine strategy for research reproducibility. Saving a strategy does not imply endorsement, live deployment, or expected profitability.
Past simulated performance does not guarantee future results. Markets change, edges decay, and strategy parameters that worked historically often stop working out-of-sample. Any use of the strategies described here is at the user's own risk.
This report is generated from historical simulations. Backtests can be wrong or incomplete, and live trading can differ materially because of liquidity, fees, slippage, latency, market resolution, outages, and data quality. Do your own review before running any strategy.