Momentum v4 Robustness
The v3 dual-regime momentum strategy found a profitable cluster, but the permutation p-value of 0.292 means we cannot rule out luck across the 100-parameter sweep. This v4 sweep narrows the price band and varies polling frequency to test whether the cluster is stable in its neighborhood — prioritizing deflated Sharpe and low sensitivity to small parameter shifts over chasing a single high-PnL cell.
Historical research only. Not investment advice.
Top strategy variants
Bottom strategy variants
Kalshi BTC Momentum Strategy — Neighborhood Robustness Sweep
Strategy Family: KXBTC15M dual-regime momentum
Sweep: price floor × price ceiling (10×10 grid, 100 variants completed)
Date: Historical simulation only
1. Short disclaimer
This is a historical simulation research report. Past performance does not guarantee future results. Nothing in this report constitutes trading advice. All metrics are derived from backtests and are subject to overfitting, lookahead bias, and fill-assumption error. The strategies described may lose money in live trading.
2. Intro / thesis
The v3 dual-regime momentum strategy on KXBTC15M turned up a cluster of profitable cells, but the permutation p-value of 0.292 across the 100-parameter sweep meant we couldn’t rule out selection noise. Rather than celebrating a single high-PnL cell, this v4 sweep was designed as a focused neighborhood check: we fixed the ceiling at 0.55 (the v3 winner’s value) in most runs and varied the price floor in small steps, while also testing broader ceiling bands to measure how quickly performance degrades when we walk away from the original sweet spot.
The goal was simple: find configurations where PnL holds up across several adjacent parameter values (low neighborhood degradation) and produce a meaningfully positive deflated Sharpe — not just a hot cell that collapses one tick over.
The results are mixed. We found a band of profitable floor values near the v3 winner, but the deflated Sharpe remains modest, the winner sits on thin daily data, and fill assumptions at extremes are stretched. The top-performing cells are still consistent with a lucky draw from the sweep.
3. Variant and strategy explanation
Market traded: Kalshi’s KXBTC15M series — 15-minute binary options on Bitcoin above/below a strike. The strategy trades YES and NO contracts based on short-term BTC momentum signals, with hard risk boundaries on entry price.
What we varied (v4 sweep):
| Parameter | Values tested |
|---|---|
risk.price_floor | 0.05, 0.09, 0.14, 0.18, 0.23, 0.27, 0.32, 0.36, 0.41, 0.45 |
risk.price_ceiling | 0.55, 0.59, 0.64, 0.68, 0.73, 0.77, 0.82, 0.86, 0.91, 0.95 |
All 100 combinations ran to completion. The base strategy logic was held constant: the same momentum entry rules, same exit rules (2-minute near-expiry sell, 15% take-profit, -25% stop-loss), same position sizing (25 contracts per trade, max 50), and the same 10-second loop interval with Coinbase BTC data refreshed every 5 seconds.
What a “variant” is: Each row in the results is a complete backtest of the base strategy with one (floor, ceiling) pair. The parameter changes only affect which contract prices the strategy is allowed to trade — contracts priced below the floor or above the ceiling are simply ignored. This is a risk filter, not a signal change.
Why these axes: The v3 winner sat at floor 0.05 / ceiling 0.55. By keeping ceiling fixed and walking the floor upward, we can see whether the profitability is concentrated in extremely cheap contracts (where fill assumptions are least reliable) or whether it survives into more moderate price bands. The broader ceiling sweep tests how much the strategy depends on avoiding expensive contracts entirely.
4. Top results
Eight variants produced positive net PnL, all with price_ceiling = 0.55. The top five, ranked by total PnL:
| Rank | Floor | Total PnL | ROI % | Sharpe | Max DD | Trades | Win Rate |
|---|---|---|---|---|---|---|---|
| 1 | 0.05 | $33.01 | 66.0% | 0.72 | -$30.91 | 52 | 37.5% |
| 2 | 0.32 | $29.74 | 59.5% | 0.67 | -$19.60 | 14 | 71.4% |
| 3 | 0.36 | $29.74 | 59.5% | 0.67 | -$19.60 | 14 | 71.4% |
| 4 | 0.41 | $29.74 | 59.5% | 0.67 | -$19.60 | 14 | 71.4% |
| 5 | 0.45 | $28.13 | 56.3% | 0.97 | -$1.77 | 8 | 75.0% |
What the top cluster tells us:
The profitability band is real but narrow. Every positive variant requires the ceiling at 0.55 — raise it even to 0.59 and every cell turns negative. The floor, however, shows some flexibility: values from 0.32 through 0.45 produce nearly identical PnL (~$29–$28) with far fewer trades (8–14 trades vs. 52 for the 0.05-floor winner), much lower drawdown, and higher win rates.
The rank-1 cell (floor 0.05) has the highest raw PnL but also carries three specific warnings from the robustness framework:
- Thin daily data: Only 7 distinct PnL days. Daily Sharpe estimates on that few observations are unreliable.
- Extreme fills: 33% of fills are at prices below $0.10 or above $0.90 — exactly where Kalshi’s order book is thinnest and backtest fill assumptions are least trustworthy.
- Deflated Sharpe of 0.17: After adjusting for the number of trials in the sweep (100 cells), the top Sharpe of 0.72 deflates to 0.17. That means the winner is not distinguishable from the luckiest cell you’d expect to see in 100 random, skill-less trials.
The rank-5 cell (floor 0.45) is arguably the more interesting configuration: $28.13 PnL on just 8 trades, -$1.77 max drawdown, 75% win rate, and a raw Sharpe of 0.97. But with only 8 trades, confidence intervals on all those metrics are wide, and the deflated Sharpe concern applies to the whole sweep, not just the rank-1 cell.
Each successful variant is saved as a runnable Turbine strategy under its unique slug (e.g., momentum-v4-robustness-61d3ce259ddb for the rank-1 cell). None are live or recommended; they are research artifacts.
5. Bottom results
The worst performers cluster at price_ceiling = 0.95 — effectively no ceiling at all. Once the strategy is allowed to buy contracts up to $0.95, losses are deep and consistent across all floor values:
| Rank | Floor | Ceiling | Total PnL | ROI % | Sharpe | Max DD | Trades | Win Rate |
|---|---|---|---|---|---|---|---|---|
| 94 | 0.45 | 0.95 | -$57.26 | -114.5% | -0.46 | -$108.86 | 150 | 65.3% |
| 95 | 0.23 | 0.95 | -$61.36 | -122.7% | -0.66 | -$107.34 | 154 | 66.2% |
| 97 | 0.18 | 0.95 | -$67.15 | -134.3% | -0.75 | -$113.13 | 156 | 65.4% |
| 98 | 0.05 | 0.95 | -$68.32 | -136.6% | -0.75 | -$114.30 | 158 | 64.6% |
What the bottom cluster tells us:
The ceiling parameter dominates. Opening the strategy to expensive contracts (above ~$0.55) destroys performance regardless of floor. Even with high win rates (64–67%), the size of the losses on losing trades overwhelms the winners — a classic negative expectancy profile dressed in a respectable win rate. Trade frequency explodes (150+ trades vs. 8–52 in the winning cells), suggesting the strategy is firing on marginal signals when it doesn’t have the price filter to protect it.
The marginal means confirm this: average PnL across all floor values is +$24.01 at ceiling 0.55, drops to -$7.73 at ceiling 0.59, and craters to -$61.26 at ceiling 0.95. The strategy’s edge — if there is one — lives entirely in the sub-$0.55 contract space.
6. Conclusion
This v4 sweep achieved its objective: it mapped the neighborhood around the v3 winner and confirmed that the profitable region is real, narrow, and tied to a strict price ceiling of 0.55. Within that band, floor values of 0.32–0.45 produce similar PnL with better risk characteristics and fewer extreme-price fills than the raw PnL leader at floor 0.05.
However, three cautions override any temptation to call these results validated:
The deflated Sharpe is 0.17. After accounting for the 100 cells swept, the top result is statistically indistinguishable from selection noise. This is the single most important number in the report. The raw Sharpe of 0.97 on the floor-0.45 cell looks impressive, but with only 8 trades across 7 PnL days, it simply doesn’t survive multiplicity adjustment.
The extreme-price fill warning on the top cell (33% of fills at <$0.10 or >$0.90) means real-world execution would almost certainly be worse than the backtest shows. Kalshi’s liquidity at those price extremes is sparse; assuming fills at back
This report is generated from historical simulations. Backtests can be wrong or incomplete, and live trading can differ materially because of liquidity, fees, slippage, latency, market resolution, outages, and data quality. Do your own review before running any strategy.