We Tested 5,732 Kalshi Bitcoin 15-Minute Bot Variants
Our updated Kalshi BTC 15-minute strategy research covers 5,732 backtests. Compare momentum, panic fade, MACD, entry bands, fees, and drawdowns.
Across 5,732 completed Kalshi Bitcoin 15-minute backtests, 2,154 bot variants finished with positive simulated P&L. Another 2,496 lost money. The remaining 1,082 finished at zero. The newer research changes the picture from our April study: panic fade lost in a later test, momentum results depended heavily on the entry and exit rules, and the latest batches show why win rate alone is a poor way to pick a BTC trading bot.
Research snapshot: September 29, 2026. This update draws on all 63 completed public research reports whose base strategy trades Kalshi's KXBTC15M Bitcoin 15-minute series. Those reports completed between June 9 and September 24. We compiled existing results; we did not run a new 5,732-variant batch for this article. The actual market-data windows differ and are discussed below.
Historical simulation research, not investment advice. These are backtests, not live trading results. Prediction-market trading carries substantial risk.
Our April study of 4,904 Kalshi Bitcoin strategies found only 102 profitable variants. Panic fade accounted for 93 of the top 100 results. That was one parameter sweep, one window, and one execution model.
The public Turbine research library now gives us a broader set of questions to examine: does that winner survive another test? Which momentum rules actually contribute? Does a tighter price ceiling help? What changes when the bot trades five contracts instead of ten?
There are useful findings here. There are also repeated tests, inactive strategies, partial data windows, and unresolved discrepancies. The count measures completed backtests, not 5,732 independent discoveries.
Results at a glance
| Measure | Updated BTC 15-minute research |
|---|---|
| Completed public reports | 63 |
| Completed bot-variant backtests | 5,732 |
| Positive simulated P&L | 2,154 |
| Negative simulated P&L | 2,496 |
| Zero simulated P&L | 1,082 |
| Variants with no trades | 972 |
| 100-variant reports | 57 |
| Smaller comparison batches | 6 reports, 32 variants |
The 972 no-trade variants are part of the total, not an additional category. A completed backtest can successfully evaluate its rules without ever entering a position. Zero P&L and zero trades are also different: some active variants finished at zero.
About 37.6% of the completed variants recorded positive P&L. That is a description of this research collection. It is not the probability that a new bot will make money. The reports overlap in market data and strategy ideas, and some configurations produce identical trades.
We have not added the old 4,904 runs to this headline. The updated count comes from the public research reports listed at the end of this article.
The April panic-fade winner lost in a later test
In April, 93 of 96 panic-fade variants made money. The strongest returned +18.32% on that study's fixed $10,000 notional.
The later BTC Panic Fade report tested sharp YES-price declines, a recovery exit, and different fade sizes. All 100 completed variants lost money. Recorded P&L ranged from -$27.23 to -$5.39.
That does not prove the market stopped reversing. The later test used a different configuration and data window; it is not a controlled rerun of every April strategy. It does establish something narrower and useful: the April leaderboard was not enough to justify treating panic fade as a dependable Kalshi 15-minute Bitcoin strategy.
The same caution applies to mean reversion. April's 432 variants all lost. Later early-window reversal studies were mixed: one report had 88 positive variants out of 100, while another had 25 out of 100. A separate price-threshold re-verification produced only four positive variants out of 100, with P&L ranging from -$155.07 to +$10.27.
“Mean reversion” is too broad a label to settle the question. Entry timing, price bands, exits, and execution assumptions determine which strategy was actually tested.
Momentum worked in some families and failed in others
Several research reports test Coinbase BTC-USD signals against Kalshi's Bitcoin contract prices. Those signals include short-term returns, velocity, moving averages, and VWAP: the volume-weighted average spot price.
The results do not support buying every contract that agrees with a spot-price trend.
| Research family | Positive / completed variants | Recorded simulated P&L range |
|---|---|---|
| Momentum continuation | 100 / 100 | +$1,131.82 to +$3,540.82 |
| Selective EMA momentum | 99 / 100 | -$4.94 to +$122.86 |
| BTC 15-minute spot momentum | 99 / 100 | -$1.70 to +$24.40 |
| Five-minute momentum with decay exits | 26 / 100 | -$29.60 to +$18.39 |
| Five-minute and one-hour trend confirmation | 0 / 100 | -$27.51 to -$24.11 |
These are within-report results, not a ranking of equally funded accounts. Position sizes and risk settings differ, so comparing the dollar ranges across rows does not identify the best bot.
The continuation study entered near expiry when spot price, SMA, VWAP, five-minute change, and one-minute velocity aligned. Its requested 30-day test had recorded coverage from August 8 to September 7, marked partial. The positive results spanned the entire tested price-band grid. That makes it more interesting than a single winning cell, but it remains one selected family on one historical window.
By contrast, the five-minute/one-hour trend-confirmation family lost throughout its grid. Its artifact warns that fees consumed 146% of the winner's gross P&L and that changing two entry-band parameters produced identical trades. Four later validation reports of that thesis also had no positive variants.
A parameter grid that never changes the trades is not evidence that a strategy is robust. It means those parameters were not exercised in that window.
The September update: price bands, MACD, and smaller entries
The newest reports are smaller experiments with more specific questions. These are the clearest additions to the April article.
A lower price ceiling beat a higher win rate
The September 16 Coinbase VWAP Momentum comparison changed only the entry price ceiling across four configurations.
| Entry ceiling | Simulated P&L | Trades | Win rate | Max drawdown |
|---|---|---|---|---|
| $0.65 | +$36.28 | 3,448 | 50.4% | -$28.99 |
| $0.70 | +$21.58 | 3,952 | 51.1% | -$32.41 |
| $0.75, saved baseline | +$23.79 | 4,407 | 52.3% | -$30.19 |
| $0.85 | +$17.18 | 5,200 | 55.5% | -$36.85 |
The highest ceiling produced the most trades and the highest win rate. It also produced the lowest P&L and the largest drawdown. The $0.65 ceiling finished ahead on both P&L and drawdown despite winning barely half its trades.
This does not establish $0.65 as a universal entry limit. It shows that buying a more likely winner at a higher price can leave less room to earn enough on each win. The batch contained four comparisons and no permutation test.
A September 11 VWAP comparison also found a tighter band ahead of its baseline: +$50.85 versus +$44.12, with fewer trades and a smaller drawdown. Tightening its stop produced identical recorded results to the baseline. Changing a setting is not the same thing as demonstrating that it affected behavior.
Stricter MACD confirmation improved win rate but reduced dollars earned
The September 22 MACD threshold report crossed five one-minute histogram thresholds with two five-minute thresholds. All ten variants recorded positive P&L.
| MACD histogram thresholds, 1m / 5m | Simulated P&L | Trades | Win rate | Max drawdown |
|---|---|---|---|---|
| 0 / 0 | +$54.61 | 1,697 | 52.8% | -$24.33 |
| 2.5 / 0 | +$75.38 | 1,138 | 57.8% | -$11.51 |
| 5 / 0 | +$69.07 | 774 | 60.7% | -$8.52 |
| 20 / 10 | +$18.08 | 78 | 76.9% | -$3.08 |
The strictest configuration won 76.9% of its trades and had the smallest drawdown of the ten. It also earned the least. The highest P&L came from a modest one-minute threshold with no additional positive five-minute threshold.
That is the tradeoff a BTC 15-minute bot needs to expose: frequency, net P&L, and drawdown alongside accuracy. A higher win rate does not automatically produce a better outcome.
We use recorded P&L here because the report's percentage returns do not match a simple return on its stated $1,000 starting portfolio. For example, +$75.38 is listed as 150.76% ROI. Without reconciling that denominator, the percentage should not be presented as an account return. This batch also had no sweep artifact or permutation test.
Wider bands mattered more than smaller entries in the latest batch
The September 24 momentum-alignment report crossed two entry sizes with two price bands, keeping the stated signals and exits fixed.
| Entry size / price band | Simulated P&L | Trades | Win rate | Max drawdown |
|---|---|---|---|---|
| 10 contracts / $0.45–$0.55 | +$2,781.79 | 2,703 | 59.9% | -$160.69 |
| 5 contracts / $0.45–$0.55 | +$2,252.87 | 3,341 | 60.7% | -$91.40 |
| 10 contracts / $0.35–$0.65 | +$4,516.86 | 6,218 | 55.0% | -$312.99 |
| 5 contracts / $0.35–$0.65 | +$4,298.37 | 7,666 | 59.7% | -$148.70 |
Both wider-band variants earned more than both narrow-band variants. The five-contract wide-band configuration recorded slightly less P&L than the ten-contract version with less than half its drawdown.
But the report explicitly leaves two issues unresolved. Changing size also changed trade counts substantially, without a confirmed explanation. The base risk ceiling and wider rule-level band differ, so the effective entry boundary needs verification. These findings are worth investigating; they are not a clean demonstration that reducing size lowers fee drag.
This wider-band result and the tighter-band VWAP result are not contradictory universal rules. They come from different strategies. A price band determines which opportunities a particular signal is allowed to trade.
Exits changed the entry-signal leaderboard
Two September 11 reports isolated the entry families in a saved BTC bot. One used take-profit and stop rules. The other flattened near settlement.
With profit/stop exits, early entry led the six tested families at +$1,166.40. The late-momentum candidate lost $39.52 despite an 85.0% win rate. That candidate used smaller entries and a different spread gate, so its position in the table is not a comparison at equal size.
With near-settlement exits, acceleration momentum led the five tested families at +$1,894.39. The VWAP family earned +$1,105.50 but recorded a -$1,751.92 drawdown. Positive final P&L hid a much rougher path.
These are separate batches, not a perfectly matched experiment proving one exit is superior. They show why a bot's entry signal cannot be evaluated separately from its exit and risk rules.
What the full research collection does not establish
One shared 30-day market window. Many reports requested 30 days but recorded partial or mixed coverage. Some later publications replay earlier data. For example, the July early-reversal report recorded data ending June 11; the August tight-band report recorded August 3–9. Publication date does not tell you when the simulated trades happened.
One consistent execution model. The artifacts include official API data, archived order books, and Turbine-derived execution events. Maker studies use different assumptions from directional taker strategies. We do not sum dollar profits or average ROI across those configurations.
Thousands of independent strategies. The count includes related parameter cells, repeated validation runs, and inactive or inert configurations. Distinct saved strategy IDs do not establish independent trade histories.
Live fill quality or a deployable edge. Several reports use extreme-price entries or small trade samples. The BTC maker sweep had 26 positive variants out of 100, but its artifact warns that the best result is not distinguishable from a lucky selection. Its permutation test failed, so a default numeric field cannot be treated as a valid p-value.
A holdout test for every winner. Some reports scramble external-feed timing while keeping market prices anchored. That asks whether feed timing matters under the recorded setup. It does not independently validate price filters, fill assumptions, or performance on an untouched future window.
The late-conviction study illustrates the distinction: 40 variants had positive P&L, but its feed-timing permutation result was p = 0.248 and its artifact flagged extreme-price fills. A green leaderboard alone did not settle the research question.
How we counted the Kalshi BTC 15-minute backtests
We read the paginated public research directory and the underlying public report data on September 29. We selected complete reports whose base strategy names KXBTC15M, then counted completed variants. The result is 57 reports with 100 variants each, plus six smaller batches totaling 32: 5,700 + 32 = 5,732.
Four additional reports relate directly to BTC but are outside this count. Hourly BTC early-high entries placed no trades. Daily BTC market making had 61 positive variants out of 100 but weak selection-adjusted evidence. Two studies used BTC velocity to trade ETH: one recorded positive outcomes on small samples, while the other lost across all 100 variants. Those findings do not establish performance on Bitcoin 15-minute contracts.
What changed since the April Bitcoin strategy study
The update is more specific than a new winner's name. Panic fade lost in a later configuration. Momentum had positive and negative families. September's smaller experiments made entry price, signal selectivity, and sizing easier to inspect. High win rates repeatedly failed to identify the highest P&L or the safest path.
The next useful test is a locked configuration on a separate window, with an explicit capital denominator, verified trade counts, fees, data coverage, and fill assumptions. That is how a historical result earns more weight than its place on a leaderboard.
To inspect your own rules, open Turbine Studio, define the market, signal, entry band, exit, and position limit, and review the trade log alongside P&L and drawdown. Our prediction-market backtesting guide explains the workflow; the overfitting guide explains why selecting the best historical result is only the beginning.
Full BTC 15-minute research inventory
Every report counted in the headline appears below. Dates are report completion dates, not the end of market-data coverage. Positive means recorded P&L greater than zero; it does not mean validated or recommended.
Disclaimer
This article is for informational and educational purposes only. It is not investment advice, a recommendation to trade, or a promise of future performance. Turbine publishes the research and provides the backtesting software discussed here.
All reported outcomes are historical simulations. They do not represent actual trading. Data gaps, execution assumptions, fees, slippage, latency, partial fills, selection bias, and changes in market conditions can materially change results. A simulated stop or position limit does not guarantee the same outcome in live trading.
Past performance does not predict future results. Trading prediction-market contracts involves substantial risk, including loss of the entire amount invested. Strategies and research links are provided for inspection and verification, not as endorsements.
