Does an AI model predict Bitcoin?
Kronos is an open foundation model trained only on price candles from more than 45 exchanges. I connected it to StratForge, ran it on a GPU and tested it on a year of Bitcoin prices it had never seen. It showed no reliable edge, and its uncertainty band was far too narrow.
Method
Set the test up so it can fail.
- Unseen data only: Bitcoin from 1 September 2025, after the Kronos paper (August 2025).
- Non-overlapping forecasts, 400 candles of context each, 20 sampled future paths per forecast.
- Simple baselines: "always the more common direction", "the next move repeats the last one" (momentum), and "no change".
- Calibration: the middle 80% of paths should hold the real price 80% of the time.
- Many comparisons, one correction: with about ten tests, one result below p = 0.05 is expected by chance.
The library averages its sampled paths internally, so I wrote a small wrapper that keeps every path. Without it there is no band to test.
Results
Right about the direction half the time.
| Candles | Ahead | Forecasts | Kronos direction | Base rate | Momentum | Error vs "no change" |
|---|---|---|---|---|---|---|
| 1 h | 1 h | 402 | 48.3% | 51.5% | 47.3% | 0.33% vs 0.29% |
| 1 h | 6 h | 402 | 55.5% | 50.5% | 48.3% | 0.76% vs 0.69% |
| 1 h | 24 h | 402 | 48.8% | 51.0% | 51.2% | 2.67% vs 1.58% |
| 4 h | 4 h | 401 | 55.1% | 50.1% | 46.9% | 0.71% vs 0.57% |
| 4 h | 24 h | 401 | 46.4% | 50.9% | 51.1% | 1.87% vs 1.58% |
The best row (4 h candles, next candle, p = 0.026) does not survive a correction for about ten comparisons.
How often the 80% band held the real price
The white line is the 80% target. The paths moved about 0.35% an hour against 0.57% in the real data, so the band gets worse the further ahead it looks.
And as a trading signal
Going long or short on the sign of the median path, with 0.1% costs per round trip, lost 55.1% against 23.0% for simply holding (1 h candles, 24 h ahead) and 63.3% against 21.0% (4 h candles, 24 h ahead).
What I did with it
Ship the model, and the warning with it.
Kronos stays in StratForge as one more input: a dashed median line and a band on the chart, and a tool the analyst can call. The tooltip, the analyst's instructions and the README all say what this test found, so nobody reads the line as a signal.
The negative result is the useful one: it keeps anyone using the app from trusting a confident-looking forecast.
- Python
- PyTorch
- CUDA
- Hugging Face
- pandas
- Statistics
- Binance API