Pedro Mariano / portfolio

Does an AI model predict Bitcoin?

Kronos is an open foundation model trained only on price candles from more than 45 exchanges. I connected it to StratForge, ran it on a GPU and tested it on a year of Bitcoin prices it had never seen. It showed no reliable edge, and its uncertainty band was far too narrow.

2,008forecasts, at five combinations of candle size and horizon
13 monthsof BTCUSDT after the model's paper, so unseen data
~3 sfor 20 sampled paths on an RTX 5060, against ~40 s for 5 on a CPU
No edge"no change" beat it on the size of the move at every horizon

Method

Set the test up so it can fail.

  • Unseen data only: Bitcoin from 1 September 2025, after the Kronos paper (August 2025).
  • Non-overlapping forecasts, 400 candles of context each, 20 sampled future paths per forecast.
  • Simple baselines: "always the more common direction", "the next move repeats the last one" (momentum), and "no change".
  • Calibration: the middle 80% of paths should hold the real price 80% of the time.
  • Many comparisons, one correction: with about ten tests, one result below p = 0.05 is expected by chance.

The library averages its sampled paths internally, so I wrote a small wrapper that keeps every path. Without it there is no band to test.

Results

Right about the direction half the time.

CandlesAheadForecastsKronos directionBase rateMomentumError vs "no change"
1 h1 h40248.3%51.5%47.3%0.33% vs 0.29%
1 h6 h40255.5%50.5%48.3%0.76% vs 0.69%
1 h24 h40248.8%51.0%51.2%2.67% vs 1.58%
4 h4 h40155.1%50.1%46.9%0.71% vs 0.57%
4 h24 h40146.4%50.9%51.1%1.87% vs 1.58%

The best row (4 h candles, next candle, p = 0.026) does not survive a correction for about ten comparisons.

How often the 80% band held the real price

1 h → 1 h
57%
1 h → 6 h
53%
1 h → 24 h
28%
4 h → 4 h
61%
4 h → 24 h
46%

The white line is the 80% target. The paths moved about 0.35% an hour against 0.57% in the real data, so the band gets worse the further ahead it looks.

And as a trading signal

Going long or short on the sign of the median path, with 0.1% costs per round trip, lost 55.1% against 23.0% for simply holding (1 h candles, 24 h ahead) and 63.3% against 21.0% (4 h candles, 24 h ahead).

What I did with it

Ship the model, and the warning with it.

Kronos stays in StratForge as one more input: a dashed median line and a band on the chart, and a tool the analyst can call. The tooltip, the analyst's instructions and the README all say what this test found, so nobody reads the line as a signal.

The negative result is the useful one: it keeps anyone using the app from trusting a confident-looking forecast.

  • Python
  • PyTorch
  • CUDA
  • Hugging Face
  • pandas
  • Statistics
  • Binance API