File size: 5,792 Bytes
3d4420a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84f7958
3d4420a
84f7958
3d4420a
84f7958
3d4420a
84f7958
 
 
 
 
 
3d4420a
84f7958
3d4420a
84f7958
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3d4420a
84f7958
 
 
3d4420a
 
 
ed2b394
84f7958
3d4420a
ed2b394
84f7958
ed2b394
 
 
 
84f7958
ed2b394
 
84f7958
ed2b394
 
3d4420a
ed2b394
3d4420a
 
84f7958
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3d4420a
84f7958
3d4420a
84f7958
3d4420a
84f7958
3d4420a
84f7958
3d4420a
84f7958
 
3d4420a
84f7958
 
 
ed2b394
84f7958
 
 
ed2b394
84f7958
ed2b394
84f7958
 
 
 
ed2b394
84f7958
ed2b394
84f7958
 
 
 
 
 
 
 
 
 
 
 
ed2b394
84f7958
ed2b394
f288c67
ed2b394
f288c67
3d4420a
84f7958
 
 
3d4420a
84f7958
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3d4420a
 
 
84f7958
 
 
 
 
 
3d4420a
84f7958
3d4420a
84f7958
 
3d4420a
84f7958
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
---
license: apache-2.0
base_model: amazon/chronos-2
tags:
  - time-series
  - forecasting
  - chronos
  - chronos-2
  - quantization
  - torchao
  - int8
pipeline_tag: time-series-forecasting
library_name: chronos-forecasting
---

# Smaller Chronos-2 (INT8)

[Chronos-2](https://huggingface.co/amazon/chronos-2) is Amazon's 120M time-series model: you give it history, it forecasts the next hours/days, no extra training.

This repo is **the same model with smaller files**. Weights are stored as 8-bit integers (INT8). It is **not** a fine-tune and **not** a new architecture.

| | Amazon original | this repo |
|--|--|--|
| download | 478 MB | **131 MB** |
| GPU memory while running (3090, one series) | ~0.56 GB | **~0.25 GB** |
| Python API | `Chronos2Pipeline.from_pretrained` | `load.py` (this repo) |
| license | Apache-2.0 | Apache-2.0 |

**Pick this** if you want Chronos-2 in Python, but a lighter download and less VRAM.

**Pick [amazon/chronos-2](https://huggingface.co/amazon/chronos-2)** if you want the one-liner load and do not care about 350 MB.

**Pick an ONNX / TensorRT Chronos-2** if you only care about production latency. Those are a different runtime, not this Python path.

---

## Install

```bash
pip install "chronos-forecasting>=2.0" torchao safetensors huggingface_hub pandas pyarrow
```

GPU: a recent PyTorch with CUDA. CPU works; it will be slower.

---

## Load (required)

Hugging Face's usual `from_pretrained` **cannot** read these INT8 files. Use `load.py` from this repo:

```python
from pathlib import Path
import importlib.util
from huggingface_hub import snapshot_download

repo = snapshot_download("oxfrug/chronos-2-int8-torchao")

spec = importlib.util.spec_from_file_location("c2int8", Path(repo) / "load.py")
mod = importlib.util.module_from_spec(spec)
spec.loader.exec_module(mod)

pipe = mod.load(repo, device="cuda")  # "cpu" if you have no GPU
```

If you already cloned the folder:

```python
from load import load
pipe = load("/path/to/chronos-2-int8-torchao", device="cuda")
```

You get a normal `Chronos2Pipeline`. After this, Amazon's docs apply.

---

## Forecast

One numpy series:

```python
import numpy as np

history = np.array([12.0, 12.4, 11.9, ...], dtype=np.float32)  # oldest → newest
quantiles, mean = pipe.predict_quantiles(
    inputs=[history],
    prediction_length=24,
    quantile_levels=[0.1, 0.5, 0.9],
)
# median forecast: quantiles[0][:, 1]  (shape: 1 × horizon × 3)
```

Many series / covariates — same as upstream, with `pipe.predict_df(...)`.

---

## Make it faster

INT8 here is **smaller, not faster**, on an RTX 3090. The model is small; most of the time is starting GPU kernels, not moving weights.

Two switches help **both** this INT8 and the original FP32 model:

1. **TF32** — tell PyTorch to use the GPU's fast float path (it is off by default).
2. **`torch.compile`** — fuse those kernels. First call is slow (compile); later calls drop a lot.

```python
import torch
from fast_infer import speedup   # file in this repo

torch.set_float32_matmul_precision("high")
pipe = speedup(pipe)  # TF32 + torch.compile
```

Or from a terminal, after downloading this folder:

```bash
python fast_infer.py                         # this INT8
python fast_infer.py --fp32 amazon/chronos-2 # original model, same knobs
```

Timed on this machine (RTX 3090, 512 past points, forecast 24, **one** series):

| | time per call | notes |
|--|--|--|
| original FP32 | ~7 ms | no extra knobs |
| this INT8 | ~9 ms | smaller, slightly slower |
| FP32 + compile + TF32 | **~3 ms** | best speed here |
| INT8 + compile + TF32 | ~3.6 ms | still a bit behind compiled FP32 |

Forecasting **many series in one call** (`predict_df` with several ids) is the other real win. Amazon's “hundreds of series per second” numbers are batched, not one sine wave.

Half-precision (FP16 / BF16) did **not** help this 120M model on a 3090.

---

## Did INT8 change the forecasts?

This pack only compresses weights.

What i did: take a few public series, hide the last 12–168 points, forecast them with the original model and with this INT8, compare.

- **German electricity (hourly)** — INT8 stays close. 24h median error vs original +1.8%; 168h actually −3%. Correlation of the two forecasts ≈ 0.99.
- **M4 hourly / daily / monthly** — same story on shape (high correlation), median error a bit worse (about +9% to +15% MASE on those short holds).
- **M4 weekly (13 steps)** — the two forecasts **diverged** (correlation near zero). Do not read that row as “INT8 is better.” Short weekly holds are noisy.

P10–P90 intervals: on 12–24 step holds, *both* models often cover ~50–60% of points instead of 80%. Compare INT8 to the original on *your* series if you use the bands, not to the textbook 80%.

Raw dumps: `eval/series.json`, `eval/coverage.json`, `eval/speed.json`.

---

## How the file was made

- Start from `amazon/chronos-2` (full float weights).
- torchao **weight-only INT8**: each big linear layer stores integers + a scale. Activations stay float.
- The small **quantile head** (the part that turns hidden states into P10/P50/P90) is left in float, on purpose.
- No calibration data, no extra training.

That is why `from_pretrained` fails: Hugging Face does not know this packing. `load.py` rebuilds the layers and fills them from `model.safetensors`.

---

## Limits

- Smaller files, less GPU RAM — **not** a speed-up by itself on a 3090.
- Not a domain fine-tune (no Nordic energy LoRA, etc.).
- Not GIFT-Eval / fev-bench.
- If you need intervals, check coverage on your data.

---

## Cite the base model

Ansari et al., *Chronos-2: From Univariate to Universal Forecasting*, 2025.  
https://arxiv.org/abs/2510.15821

Quant pack by [oxfrug](https://huggingface.co/oxfrug).
"""