Shadowell commited on
Commit
2e8cd36
·
verified ·
1 Parent(s): f6630b7

Upload Kairos checkpoint

Browse files
Files changed (3) hide show
  1. README.md +52 -0
  2. config.json +18 -0
  3. model.safetensors +3 -0
README.md ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags: [time-series, finance, kronos, kairos, crypto]
4
+ library_name: pytorch
5
+ ---
6
+
7
+ # Shadowell/Kairos-base-crypto
8
+
9
+ Fine-tuned **Kronos-Tokenizer-base** (crypto (BTC/USDT + ETH/USDT 1-min)), produced by
10
+ **[Kairos](https://github.com/Shadowell/Kairos)**.
11
+
12
+ The tokenizer encodes a rolling 6-dim OHLCV+amount window into two streams of
13
+ discrete tokens (BSQ — Binary Spherical Quantization) that Kronos predictors
14
+ consume downstream. Only the OHLCV inputs are used for training; the 32-dim
15
+ exogenous channel required by `KronosWithExogenous` is not needed for the
16
+ tokenizer itself.
17
+
18
+ ## Results on test set (304,770 windows, 2 symbols)
19
+
20
+ | metric | baseline | finetuned | Δ |
21
+ |---|---|---|---|
22
+ | recon_mse_full | 0.005504 | 0.004141 | -24.8% |
23
+ | recon_mae_full | 0.055000 | 0.047185 | -14.2% |
24
+ | bsq_loss_mean | -0.070294 | -0.070731 | +0.6% |
25
+ | s1 util | 0.6904 | 0.7314 | +5.9% |
26
+ | s2 util | 0.4326 | 0.4277 | -1.1% |
27
+ | s1 entropy (bits) | 5.9374 | 6.1164 | +3.0% |
28
+ | s2 entropy (bits) | 4.6063 | 4.7127 | +2.3% |
29
+
30
+ | channel | baseline | finetuned | Δ |
31
+ |---|---|---|---|
32
+ | open | 0.00528 | 0.00332 | -37.0% |
33
+ | high | 0.00556 | 0.00504 | -9.5% |
34
+ | low | 0.00576 | 0.00502 | -12.9% |
35
+ | close | 0.00437 | 0.00308 | -29.6% |
36
+ | vol | 0.00615 | 0.00420 | -31.8% |
37
+ | amt | 0.00589 | 0.00420 | -28.8% |
38
+
39
+ ## Usage
40
+
41
+ ```python
42
+ from kairos import KronosTokenizer
43
+ tok = KronosTokenizer.from_pretrained("Shadowell/Kairos-base-crypto")
44
+ # Encode a [B, T, 6] OHLCV tensor into (s1_ids, s2_ids)
45
+ s1, s2 = tok.encode(x, half=True)
46
+ ```
47
+
48
+ ## Training recipe
49
+
50
+ See `docs/CRYPTO_TOKENIZER_RUN.md` in the upstream repo for the full
51
+ reproduction checklist (data collection, preparation, 15-epoch + patience 3
52
+ fine-tune, reconstruction evaluation).
config.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "attn_dropout_p": 0.0,
3
+ "beta": 0.05,
4
+ "d_in": 6,
5
+ "d_model": 256,
6
+ "ff_dim": 512,
7
+ "ffn_dropout_p": 0.0,
8
+ "gamma": 1.1,
9
+ "gamma0": 1.0,
10
+ "group_size": 4,
11
+ "n_dec_layers": 4,
12
+ "n_enc_layers": 4,
13
+ "n_heads": 4,
14
+ "resid_dropout_p": 0.0,
15
+ "s1_bits": 10,
16
+ "s2_bits": 10,
17
+ "zeta": 0.05
18
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:9207551f3a7c1b66e5517602c3b7becbfd915acbb2dad9c96fd13812a1d93779
3
+ size 15842368