Abhaykoul commited on
Commit
3ddb1a3
·
verified ·
1 Parent(s): d233641

Initial release of native LF2 2-bit quantized embedding model

Browse files
Files changed (5) hide show
  1. README.md +113 -0
  2. config.json +12 -0
  3. lf2_native.py +1302 -0
  4. model.safetensors +3 -0
  5. tokenizer.json +0 -0
README.md ADDED
@@ -0,0 +1,113 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: mit
5
+ library_name: tokenizers
6
+ tags:
7
+ - sentence-similarity
8
+ - feature-extraction
9
+ - embeddings
10
+ - rag
11
+ - quantized
12
+ - 2-bit
13
+ - lf2
14
+ - matryoshka
15
+ - ultra-lightweight
16
+ - code-search
17
+ - retrieval
18
+ - vortexa
19
+ pipeline_tag: feature-extraction
20
+ ---
21
+
22
+ <div align="center">
23
+
24
+ # 🚀 vtx-embed-1M-lf2 (`nano-2bit`)
25
+
26
+ **The world's most memory-efficient native 2-bit static embedding model powering [vortexa](https://github.com/OEvortex/vortexa).**
27
+ Native 2-Bit LF2 integer quantization · **0.39 MB RAM** · 1.05M Parameters · Fused Numba dequant & mean-pooling · Sub-millisecond CPU latency
28
+
29
+ [![HuggingFace](https://img.shields.io/badge/🤗%20HuggingFace-VTXAI%2Fvtx-embed-1M-blue)](https://huggingface.co/VTXAI/vtx-embed-1M)
30
+ [![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](https://opensource.org/licenses/MIT)
31
+ [![Python 3.8+](https://img.shields.io/badge/Python-3.8%2B-blue)](https://python.org)
32
+
33
+ </div>
34
+
35
+ ---
36
+
37
+ ## ⚡ What is LF2?
38
+
39
+ **LF2** is an ultra-compact, integer-native 2-bit quantization format for embedding matrices:
40
+ - **Zero FP32 Parameter Tables**: Parameters are stored entirely as packed 2-bit levels (4 weights per `uint8` byte) and double-quantized `uint8` scales and minimums.
41
+ - **Extreme Compression**: Memory drops from 0.57 MB (LF4 4-bit) down to **0.39 MB** (2-bit), retaining **95.40% mean token cosine similarity** to full precision.
42
+ - **Fused Dequantization & Pooling**: Evaluated on-the-fly using Numba JIT kernels directly into registers / L1 cache without allocating full FP32 token tables in RAM.
43
+
44
+ ---
45
+
46
+ ## 📄 Model Details
47
+
48
+ | Property | Value |
49
+ | :--- | :--- |
50
+ | **Model Name / Tier** | **vtx-embed-1M-lf2** (`"nano-2bit"`) |
51
+ | **Total Parameters** | **1.05M** |
52
+ | **Quantization Format** | `lf2` (Native 2-bit integer block quantization) |
53
+ | **In-RAM Memory** | **0.39 MB** |
54
+ | **On-Disk Size** | **0.39 MB** |
55
+ | **Embedding Dimension** | 64 |
56
+ | **Vocabulary Size** | 16,384 |
57
+ | **Block Size** | 16 |
58
+ | **Mean Cosine Similarity vs Orig** | **0.9540** |
59
+ | **License** | MIT |
60
+
61
+ ---
62
+
63
+ ## 💻 Quickstart Usage
64
+
65
+ ### Standalone Inference with `lf2_native.py`
66
+
67
+ This repository includes [`lf2_native.py`](lf2_native.py) directly for zero-dependency inference (only requires `numpy`, `safetensors`, and `tokenizers`; `numba` optional for maximum speed):
68
+
69
+ ```python
70
+ from huggingface_hub import snapshot_download
71
+ import sys
72
+
73
+ # 1. Download model repository
74
+ model_path = snapshot_download(repo_id="VTXAI/vtx-embed-1M-lf2")
75
+ sys.path.append(model_path)
76
+
77
+ from lf2_native import VortexEmbedLF2
78
+
79
+ # 2. Load model directly from local directory
80
+ model = VortexEmbedLF2.from_pretrained(model_path)
81
+ print(f"Model In-RAM size: {model.model_size_mb:.2f} MB")
82
+
83
+ # 3. Encode sentences
84
+ texts = [
85
+ "What is the capital of India?",
86
+ "Explain gravity and general relativity",
87
+ ]
88
+ embeddings = model.encode(texts)
89
+ print("Embeddings shape:", embeddings.shape) # (2, 64)
90
+
91
+ # 4. Semantic similarity
92
+ sim = embeddings[0] @ embeddings[1]
93
+ print("Cosine Similarity:", sim)
94
+ ```
95
+
96
+ ---
97
+
98
+ ## 📜 Citation
99
+
100
+ ```bibtex
101
+ @misc{vtx-embed-1m-lf2,
102
+ title = {vtx-embed-1M-lf2: Native 2-Bit Embeddings for Ultra-Low Footprint Semantic Search},
103
+ author = {VTXAI},
104
+ year = {2026},
105
+ url = {https://huggingface.co/VTXAI/vtx-embed-1M-lf2}
106
+ }
107
+ ```
108
+
109
+ ---
110
+
111
+ ## 📄 License
112
+
113
+ MIT License — free for commercial and research use.
config.json ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab_size": 16384,
3
+ "embedding_dim": 64,
4
+ "block_size": 16,
5
+ "num_blocks": 4,
6
+ "global_min": -49.90625,
7
+ "global_max": 47.2734375,
8
+ "global_scale_max": 25.01953125,
9
+ "matryoshka_dim": null,
10
+ "quantization": "lf2",
11
+ "bits": 2
12
+ }
lf2_native.py ADDED
@@ -0,0 +1,1302 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Vortex-Embed LF2 — Native 2-Bit Embedding Engine.
2
+
3
+ LF2 format (per weight-block of `block_size` fp32 weights):
4
+ - codes: 2-bit levels {0,1,2,3}, 4 weights packed per uint8 byte.
5
+ - scale_u8: per-block uint8, double-quantized step
6
+ (step = global_scale_max * scale_u8 / 255, step_fp = span / 3).
7
+ - min_u8: per-block uint8, double-quantized block minimum
8
+ (bmin = global_min + min_u8 / 255 * (global_max - global_min)).
9
+
10
+ Integer-native guarantee: RAM holds ONLY uint8 bytes
11
+ (packed codes + int8 metadata). The only floats in the whole checkpoint
12
+ are 3 fp32 scalars per tensor (global_min/max/scale_max, 12 bytes) —
13
+ zero FP32/FP16 parameter tables. Dequant happens on-the-fly for the
14
+ unique token IDs of each encode batch (registers/L1 temp buffer only),
15
+ mirroring the LF4 native engine's pooling/normalize path exactly.
16
+
17
+ Realtime paths (research: Model2Vec static-lookup + mean-pool O(n*d);
18
+ SwiftEmbed SIMD/prefetch/zero-copy; QuIP#/QTIP L1-resident codebooks +
19
+ bandwidth-bound fused dequant; AQLM additive LUTs):
20
+ - fused numba dequant+mean-pool (no (N,dim) temp, no np.unique, no
21
+ torch construction) — default fast path for index loops.
22
+ - opt-in preloaded fp32 table (Model2Vec/SwiftEmbed row-index mode).
23
+ - opt-in precomputed fp32 step/min meta (skip double-quant per batch).
24
+ - truncated-dim early exit (matryoshka needs only leading blocks).
25
+ - streaming indexer (tokenize-once, chunked, parallel).
26
+ """
27
+ from __future__ import annotations
28
+
29
+ import json
30
+ import os
31
+ from pathlib import Path
32
+ from typing import Iterator, List, Optional, Sequence, Union
33
+
34
+ import numpy as np
35
+ from safetensors.numpy import load_file, save_file
36
+
37
+ try:
38
+ from tokenizers import Tokenizer
39
+ except ImportError: # pragma: no cover
40
+ Tokenizer = None
41
+
42
+ try:
43
+ import numba
44
+
45
+ _NUMBA_OK = True
46
+ except ImportError: # pragma: no cover
47
+ numba = None # type: ignore
48
+ _NUMBA_OK = False
49
+
50
+ LEVELS = 3 # 2-bit -> levels {0,1,2,3}, step = span / 3
51
+ VALS_PER_BYTE = 4
52
+
53
+ # 256-entry byte->4x2bit LUT (1KB, L1-resident ala QuIP# E8P codebook).
54
+ _LUT4 = np.empty((256, VALS_PER_BYTE), dtype=np.uint8)
55
+ for _b in range(256):
56
+ _LUT4[_b, 0] = (_b >> 0) & 0x03
57
+ _LUT4[_b, 1] = (_b >> 2) & 0x03
58
+ _LUT4[_b, 2] = (_b >> 4) & 0x03
59
+ _LUT4[_b, 3] = (_b >> 6) & 0x03
60
+
61
+
62
+ if _NUMBA_OK:
63
+
64
+ @numba.njit(cache=True, fastmath=True)
65
+ def _fused_seq_full(
66
+ packed, scale_u8, min_u8, flat, starts, out,
67
+ gmin, grange, smax, nb, bs,
68
+ ):
69
+ # Full-dim specialization: no dd>=dim branch (out_dim == nb*bs).
70
+ n = starts.shape[0] - 1
71
+ pb = bs // 4
72
+ inv255 = 1.0 / 255.0
73
+ for d in range(n):
74
+ s0 = starts[d]
75
+ s1 = starts[d + 1]
76
+ L = s1 - s0
77
+ if L <= 0:
78
+ continue
79
+ inv = 1.0 / L
80
+ for ti in range(s0, s1):
81
+ tid = flat[ti]
82
+ dd = 0
83
+ for b in range(nb):
84
+ step = smax * (scale_u8[tid, b] * inv255)
85
+ bmin = gmin + (min_u8[tid, b] * inv255) * grange
86
+ poff = b * pb
87
+ for k in range(bs):
88
+ byte = packed[tid, poff + (k >> 2)]
89
+ code = (byte >> ((k & 3) * 2)) & 3
90
+ out[d, dd] += (code * step + bmin) * inv
91
+ dd += 1
92
+
93
+ @numba.njit(cache=True, fastmath=True)
94
+ def _fused_pool_seq(
95
+ packed, scale_u8, min_u8, flat, starts, out,
96
+ gmin, grange, smax, dim, nb, bs,
97
+ ):
98
+ n = starts.shape[0] - 1
99
+ pb = bs // 4
100
+ inv255 = 1.0 / 255.0
101
+ for d in range(n):
102
+ s0 = starts[d]
103
+ s1 = starts[d + 1]
104
+ L = s1 - s0
105
+ if L <= 0:
106
+ continue
107
+ inv = 1.0 / L
108
+ for ti in range(s0, s1):
109
+ tid = flat[ti]
110
+ base = 0
111
+ for b in range(nb):
112
+ step = smax * (scale_u8[tid, b] * inv255)
113
+ bmin = gmin + (min_u8[tid, b] * inv255) * grange
114
+ poff = b * pb
115
+ for k in range(bs):
116
+ dd = base + k
117
+ if dd >= dim:
118
+ break
119
+ byte = packed[tid, poff + (k >> 2)]
120
+ code = (byte >> ((k & 3) * 2)) & 3
121
+ out[d, dd] += (code * step + bmin) * inv
122
+ base += bs
123
+
124
+ @numba.njit(cache=True, fastmath=True, parallel=True)
125
+ def _fused_pool_nb(
126
+ packed, scale_u8, min_u8, flat, starts, out,
127
+ gmin, grange, smax, dim, nb, bs,
128
+ ):
129
+ n = starts.shape[0] - 1
130
+ pb = bs // 4
131
+ inv255 = 1.0 / 255.0
132
+ for d in numba.prange(n):
133
+ s0 = starts[d]
134
+ s1 = starts[d + 1]
135
+ L = s1 - s0
136
+ if L <= 0:
137
+ continue
138
+ inv = 1.0 / L
139
+ for ti in range(s0, s1):
140
+ tid = flat[ti]
141
+ base = 0
142
+ for b in range(nb):
143
+ step = smax * (scale_u8[tid, b] * inv255)
144
+ bmin = gmin + (min_u8[tid, b] * inv255) * grange
145
+ poff = b * pb
146
+ for k in range(bs):
147
+ dd = base + k
148
+ if dd >= dim:
149
+ break
150
+ byte = packed[tid, poff + (k >> 2)]
151
+ code = (byte >> ((k & 3) * 2)) & 3
152
+ out[d, dd] += (code * step + bmin) * inv
153
+ base += bs
154
+
155
+ @numba.njit(cache=True, fastmath=True)
156
+ def _fused_pool_w_seq(
157
+ packed, scale_u8, min_u8, flat, starts, out, wrow,
158
+ gmin, grange, smax, dim, nb, bs,
159
+ ):
160
+ n = starts.shape[0] - 1
161
+ pb = bs // 4
162
+ inv255 = 1.0 / 255.0
163
+ for d in range(n):
164
+ s0 = starts[d]
165
+ s1 = starts[d + 1]
166
+ wsum = 0.0
167
+ for ti in range(s0, s1):
168
+ wsum += wrow[ti]
169
+ if wsum < 1e-12:
170
+ continue
171
+ inv = 1.0 / wsum
172
+ for ti in range(s0, s1):
173
+ tid = flat[ti]
174
+ w = wrow[ti] * inv
175
+ base = 0
176
+ for b in range(nb):
177
+ step = smax * (scale_u8[tid, b] * inv255)
178
+ bmin = gmin + (min_u8[tid, b] * inv255) * grange
179
+ poff = b * pb
180
+ for k in range(bs):
181
+ dd = base + k
182
+ if dd >= dim:
183
+ break
184
+ byte = packed[tid, poff + (k >> 2)]
185
+ code = (byte >> ((k & 3) * 2)) & 3
186
+ out[d, dd] += (code * step + bmin) * w
187
+ base += bs
188
+
189
+ @numba.njit(cache=True, fastmath=True, parallel=True)
190
+ def _fused_pool_w_nb(
191
+ packed, scale_u8, min_u8, flat, starts, out, wrow,
192
+ gmin, grange, smax, dim, nb, bs,
193
+ ):
194
+ n = starts.shape[0] - 1
195
+ pb = bs // 4
196
+ inv255 = 1.0 / 255.0
197
+ for d in numba.prange(n):
198
+ s0 = starts[d]
199
+ s1 = starts[d + 1]
200
+ wsum = 0.0
201
+ for ti in range(s0, s1):
202
+ wsum += wrow[ti]
203
+ if wsum < 1e-12:
204
+ continue
205
+ inv = 1.0 / wsum
206
+ for ti in range(s0, s1):
207
+ tid = flat[ti]
208
+ w = wrow[ti] * inv
209
+ base = 0
210
+ for b in range(nb):
211
+ step = smax * (scale_u8[tid, b] * inv255)
212
+ bmin = gmin + (min_u8[tid, b] * inv255) * grange
213
+ poff = b * pb
214
+ for k in range(bs):
215
+ dd = base + k
216
+ if dd >= dim:
217
+ break
218
+ byte = packed[tid, poff + (k >> 2)]
219
+ code = (byte >> ((k & 3) * 2)) & 3
220
+ out[d, dd] += (code * step + bmin) * w
221
+ base += bs
222
+
223
+ @numba.njit(cache=True, fastmath=True)
224
+ def _table_pool_seq(table, flat, starts, out):
225
+ n = starts.shape[0] - 1
226
+ dim = out.shape[1]
227
+ for d in range(n):
228
+ s0 = starts[d]
229
+ s1 = starts[d + 1]
230
+ L = s1 - s0
231
+ if L <= 0:
232
+ continue
233
+ inv = 1.0 / L
234
+ for ti in range(s0, s1):
235
+ tid = flat[ti]
236
+ for j in range(dim):
237
+ out[d, j] += table[tid, j] * inv
238
+
239
+ @numba.njit(cache=True, fastmath=True, parallel=True)
240
+ def _table_pool_nb(table, flat, starts, out):
241
+ n = starts.shape[0] - 1
242
+ dim = out.shape[1]
243
+ for d in numba.prange(n):
244
+ s0 = starts[d]
245
+ s1 = starts[d + 1]
246
+ L = s1 - s0
247
+ if L <= 0:
248
+ continue
249
+ inv = 1.0 / L
250
+ for ti in range(s0, s1):
251
+ tid = flat[ti]
252
+ for j in range(dim):
253
+ out[d, j] += table[tid, j] * inv
254
+
255
+ @numba.njit(cache=True, fastmath=True)
256
+ def _norm_seq(x):
257
+ n = x.shape[0]
258
+ dim = x.shape[1]
259
+ for i in range(n):
260
+ s = 0.0
261
+ for j in range(dim):
262
+ s += x[i, j] * x[i, j]
263
+ inv = 1.0 / (np.sqrt(s) + 1e-12)
264
+ for j in range(dim):
265
+ x[i, j] *= inv
266
+
267
+ @numba.njit(cache=True, fastmath=True, parallel=True)
268
+ def _norm_nb(x):
269
+ n = x.shape[0]
270
+ dim = x.shape[1]
271
+ for i in numba.prange(n):
272
+ s = 0.0
273
+ for j in range(dim):
274
+ s += x[i, j] * x[i, j]
275
+ inv = 1.0 / (np.sqrt(s) + 1e-12)
276
+ for j in range(dim):
277
+ x[i, j] *= inv
278
+
279
+ # Parallel crossover: prange thread-spawn costs ~ms cold but the pool
280
+ # stays hot after warm_kernels (which warms with a realistic-size
281
+ # batch); above this many docs parallel wins by 4-6x. Measured:
282
+ # n=256: seq 1.79ms vs par 0.30ms; n=1379: seq 14.3ms vs par 2.9ms.
283
+ _PAR_MIN_DOCS = 128
284
+ _PAR_MIN_DOCS_NORM = 1024
285
+ else: # pragma: no cover
286
+ _fused_pool_nb = None
287
+ _fused_pool_seq = None
288
+ _fused_seq_full = None
289
+ _fused_pool_w_nb = None
290
+ _fused_pool_w_seq = None
291
+ _table_pool_nb = None
292
+ _table_pool_seq = None
293
+ _norm_nb = None
294
+ _norm_seq = None
295
+ _PAR_MIN_DOCS = 10**9
296
+ _PAR_MIN_DOCS_NORM = 10**9
297
+
298
+
299
+ def quantize_lf2_block(
300
+ x: np.ndarray, block_size: int
301
+ ) -> tuple[np.ndarray, np.ndarray, np.ndarray, float, float, float]:
302
+ """Quantize fp32 matrix to LF2 integer-native format.
303
+
304
+ Returns (packed_uint8, scale_u8, min_u8, global_min, global_max,
305
+ global_scale_max).
306
+ """
307
+ x = np.asarray(x, dtype=np.float32)
308
+ n, d = x.shape
309
+ assert d % block_size == 0, f"dim {d} not divisible by block {block_size}"
310
+ n_blocks = d // block_size
311
+ xb = x.reshape(n, n_blocks, block_size)
312
+ bmin = xb.min(axis=2)
313
+ bmax = xb.max(axis=2)
314
+ span = (bmax - bmin) / float(LEVELS)
315
+ span = np.where(span == 0, 1.0, span)
316
+ q = np.clip(np.round((xb - bmin[:, :, None]) / span[:, :, None]), 0, LEVELS)
317
+ q = q.astype(np.uint8).reshape(n, d)
318
+ step_fp = span # (n, n_blocks) float32
319
+
320
+ gmin = float(x.min())
321
+ gmax = float(x.max())
322
+ grange = (gmax - gmin) if gmax > gmin else 1.0
323
+ min_u8 = np.clip(
324
+ np.round(255.0 * (bmin - gmin) / grange), 0, 255
325
+ ).astype(np.uint8)
326
+ smax = float(step_fp.max())
327
+ scale_u8 = np.clip(np.round(255.0 * step_fp / smax), 0, 255).astype(np.uint8)
328
+
329
+ # Pack 4x 2-bit codes per byte, block-aligned (block_size % 4 == 0 required)
330
+ assert block_size % VALS_PER_BYTE == 0
331
+ qb = q.reshape(n, -1, VALS_PER_BYTE)
332
+ shifts = np.array([0, 2, 4, 6], dtype=np.uint8)
333
+ packed = np.zeros((n, q.shape[1] // VALS_PER_BYTE), dtype=np.uint8)
334
+ for i in range(VALS_PER_BYTE):
335
+ packed |= (qb[:, :, i] << shifts[i]).astype(np.uint8)
336
+ return packed, scale_u8, min_u8, gmin, gmax, smax
337
+
338
+
339
+ def dequantize_lf2_meta(
340
+ scale_u8: np.ndarray,
341
+ min_u8: np.ndarray,
342
+ gmin: float,
343
+ gmax: float,
344
+ smax: float,
345
+ ) -> tuple[np.ndarray, np.ndarray]:
346
+ """Double-quant metadata -> per-block float step/min (temp buffers only)."""
347
+ grange = (gmax - gmin) if gmax > gmin else 1.0
348
+ step = smax * scale_u8.astype(np.float32) / 255.0
349
+ bmin = gmin + min_u8.astype(np.float32) / 255.0 * grange
350
+ return step, bmin
351
+
352
+
353
+ class LF2Config:
354
+ def __init__(
355
+ self,
356
+ vocab_size: int = 29528,
357
+ embedding_dim: int = 256,
358
+ block_size: int = 32,
359
+ num_blocks: int = 8,
360
+ global_min: float = 0.0,
361
+ global_max: float = 0.0,
362
+ global_scale_max: float = 1.0,
363
+ matryoshka_dim: Optional[int] = None,
364
+ **kwargs,
365
+ ):
366
+ self.vocab_size = vocab_size
367
+ self.embedding_dim = embedding_dim
368
+ self.block_size = block_size
369
+ self.num_blocks = num_blocks
370
+ self.global_min = global_min
371
+ self.global_max = global_max
372
+ self.global_scale_max = global_scale_max
373
+ self.matryoshka_dim = matryoshka_dim
374
+
375
+ @classmethod
376
+ def from_dict(cls, d: dict) -> "LF2Config":
377
+ return cls(**d)
378
+
379
+ def to_dict(self) -> dict:
380
+ return {
381
+ "vocab_size": self.vocab_size,
382
+ "embedding_dim": self.embedding_dim,
383
+ "block_size": self.block_size,
384
+ "num_blocks": self.num_blocks,
385
+ "global_min": self.global_min,
386
+ "global_max": self.global_max,
387
+ "global_scale_max": self.global_scale_max,
388
+ "matryoshka_dim": self.matryoshka_dim,
389
+ "quantization": "lf2",
390
+ "bits": 2,
391
+ }
392
+
393
+
394
+ class VortexEmbedLF2:
395
+ """Native 2-bit sentence embedding model. Same encode path as LF4 engine."""
396
+
397
+ def __init__(
398
+ self,
399
+ packed: np.ndarray,
400
+ scale_u8: np.ndarray,
401
+ min_u8: np.ndarray,
402
+ tokenizer_data: Union[str, Path],
403
+ config: Union[dict, LF2Config],
404
+ *,
405
+ matryoshka_dim: Optional[int] = None,
406
+ ) -> None:
407
+ self.packed = np.asarray(packed, dtype=np.uint8)
408
+ self.scale_u8 = np.asarray(scale_u8, dtype=np.uint8)
409
+ self.min_u8 = np.asarray(min_u8, dtype=np.uint8)
410
+ self.tokenizer_data = str(tokenizer_data)
411
+ self.config = (
412
+ config if isinstance(config, LF2Config) else LF2Config.from_dict(config)
413
+ )
414
+ self.vocab_size = int(self.config.vocab_size)
415
+ self.dim = int(self.config.embedding_dim)
416
+ self.block_size = int(self.config.block_size)
417
+ self.num_blocks = int(self.config.num_blocks)
418
+ self.matryoshka_dim = matryoshka_dim or self.config.matryoshka_dim
419
+ self._tokenizer: Optional[Tokenizer] = None
420
+ self._sif_weights: Optional[np.ndarray] = None
421
+ self._pc_directions: Optional[np.ndarray] = None
422
+ # SIF 'a' mirrors the LF4 single-file engine default.
423
+ self.sif_a: float = 0.05
424
+ self.sif_pc: float = 1.0
425
+ self.pc_k: int = 1
426
+ # Hot-token row cache (opt-in, runtime only — not a parameter table).
427
+ # Maps token id -> dequantized fp32 row. Disabled by default to keep
428
+ # the native guarantee; enable with enable_cache() for indexing loops
429
+ # over skewed corpora. Not shared across threads (use shallow_clone).
430
+ self._row_cache: Optional[dict] = None
431
+ self._row_cache_cap: int = 0
432
+ # Opt-in runtime accelerators (not parameter tables; transient temp).
433
+ # preload_meta(): fp32 step/min per (vocab, block) — skips per-batch
434
+ # double-quant (Model2Vec-style precompute, ~2x240KB for 30k vocab).
435
+ # preload_table(): full fp32 table (V,D) — SwiftEmbed row-index mode,
436
+ # max tok/s at ~30MB transient (freed with unload_table()).
437
+ self._step_f32: Optional[np.ndarray] = None
438
+ self._bmin_f32: Optional[np.ndarray] = None
439
+ self._fp_table: Optional[np.ndarray] = None
440
+ # Warm numba kernels at construction (compile once, not per batch).
441
+ self._nb_warmed: bool = False
442
+
443
+ def enable_cache(self, cap: int = 4096) -> "VortexEmbedLF2":
444
+ """Opt-in LRU cache of dequantized token rows (runtime temp only)."""
445
+ from collections import OrderedDict
446
+
447
+ self._row_cache = OrderedDict()
448
+ self._row_cache_cap = max(int(cap), 1)
449
+ return self
450
+
451
+ def disable_cache(self) -> "VortexEmbedLF2":
452
+ self._row_cache = None
453
+ self._row_cache_cap = 0
454
+ return self
455
+
456
+ def shallow_clone(self) -> "VortexEmbedLF2":
457
+ """Share read-only params, fresh fit-state and empty cache config."""
458
+ c = VortexEmbedLF2(
459
+ self.packed, self.scale_u8, self.min_u8,
460
+ self.tokenizer_data, self.config,
461
+ matryoshka_dim=self.matryoshka_dim,
462
+ )
463
+ c.sif_a, c.sif_pc, c.pc_k = self.sif_a, self.sif_pc, self.pc_k
464
+ if self._row_cache is not None:
465
+ c.enable_cache(self._row_cache_cap)
466
+ # Share accelerators read-only (they are deterministic of params).
467
+ c._step_f32, c._bmin_f32, c._fp_table = (
468
+ self._step_f32, self._bmin_f32, self._fp_table)
469
+ return c
470
+
471
+ # -- realtime accelerators (opt-in, runtime temp only) ---------------
472
+ def preload_meta(self) -> "VortexEmbedLF2":
473
+ """Precompute fp32 step/min tables (skips per-batch double-quant)."""
474
+ step, bmin = dequantize_lf2_meta(
475
+ np.arange(256, dtype=np.uint8)[self.scale_u8.ravel()].reshape(
476
+ self.scale_u8.shape) * 0 + self.scale_u8,
477
+ self.min_u8,
478
+ self.config.global_min, self.config.global_max,
479
+ self.config.global_scale_max,
480
+ ) if False else dequantize_lf2_meta(
481
+ self.scale_u8, self.min_u8,
482
+ self.config.global_min, self.config.global_max,
483
+ self.config.global_scale_max,
484
+ )
485
+ self._step_f32 = np.ascontiguousarray(step, dtype=np.float32)
486
+ self._bmin_f32 = np.ascontiguousarray(bmin, dtype=np.float32)
487
+ return self
488
+
489
+ def unload_meta(self) -> "VortexEmbedLF2":
490
+ self._step_f32 = None
491
+ self._bmin_f32 = None
492
+ return self
493
+
494
+ def preload_table(self, dim: Optional[int] = None) -> "VortexEmbedLF2":
495
+ """Materialize full fp32 table transiently (SwiftEmbed row-index mode).
496
+
497
+ `dim` truncates columns (matryoshka early-exit). Call unload_table()
498
+ to restore the integer-native guarantee.
499
+ """
500
+ d = dim or self.dim
501
+ full = self._dequantize_fresh(
502
+ np.arange(self.vocab_size, dtype=np.int64))[:, :d]
503
+ self._fp_table = np.ascontiguousarray(full, dtype=np.float32)
504
+ return self
505
+
506
+ def unload_table(self) -> "VortexEmbedLF2":
507
+ self._fp_table = None
508
+ return self
509
+
510
+ def warm_kernels(self) -> "VortexEmbedLF2":
511
+ """Compile numba kernels once (avoid first-batch compile stall)."""
512
+ if not _NUMBA_OK or self._nb_warmed:
513
+ return self
514
+ try:
515
+ pk = np.ascontiguousarray(self.packed[:8])
516
+ sc = np.ascontiguousarray(self.scale_u8[:8])
517
+ mn = np.ascontiguousarray(self.min_u8[:8])
518
+ gmin = float(self.config.global_min)
519
+ gr = float((self.config.global_max - self.config.global_min)
520
+ if self.config.global_max > self.config.global_min else 1.0)
521
+ smax = float(self.config.global_scale_max)
522
+ nb, bs = self.num_blocks, self.block_size
523
+ # Warm BOTH seq + par variants (dispatch picks by batch size).
524
+ # The par warm uses a realistic-size batch so the numba thread
525
+ # pool is hot before serving (cold spawn costs ~ms).
526
+ nW = 512
527
+ flW = np.random.default_rng(0).integers(
528
+ 0, 8, size=2048).astype(np.int64)
529
+ stW = np.linspace(0, 2048, nW + 1).astype(np.int64)
530
+ ouW = np.zeros((nW, self.dim), dtype=np.float32)
531
+ _fused_pool_nb(pk, sc, mn, flW, stW, ouW, gmin, gr, smax,
532
+ self.dim, nb, bs)
533
+ fl = np.array([0, 1, 2], dtype=np.int64)
534
+ st = np.array([0, 2, 3], dtype=np.int64)
535
+ ou = np.zeros((2, self.dim), dtype=np.float32)
536
+ _fused_seq_full(pk, sc, mn, fl, st, ou, gmin, gr, smax, nb, bs)
537
+ ou[:] = 0
538
+ _fused_pool_seq(pk, sc, mn, fl, st, ou, gmin, gr, smax,
539
+ self.dim, nb, bs)
540
+ _norm_seq(ou)
541
+ _norm_nb(ou)
542
+ self._nb_warmed = True
543
+ except Exception:
544
+ pass
545
+ return self
546
+
547
+ # -- properties -----------------------------------------------------
548
+ @property
549
+ def tokenizer(self) -> Tokenizer:
550
+ if self._tokenizer is None:
551
+ if Tokenizer is None: # pragma: no cover
552
+ raise RuntimeError("tokenizers required: pip install tokenizers")
553
+ self._tokenizer = Tokenizer.from_file(self.tokenizer_data)
554
+ return self._tokenizer
555
+
556
+ @property
557
+ def int_bytes(self) -> int:
558
+ """Integer parameter bytes in RAM (codes + int8 meta)."""
559
+ return (
560
+ int(self.packed.nbytes)
561
+ + int(self.scale_u8.nbytes)
562
+ + int(self.min_u8.nbytes)
563
+ )
564
+
565
+ @property
566
+ def model_size_mb(self) -> float:
567
+ return (self.int_bytes + 12) / 1e6 # +3 fp32 global scalars
568
+
569
+ @property
570
+ def on_disk_size_mb(self) -> float:
571
+ return (self.int_bytes + 12) / 1e6
572
+
573
+ # -- io --------------------------------------------------------------
574
+ @classmethod
575
+ def quantize_from_matrix(
576
+ cls,
577
+ w_fp32: np.ndarray,
578
+ tokenizer_data: Union[str, Path],
579
+ block_size: int = 32,
580
+ matryoshka_dim: Optional[int] = None,
581
+ ) -> "VortexEmbedLF2":
582
+ packed, scale_u8, min_u8, gmin, gmax, smax = quantize_lf2_block(
583
+ w_fp32, block_size
584
+ )
585
+ n, d = w_fp32.shape
586
+ cfg = LF2Config(
587
+ vocab_size=n,
588
+ embedding_dim=d,
589
+ block_size=block_size,
590
+ num_blocks=d // block_size,
591
+ global_min=gmin,
592
+ global_max=gmax,
593
+ global_scale_max=smax,
594
+ matryoshka_dim=matryoshka_dim,
595
+ )
596
+ return cls(packed, scale_u8, min_u8, tokenizer_data, cfg,
597
+ matryoshka_dim=matryoshka_dim)
598
+
599
+ def save_pretrained(self, path: Union[str, Path]) -> None:
600
+ out = Path(path)
601
+ out.mkdir(parents=True, exist_ok=True)
602
+ save_file(
603
+ {
604
+ "embedding_packed": self.packed,
605
+ "embedding_scale_u8": self.scale_u8,
606
+ "embedding_min_u8": self.min_u8,
607
+ },
608
+ str(out / "model.safetensors"),
609
+ )
610
+ (out / "config.json").write_text(json.dumps(self.config.to_dict(), indent=2))
611
+ if not (out / "tokenizer.json").exists():
612
+ (out / "tokenizer.json").write_text(Path(self.tokenizer_data).read_text())
613
+
614
+ @classmethod
615
+ def from_pretrained(
616
+ cls, path: Union[str, Path], matryoshka_dim: Optional[int] = None
617
+ ) -> "VortexEmbedLF2":
618
+ path = Path(path)
619
+ tensors = load_file(str(path / "model.safetensors"))
620
+ config = json.loads((path / "config.json").read_text())
621
+ return cls(
622
+ packed=tensors["embedding_packed"],
623
+ scale_u8=tensors["embedding_scale_u8"],
624
+ min_u8=tensors["embedding_min_u8"],
625
+ tokenizer_data=str(path / "tokenizer.json"),
626
+ config=config,
627
+ matryoshka_dim=matryoshka_dim,
628
+ )
629
+
630
+ def dequantize_all(self) -> np.ndarray:
631
+ """Full fp32 table (offline analysis only — never held in RAM at runtime)."""
632
+ return self.dequantize_ids(
633
+ np.arange(self.vocab_size, dtype=np.int64))
634
+
635
+ # -- SIF-IDF + PC removal (mirrors LF4 engine) -------------------------
636
+ def fit_idf(self, corpus_token_lists: Sequence[Sequence[int]]) -> "VortexEmbedLF2":
637
+ flat = (
638
+ np.concatenate(corpus_token_lists)
639
+ if corpus_token_lists
640
+ else np.empty(0, dtype=np.int64)
641
+ )
642
+ total = max(int(flat.size), 1)
643
+ counts = np.bincount(flat, minlength=self.vocab_size).astype(np.float64)
644
+ p = counts / total
645
+ denom = self.sif_a + p
646
+ with np.errstate(divide="ignore", invalid="ignore"):
647
+ weights = np.where(p > 0, self.sif_a / denom, 1.0)
648
+ self._sif_weights = weights.astype(np.float32)
649
+ return self
650
+
651
+ def fit_pc(
652
+ self, corpus_embeddings: np.ndarray, k: Optional[int] = None
653
+ ) -> "VortexEmbedLF2":
654
+ if k is None:
655
+ k = self.pc_k
656
+ if corpus_embeddings.size == 0 or k <= 0:
657
+ return self
658
+ x = corpus_embeddings.astype(np.float32)
659
+ x = x - x.mean(axis=0, keepdims=True)
660
+ try:
661
+ _, _, vt = np.linalg.svd(x, full_matrices=False)
662
+ pcs = vt[:k].astype(np.float32)
663
+ pcs = pcs / (np.linalg.norm(pcs, axis=1, keepdims=True) + 1e-12)
664
+ self._pc_directions = pcs
665
+ except np.linalg.LinAlgError:
666
+ self._pc_directions = None
667
+ return self
668
+
669
+ def _apply_pc(self, x: np.ndarray) -> np.ndarray:
670
+ if self.sif_pc <= 0 or self._pc_directions is None:
671
+ return x
672
+ out = x
673
+ for pc in self._pc_directions:
674
+ proj = (out @ pc)[:, None] * pc[None, :]
675
+ out = out - self.sif_pc * proj
676
+ return out
677
+
678
+ def reset_fit(self) -> "VortexEmbedLF2":
679
+ self._sif_weights = None
680
+ self._pc_directions = None
681
+ return self
682
+
683
+ # -- native on-the-fly dequant (H17: planar-fill + single astype) ----
684
+ def dequantize_ids(self, token_ids: np.ndarray) -> np.ndarray:
685
+ """2-pass dequant: strided fills happen on a uint8 temp (1/4 the
686
+ traffic), then ONE u8->f32 cast + ONE contiguous blocked fmadd.
687
+ No (N, dim) float temp, no per-stream astype."""
688
+ if token_ids.size == 0:
689
+ return np.empty((0, self.dim), dtype=np.float32)
690
+ n = len(token_ids)
691
+ nb, bs = self.num_blocks, self.block_size
692
+ cache = self._row_cache
693
+ if cache is not None and n <= 512:
694
+ # Hot-token path: reuse cached rows, dequantize misses only.
695
+ rows: List[Optional[np.ndarray]] = [cache.get(int(t)) for t in token_ids]
696
+ # LRU touch on hits
697
+ for t, r in zip(token_ids, rows):
698
+ if r is not None:
699
+ cache.move_to_end(int(t))
700
+ missing = np.array(
701
+ [t for t, r in zip(token_ids, rows) if r is None], dtype=np.int64
702
+ )
703
+ if missing.size:
704
+ got = self._dequantize_fresh(missing)
705
+ for t, r in zip(missing, got):
706
+ cache[int(t)] = r
707
+ if len(cache) > self._row_cache_cap:
708
+ cache.popitem(last=False)
709
+ it = iter(zip(missing, got))
710
+ lut = {int(t): r for t, r in it}
711
+ rows = [r if r is not None else lut[int(t)]
712
+ for t, r in zip(token_ids, rows)]
713
+ return np.stack(list(rows)).astype(np.float32)
714
+ return self._dequantize_fresh(np.asarray(token_ids, dtype=np.int64))
715
+
716
+ def _dequantize_fresh(self, token_ids: np.ndarray,
717
+ out_dim: Optional[int] = None) -> np.ndarray:
718
+ """2-pass dequant: strided fills happen on a uint8 temp (1/4 the
719
+ traffic), then ONE u8->f32 cast + ONE contiguous blocked fmadd.
720
+ No (N, dim) float temp, no per-stream astype.
721
+
722
+ `out_dim` enables matryoshka early-exit (leading blocks only).
723
+ Uses preloaded fp32 meta when available (skips double-quant).
724
+ """
725
+ if token_ids.size == 0:
726
+ d = out_dim or self.dim
727
+ return np.empty((0, d), dtype=np.float32)
728
+ n = len(token_ids)
729
+ nb, bs = self.num_blocks, self.block_size
730
+ dim = out_dim or self.dim
731
+ nb_need = min(nb, (dim + bs - 1) // bs)
732
+ p = self.packed[token_ids]
733
+ if nb_need < nb:
734
+ # Slice leading bytes/blocks only (truncated-dim early exit).
735
+ pb = bs // VALS_PER_BYTE
736
+ p = p[:, : nb_need * pb]
737
+ p = p.reshape(n, nb_need, bs // VALS_PER_BYTE)
738
+ t = np.empty((n, nb_need, bs), dtype=np.uint8)
739
+ t[:, :, 0::4] = p & 0x03
740
+ t[:, :, 1::4] = (p >> 2) & 0x03
741
+ t[:, :, 2::4] = (p >> 4) & 0x03
742
+ t[:, :, 3::4] = (p >> 6) & 0x03
743
+ if self._step_f32 is not None and self._bmin_f32 is not None:
744
+ step = self._step_f32[token_ids][:, :nb_need]
745
+ bmin = self._bmin_f32[token_ids][:, :nb_need]
746
+ else:
747
+ step, bmin = dequantize_lf2_meta(
748
+ self.scale_u8[token_ids][:, :nb_need]
749
+ if nb_need < nb else self.scale_u8[token_ids],
750
+ self.min_u8[token_ids][:, :nb_need]
751
+ if nb_need < nb else self.min_u8[token_ids],
752
+ self.config.global_min,
753
+ self.config.global_max,
754
+ self.config.global_scale_max,
755
+ )
756
+ f = t.astype(np.float32)
757
+ f *= step[:, :, None]
758
+ f += bmin[:, :, None]
759
+ return f.reshape(n, nb_need * bs)[:, :dim]
760
+
761
+ def _dequantize_lut(self, token_ids: np.ndarray,
762
+ out_dim: Optional[int] = None) -> np.ndarray:
763
+ """LUT-gather variant (256x4 table, QuIP#-style L1 codebook)."""
764
+ if token_ids.size == 0:
765
+ return np.empty((0, out_dim or self.dim), dtype=np.float32)
766
+ n = len(token_ids)
767
+ nb, bs = self.num_blocks, self.block_size
768
+ dim = out_dim or self.dim
769
+ nb_need = min(nb, (dim + bs - 1) // bs)
770
+ pb = bs // VALS_PER_BYTE
771
+ p = self.packed[token_ids]
772
+ if nb_need < nb:
773
+ p = p[:, : nb_need * pb]
774
+ t = _LUT4[p] # (n, bytes, 4) uint8, single gather
775
+ t = t.reshape(n, nb_need, bs)
776
+ if self._step_f32 is not None and self._bmin_f32 is not None:
777
+ step = self._step_f32[token_ids][:, :nb_need]
778
+ bmin = self._bmin_f32[token_ids][:, :nb_need]
779
+ else:
780
+ step, bmin = dequantize_lf2_meta(
781
+ self.scale_u8[token_ids][:, :nb_need]
782
+ if nb_need < nb else self.scale_u8[token_ids],
783
+ self.min_u8[token_ids][:, :nb_need]
784
+ if nb_need < nb else self.min_u8[token_ids],
785
+ self.config.global_min,
786
+ self.config.global_max,
787
+ self.config.global_scale_max,
788
+ )
789
+ f = t.astype(np.float32)
790
+ f *= step[:, :, None]
791
+ f += bmin[:, :, None]
792
+ return f.reshape(n, nb_need * bs)[:, :dim]
793
+
794
+ # -- encode (same segment-sum path as LF4 engine) ----------------------
795
+ def _tokenize_batch(self, texts: Sequence[str]) -> List[List[int]]:
796
+ encoded = self.tokenizer.encode_batch(list(texts))
797
+ return [
798
+ [tid for tid in item.ids if 0 <= int(tid) < self.vocab_size]
799
+ for item in encoded
800
+ ]
801
+
802
+ @staticmethod
803
+ def _normalize_inplace(x: np.ndarray) -> None:
804
+ norms = np.linalg.norm(x, axis=1, keepdims=True)
805
+ np.divide(x, np.maximum(norms, 1e-12), out=x)
806
+
807
+ @staticmethod
808
+ def _flat_starts(token_lists: Sequence[Sequence[int]],
809
+ max_tokens: int = 0):
810
+ n = len(token_lists)
811
+ if n == 0:
812
+ return (np.empty(0, dtype=np.int64), np.zeros(1, dtype=np.int64),
813
+ np.empty(0, dtype=np.int64))
814
+ if max_tokens and max_tokens > 0:
815
+ trunc = [ids[:max_tokens] if len(ids) > max_tokens else ids
816
+ for ids in token_lists]
817
+ else:
818
+ trunc = list(token_lists)
819
+ # One Python-level pass with C-speed list.extend + a single
820
+ # array build (2x faster than np.concatenate's per-list convert).
821
+ big: List[int] = []
822
+ ap = big.extend
823
+ for t in trunc:
824
+ ap(t)
825
+ lens = np.fromiter((len(t) for t in trunc), dtype=np.int64, count=n)
826
+ if big:
827
+ flat = np.asarray(big, dtype=np.int64)
828
+ else:
829
+ flat = np.empty(0, dtype=np.int64)
830
+ starts = np.empty(n + 1, dtype=np.int64)
831
+ starts[0] = 0
832
+ np.cumsum(lens, out=starts[1:])
833
+ return flat, starts, lens
834
+
835
+ def _encode_fused(self, token_lists, *, normalize: bool,
836
+ out_dim: int, max_tokens: int = 0) -> Optional[np.ndarray]:
837
+ """Fused numba dequant+pool: no unique, no (T,dim) temp, no torch."""
838
+ if not _NUMBA_OK or self._pc_directions is not None:
839
+ return None
840
+ flat, starts, lens = self._flat_starts(token_lists, max_tokens)
841
+ n = len(token_lists)
842
+ if flat.size == 0:
843
+ return np.zeros((n, out_dim), dtype=np.float32)
844
+ out = np.zeros((n, out_dim), dtype=np.float32)
845
+ nb_need = min(self.num_blocks, (out_dim + self.block_size - 1) // self.block_size)
846
+ gmin = float(self.config.global_min)
847
+ gmax = float(self.config.global_max)
848
+ grange = (gmax - gmin) if gmax > gmin else 1.0
849
+ smax = float(self.config.global_scale_max)
850
+ par = n >= _PAR_MIN_DOCS
851
+ try:
852
+ if self._sif_weights is not None:
853
+ wrow = self._sif_weights[flat].astype(np.float32)
854
+ kern = _fused_pool_w_nb if par else _fused_pool_w_seq
855
+ kern(self.packed, self.scale_u8, self.min_u8,
856
+ flat, starts, out, wrow,
857
+ gmin, grange, smax, out_dim,
858
+ nb_need, self.block_size)
859
+ elif out_dim == nb_need * self.block_size and not par:
860
+ # Fast lane: full-dim seq kernel, no bounds branch.
861
+ _fused_seq_full(self.packed, self.scale_u8, self.min_u8,
862
+ flat, starts, out,
863
+ gmin, grange, smax,
864
+ nb_need, self.block_size)
865
+ else:
866
+ kern = _fused_pool_nb if par else _fused_pool_seq
867
+ kern(self.packed, self.scale_u8, self.min_u8,
868
+ flat, starts, out,
869
+ gmin, grange, smax, out_dim,
870
+ nb_need, self.block_size)
871
+ except Exception:
872
+ return None
873
+ if normalize and n:
874
+ if _norm_nb is not None:
875
+ try:
876
+ (_norm_nb if n >= _PAR_MIN_DOCS_NORM else _norm_seq)(out)
877
+ except Exception:
878
+ self._normalize_inplace(out)
879
+ else:
880
+ self._normalize_inplace(out)
881
+ return out
882
+
883
+ def _encode_table(self, token_lists, *, normalize: bool,
884
+ out_dim: int, max_tokens: int = 0) -> Optional[np.ndarray]:
885
+ """Preloaded-fp32 row-index pool (Model2Vec/SwiftEmbed mode)."""
886
+ if self._fp_table is None:
887
+ return None
888
+ if self._pc_directions is not None:
889
+ return None
890
+ tab = self._fp_table
891
+ if tab.shape[1] < out_dim:
892
+ return None
893
+ tab = tab[:, :out_dim]
894
+ flat, starts, _ = self._flat_starts(token_lists, max_tokens)
895
+ n = len(token_lists)
896
+ if flat.size == 0:
897
+ return np.zeros((n, out_dim), dtype=np.float32)
898
+ out = np.zeros((n, out_dim), dtype=np.float32)
899
+ try:
900
+ if self._sif_weights is not None or not _NUMBA_OK:
901
+ raise RuntimeError("fallback")
902
+ kern = _table_pool_nb if n >= _PAR_MIN_DOCS else _table_pool_seq
903
+ kern(np.ascontiguousarray(tab), flat, starts, out)
904
+ except Exception:
905
+ # Numpy fallback: unique-dedup gather + reduceat (still no dequant).
906
+ uq, inv = np.unique(flat, return_inverse=True)
907
+ te = np.ascontiguousarray(tab[uq])[inv]
908
+ if self._sif_weights is not None:
909
+ w = self._sif_weights[flat].astype(np.float32)[:, None]
910
+ te = te * w
911
+ ends = starts[1:]
912
+ bounds = starts[:-1]
913
+ # guard empty docs: reduceat needs valid indices; handle via mask
914
+ sums = np.add.reduceat(te, bounds, axis=0) if te.size else out
915
+ lens = np.diff(starts).astype(np.float32)
916
+ if self._sif_weights is not None:
917
+ wf = self._sif_weights[flat].astype(np.float32)
918
+ wpr = np.add.reduceat(wf, bounds)
919
+ wpr = np.maximum(wpr, 1e-12)
920
+ else:
921
+ wpr = np.maximum(lens, 1.0)
922
+ # Fix rows for empty docs (reduceat wraps around): zero them.
923
+ out = sums / wpr[:, None]
924
+ out[lens == 0] = 0.0
925
+ if normalize and n:
926
+ self._normalize_inplace(out)
927
+ return out.astype(np.float32)
928
+ if normalize and n:
929
+ self._normalize_inplace(out)
930
+ return out
931
+
932
+ def _encode_subbatch(
933
+ self, token_lists: Sequence[Sequence[int]], *, normalize: bool,
934
+ fast: bool = True,
935
+ ) -> np.ndarray:
936
+ n = len(token_lists)
937
+ if fast:
938
+ # Fast dispatch order: table (fastest) -> fused numba (no temp)
939
+ # -> legacy unique+torch (exact legacy numerics, SIF/PC-safe).
940
+ got = self._encode_table(token_lists, normalize=False,
941
+ out_dim=self.dim)
942
+ if got is not None:
943
+ embs = self._apply_pc(got)
944
+ if normalize:
945
+ self._normalize_inplace(embs)
946
+ return embs
947
+ got = self._encode_fused(token_lists, normalize=False,
948
+ out_dim=self.dim)
949
+ if got is not None:
950
+ embs = self._apply_pc(got)
951
+ if normalize:
952
+ self._normalize_inplace(embs)
953
+ return embs
954
+ return self._encode_legacy(token_lists, normalize=normalize)
955
+
956
+ def _encode_legacy(self, token_lists, *, normalize: bool) -> np.ndarray:
957
+ n = len(token_lists)
958
+ flat = (
959
+ np.concatenate(token_lists)
960
+ if token_lists
961
+ else np.empty(0, dtype=np.int64)
962
+ )
963
+ if flat.size == 0:
964
+ return np.zeros((n, self.dim), dtype=np.float32)
965
+ unique_ids, inverse = np.unique(flat, return_inverse=True)
966
+ token_embs = self.dequantize_ids(unique_ids)[inverse]
967
+ if self._sif_weights is not None:
968
+ w = self._sif_weights[flat].astype(np.float32)[:, None]
969
+ token_embs = token_embs * w
970
+ try:
971
+ import torch
972
+
973
+ ro = torch.from_numpy(
974
+ np.repeat(
975
+ np.arange(n, dtype=np.int64),
976
+ [len(ids) for ids in token_lists],
977
+ )
978
+ )
979
+ em = torch.from_numpy(np.ascontiguousarray(token_embs))
980
+ sums = torch.zeros((n, self.dim), dtype=torch.float32)
981
+ sums.index_add_(0, ro, em)
982
+ sums = sums.numpy()
983
+ except ImportError:
984
+ chunk_lens = np.array(
985
+ [len(ids) for ids in token_lists], dtype=np.int64
986
+ )
987
+ ends = np.cumsum(chunk_lens)
988
+ bounds = np.empty(n + 1, dtype=np.int64)
989
+ bounds[0] = 0
990
+ bounds[1:] = ends
991
+ sums = np.add.reduceat(token_embs, bounds[:-1], axis=0)
992
+ chunk_lens = np.array([len(ids) for ids in token_lists], dtype=np.int64)
993
+ if self._sif_weights is not None:
994
+ w_full = self._sif_weights[flat].astype(np.float32)
995
+ ends = np.cumsum(chunk_lens)
996
+ bounds = np.empty(n + 1, dtype=np.int64)
997
+ bounds[0] = 0
998
+ bounds[1:] = ends
999
+ w_per_row = np.add.reduceat(w_full, bounds[:-1])
1000
+ w_per_row = np.maximum(w_per_row, 1e-12)
1001
+ else:
1002
+ w_per_row = np.maximum(chunk_lens.astype(np.float32), 1.0)
1003
+ embs = sums / w_per_row[:, None]
1004
+ embs = self._apply_pc(embs)
1005
+ if normalize:
1006
+ self._normalize_inplace(embs)
1007
+ return embs
1008
+
1009
+ def encode_batch(
1010
+ self,
1011
+ texts: Sequence[str],
1012
+ *,
1013
+ normalize: bool = True,
1014
+ truncate_dim: Optional[int] = None,
1015
+ ) -> np.ndarray:
1016
+ if not texts:
1017
+ return np.zeros((0, self.dim), dtype=np.float32)
1018
+ if len(texts) == 1:
1019
+ # H17 latency path: single text needs no segment sum — plain
1020
+ # (weighted) mean skips torch construction + index_add entirely.
1021
+ embs = self._encode_single(texts[0])
1022
+ embs = embs[None, :]
1023
+ else:
1024
+ embs = self._encode_subbatch(self._tokenize_batch(list(texts)),
1025
+ normalize=False)
1026
+ dim = truncate_dim if truncate_dim is not None else self.matryoshka_dim
1027
+ if dim is not None and 0 < dim < self.dim:
1028
+ embs = embs[:, :dim]
1029
+ if normalize and embs.shape[0] > 0:
1030
+ self._normalize_inplace(embs)
1031
+ return embs
1032
+
1033
+ def _encode_single(self, text: str) -> np.ndarray:
1034
+ ids = self._tokenize_batch([text])[0]
1035
+ if not ids:
1036
+ return np.zeros((self.dim,), dtype=np.float32)
1037
+ flat = np.asarray(ids, dtype=np.int64)
1038
+ if flat.size <= 64:
1039
+ # Short-text path: mean over occurrences needs no dedup math —
1040
+ # dequantize flat directly, skipping unique + inverse gather.
1041
+ # (Identical to dedup-then-mean up to fp summation order.)
1042
+ token_embs = self.dequantize_ids(flat)
1043
+ else:
1044
+ unique_ids, inverse = np.unique(flat, return_inverse=True)
1045
+ token_embs = self.dequantize_ids(unique_ids)[inverse]
1046
+ if self._sif_weights is not None:
1047
+ w = self._sif_weights[flat].astype(np.float32)
1048
+ embs = (token_embs * w[:, None]).sum(axis=0) / max(float(w.sum()), 1e-12)
1049
+ else:
1050
+ embs = token_embs.mean(axis=0)
1051
+ return self._apply_pc(embs[None, :])[0]
1052
+
1053
+ def encode(
1054
+ self,
1055
+ texts: Union[str, Sequence[str]],
1056
+ *,
1057
+ normalize: bool = True,
1058
+ truncate_dim: Optional[int] = None,
1059
+ ) -> np.ndarray:
1060
+ if isinstance(texts, str):
1061
+ return self.encode_batch(
1062
+ [texts], normalize=normalize, truncate_dim=truncate_dim
1063
+ )[0]
1064
+ return self.encode_batch(
1065
+ list(texts), normalize=normalize, truncate_dim=truncate_dim
1066
+ )
1067
+
1068
+ # -- realtime indexing API (pre-tokenized + parallel) ------------------
1069
+ def encode_ids(
1070
+ self,
1071
+ token_lists: Sequence[Sequence[int]],
1072
+ *,
1073
+ normalize: bool = True,
1074
+ truncate_dim: Optional[int] = None,
1075
+ max_tokens: Optional[int] = None,
1076
+ fast: bool = True,
1077
+ ) -> np.ndarray:
1078
+ """Encode pre-tokenized id lists (skips the tokenizer entirely).
1079
+
1080
+ `max_tokens` truncates each list (Model2Vec-style max_length).
1081
+ Reactive indexers tokenize once, then call this per batch.
1082
+ `fast=True` (default) uses table/fused kernels with truncated-dim
1083
+ early-exit; `fast=False` forces the legacy unique+torch path.
1084
+ """
1085
+ lists: List[Sequence[int]] = list(token_lists)
1086
+ mt = int(max_tokens) if max_tokens is not None and max_tokens > 0 else 0
1087
+ if not lists:
1088
+ d0 = truncate_dim or self.matryoshka_dim or self.dim
1089
+ return np.zeros((0, d0), dtype=np.float32)
1090
+ dim = truncate_dim if truncate_dim is not None else self.matryoshka_dim
1091
+ out_dim = dim if dim is not None and 0 < dim < self.dim else self.dim
1092
+ if fast:
1093
+ got = self._encode_table(lists, normalize=False,
1094
+ out_dim=out_dim, max_tokens=mt)
1095
+ if got is None:
1096
+ got = self._encode_fused(lists, normalize=False,
1097
+ out_dim=out_dim, max_tokens=mt)
1098
+ if got is not None:
1099
+ got = self._apply_pc(got)
1100
+ if normalize and got.shape[0]:
1101
+ self._normalize_inplace(got)
1102
+ return got
1103
+ if mt:
1104
+ lists = [ids[:mt] for ids in lists]
1105
+ if len(lists) == 1:
1106
+ flat = np.asarray(lists[0], dtype=np.int64)
1107
+ if flat.size == 0:
1108
+ embs = np.zeros((1, self.dim), dtype=np.float32)
1109
+ else:
1110
+ embs = self._mean_pool(flat)[None, :]
1111
+ else:
1112
+ embs = self._encode_subbatch(lists, normalize=False, fast=fast)
1113
+ if out_dim < self.dim:
1114
+ embs = embs[:, :out_dim]
1115
+ if normalize and embs.shape[0] > 0:
1116
+ self._normalize_inplace(embs)
1117
+ return embs
1118
+
1119
+ def _mean_pool(self, flat: np.ndarray) -> np.ndarray:
1120
+ if flat.size <= 64:
1121
+ token_embs = self.dequantize_ids(flat)
1122
+ else:
1123
+ unique_ids, inverse = np.unique(flat, return_inverse=True)
1124
+ token_embs = self.dequantize_ids(unique_ids)[inverse]
1125
+ if self._sif_weights is not None:
1126
+ w = self._sif_weights[flat].astype(np.float32)
1127
+ embs = (token_embs * w[:, None]).sum(axis=0) / max(float(w.sum()), 1e-12)
1128
+ else:
1129
+ embs = token_embs.mean(axis=0)
1130
+ return self._apply_pc(embs[None, :])[0]
1131
+
1132
+ def _fast_usable(self) -> bool:
1133
+ """True when a single-shot numba path handles this config.
1134
+
1135
+ The fused/table kernels already parallelize over docs internally
1136
+ (prange), so ThreadPool sharding on top only adds spawn + order
1137
+ restore overhead. Single-shot is the fastest option.
1138
+ """
1139
+ return bool(_NUMBA_OK) and self._pc_directions is None
1140
+
1141
+ def encode_parallel(
1142
+ self,
1143
+ token_lists: Sequence[Sequence[int]],
1144
+ *,
1145
+ n_jobs: int = 8,
1146
+ batch: int = 256,
1147
+ normalize: bool = True,
1148
+ truncate_dim: Optional[int] = None,
1149
+ max_tokens: Optional[int] = None,
1150
+ fast: bool = True,
1151
+ ) -> np.ndarray:
1152
+ """Shard pre-tokenized id lists across worker threads.
1153
+
1154
+ Tokenize ONCE in the caller, then fan out pure vector math.
1155
+ When the single-shot numba fast path applies (default: no PC fit),
1156
+ `n_jobs`/`batch` are bypassed — one call already saturates cores.
1157
+ Threaded sharding remains for the legacy torch path (`fast=False`)
1158
+ or PC-fitted models.
1159
+ """
1160
+ if fast and self._fast_usable():
1161
+ return self.encode_ids(
1162
+ token_lists, normalize=normalize, truncate_dim=truncate_dim,
1163
+ max_tokens=max_tokens, fast=True)
1164
+ from concurrent.futures import ThreadPoolExecutor
1165
+
1166
+ lists: List[Sequence[int]] = list(token_lists)
1167
+ if max_tokens is not None and max_tokens > 0:
1168
+ lists = [ids[:max_tokens] for ids in lists]
1169
+ n = len(lists)
1170
+ if n == 0:
1171
+ d = truncate_dim or self.matryoshka_dim or self.dim
1172
+ return np.zeros((0, d), dtype=np.float32)
1173
+ if n_jobs <= 0:
1174
+ n_jobs = max(1, os.cpu_count() or 1)
1175
+ n_jobs = max(1, min(int(n_jobs), n))
1176
+ if n_jobs == 1:
1177
+ return self.encode_ids(
1178
+ lists, normalize=normalize, truncate_dim=truncate_dim,
1179
+ fast=fast,
1180
+ )
1181
+ chunks = [lists[i::n_jobs] for i in range(n_jobs)]
1182
+ workers = [self.shallow_clone() for _ in range(n_jobs)]
1183
+
1184
+ def _run(args) -> np.ndarray:
1185
+ w, ch = args
1186
+ out = []
1187
+ for i in range(0, len(ch), batch):
1188
+ out.append(
1189
+ w.encode_ids(
1190
+ ch[i:i + batch],
1191
+ normalize=False,
1192
+ truncate_dim=truncate_dim,
1193
+ fast=fast,
1194
+ )
1195
+ )
1196
+ return np.vstack(out) if out else np.zeros((0, self.dim))
1197
+
1198
+ with ThreadPoolExecutor(max_workers=n_jobs) as ex:
1199
+ parts = list(ex.map(_run, zip(workers, chunks)))
1200
+ # Restore original order (round-robin interleave invert)
1201
+ order = np.argsort(
1202
+ np.concatenate([np.arange(i, n, n_jobs) for i in range(n_jobs)])
1203
+ )
1204
+ embs = np.vstack(parts)[order]
1205
+ dim = truncate_dim if truncate_dim is not None else self.matryoshka_dim
1206
+ if dim is not None and 0 < dim < self.dim and embs.shape[1] > dim:
1207
+ embs = embs[:, :dim]
1208
+ if normalize and embs.shape[0] > 0:
1209
+ self._normalize_inplace(embs)
1210
+ return embs
1211
+
1212
+ # -- streaming realtime indexer -------------------------------------
1213
+ def tokenize_texts(self, texts: Sequence[str],
1214
+ max_tokens: int = 0) -> List[List[int]]:
1215
+ """Tokenize once (main thread); reuse lists for encode_ids*."""
1216
+ lists = self._tokenize_batch(list(texts))
1217
+ if max_tokens and max_tokens > 0:
1218
+ lists = [ids[:max_tokens] for ids in lists]
1219
+ return lists
1220
+
1221
+ def index_texts(self, texts: Sequence[str], *,
1222
+ batch: int = 512, n_jobs: int = 0,
1223
+ normalize: bool = True,
1224
+ truncate_dim: Optional[int] = None,
1225
+ max_tokens: Optional[int] = None,
1226
+ fast: bool = True,
1227
+ show_progress: bool = False) -> np.ndarray:
1228
+ """End-to-end realtime index: tokenize-once + chunked parallel pool.
1229
+
1230
+ Tokenizes the whole input in ONE tokenizer call (Rust-batched),
1231
+ then pools each 50k-doc chunk in a single numba shot. `n_jobs`/
1232
+ `batch` only affect the legacy path; the fast path ignores them
1233
+ (internal prange already saturates cores).
1234
+ """
1235
+ texts = list(texts)
1236
+ n = len(texts)
1237
+ d = truncate_dim or self.matryoshka_dim or self.dim
1238
+ if n == 0:
1239
+ return np.zeros((0, d), dtype=np.float32)
1240
+ mt = int(max_tokens or 0)
1241
+ if n <= 50000:
1242
+ lists = self.tokenize_texts(texts, max_tokens=mt)
1243
+ return self.encode_ids(
1244
+ lists, normalize=normalize, truncate_dim=truncate_dim,
1245
+ fast=fast)
1246
+ out_parts: List[np.ndarray] = []
1247
+ step = 50000
1248
+ it = range(0, n, step)
1249
+ if show_progress:
1250
+ try:
1251
+ from tqdm import tqdm # type: ignore
1252
+ it = tqdm(it, desc="index")
1253
+ except Exception:
1254
+ pass
1255
+ for s in it:
1256
+ lists = self.tokenize_texts(texts[s:s + step], max_tokens=mt)
1257
+ out_parts.append(self.encode_ids(
1258
+ lists, normalize=normalize, truncate_dim=truncate_dim,
1259
+ fast=fast))
1260
+ return np.vstack(out_parts) if out_parts else np.zeros((0, d))
1261
+
1262
+ def index_stream(self, texts: Iterator[str], *,
1263
+ batch: int = 512, n_jobs: int = 0,
1264
+ normalize: bool = True,
1265
+ truncate_dim: Optional[int] = None,
1266
+ max_tokens: Optional[int] = None,
1267
+ fast: bool = True) -> Iterator[np.ndarray]:
1268
+ """Yield embedding chunks for an unbounded text iterator."""
1269
+ buf: List[str] = []
1270
+ width = max(int(batch) * max(int(n_jobs or 1), 1), 512)
1271
+ for t in texts:
1272
+ buf.append(t)
1273
+ if len(buf) >= width:
1274
+ lists = self.tokenize_texts(buf, int(max_tokens or 0))
1275
+ yield self.encode_parallel(
1276
+ lists, n_jobs=n_jobs or 1, batch=int(batch),
1277
+ normalize=normalize, truncate_dim=truncate_dim,
1278
+ fast=fast)
1279
+ buf = []
1280
+ if buf:
1281
+ lists = self.tokenize_texts(buf, int(max_tokens or 0))
1282
+ yield self.encode_parallel(
1283
+ lists, n_jobs=n_jobs or 1, batch=int(batch),
1284
+ normalize=normalize, truncate_dim=truncate_dim,
1285
+ fast=fast)
1286
+
1287
+ def search(self, queries: np.ndarray, index: np.ndarray, top_k: int = 10,
1288
+ index_normalized: bool = False) -> tuple[np.ndarray, np.ndarray]:
1289
+ """Cosine top-k (queries assumed L2-normalized)."""
1290
+ q = np.asarray(queries, dtype=np.float32)
1291
+ if q.ndim == 1:
1292
+ q = q[None, :]
1293
+ idx = index if index_normalized else (
1294
+ index / np.maximum(np.linalg.norm(index, axis=1, keepdims=True), 1e-12))
1295
+ sims = q @ idx.T
1296
+ k = max(1, min(int(top_k), index.shape[0]))
1297
+ part = np.argpartition(-sims, k - 1, axis=1)[:, :k]
1298
+ row = np.take_along_axis(sims, part, axis=1)
1299
+ order = np.argsort(-row, axis=1)
1300
+ idx_out = np.take_along_axis(part, order, axis=1)
1301
+ sco_out = np.take_along_axis(row, order, axis=1)
1302
+ return sco_out, idx_out
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d2e24e22ff8f952811c443d56ce613e3933fb955a34a03edb1ded71bfe7b9a42
3
+ size 393472
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff