File size: 6,898 Bytes
2a66e90
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
---
license: apache-2.0
library_name: sentence-transformers
pipeline_tag: feature-extraction
base_model: ThakiCloud/SKILLRET-Edge-22M
base_model_relation: quantized
datasets:
  - ThakiCloud/SKILLRET
language:
  - en
tags:
  - sentence-transformers
  - feature-extraction
  - retrieval
  - skill-retrieval
  - agent
  - on-device
  - mixed-precision
  - quantization
---

# SKILLRET-Edge-22M-drift-mixed

The **22.7M-parameter** SKILLRET-Edge bi-encoder for **agent skill retrieval**, quantized to a
**mixed-precision** layout chosen by a **label-free drift signal** instead of one uniform bit width.
Same architecture, same tokenizer, same query contract as `ThakiCloud/SKILLRET-Edge-22M-int4` β€” this
repo exists to show a *different allocation technique*, not to replace the int4/int3 variants.

Scored on the public `ThakiCloud/SKILLRET` **test split** (4,392 queries / 6,006 skills).

**NDCG@10 74.57** Β· **14.1 MB packed** (the int3 file size, 99.1% of fp16)

## Where it sits

| Variant | Packed size | NDCG@10 | Note |
|---|---|---|---|
| fp16 | 45.4 MB | 75.25 | reference |
| int4 / g16 (uniform) | 17.0 MB | 75.14 | best quality per the int4 card |
| **drift-mixed / g16 (this repo)** | **14.1 MB** | **74.57** | int3 bytes, +0.65pp over int3 |
| int3 / g16 (uniform) | 14.2 MB | 73.82 | |
| int2 / g16 (uniform) | 11.9 MB | 59.73 | 4.18 real bits β€” no size win over int4 g128 |

Paired query-level bootstrap on the same 4,392 queries (5,000 resamples), all arms built and scored
under one protocol (int3 uniform comes from the 5.2-bit run of the same series):

| Contrast | Ξ” NDCG@10 [95% CI] |
|---|---|
| drift-mixed (14.1 MB) βˆ’ uniform int3 g16 (14.65 MB, same protocol) | +0.65pp [+0.35, +0.96] |
| drift-mixed at the *same* 5.2-real-bit budget βˆ’ uniform int3 | +0.98pp [+0.64, +1.33] |
| drift-mixed βˆ’ uniform int2 g16 | +14.83pp [+14.08, +15.59] |
| drift-mixed βˆ’ label-supervised allocation (per-module Ξ”NDCG, same budget) | +0.71pp [+0.43, +0.99] |
| drift-mixed βˆ’ 5 random allocations at the same bytes | +7.7 to +9.6pp (every CI excludes 0) |

Read the first row honestly: the gain over uniform int3 is real but under one point. The value of the
technique is the **ordering** (random allocations at the same bytes lose 8–9 points), and the budgets
uniform cannot express β€” at 4.5 real bits (~12.7 MB) the drift allocation keeps **70.42**
where uniform int2 gives 59.73 and uniform int3 does not fit.

## How the bits were allocated

Each of 12 allocation units was quantized **alone** at 2/3/4 bits while the rest stayed fp32, and the
resulting *drift* of the output embeddings β€” mean (1 βˆ’ cosine) between full-precision and quantized
embeddings, averaged over 1,500 calibration skills and 150 calibration queries, balanced across the two
β€” was recorded. No relevance labels are used. An exhaustive integer program then picks the per-unit
bit width that minimises predicted drift under a hard byte budget (fp16 scale + zero per group of 16
counted). The drift ordering reproduced across disjoint calibration halves at every size tried
(Spearman 1.000 at 75/150/300/600 queries); the label-based ordering on this model did **not**
(βˆ’0.29 to +0.34), which is why the label-supervised row above is reported but not recommended.

| Unit | What it covers | Params | Bits |
|---|---|---|---|
| `embedding.v0` | word embedding rows 0–25% | 2.93M | **3** |
| `embedding.v1` | rows 25–50% | 2.93M | **2** |
| `embedding.v2` | rows 50–75% | 2.93M | **2** |
| `embedding.v3` | rows 75–100% | 2.93M | **2** |
| `attention.b0` | attention, layers 0–1 | 1.18M | **3** |
| `attention.b1` | attention, layer 2 | 0.59M | **4** |
| `attention.b2` | attention, layers 3–4 | 1.18M | **4** |
| `attention.b3` | attention, layer 5 | 0.59M | **4** |
| `ffn.b0` | FFN, layers 0–1 | 2.36M | **3** |
| `ffn.b1` | FFN, layer 2 | 1.18M | **3** |
| `ffn.b2` | FFN, layers 3–4 | 2.36M | **4** |
| `ffn.b3` | FFN, layer 5 | 1.18M | **3** |

The picture is the one you would guess after seeing the sensitivities: the word-embedding table is
51.6% of the parameters and its lower-frequency rows barely move the output at 2 bits, while the
middle attention/FFN blocks are where 2-bit damage concentrates.

## Files

| File | What it is |
|---|---|
| `model.safetensors` | The quantized values **de-quantized back into fp16** so `sentence-transformers` loads it today. This file produces the score above (re-evaluated after packing: 74.569 vs the ledger's 74.559). |
| `model-drift-mixed-g16.bin` | The **actual packed payload β€” 14.140 MB**: per-unit bit-packed indices + fp16 scale/zero per 16 weights, plus the two small fp16 position/token-type tables. Predicted 14.140 MB, 0.0% error. |
| `quantization.json` | The plan (`plan`), per-unit parameter counts, and the per-tensor layout (name, rows slice, bits, offset) if you want to write a kernel. |
| `query_prefix.json` | `resolved: ""` β€” encode queries **bare**, no instruction prefix. |

The safetensors path runs fp32 arithmetic; latency is identical to the int4/int3 cards (same
architecture, same dequantized format) and is *not* an integer-kernel speed.

## Usage

```python
from sentence_transformers import SentenceTransformer
m = SentenceTransformer("ThakiCloud/SKILLRET-Edge-22M-drift-mixed")
q = m.encode(["I need to convert a spreadsheet into a chart"], normalize_embeddings=True)
d = m.encode(["Chart Builder β€” turn tabular data into bar/line charts"], normalize_embeddings=True)
print(q @ d.T)
```

## Caveats

- Standard error on this split is about Β±0.45; the drift-vs-int3 gap is significant only because it is
  paired on the same queries. Do not read it as a ranking between cards scored in different runs.
- The protocol matches the other Edge cards: CLS pooling, L2-normalised, `max_length=256`, bare queries.
- Mixed precision means a real integer kernel has to handle per-tensor (and, for the embedding table,
  per-row-band) bit widths. That is the cost you pay for +0.65pp at int3 bytes.
- The base student, the distillation recipe and the int4/int3 measurements are documented on
  `ThakiCloud/SKILLRET-Edge-22M-int4`.

## License

Apache-2.0, inherited from the base model.

## Citation

Benchmark: [SkillRet, arXiv:2605.05726](https://arxiv.org/abs/2605.05726)

Student and uniform-PTQ measurements:
[Where Post-Training Quantization Breaks Text Embedders, arXiv:2609.16391](https://arxiv.org/abs/2609.16391)

Allocation signal (the drift criterion and the allocator used here):
*Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders*, Han, 2026 β€”
arXiv preprint, identifier to be added on announcement.

```bibtex
@article{han2026quantizebydrift,
  title   = {Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders},
  author  = {Han, Hyojung},
  journal = {arXiv preprint},
  year    = {2026}
}
```