--- license: apache-2.0 library_name: sentence-transformers pipeline_tag: feature-extraction base_model: ThakiCloud/SKILLRET-Edge-109M base_model_relation: quantized datasets: - ThakiCloud/SKILLRET language: - en tags: - sentence-transformers - feature-extraction - retrieval - skill-retrieval - agent - on-device --- # SKILLRET-Edge-109M-int3 A **109.5M-parameter** bi-encoder for **agent skill retrieval** — picking the right skill out of a catalogue for a natural-language request. Distilled from `ThakiCloud/SKILLRET-Embedding-0.6B` and small enough to run on a CPU next to the agent. Scored on the public `ThakiCloud/SKILLRET` **test split** (4,392 queries / 6,006 skills), the metric the SkillRet paper uses as its headline. **NDCG@10 78.04** · **68.4 MB packed** > **Correction — September 2026.** Earlier versions of this card reported the SKILLRET-Embedding-0.6B teacher as 80.82 NDCG@10 on the current 4,392-query / 6,006-skill evaluation split. That score was measured before the train/evaluation query-prefix contract was corrected. Re-evaluation on `ThakiCloud/SKILLRET` revision `a050ad2` under the canonical query contract gives **78.48** NDCG@10. The model weights are unchanged. The earlier 605 MB size was also not the serialized artifact size; the distributed BF16 checkpoint is 1191.6 MB (decimal MB). We retain this note so results quoted from earlier versions of the card can be interpreted correctly. ## Results | Variant | Size on disk | NDCG@10 | vs teacher | |---|---|---|---| | SKILLRET-Embedding-0.6B (teacher) | 1191.6 MB | 78.48 | — | | fp16 | 219.0 MB | **79.18** ± 0.42 | 98.0% | | int8 / g16 | 136.9 MB | 79.18 ± 0.42 | 98.0% | | int4 / g16 | 82.0 MB | 79.21 ± 0.42 | 98.0% | | **int3 / g16** | **68.4 MB** | **78.04** ± 0.44 | 96.6% | This model recovers **98% of a 26x larger teacher**. int8 and int4 are free; int3 costs 1.14pp and is the point where quantization starts to be a real trade. ## Usage ```python from sentence_transformers import SentenceTransformer m = SentenceTransformer("ThakiCloud/SKILLRET-Edge-109M-int3") # ⛔ Encode queries BARE — no instruction prefix. See query_prefix.json. q = m.encode(["I need to convert a spreadsheet into a chart"], normalize_embeddings=True) d = m.encode(["Chart Builder — turn tabular data into bar/line charts"], normalize_embeddings=True) print(q @ d.T) ``` ## What "int3 / g16" means here Weights are quantized **group-wise, asymmetric min/max**, group size 16, with scales and zero-points stored in **fp16** — the same shape as a GGUF Q4_K/Q8_0 block. This repo ships two things, deliberately: | File | What it is | |---|---| | `model.safetensors` | The quantized values, **de-quantized back into fp16** so any `transformers` / `sentence-transformers` install loads it today. This is the file that produces the score above. | | `model-int3-g16.bin` | The **actual packed payload — 68.350 MB**. This is what the size claim means. | **Be clear about what this does and does not buy you.** The packed file is the real storage footprint and it is verified: predicted 68.594 MB vs 68.350 MB on disk (0.36% error). But the safetensors path runs **fp32 arithmetic** — the latency figures below were measured that way, and they are *not* integer-kernel speeds. Routing the packed weights through a real INT kernel (llama.cpp, ONNX Runtime) is unmeasured and is where further speedup would come from. `quantization.json` carries the per-tensor layout if you want to write that kernel. ## How it was trained Knowledge distillation from `ThakiCloud/SKILLRET-Embedding-0.6B` (NDCG@10 78.48 on the same split), `kd_weight=0.7`, multi-positive InfoNCE over 1–3 golds per query, 12 epochs, cosine schedule. The epoch was chosen on a **skill-disjoint holdout** — the test split was never used for selection. ## ⚠️ Read this before you compare numbers **Use the same query prefix at train and eval time.** This repo ships `query_prefix.json` recording the contract (`resolved: ""`, i.e. queries are encoded bare, with no instruction prefix). Which prefix you pick barely matters after fine-tuning — none 79.18 vs an instruction prefix 78.50, inside the standard error — but **keeping it consistent matters a great deal.** We measured a **+8.84pp** swing on a single checkpoint purely from a train/eval prefix mismatch, and that mismatch also manufactured a fake result in which quantization appeared to *beat* fp16. It does not. Other caveats worth stating plainly: - The model card of the SkillRet reference models reports a 4,997-query / 6,660-skill split. That split is **not in the currently published dataset** — we verified the public files are hash-identical to ours at 4,392 / 6,006. Do not convert between the two. - Standard error on this split is about ±0.45. Differences under ~1pp are not rankings. - Pooling is CLS, embeddings are L2-normalized, `max_length=256` at evaluation. ## What did NOT work (so you don't repeat it) | Attempt | Result | |---|---| | Distilling from the **8B** teacher instead of 0.6B | **−2.40pp** (capacity gap) | | INT2 / ternary post-training quantization | Total collapse (~0.1 NDCG@10) — GPTQ and QuIP do not rescue it | | Hard-negative mining (our mined set) | −1.0pp | | LEAF auxiliary projection loss (w=0.3) | −1.51pp (paired t = −7.78) | ## License Apache-2.0, inherited from the base model. ## Citation Benchmark: [SkillRet, arXiv:2605.05726](https://arxiv.org/abs/2605.05726) Paper (measurements behind the quantized variants and the student frontier): [Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families, arXiv:2609.16391](https://arxiv.org/abs/2609.16391) ```bibtex @article{han2026ptqembedders, title = {Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families}, author = {Han, Hyojung}, journal = {arXiv preprint arXiv:2609.16391}, year = {2026} } ```