SKILLRET-Edge-22M-int4

A 22.7M-parameter bi-encoder for agent skill retrieval — picking the right skill out of a catalogue for a natural-language request. Distilled from ThakiCloud/SKILLRET-Embedding-0.6B and small enough to run on a CPU next to the agent.

Scored on the public ThakiCloud/SKILLRET test split (4,392 queries / 6,006 skills), the metric the SkillRet paper uses as its headline.

NDCG@10 75.14 · 17.0 MB packed

Correction — September 2026. Earlier versions of this card reported the SKILLRET-Embedding-0.6B teacher as 80.82 NDCG@10 on the current 4,392-query / 6,006-skill evaluation split. That score was measured before the train/evaluation query-prefix contract was corrected. Re-evaluation on ThakiCloud/SKILLRET revision a050ad2 under the canonical query contract gives 78.48 NDCG@10. The model weights are unchanged. The earlier 605 MB size was also not the serialized artifact size; the distributed BF16 checkpoint is 1191.6 MB (decimal MB). We retain this note so results quoted from earlier versions of the card can be interpreted correctly.

Results

Variant Size on disk NDCG@10 vs teacher
SKILLRET-Embedding-0.6B (teacher) 1191.6 MB 78.48
fp16 45.4 MB 75.26 ± 0.45 93.1%
int8 / g16 28.4 MB 75.27 ± 0.45 93.1%
int4 / g16 17.0 MB 75.14 ± 0.45 93.0%
int3 / g16 14.2 MB 73.82 ± 0.46 91.3%
untrained base 90.9 MB 50.07 ± 0.58 61.9%

int4 costs 0.12pp against fp16 — a paired t of 1.15, i.e. a statistical tie — while cutting the file 2.7x.

CPU latency

Machine single query p50 batch-32 per item
Apple M4 Pro, 4 threads 6.5 ms 5.11 ms
AMD EPYC 9355, 4 threads 3.1 ms 1.25 ms
0.6B teacher, same bench 518–650 ms

⚠️ Two different machines, so this is not an isolated ISA comparison — memory and clock differ too. fp32 arithmetic, weights-only quantization.

Usage

from sentence_transformers import SentenceTransformer
m = SentenceTransformer("ThakiCloud/SKILLRET-Edge-22M-int4")

# ⛔ Encode queries BARE — no instruction prefix. See query_prefix.json.
q = m.encode(["I need to convert a spreadsheet into a chart"], normalize_embeddings=True)
d = m.encode(["Chart Builder — turn tabular data into bar/line charts"], normalize_embeddings=True)
print(q @ d.T)

What "int4 / g16" means here

Weights are quantized group-wise, asymmetric min/max, group size 16, with scales and zero-points stored in fp16 — the same shape as a GGUF Q4_K/Q8_0 block.

This repo ships two things, deliberately:

File What it is
model.safetensors The quantized values, de-quantized back into fp16 so any transformers / sentence-transformers install loads it today. This is the file that produces the score above.
model-int4-g16.bin The actual packed payload — 17.012 MB. This is what the size claim means.

Be clear about what this does and does not buy you. The packed file is the real storage footprint and it is verified: predicted 17.074 MB vs 17.012 MB on disk (0.36% error). But the safetensors path runs fp32 arithmetic — the latency figures below were measured that way, and they are not integer-kernel speeds. Routing the packed weights through a real INT kernel (llama.cpp, ONNX Runtime) is unmeasured and is where further speedup would come from.

quantization.json carries the per-tensor layout if you want to write that kernel.

How it was trained

Knowledge distillation from ThakiCloud/SKILLRET-Embedding-0.6B (NDCG@10 78.48 on the same split), kd_weight=0.7, multi-positive InfoNCE over 1–3 golds per query, 12 epochs, cosine schedule. The epoch was chosen on a skill-disjoint holdout — the test split was never used for selection.

⚠️ Read this before you compare numbers

Use the same query prefix at train and eval time. This repo ships query_prefix.json recording the contract (resolved: "", i.e. queries are encoded bare, with no instruction prefix). Which prefix you pick barely matters after fine-tuning — none 79.18 vs an instruction prefix 78.50, inside the standard error — but keeping it consistent matters a great deal. We measured a +8.84pp swing on a single checkpoint purely from a train/eval prefix mismatch, and that mismatch also manufactured a fake result in which quantization appeared to beat fp16. It does not.

Other caveats worth stating plainly:

  • The model card of the SkillRet reference models reports a 4,997-query / 6,660-skill split. That split is not in the currently published dataset — we verified the public files are hash-identical to ours at 4,392 / 6,006. Do not convert between the two.
  • Standard error on this split is about ±0.45. Differences under ~1pp are not rankings.
  • Pooling is CLS, embeddings are L2-normalized, max_length=256 at evaluation.

What did NOT work (so you don't repeat it)

Attempt Result
Distilling from the 8B teacher instead of 0.6B −2.40pp (capacity gap)
INT2 / ternary post-training quantization Total collapse (~0.1 NDCG@10) — GPTQ and QuIP do not rescue it
Hard-negative mining (our mined set) −1.0pp
LEAF auxiliary projection loss (w=0.3) −1.51pp (paired t = −7.78)

License

Apache-2.0, inherited from the base model.

Citation

Benchmark: SkillRet, arXiv:2605.05726

Downloads last month
41
Safetensors
Model size
22.7M params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ThakiCloud/SKILLRET-Edge-22M-int4

Quantized
(2)
this model

Dataset used to train ThakiCloud/SKILLRET-Edge-22M-int4

Collection including ThakiCloud/SKILLRET-Edge-22M-int4

Paper for ThakiCloud/SKILLRET-Edge-22M-int4