Sentence Similarity
Transformers
Safetensors
English
Arabic
Urdu
mentee_embed
feature-extraction
embeddings
retrieval
contrastive-learning
multilingual
from-scratch
custom_code
Instructions to use MenteEAI/mentee-embed-v4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MenteEAI/mentee-embed-v4 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("MenteEAI/mentee-embed-v4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 14,109 Bytes
52fea85 4aaa9af c438e94 4aaa9af 26f9e10 4aaa9af 26f9e10 52fea85 c438e94 4aaa9af c438e94 4aaa9af c438e94 4aaa9af c438e94 26f9e10 c438e94 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af 26f9e10 4aaa9af c438e94 4aaa9af c438e94 4aaa9af c438e94 4aaa9af c438e94 4aaa9af 0495079 4aaa9af 0495079 4aaa9af | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 | ---
license: apache-2.0
pipeline_tag: sentence-similarity
library_name: transformers
language:
- en
- ar
- ur
tags:
- embeddings
- sentence-similarity
- retrieval
- contrastive-learning
- multilingual
- from-scratch
datasets:
- sentence-transformers/all-NLI
- facebook/xnli
- Helsinki-NLP/opus-100
- sentence-transformers/msmarco-bm25
- nlpai-lab/miracl-multilingual-triplets
metrics:
- mrr@10
- ndcg@10
- spearman_correlation
- recall@10
---
# mentee-embed-v4
**A 41M-parameter trilingual text embedding model trained entirely from scratch — no pretrained backbone, no BERT, no RoBERTa. Random initialization only.**
v4 adds mMARCO Arabic retrieval data (500K triplets) and a third distillation round with hard negatives, yielding significant improvements over v3 across all languages and benchmarks.
Developed by **Syed Syab Ahmad Shah** and **Team MenteE AI** (menteeai.org) as part of an ongoing research effort to build competitive multilingual embeddings from the ground up for Arabic, English, and Urdu.
---
## Key Facts
| Property | Value |
|----------|-------|
| Parameters | **41M** |
| Embedding dimension | 384 |
| Max sequence length | 128 tokens |
| Languages | Arabic · English · Urdu |
| Architecture | 12-layer Transformer, custom BPE tokenizer (50K vocab) |
| Initialization | **Random** (trained from scratch) |
| Training data | ~2.6M triplets (NLI + MS-MARCO + OPUS + MIRACL + mMARCO Arabic) |
| Pooling | Mean pooling |
| Training objective | Relational distillation + InfoNCE contrastive (teacher: `multilingual-e5-base`) |
---
## What "From Scratch" Means
Most embedding models fine-tune an existing pretrained encoder (BERT, RoBERTa, MPNet). **mentee-embed-v4 does not.** We:
1. Trained a **custom BPE tokenizer** on Arabic, English, and Urdu text
2. Initialized a **12-layer Transformer with random weights**
3. Ran **masked language modeling** pretraining (50K steps on trilingual corpus)
4. Applied **three-round contrastive distillation** with hard negative mining
No pretrained checkpoint was used at any stage.
---
## v4 vs v3 Improvements
| Benchmark | v3 | v4 | Change |
|-----------|----|----|--------|
| Bench MRR@10 (all) | 0.103 | **0.252** | +146% |
| MIRACL EN MRR@10 | 0.768 | **0.916** | +19% |
| MIRACL AR MRR@10 | 0.447 | **0.874** | +96% |
| MIRACL UR MRR@10 | 0.354 | **0.572** | +62% |
| MS-MARCO MRR@10 | 0.592 | **0.706** | +19% |
| STS-B Spearman | 0.683 | **0.725** | +6% |
Key improvements come from:
- **mMARCO Arabic** — 500K Arabic retrieval triplets added to training
- **50K MLM steps** — 6x more pretraining than v3 (8K steps)
- **3 distillation rounds** — extra hard-negative round for better discrimination
---
## Full Benchmark Results
All baselines evaluated under identical conditions on the same hardware (RTX 5090).
### MenteE Embed Bench — Custom EN/AR/UR Pilot
*121 queries · 18 domains · dialect Arabic · Roman Urdu · hard negatives*
| Model | Params | EN MRR@10 | AR MRR@10 | UR MRR@10 | Dialect AR | Roman UR | All MRR@10 |
|---|---|---|---|---|---|---|---|
| **mentee-embed-v4** | **41M** | 0.369 | 0.120 | 0.261 | 0.095 | 0.399 | **0.252** |
| paraphrase-mpnet-base-v2 | 278M | 0.764 | 0.602 | 0.611 | 0.445 | 0.661 | 0.661 |
| paraphrase-MiniLM-L12-v2 | 118M | 0.682 | 0.568 | 0.519 | 0.385 | 0.480 | 0.592 |
| multilingual-e5-base | 278M | 0.667 | 0.352 | 0.413 | 0.310 | 0.430 | 0.481 |
| all-MiniLM-L6-v2 | 23M | 0.873 | 0.087 | 0.289 | 0.084 | 0.486 | 0.425 |

#### Per-Language Breakdown

| Metric | v4 | mpnet | e5-base | MiniLM-L6 |
|--------|-----|-------|---------|-----------|
| English MRR@10 | 0.369 | 0.764 | 0.667 | 0.873 |
| Arabic MRR@10 | 0.120 | 0.602 | 0.352 | 0.087 |
| Urdu MRR@10 | 0.261 | 0.611 | 0.413 | 0.289 |
| Dialect AR MRR@10 | 0.095 | 0.445 | 0.310 | 0.084 |
| Roman UR MRR@10 | 0.399 | 0.661 | 0.430 | 0.486 |
#### Detailed Metrics (NDCG@10, Recall@1, Recall@5, Recall@10)
| Model | All NDCG | All R@1 | All R@5 | All R@10 | EN NDCG | AR NDCG | UR NDCG |
|---|---|---|---|---|---|---|---|
| **mentee-embed-v4** | 0.238 | 0.083 | 0.203 | 0.273 | 0.331 | 0.166 | 0.212 |
| mpnet-base | 0.650 | 0.438 | 0.620 | 0.736 | 0.757 | 0.589 | 0.595 |
| MiniLM-L12 | 0.580 | 0.355 | 0.562 | 0.661 | 0.674 | 0.556 | 0.504 |
| e5-base | 0.468 | 0.248 | 0.446 | 0.549 | 0.653 | 0.341 | 0.398 |
| MiniLM-L6 | 0.412 | 0.231 | 0.376 | 0.483 | 0.868 | 0.104 | 0.266 |
#### Per-Domain MRR@10 (v4)

| Domain | MRR@10 | Domain | MRR@10 |
|---|---|---|---|
| Travel | 0.643 | Fashion | 0.200 |
| Hotels | 0.378 | History | 0.195 |
| Restaurants | 0.357 | Sports | 0.179 |
| Technology | 0.333 | Health | 0.164 |
| Finance | 0.286 | Entertainment | 0.143 |
| Automotive | 0.222 | Social Media | 0.029 |
---
### Protocol A — In-batch Retrieval (MIRACL)
*acc@1 / MRR@10*
| Model | MIRACL-EN | MIRACL-AR | MIRACL-UR |
|-------|-----------|-----------|-----------|
| **mentee-embed-v4** | 0.870 / **0.916** | 0.825 / **0.874** | 0.475 / **0.572** |
| paraphrase-mpnet-base-v2 | 0.965 / 0.982 | 0.860 / 0.898 | 0.755 / 0.824 |
| paraphrase-MiniLM-L12-v2 | 0.965 / 0.979 | 0.815 / 0.868 | 0.720 / 0.785 |
| multilingual-e5-base | 0.980 / 0.990 | 0.940 / 0.958 | 0.960 / 0.970 |
| all-MiniLM-L6-v2 | 0.980 / 0.990 | 0.005 / 0.027 | 0.015 / 0.031 |

> v4 achieves strong MIRACL scores despite being 6.8x smaller than mpnet-base. MIRACL AR (0.874) is within 3% of mpnet-base (0.898).
---
### Protocol C — MS-MARCO Corpus Retrieval (10K passages)
| Model | Params | MRR@10 | R@5 | R@100 |
|-------|--------|--------|-----|-------|
| all-MiniLM-L6-v2 | 23M | 0.877 | 0.992 | 0.992 |
| multilingual-e5-base | 278M | 0.856 | 0.992 | 0.996 |
| paraphrase-mpnet-base-v2 | 278M | 0.803 | 0.964 | 0.988 |
| **mentee-embed-v4** | **41M** | **0.706** | **0.913** | **0.980** |

---
### STS-B (Semantic Textual Similarity)
| Model | Spearman | Pearson | MAE | RMSE |
|-------|----------|---------|-----|------|
| paraphrase-mpnet-base-v2 | 0.868 | 0.870 | — | — |
| bge-small-en-v1.5 | 0.859 | — | — | — |
| paraphrase-MiniLM-L12-v2 | 0.844 | 0.849 | — | — |
| multilingual-e5-base | 0.842 | 0.847 | — | — |
| **mentee-embed-v4** | **0.725** | **0.730** | **2.059** | **2.458** |

---
## Speed & Efficiency
Measured on RTX 5090.
| Model | Params | Sents/sec (bs=128) | Latency@1 (ms) | VRAM (MB) | Cost/1B sents |
|---|---|---|---|---|---|
| **mentee-embed-v4** | **41M** | **18,115** | **5.15** | **175** | **$7.15** |
| all-MiniLM-L6-v2 | 23M | 12,445 | 4.71 | 103 | $10.40 |
| multilingual-e5-small | 118M | 9,749 | 7.88 | 482 | $13.28 |
| paraphrase-MiniLM-L12-v2 | 118M | 9,456 | 8.17 | 573 | $13.69 |
| multilingual-e5-base | 278M | 5,379 | 8.07 | 1126 | $24.07 |
| paraphrase-mpnet-base-v2 | 278M | 5,158 | 8.04 | 1126 | $25.10 |

> v4 is the **fastest model tested** — 3.5x faster than mpnet-base, 45% faster than MiniLM-L12.
### Batch Size Scaling
| Batch Size | Sents/sec | Latency (ms) | Peak VRAM (MB) |
|---|---|---|---|
| 1 | 145 | 6.90 | 174 |
| 8 | 902 | 8.87 | 177 |
| 16 | 1,707 | 9.37 | 181 |
| 32 | 3,424 | 9.35 | 188 |
| 64 | 6,470 | 9.89 | 202 |
| 128 | 12,418 | 10.31 | 230 |
| 256 | 20,546 | 12.46 | 287 |
| 512 | 22,139 | 23.13 | 400 |
### Precision Formats
| Format | Sents/sec (bs=128) | STS-B Spearman | Delta vs FP32 | VRAM (MB) |
|---|---|---|---|---|
| FP32 | 12,500 | 0.7249 | — | 289 |
| FP16 | 3,632 | 0.7249 | 0.0 | 149 |
| BF16 | 2,837 | 0.7248 | -0.0001 | 149 |
### Efficiency Metrics (Per Parameter)
| Model | Bench MRR/1M params | Sents/sec per 1M params | Cost per 1B sents |
|---|---|---|---|
| **mentee-embed-v4** | **0.525** | **441.82** | **$7.15** |
| all-MiniLM-L6-v2 | 0.344 | 541.09 | $10.40 |
| paraphrase-MiniLM-L12-v2 | 0.452 | 80.34 | $13.69 |
| multilingual-e5-base | 0.166 | 19.35 | $24.07 |
| paraphrase-mpnet-base-v2 | 0.224 | 18.55 | $25.10 |


---
## Radar — v4 vs mpnet-base
Normalized comparison across all evaluation dimensions:

| Dimension | v4 (41M) | mpnet (278M) | v4 as % of mpnet |
|---|---|---|---|
| English MRR@10 | 0.369 | 0.764 | 48% |
| Arabic MRR@10 | 0.120 | 0.602 | 20% |
| Urdu MRR@10 | 0.261 | 0.611 | 43% |
| MIRACL EN | 0.916 | 0.982 | 93% |
| MIRACL AR | 0.874 | 0.898 | 97% |
| MIRACL UR | 0.572 | 0.824 | 69% |
| MS-MARCO | 0.706 | 0.803 | 88% |
| STS-B Spearman | 0.725 | 0.868 | 84% |
| Throughput | 18,115 | 5,158 | 351% |
| Cost efficiency | $7.15 | $25.10 | 351% |
---
## Honest Assessment
### Where v4 Wins
- **Speed** — fastest embedding model tested (18,115 sents/sec)
- **Cost** — cheapest per billion sentences ($7.15)
- **MIRACL AR** — 0.874 vs mpnet's 0.898 (within 3%) despite 6.8x fewer params
- **MS-MARCO** — 0.706, competitive with models 3-7x larger
- **Efficiency** — best MRR per parameter of all models tested
### Where v4 Loses
- **Custom bench** — 0.252 vs mpnet's 0.661 (2.4x gap)
- **English** — 0.369 vs MiniLM-L6's 0.873
- **STS-B** — 0.725 vs mpnet's 0.868 (16% gap)
- **Arabic dialect** — 0.095 (struggles with informal Arabic)
- **Urdu** — 0.261 vs mpnet's 0.611
### Verdict
v4 is a **strong proof-of-concept** — it proves you can train a competitive embedding model from scratch. The MIRACL AR score (0.874) is within 3% of mpnet-base (0.898) despite being 6.8x smaller.
For production use, v4 works well for:
- Arabic retrieval (MIRACL-scale tasks)
- High-throughput, cost-sensitive applications
- Resource-constrained environments (175 MB VRAM)
v4 is not yet suitable for:
- Domain-specific English retrieval
- Dialectal Arabic / informal Urdu
- Tasks requiring STS-level semantic similarity
---
## Training Data
| Source | Language | Triplets | Type |
|--------|----------|----------|------|
| all-NLI (sentence-transformers) | EN | 558K | NLI triplets |
| XNLI | AR | 128K | NLI triplets |
| XNLI | UR | 125K | NLI triplets |
| OPUS-100 EN-UR | EN/UR | 300K | Parallel translation |
| OPUS-100 AR-EN | AR/EN | 300K | Parallel translation |
| MS-MARCO BM25 triplets | EN | 500K | Passage retrieval |
| MS-MARCO hard negatives | EN | 200K | Hard retrieval |
| mMARCO Arabic | AR | 500K | Passage retrieval |
| MIRACL | EN/AR/UR | ~9K | Wikipedia retrieval |
| **Total** | | **~2.6M** | |
---
## Training Pipeline
```
Stage 1 — MLM Pretraining (50K steps)
Random init → masked language modeling on trilingual corpus
50,000 steps · batch=64 · vocab=50K BPE
Stage 2 — Teacher Encoding
Teacher: intfloat/multilingual-e5-base (768-dim)
Encode all 2.6M triplets to soft labels
Stage 3 — Distillation Round 1 (no hard negatives)
InfoNCE contrastive + relational distillation
4,000 steps · batch=512 · temp=0.05
Stage 4 — Hard Negative Mining Round 1
GPU-accelerated top-5 mining across full 2.6M corpus
Stage 5 — Distillation Round 2 (with hard negatives)
Same objective + mined hard negatives per anchor
10,000 steps · batch=512 · temp=0.05
Stage 6 — Hard Negative Mining Round 2 (on improved model)
Re-mine with the better round-2 model
Stage 7 — Distillation Round 3 (final polish)
Train on round-2 mined negatives
15,000 cumulative steps · batch=512 · temp=0.05
```
---
## Usage
```python
# pip install torch transformers tokenizers huggingface_hub
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MenteEAI/mentee-embed-v4", trust_remote_code=True)
model = AutoModel.from_pretrained("MenteEAI/mentee-embed-v4", trust_remote_code=True)
sentences = [
"Hello, how are you?",
"مرحبا، كيف حالك؟",
"ہیلو، آپ کیسے ہیں؟"
]
embeddings = model.encode(sentences, tokenizer=tok)
print(embeddings.shape) # torch.Size([3, 384])
```
> `trust_remote_code=True` is required — standard for custom-architecture models on HuggingFace. The code runs entirely on your machine.
### Similarity search
```python
from transformers import AutoModel, AutoTokenizer
tok = AutoTokenizer.from_pretrained("MenteEAI/mentee-embed-v4", trust_remote_code=True)
model = AutoModel.from_pretrained("MenteEAI/mentee-embed-v4", trust_remote_code=True)
query = model.encode(["What is machine learning?"], tokenizer=tok)
passages = model.encode([
"Machine learning is a subset of artificial intelligence.",
"The weather today is sunny.",
"تعلم الآلة هو فرع من فروع الذكاء الاصطناعي.",
], tokenizer=tok)
scores = query @ passages.T
print(scores) # tensor([[0.81, 0.60, 0.79]])
```
---
## Limitations
- **Custom benchmark gap** — v4 scores 0.252 vs 0.661 for mpnet-base on the custom bench; the gap reflects the difference between 2.6M training pairs and web-scale data
- **Arabic and Urdu** lag behind English on the custom bench, though MIRACL scores are strong (AR 0.874, UR 0.572)
- Vocabulary limited to 50K tokens trained on ~2.6M sentences — rare scripts and dialects may tokenize poorly
- Not evaluated on MTEB full suite yet
---
## Citation
```bibtex
@misc{mentee-embed-v4-2026,
title = {How Far Can Multilingual Text Embeddings Be Trained From Scratch?
A Compute-Efficient Study of Arabic, English, and Urdu},
author = {Shah, Syed Syab Ahmad and Team MenteE AI},
year = {2026},
url = {https://huggingface.co/MenteEAI/mentee-embed-v4},
note = {MenteE AI. Apache-2.0 License}
}
```
---
## About MenteE AI
Built by **Syed Syab Ahmad Shah** and Team MenteE AI.
[menteeai.org](https://menteeai.org) · syab@menteeai.org
|