Clarify complete benchmark ranks and size–quality comparisons
Browse filesAdd score gaps and model-specific benchmark figures. Preserve released weights, evaluation scores and original-model comparisons.
- .gitattributes +2 -0
- README.md +7 -7
- assets/general-audio19-nano.png +3 -0
- assets/general-audio19-nano.svg +912 -0
- assets/general-english41-nano.png +3 -0
- assets/general-english41-nano.svg +1062 -0
- benchmarks/complete-panel-ranks.md +9 -9
- benchmarks/general-rank-data.json +0 -0
- benchmarks/general-rank-figures.json +373 -0
- benchmarks/pareto-gallery.md +4 -9
- benchmarks/pareto-methodology.md +2 -2
.gitattributes
CHANGED
|
@@ -40,3 +40,5 @@ assets/pareto-vehiclesoundclustering.png filter=lfs diff=lfs merge=lfs -text
|
|
| 40 |
assets/complete-panel-audio19-mini.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
assets/complete-panel-english41-nano.png filter=lfs diff=lfs merge=lfs -text
|
| 42 |
assets/pareto-imdb.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
|
|
|
| 40 |
assets/complete-panel-audio19-mini.png filter=lfs diff=lfs merge=lfs -text
|
| 41 |
assets/complete-panel-english41-nano.png filter=lfs diff=lfs merge=lfs -text
|
| 42 |
assets/pareto-imdb.png filter=lfs diff=lfs merge=lfs -text
|
| 43 |
+
assets/general-audio19-nano.png filter=lfs diff=lfs merge=lfs -text
|
| 44 |
+
assets/general-english41-nano.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -59,10 +59,10 @@ The common protocol uses labeled TRAIN prototypes for text classification and al
|
|
| 59 |
|
| 60 |
The primary metric, **Mean(TaskType)**, weights each task type equally. Mean(Task) weights individual tasks equally and is supplementary. Rankings use complete benchmark results from the September 17, 2026 registry snapshots plus both current Vela models; “≤ size” counts models with no more total parameters than this model.
|
| 61 |
|
| 62 |
-
| Benchmark | Mean(TaskType) | Global rank | Rank at ≤ size | Mean(Task) |
|
| 63 |
-
| --- | ---: | ---: | ---: | ---: |
|
| 64 |
-
| MTEB English v2 · 41 tasks | 60.78 | 65/188 | 5/75 | 64.88 |
|
| 65 |
-
| MAEB audio-only · 19 tasks | 52.34 | 19/64 | 6/27 | 43.59 |
|
| 66 |
|
| 67 |
Nano ranks **5/75** on the English panel and **6/27** on the audio panel among models with no more than its 163.8M total parameters. These are snapshot-relative comparisons across reported protocols, including single-modality specialists. The original small has not been evaluated on these complete panels. [Full rankings and both aggregate metrics](benchmarks/complete-panel-ranks.md) · [All task scores and methods](benchmarks/EVALUATION.md).
|
| 68 |
|
|
@@ -70,12 +70,12 @@ Audio Mean(TaskType) improves from 46.97 to 52.34 over the previous Nano. Parame
|
|
| 70 |
|
| 71 |
### Quality and model size
|
| 72 |
|
| 73 |
-
|
| 74 |
|
| 75 |
<table>
|
| 76 |
<tr>
|
| 77 |
-
<td width="50%"><a href="assets/
|
| 78 |
-
<td width="50%"><a href="assets/
|
| 79 |
</tr>
|
| 80 |
</table>
|
| 81 |
|
|
|
|
| 59 |
|
| 60 |
The primary metric, **Mean(TaskType)**, weights each task type equally. Mean(Task) weights individual tasks equally and is supplementary. Rankings use complete benchmark results from the September 17, 2026 registry snapshots plus both current Vela models; “≤ size” counts models with no more total parameters than this model.
|
| 61 |
|
| 62 |
+
| Benchmark | Mean(TaskType) | Global rank | Rank at ≤ size | Gap to best at ≤ size | Mean(Task) |
|
| 63 |
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
| 64 |
+
| MTEB English v2 · 41 tasks | 60.78 | 65/188 | 5/75 | 0.61 pp | 64.88 |
|
| 65 |
+
| MAEB audio-only · 19 tasks | 52.34 | 19/64 | 6/27 | 3.51 pp | 43.59 |
|
| 66 |
|
| 67 |
Nano ranks **5/75** on the English panel and **6/27** on the audio panel among models with no more than its 163.8M total parameters. These are snapshot-relative comparisons across reported protocols, including single-modality specialists. The original small has not been evaluated on these complete panels. [Full rankings and both aggregate metrics](benchmarks/complete-panel-ranks.md) · [All task scores and methods](benchmarks/EVALUATION.md).
|
| 68 |
|
|
|
|
| 70 |
|
| 71 |
### Quality and model size
|
| 72 |
|
| 73 |
+
Each plot combines the full benchmark score, global and size-constrained ranks, and the gap to the best model at no greater total size. Every complete model with a known size is plotted, including models below the observed Pareto frontier. Highlighting Vela does not imply frontier membership. Click either figure for full resolution.
|
| 74 |
|
| 75 |
<table>
|
| 76 |
<tr>
|
| 77 |
+
<td width="50%"><a href="assets/general-english41-nano.png"><img src="assets/general-english41-nano.png" alt="Vela Nano: complete English benchmark ranking and size–quality comparison" /></a></td>
|
| 78 |
+
<td width="50%"><a href="assets/general-audio19-nano.png"><img src="assets/general-audio19-nano.png" alt="Vela Nano: complete audio benchmark ranking and size–quality comparison" /></a></td>
|
| 79 |
</tr>
|
| 80 |
</table>
|
| 81 |
|
assets/general-audio19-nano.png
ADDED
|
Git LFS Details
|
assets/general-audio19-nano.svg
ADDED
|
|
assets/general-english41-nano.png
ADDED
|
Git LFS Details
|
assets/general-english41-nano.svg
ADDED
|
|
benchmarks/complete-panel-ranks.md
CHANGED
|
@@ -1,12 +1,12 @@
|
|
| 1 |
-
# Complete
|
| 2 |
|
| 3 |
-
|
| 4 |
|
| 5 |
-
|
|
| 6 |
-
|---|---|---:|---|---|---:|
|
| 7 |
-
|
|
| 8 |
-
|
|
| 9 |
-
| Vela Omni Mini | English41 | 58.7867 | 92/188 | 50/134 | 63.4923 | 82/188 | 42/134 |
|
| 10 |
-
| Vela Omni Mini | audio19 | 54.8679 | 12/64 | 5/50 | 47.7720 | 9/64 | 3/50 |
|
| 11 |
|
| 12 |
-
|
|
|
|
|
|
|
|
|
| 1 |
+
# Complete benchmark rankings
|
| 2 |
|
| 3 |
+
Scores are 0–100. Mean(TaskType) weights task types equally; Mean(Task) weights tasks equally. The gap is measured in score points against the best complete model with no more total parameters. All comparisons use unrounded scores.
|
| 4 |
|
| 5 |
+
| Benchmark | Mean(TaskType) | Global rank | Rank at ≤ size | Gap to best at ≤ size | Mean(Task) |
|
| 6 |
+
| --- | ---: | ---: | ---: | ---: | ---: |
|
| 7 |
+
| MTEB English v2 · 41 tasks | 60.78 | 65/188 | 5/75 | 0.61 pp | 64.88 |
|
| 8 |
+
| MAEB audio-only · 19 tasks | 52.34 | 19/64 | 6/27 | 3.51 pp | 43.59 |
|
|
|
|
|
|
|
| 9 |
|
| 10 |
+
The September 17, 2026 registry snapshots contain 186 complete English models and 62 complete audio models; both Vela models are added to each panel. The 13 English peers without a known parameter count remain in global ranks but cannot enter size comparisons. Vela size includes text, image and audio components; peer sizes are registry-reported. Recorded evaluation protocols can differ. These panels have separate cohorts and are not combined into an overall multimodal rank.
|
| 11 |
+
|
| 12 |
+
[Comparison data and source URLs](general-rank-data.json) · [Methodology](pareto-methodology.md)
|
benchmarks/general-rank-data.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
benchmarks/general-rank-figures.json
ADDED
|
@@ -0,0 +1,373 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"snapshot": "2026-09-17",
|
| 3 |
+
"data_sha256": "b2020a288d71fd61ca3dd56e4bcd95f3a76bb5d4ea57f1eff379c414b4519215",
|
| 4 |
+
"figures": {
|
| 5 |
+
"English41-nano": {
|
| 6 |
+
"plotted_ids": [
|
| 7 |
+
"scores_MTEB_eng_v2:row:0",
|
| 8 |
+
"scores_MTEB_eng_v2:row:1",
|
| 9 |
+
"scores_MTEB_eng_v2:row:2",
|
| 10 |
+
"scores_MTEB_eng_v2:row:4",
|
| 11 |
+
"scores_MTEB_eng_v2:row:5",
|
| 12 |
+
"scores_MTEB_eng_v2:row:6",
|
| 13 |
+
"scores_MTEB_eng_v2:row:9",
|
| 14 |
+
"scores_MTEB_eng_v2:row:10",
|
| 15 |
+
"scores_MTEB_eng_v2:row:11",
|
| 16 |
+
"scores_MTEB_eng_v2:row:12",
|
| 17 |
+
"scores_MTEB_eng_v2:row:13",
|
| 18 |
+
"scores_MTEB_eng_v2:row:14",
|
| 19 |
+
"scores_MTEB_eng_v2:row:15",
|
| 20 |
+
"scores_MTEB_eng_v2:row:16",
|
| 21 |
+
"scores_MTEB_eng_v2:row:17",
|
| 22 |
+
"scores_MTEB_eng_v2:row:18",
|
| 23 |
+
"scores_MTEB_eng_v2:row:19",
|
| 24 |
+
"scores_MTEB_eng_v2:row:20",
|
| 25 |
+
"scores_MTEB_eng_v2:row:21",
|
| 26 |
+
"scores_MTEB_eng_v2:row:22",
|
| 27 |
+
"scores_MTEB_eng_v2:row:23",
|
| 28 |
+
"scores_MTEB_eng_v2:row:24",
|
| 29 |
+
"scores_MTEB_eng_v2:row:25",
|
| 30 |
+
"scores_MTEB_eng_v2:row:26",
|
| 31 |
+
"scores_MTEB_eng_v2:row:28",
|
| 32 |
+
"scores_MTEB_eng_v2:row:30",
|
| 33 |
+
"scores_MTEB_eng_v2:row:31",
|
| 34 |
+
"scores_MTEB_eng_v2:row:32",
|
| 35 |
+
"scores_MTEB_eng_v2:row:33",
|
| 36 |
+
"scores_MTEB_eng_v2:row:34",
|
| 37 |
+
"scores_MTEB_eng_v2:row:35",
|
| 38 |
+
"scores_MTEB_eng_v2:row:36",
|
| 39 |
+
"scores_MTEB_eng_v2:row:37",
|
| 40 |
+
"scores_MTEB_eng_v2:row:38",
|
| 41 |
+
"scores_MTEB_eng_v2:row:39",
|
| 42 |
+
"scores_MTEB_eng_v2:row:40",
|
| 43 |
+
"scores_MTEB_eng_v2:row:41",
|
| 44 |
+
"scores_MTEB_eng_v2:row:42",
|
| 45 |
+
"scores_MTEB_eng_v2:row:43",
|
| 46 |
+
"scores_MTEB_eng_v2:row:44",
|
| 47 |
+
"scores_MTEB_eng_v2:row:46",
|
| 48 |
+
"scores_MTEB_eng_v2:row:47",
|
| 49 |
+
"scores_MTEB_eng_v2:row:48",
|
| 50 |
+
"scores_MTEB_eng_v2:row:49",
|
| 51 |
+
"scores_MTEB_eng_v2:row:50",
|
| 52 |
+
"scores_MTEB_eng_v2:row:51",
|
| 53 |
+
"scores_MTEB_eng_v2:row:52",
|
| 54 |
+
"scores_MTEB_eng_v2:row:53",
|
| 55 |
+
"scores_MTEB_eng_v2:row:60",
|
| 56 |
+
"scores_MTEB_eng_v2:row:61",
|
| 57 |
+
"scores_MTEB_eng_v2:row:63",
|
| 58 |
+
"scores_MTEB_eng_v2:row:64",
|
| 59 |
+
"scores_MTEB_eng_v2:row:65",
|
| 60 |
+
"scores_MTEB_eng_v2:row:66",
|
| 61 |
+
"scores_MTEB_eng_v2:row:67",
|
| 62 |
+
"scores_MTEB_eng_v2:row:68",
|
| 63 |
+
"scores_MTEB_eng_v2:row:70",
|
| 64 |
+
"scores_MTEB_eng_v2:row:71",
|
| 65 |
+
"scores_MTEB_eng_v2:row:72",
|
| 66 |
+
"scores_MTEB_eng_v2:row:73",
|
| 67 |
+
"scores_MTEB_eng_v2:row:74",
|
| 68 |
+
"scores_MTEB_eng_v2:row:76",
|
| 69 |
+
"scores_MTEB_eng_v2:row:77",
|
| 70 |
+
"scores_MTEB_eng_v2:row:79",
|
| 71 |
+
"scores_MTEB_eng_v2:row:80",
|
| 72 |
+
"scores_MTEB_eng_v2:row:81",
|
| 73 |
+
"scores_MTEB_eng_v2:row:82",
|
| 74 |
+
"scores_MTEB_eng_v2:row:83",
|
| 75 |
+
"scores_MTEB_eng_v2:row:84",
|
| 76 |
+
"scores_MTEB_eng_v2:row:86",
|
| 77 |
+
"scores_MTEB_eng_v2:row:87",
|
| 78 |
+
"scores_MTEB_eng_v2:row:88",
|
| 79 |
+
"scores_MTEB_eng_v2:row:89",
|
| 80 |
+
"scores_MTEB_eng_v2:row:90",
|
| 81 |
+
"scores_MTEB_eng_v2:row:91",
|
| 82 |
+
"scores_MTEB_eng_v2:row:92",
|
| 83 |
+
"scores_MTEB_eng_v2:row:93",
|
| 84 |
+
"scores_MTEB_eng_v2:row:94",
|
| 85 |
+
"scores_MTEB_eng_v2:row:95",
|
| 86 |
+
"scores_MTEB_eng_v2:row:96",
|
| 87 |
+
"scores_MTEB_eng_v2:row:97",
|
| 88 |
+
"scores_MTEB_eng_v2:row:98",
|
| 89 |
+
"scores_MTEB_eng_v2:row:99",
|
| 90 |
+
"scores_MTEB_eng_v2:row:100",
|
| 91 |
+
"scores_MTEB_eng_v2:row:101",
|
| 92 |
+
"scores_MTEB_eng_v2:row:102",
|
| 93 |
+
"scores_MTEB_eng_v2:row:103",
|
| 94 |
+
"scores_MTEB_eng_v2:row:104",
|
| 95 |
+
"scores_MTEB_eng_v2:row:106",
|
| 96 |
+
"scores_MTEB_eng_v2:row:107",
|
| 97 |
+
"scores_MTEB_eng_v2:row:109",
|
| 98 |
+
"scores_MTEB_eng_v2:row:110",
|
| 99 |
+
"scores_MTEB_eng_v2:row:112",
|
| 100 |
+
"scores_MTEB_eng_v2:row:114",
|
| 101 |
+
"scores_MTEB_eng_v2:row:115",
|
| 102 |
+
"scores_MTEB_eng_v2:row:117",
|
| 103 |
+
"scores_MTEB_eng_v2:row:118",
|
| 104 |
+
"scores_MTEB_eng_v2:row:122",
|
| 105 |
+
"scores_MTEB_eng_v2:row:123",
|
| 106 |
+
"scores_MTEB_eng_v2:row:124",
|
| 107 |
+
"scores_MTEB_eng_v2:row:125",
|
| 108 |
+
"scores_MTEB_eng_v2:row:128",
|
| 109 |
+
"scores_MTEB_eng_v2:row:129",
|
| 110 |
+
"scores_MTEB_eng_v2:row:130",
|
| 111 |
+
"scores_MTEB_eng_v2:row:132",
|
| 112 |
+
"scores_MTEB_eng_v2:row:133",
|
| 113 |
+
"scores_MTEB_eng_v2:row:134",
|
| 114 |
+
"scores_MTEB_eng_v2:row:135",
|
| 115 |
+
"scores_MTEB_eng_v2:row:136",
|
| 116 |
+
"scores_MTEB_eng_v2:row:137",
|
| 117 |
+
"scores_MTEB_eng_v2:row:139",
|
| 118 |
+
"scores_MTEB_eng_v2:row:140",
|
| 119 |
+
"scores_MTEB_eng_v2:row:143",
|
| 120 |
+
"scores_MTEB_eng_v2:row:144",
|
| 121 |
+
"scores_MTEB_eng_v2:row:148",
|
| 122 |
+
"scores_MTEB_eng_v2:row:150",
|
| 123 |
+
"scores_MTEB_eng_v2:row:151",
|
| 124 |
+
"scores_MTEB_eng_v2:row:152",
|
| 125 |
+
"scores_MTEB_eng_v2:row:153",
|
| 126 |
+
"scores_MTEB_eng_v2:row:154",
|
| 127 |
+
"scores_MTEB_eng_v2:row:160",
|
| 128 |
+
"scores_MTEB_eng_v2:row:162",
|
| 129 |
+
"scores_MTEB_eng_v2:row:165",
|
| 130 |
+
"scores_MTEB_eng_v2:row:166",
|
| 131 |
+
"scores_MTEB_eng_v2:row:169",
|
| 132 |
+
"scores_MTEB_eng_v2:row:171",
|
| 133 |
+
"scores_MTEB_eng_v2:row:172",
|
| 134 |
+
"scores_MTEB_eng_v2:row:173",
|
| 135 |
+
"scores_MTEB_eng_v2:row:177",
|
| 136 |
+
"scores_MTEB_eng_v2:row:186",
|
| 137 |
+
"scores_MTEB_eng_v2:row:187",
|
| 138 |
+
"scores_MTEB_eng_v2:row:191",
|
| 139 |
+
"scores_MTEB_eng_v2:row:192",
|
| 140 |
+
"scores_MTEB_eng_v2:row:195",
|
| 141 |
+
"scores_MTEB_eng_v2:row:197",
|
| 142 |
+
"scores_MTEB_eng_v2:row:198",
|
| 143 |
+
"scores_MTEB_eng_v2:row:199",
|
| 144 |
+
"scores_MTEB_eng_v2:row:200",
|
| 145 |
+
"scores_MTEB_eng_v2:row:202",
|
| 146 |
+
"scores_MTEB_eng_v2:row:203",
|
| 147 |
+
"scores_MTEB_eng_v2:row:212",
|
| 148 |
+
"scores_MTEB_eng_v2:row:213",
|
| 149 |
+
"scores_MTEB_eng_v2:row:217",
|
| 150 |
+
"scores_MTEB_eng_v2:row:218",
|
| 151 |
+
"scores_MTEB_eng_v2:row:220",
|
| 152 |
+
"scores_MTEB_eng_v2:row:221",
|
| 153 |
+
"scores_MTEB_eng_v2:row:230",
|
| 154 |
+
"scores_MTEB_eng_v2:row:231",
|
| 155 |
+
"scores_MTEB_eng_v2:row:232",
|
| 156 |
+
"scores_MTEB_eng_v2:row:233",
|
| 157 |
+
"scores_MTEB_eng_v2:row:234",
|
| 158 |
+
"scores_MTEB_eng_v2:row:235",
|
| 159 |
+
"scores_MTEB_eng_v2:row:236",
|
| 160 |
+
"scores_MTEB_eng_v2:row:238",
|
| 161 |
+
"scores_MTEB_eng_v2:row:239",
|
| 162 |
+
"scores_MTEB_eng_v2:row:241",
|
| 163 |
+
"scores_MTEB_eng_v2:row:242",
|
| 164 |
+
"scores_MTEB_eng_v2:row:243",
|
| 165 |
+
"scores_MTEB_eng_v2:row:245",
|
| 166 |
+
"scores_MTEB_eng_v2:row:246",
|
| 167 |
+
"scores_MTEB_eng_v2:row:252",
|
| 168 |
+
"scores_MTEB_eng_v2:row:253",
|
| 169 |
+
"scores_MTEB_eng_v2:row:254",
|
| 170 |
+
"scores_MTEB_eng_v2:row:255",
|
| 171 |
+
"scores_MTEB_eng_v2:row:257",
|
| 172 |
+
"scores_MTEB_eng_v2:row:258",
|
| 173 |
+
"scores_MTEB_eng_v2:row:260",
|
| 174 |
+
"scores_MTEB_eng_v2:row:263",
|
| 175 |
+
"scores_MTEB_eng_v2:row:264",
|
| 176 |
+
"scores_MTEB_eng_v2:row:265",
|
| 177 |
+
"scores_MTEB_eng_v2:row:266",
|
| 178 |
+
"scores_MTEB_eng_v2:row:267",
|
| 179 |
+
"scores_MTEB_eng_v2:row:268",
|
| 180 |
+
"current:nano",
|
| 181 |
+
"current:mini"
|
| 182 |
+
],
|
| 183 |
+
"labeled_ids": [
|
| 184 |
+
"current:nano",
|
| 185 |
+
"current:mini",
|
| 186 |
+
"scores_MTEB_eng_v2:row:61",
|
| 187 |
+
"scores_MTEB_eng_v2:row:26",
|
| 188 |
+
"scores_MTEB_eng_v2:row:33",
|
| 189 |
+
"scores_MTEB_eng_v2:row:38",
|
| 190 |
+
"scores_MTEB_eng_v2:row:18",
|
| 191 |
+
"scores_MTEB_eng_v2:row:241",
|
| 192 |
+
"scores_MTEB_eng_v2:row:230",
|
| 193 |
+
"scores_MTEB_eng_v2:row:218",
|
| 194 |
+
"scores_MTEB_eng_v2:row:166",
|
| 195 |
+
"scores_MTEB_eng_v2:row:151",
|
| 196 |
+
"scores_MTEB_eng_v2:row:79",
|
| 197 |
+
"scores_MTEB_eng_v2:row:66"
|
| 198 |
+
],
|
| 199 |
+
"omitted_labels_but_points_retained": [],
|
| 200 |
+
"table_ids": [
|
| 201 |
+
"scores_MTEB_eng_v2:row:61",
|
| 202 |
+
"scores_MTEB_eng_v2:row:68",
|
| 203 |
+
"scores_MTEB_eng_v2:row:65",
|
| 204 |
+
"scores_MTEB_eng_v2:row:74",
|
| 205 |
+
"current:nano",
|
| 206 |
+
"scores_MTEB_eng_v2:row:71",
|
| 207 |
+
"scores_MTEB_eng_v2:row:66",
|
| 208 |
+
"scores_MTEB_eng_v2:row:72"
|
| 209 |
+
],
|
| 210 |
+
"rank": {
|
| 211 |
+
"score_100": 60.78184975654924,
|
| 212 |
+
"global_rank": 65,
|
| 213 |
+
"global_population": 188,
|
| 214 |
+
"known_size_frontier": false,
|
| 215 |
+
"at_most_size": {
|
| 216 |
+
"rank": 5,
|
| 217 |
+
"population": 75,
|
| 218 |
+
"gap_to_best_pp": 0.6140745488965593,
|
| 219 |
+
"best_ids": [
|
| 220 |
+
"scores_MTEB_eng_v2:row:61"
|
| 221 |
+
]
|
| 222 |
+
}
|
| 223 |
+
},
|
| 224 |
+
"known_size_frontier": [
|
| 225 |
+
"scores_MTEB_eng_v2:row:241",
|
| 226 |
+
"scores_MTEB_eng_v2:row:230",
|
| 227 |
+
"scores_MTEB_eng_v2:row:218",
|
| 228 |
+
"scores_MTEB_eng_v2:row:166",
|
| 229 |
+
"scores_MTEB_eng_v2:row:151",
|
| 230 |
+
"scores_MTEB_eng_v2:row:79",
|
| 231 |
+
"scores_MTEB_eng_v2:row:66",
|
| 232 |
+
"scores_MTEB_eng_v2:row:61",
|
| 233 |
+
"scores_MTEB_eng_v2:row:19",
|
| 234 |
+
"scores_MTEB_eng_v2:row:21",
|
| 235 |
+
"scores_MTEB_eng_v2:row:14",
|
| 236 |
+
"scores_MTEB_eng_v2:row:1",
|
| 237 |
+
"scores_MTEB_eng_v2:row:2",
|
| 238 |
+
"scores_MTEB_eng_v2:row:0"
|
| 239 |
+
],
|
| 240 |
+
"all_points_inside_axes": true,
|
| 241 |
+
"table_text_collisions": [],
|
| 242 |
+
"exports_sha256": {
|
| 243 |
+
"general-english41-nano.png": "6f874cb13d8b5fcc886fe6ec99024fef1de63e8c5816ab062a455c3f6c0ef908",
|
| 244 |
+
"general-english41-nano.svg": "cdb37918dd5417f11df19532ecab54b55d5032b36f895dcc9d0f58ecd9fe21bd"
|
| 245 |
+
}
|
| 246 |
+
},
|
| 247 |
+
"audio19-nano": {
|
| 248 |
+
"plotted_ids": [
|
| 249 |
+
"scores_MAEB_beta_audio-only:row:0",
|
| 250 |
+
"scores_MAEB_beta_audio-only:row:1",
|
| 251 |
+
"scores_MAEB_beta_audio-only:row:2",
|
| 252 |
+
"scores_MAEB_beta_audio-only:row:3",
|
| 253 |
+
"scores_MAEB_beta_audio-only:row:4",
|
| 254 |
+
"scores_MAEB_beta_audio-only:row:5",
|
| 255 |
+
"scores_MAEB_beta_audio-only:row:6",
|
| 256 |
+
"scores_MAEB_beta_audio-only:row:7",
|
| 257 |
+
"scores_MAEB_beta_audio-only:row:8",
|
| 258 |
+
"scores_MAEB_beta_audio-only:row:9",
|
| 259 |
+
"scores_MAEB_beta_audio-only:row:10",
|
| 260 |
+
"scores_MAEB_beta_audio-only:row:11",
|
| 261 |
+
"scores_MAEB_beta_audio-only:row:12",
|
| 262 |
+
"scores_MAEB_beta_audio-only:row:13",
|
| 263 |
+
"scores_MAEB_beta_audio-only:row:14",
|
| 264 |
+
"scores_MAEB_beta_audio-only:row:15",
|
| 265 |
+
"scores_MAEB_beta_audio-only:row:16",
|
| 266 |
+
"scores_MAEB_beta_audio-only:row:17",
|
| 267 |
+
"scores_MAEB_beta_audio-only:row:18",
|
| 268 |
+
"scores_MAEB_beta_audio-only:row:19",
|
| 269 |
+
"scores_MAEB_beta_audio-only:row:20",
|
| 270 |
+
"scores_MAEB_beta_audio-only:row:21",
|
| 271 |
+
"scores_MAEB_beta_audio-only:row:22",
|
| 272 |
+
"scores_MAEB_beta_audio-only:row:23",
|
| 273 |
+
"scores_MAEB_beta_audio-only:row:24",
|
| 274 |
+
"scores_MAEB_beta_audio-only:row:25",
|
| 275 |
+
"scores_MAEB_beta_audio-only:row:26",
|
| 276 |
+
"scores_MAEB_beta_audio-only:row:27",
|
| 277 |
+
"scores_MAEB_beta_audio-only:row:28",
|
| 278 |
+
"scores_MAEB_beta_audio-only:row:29",
|
| 279 |
+
"scores_MAEB_beta_audio-only:row:30",
|
| 280 |
+
"scores_MAEB_beta_audio-only:row:31",
|
| 281 |
+
"scores_MAEB_beta_audio-only:row:32",
|
| 282 |
+
"scores_MAEB_beta_audio-only:row:33",
|
| 283 |
+
"scores_MAEB_beta_audio-only:row:35",
|
| 284 |
+
"scores_MAEB_beta_audio-only:row:36",
|
| 285 |
+
"scores_MAEB_beta_audio-only:row:37",
|
| 286 |
+
"scores_MAEB_beta_audio-only:row:38",
|
| 287 |
+
"scores_MAEB_beta_audio-only:row:39",
|
| 288 |
+
"scores_MAEB_beta_audio-only:row:40",
|
| 289 |
+
"scores_MAEB_beta_audio-only:row:41",
|
| 290 |
+
"scores_MAEB_beta_audio-only:row:42",
|
| 291 |
+
"scores_MAEB_beta_audio-only:row:43",
|
| 292 |
+
"scores_MAEB_beta_audio-only:row:44",
|
| 293 |
+
"scores_MAEB_beta_audio-only:row:45",
|
| 294 |
+
"scores_MAEB_beta_audio-only:row:46",
|
| 295 |
+
"scores_MAEB_beta_audio-only:row:47",
|
| 296 |
+
"scores_MAEB_beta_audio-only:row:49",
|
| 297 |
+
"scores_MAEB_beta_audio-only:row:50",
|
| 298 |
+
"scores_MAEB_beta_audio-only:row:51",
|
| 299 |
+
"scores_MAEB_beta_audio-only:row:52",
|
| 300 |
+
"scores_MAEB_beta_audio-only:row:53",
|
| 301 |
+
"scores_MAEB_beta_audio-only:row:55",
|
| 302 |
+
"scores_MAEB_beta_audio-only:row:56",
|
| 303 |
+
"scores_MAEB_beta_audio-only:row:57",
|
| 304 |
+
"scores_MAEB_beta_audio-only:row:58",
|
| 305 |
+
"scores_MAEB_beta_audio-only:row:60",
|
| 306 |
+
"scores_MAEB_beta_audio-only:row:61",
|
| 307 |
+
"scores_MAEB_beta_audio-only:row:62",
|
| 308 |
+
"scores_MAEB_beta_audio-only:row:63",
|
| 309 |
+
"scores_MAEB_beta_audio-only:row:64",
|
| 310 |
+
"scores_MAEB_beta_audio-only:row:65",
|
| 311 |
+
"current:nano",
|
| 312 |
+
"current:mini"
|
| 313 |
+
],
|
| 314 |
+
"labeled_ids": [
|
| 315 |
+
"current:nano",
|
| 316 |
+
"current:mini",
|
| 317 |
+
"scores_MAEB_beta_audio-only:row:17",
|
| 318 |
+
"scores_MAEB_beta_audio-only:row:10",
|
| 319 |
+
"scores_MAEB_beta_audio-only:row:18",
|
| 320 |
+
"scores_MAEB_beta_audio-only:row:1",
|
| 321 |
+
"scores_MAEB_beta_audio-only:row:13",
|
| 322 |
+
"scores_MAEB_beta_audio-only:row:26",
|
| 323 |
+
"scores_MAEB_beta_audio-only:row:28",
|
| 324 |
+
"scores_MAEB_beta_audio-only:row:25",
|
| 325 |
+
"scores_MAEB_beta_audio-only:row:20",
|
| 326 |
+
"scores_MAEB_beta_audio-only:row:5",
|
| 327 |
+
"scores_MAEB_beta_audio-only:row:21",
|
| 328 |
+
"scores_MAEB_beta_audio-only:row:3"
|
| 329 |
+
],
|
| 330 |
+
"omitted_labels_but_points_retained": [],
|
| 331 |
+
"table_ids": [
|
| 332 |
+
"scores_MAEB_beta_audio-only:row:17",
|
| 333 |
+
"scores_MAEB_beta_audio-only:row:20",
|
| 334 |
+
"scores_MAEB_beta_audio-only:row:25",
|
| 335 |
+
"scores_MAEB_beta_audio-only:row:14",
|
| 336 |
+
"scores_MAEB_beta_audio-only:row:19",
|
| 337 |
+
"current:nano",
|
| 338 |
+
"scores_MAEB_beta_audio-only:row:28",
|
| 339 |
+
"scores_MAEB_beta_audio-only:row:12"
|
| 340 |
+
],
|
| 341 |
+
"rank": {
|
| 342 |
+
"score_100": 52.33540613685162,
|
| 343 |
+
"global_rank": 19,
|
| 344 |
+
"global_population": 64,
|
| 345 |
+
"known_size_frontier": false,
|
| 346 |
+
"at_most_size": {
|
| 347 |
+
"rank": 6,
|
| 348 |
+
"population": 27,
|
| 349 |
+
"gap_to_best_pp": 3.512754763326642,
|
| 350 |
+
"best_ids": [
|
| 351 |
+
"scores_MAEB_beta_audio-only:row:17"
|
| 352 |
+
]
|
| 353 |
+
}
|
| 354 |
+
},
|
| 355 |
+
"known_size_frontier": [
|
| 356 |
+
"scores_MAEB_beta_audio-only:row:28",
|
| 357 |
+
"scores_MAEB_beta_audio-only:row:25",
|
| 358 |
+
"scores_MAEB_beta_audio-only:row:20",
|
| 359 |
+
"scores_MAEB_beta_audio-only:row:17",
|
| 360 |
+
"scores_MAEB_beta_audio-only:row:10",
|
| 361 |
+
"scores_MAEB_beta_audio-only:row:5",
|
| 362 |
+
"scores_MAEB_beta_audio-only:row:21",
|
| 363 |
+
"scores_MAEB_beta_audio-only:row:3"
|
| 364 |
+
],
|
| 365 |
+
"all_points_inside_axes": true,
|
| 366 |
+
"table_text_collisions": [],
|
| 367 |
+
"exports_sha256": {
|
| 368 |
+
"general-audio19-nano.png": "314a33a29d9c78be427d48219739af50ec0a9df8785124b1b6b36185da555844",
|
| 369 |
+
"general-audio19-nano.svg": "2980af98012ec439b7480a7e97a67974371bac38fa2e69c732e3e84d5716d094"
|
| 370 |
+
}
|
| 371 |
+
}
|
| 372 |
+
}
|
| 373 |
+
}
|
benchmarks/pareto-gallery.md
CHANGED
|
@@ -1,16 +1,11 @@
|
|
| 1 |
# Figure gallery
|
| 2 |
|
| 3 |
-
Complete
|
| 4 |
|
| 5 |
-
 · [Methods and provenance](pareto-methodology.md)
|
|
|
|
| 1 |
# Figure gallery
|
| 2 |
|
| 3 |
+
Complete benchmark comparisons use all eligible models with complete results. Each figure shows ranks, score gaps and the observed size–quality frontier; highlighting does not imply frontier membership.
|
| 4 |
|
| 5 |
+

|
| 6 |
|
| 7 |
+

|
| 8 |
|
| 9 |
+
Task-level capability examples remain in the model card. They do not establish overall benchmark SOTA.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 10 |
|
| 11 |
[Full ranks](complete-panel-ranks.md) · [Methods and provenance](pareto-methodology.md)
|
benchmarks/pareto-methodology.md
CHANGED
|
@@ -6,11 +6,11 @@ The primary score is **Mean(TaskType)**: average tasks within each task type, th
|
|
| 6 |
|
| 7 |
Rankings use the September 17, 2026 dedicated MTEB registry snapshots plus both current Vela releases. English has 186 complete registry models; audio has 62. Incomplete rows (151 English, seven audio) are excluded rather than filled or combined across snapshots. Global ranks include 13 English models without a known positive size. Those models cannot enter a size-filtered rank or scatter plot. Ties share competition rank: one plus the number of strictly higher scores.
|
| 8 |
|
| 9 |
-
The plots show every complete model with a known positive total parameter count. The dashed Pareto line connects nondominated observations: no other eligible model has both no more parameters and no lower score, with one strict improvement. Vela points remain visible even when they are below this line. The adjacent
|
| 10 |
|
| 11 |
Vela sizes count all stored model parameters across text, image and audio: Nano 163,771,288 and Mini 1,361,475,288. Peer sizes use registry-reported totals, which can be rounded. Peers include both single-modality specialists and multimodal encoders. Their scores were reported under differing prompts, model revisions and evaluation protocols; these are snapshot-relative comparisons, not matched reproductions or official leaderboard submissions. There is no combined text–image–audio rank. Neither current Vela model lies on these complete-panel frontiers.
|
| 12 |
|
| 13 |
-
[All ranks](complete-panel-ranks.md) · [Complete vectors, source URLs, exclusions and plotted points](
|
| 14 |
|
| 15 |
## Selected task strengths
|
| 16 |
|
|
|
|
| 6 |
|
| 7 |
Rankings use the September 17, 2026 dedicated MTEB registry snapshots plus both current Vela releases. English has 186 complete registry models; audio has 62. Incomplete rows (151 English, seven audio) are excluded rather than filled or combined across snapshots. Global ranks include 13 English models without a known positive size. Those models cannot enter a size-filtered rank or scatter plot. Ties share competition rank: one plus the number of strictly higher scores.
|
| 8 |
|
| 9 |
+
The plots show every complete model with a known positive total parameter count. The dashed Pareto line connects nondominated observations: no other eligible model has both no more parameters and no lower score, with one strict improvement. Vela points remain visible even when they are below this line. The adjacent table shows the eight highest-scoring eligible models no larger than the highlighted Vela model, plus Vela itself if it ranks outside those eight. The displayed gap subtracts Vela's score from the highest score in this same size-constrained population. All comparisons use unrounded scores.
|
| 10 |
|
| 11 |
Vela sizes count all stored model parameters across text, image and audio: Nano 163,771,288 and Mini 1,361,475,288. Peer sizes use registry-reported totals, which can be rounded. Peers include both single-modality specialists and multimodal encoders. Their scores were reported under differing prompts, model revisions and evaluation protocols; these are snapshot-relative comparisons, not matched reproductions or official leaderboard submissions. There is no combined text–image–audio rank. Neither current Vela model lies on these complete-panel frontiers.
|
| 12 |
|
| 13 |
+
[All ranks](complete-panel-ranks.md) · [Complete vectors, source URLs, exclusions and plotted points](general-rank-data.json) · [Figure metadata](general-rank-figures.json) · [Vela evaluation protocol](https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/blob/main/benchmarks/EVALUATION.md)
|
| 14 |
|
| 15 |
## Selected task strengths
|
| 16 |
|