Xunzhuo commited on
Commit
b7c579a
·
verified ·
1 Parent(s): 9e747d2

Improve Omni Nano audio embeddings and complete benchmark rankings

Browse files
LICENSE-CLAP ADDED
@@ -0,0 +1,121 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ Creative Commons Legal Code
2
+
3
+ CC0 1.0 Universal
4
+
5
+ CREATIVE COMMONS CORPORATION IS NOT A LAW FIRM AND DOES NOT PROVIDE
6
+ LEGAL SERVICES. DISTRIBUTION OF THIS DOCUMENT DOES NOT CREATE AN
7
+ ATTORNEY-CLIENT RELATIONSHIP. CREATIVE COMMONS PROVIDES THIS
8
+ INFORMATION ON AN "AS-IS" BASIS. CREATIVE COMMONS MAKES NO WARRANTIES
9
+ REGARDING THE USE OF THIS DOCUMENT OR THE INFORMATION OR WORKS
10
+ PROVIDED HEREUNDER, AND DISCLAIMS LIABILITY FOR DAMAGES RESULTING FROM
11
+ THE USE OF THIS DOCUMENT OR THE INFORMATION OR WORKS PROVIDED
12
+ HEREUNDER.
13
+
14
+ Statement of Purpose
15
+
16
+ The laws of most jurisdictions throughout the world automatically confer
17
+ exclusive Copyright and Related Rights (defined below) upon the creator
18
+ and subsequent owner(s) (each and all, an "owner") of an original work of
19
+ authorship and/or a database (each, a "Work").
20
+
21
+ Certain owners wish to permanently relinquish those rights to a Work for
22
+ the purpose of contributing to a commons of creative, cultural and
23
+ scientific works ("Commons") that the public can reliably and without fear
24
+ of later claims of infringement build upon, modify, incorporate in other
25
+ works, reuse and redistribute as freely as possible in any form whatsoever
26
+ and for any purposes, including without limitation commercial purposes.
27
+ These owners may contribute to the Commons to promote the ideal of a free
28
+ culture and the further production of creative, cultural and scientific
29
+ works, or to gain reputation or greater distribution for their Work in
30
+ part through the use and efforts of others.
31
+
32
+ For these and/or other purposes and motivations, and without any
33
+ expectation of additional consideration or compensation, the person
34
+ associating CC0 with a Work (the "Affirmer"), to the extent that he or she
35
+ is an owner of Copyright and Related Rights in the Work, voluntarily
36
+ elects to apply CC0 to the Work and publicly distribute the Work under its
37
+ terms, with knowledge of his or her Copyright and Related Rights in the
38
+ Work and the meaning and intended legal effect of CC0 on those rights.
39
+
40
+ 1. Copyright and Related Rights. A Work made available under CC0 may be
41
+ protected by copyright and related or neighboring rights ("Copyright and
42
+ Related Rights"). Copyright and Related Rights include, but are not
43
+ limited to, the following:
44
+
45
+ i. the right to reproduce, adapt, distribute, perform, display,
46
+ communicate, and translate a Work;
47
+ ii. moral rights retained by the original author(s) and/or performer(s);
48
+ iii. publicity and privacy rights pertaining to a person's image or
49
+ likeness depicted in a Work;
50
+ iv. rights protecting against unfair competition in regards to a Work,
51
+ subject to the limitations in paragraph 4(a), below;
52
+ v. rights protecting the extraction, dissemination, use and reuse of data
53
+ in a Work;
54
+ vi. database rights (such as those arising under Directive 96/9/EC of the
55
+ European Parliament and of the Council of 11 March 1996 on the legal
56
+ protection of databases, and under any national implementation
57
+ thereof, including any amended or successor version of such
58
+ directive); and
59
+ vii. other similar, equivalent or corresponding rights throughout the
60
+ world based on applicable law or treaty, and any national
61
+ implementations thereof.
62
+
63
+ 2. Waiver. To the greatest extent permitted by, but not in contravention
64
+ of, applicable law, Affirmer hereby overtly, fully, permanently,
65
+ irrevocably and unconditionally waives, abandons, and surrenders all of
66
+ Affirmer's Copyright and Related Rights and associated claims and causes
67
+ of action, whether now known or unknown (including existing as well as
68
+ future claims and causes of action), in the Work (i) in all territories
69
+ worldwide, (ii) for the maximum duration provided by applicable law or
70
+ treaty (including future time extensions), (iii) in any current or future
71
+ medium and for any number of copies, and (iv) for any purpose whatsoever,
72
+ including without limitation commercial, advertising or promotional
73
+ purposes (the "Waiver"). Affirmer makes the Waiver for the benefit of each
74
+ member of the public at large and to the detriment of Affirmer's heirs and
75
+ successors, fully intending that such Waiver shall not be subject to
76
+ revocation, rescission, cancellation, termination, or any other legal or
77
+ equitable action to disrupt the quiet enjoyment of the Work by the public
78
+ as contemplated by Affirmer's express Statement of Purpose.
79
+
80
+ 3. Public License Fallback. Should any part of the Waiver for any reason
81
+ be judged legally invalid or ineffective under applicable law, then the
82
+ Waiver shall be preserved to the maximum extent permitted taking into
83
+ account Affirmer's express Statement of Purpose. In addition, to the
84
+ extent the Waiver is so judged Affirmer hereby grants to each affected
85
+ person a royalty-free, non transferable, non sublicensable, non exclusive,
86
+ irrevocable and unconditional license to exercise Affirmer's Copyright and
87
+ Related Rights in the Work (i) in all territories worldwide, (ii) for the
88
+ maximum duration provided by applicable law or treaty (including future
89
+ time extensions), (iii) in any current or future medium and for any number
90
+ of copies, and (iv) for any purpose whatsoever, including without
91
+ limitation commercial, advertising or promotional purposes (the
92
+ "License"). The License shall be deemed effective as of the date CC0 was
93
+ applied by Affirmer to the Work. Should any part of the License for any
94
+ reason be judged legally invalid or ineffective under applicable law, such
95
+ partial invalidity or ineffectiveness shall not invalidate the remainder
96
+ of the License, and in such case Affirmer hereby affirms that he or she
97
+ will not (i) exercise any of his or her remaining Copyright and Related
98
+ Rights in the Work or (ii) assert any associated claims and causes of
99
+ action with respect to the Work, in either case contrary to Affirmer's
100
+ express Statement of Purpose.
101
+
102
+ 4. Limitations and Disclaimers.
103
+
104
+ a. No trademark or patent rights held by Affirmer are waived, abandoned,
105
+ surrendered, licensed or otherwise affected by this document.
106
+ b. Affirmer offers the Work as-is and makes no representations or
107
+ warranties of any kind concerning the Work, express, implied,
108
+ statutory or otherwise, including without limitation warranties of
109
+ title, merchantability, fitness for a particular purpose, non
110
+ infringement, or the absence of latent or other defects, accuracy, or
111
+ the present or absence of errors, whether or not discoverable, all to
112
+ the greatest extent permissible under applicable law.
113
+ c. Affirmer disclaims responsibility for clearing rights of other persons
114
+ that may apply to the Work or any use thereof, including without
115
+ limitation any person's Copyright and Related Rights in the Work.
116
+ Further, Affirmer disclaims responsibility for obtaining any necessary
117
+ consents, permissions or other rights required for any use of the
118
+ Work.
119
+ d. Affirmer understands and acknowledges that Creative Commons is not a
120
+ party to this document and has no duty or obligation with respect to
121
+ this CC0 or use of the Work.
NOTICE CHANGED
@@ -25,3 +25,6 @@ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
25
  SOFTWARE.
26
 
27
  Text component: avsolatorio/GIST-small-Embedding-v0, revision 75e62fd210b9fde790430e0b2f040b0b00a021b1. Its official model card declares MIT; the complete pinned card is retained in licenses/GIST-MODEL-CARD.md. The card identifies BAAI/bge-small-en-v1.5 as the upstream base and discloses its training data.
 
 
 
 
25
  SOFTWARE.
26
 
27
  Text component: avsolatorio/GIST-small-Embedding-v0, revision 75e62fd210b9fde790430e0b2f040b0b00a021b1. Its official model card declares MIT; the complete pinned card is retained in licenses/GIST-MODEL-CARD.md. The card identifies BAAI/bge-small-en-v1.5 as the upstream base and discloses its training data.
28
+
29
+ Additional frozen audio branch: laion/clap-htsat-unfused, revision 8fa0f1c6d0433df6e97c127f64b2a1d6c0dcda8a, Apache-2.0.
30
+ Endpoint window pooling is a Vela readout. See LICENSE-CLAP for the unchanged upstream license.
README.md CHANGED
@@ -22,7 +22,7 @@ tags:
22
 
23
  # Vela Omni Nano
24
 
25
- Text, images and speech in one embedding space. Vela Omni Nano supports multimodal search, routing and clustering with normalized vectors that can be compared directly.
26
 
27
  [Try Vela Studio](https://huggingface.co/spaces/llm-semantic-router/vela-studio) · [Vela collection](https://huggingface.co/collections/llm-semantic-router/vela-10-6aa555ba70cc6997d6d67798) · [Detailed evaluation](benchmarks/EVALUATION.md)
28
 
@@ -30,15 +30,15 @@ Text, images and speech in one embedding space. Vela Omni Nano supports multimod
30
 
31
  | Feature | Value |
32
  | --- | --- |
33
- | Modalities | Text, images, speech |
34
- | Total parameters | 135.4M (135,383,808) |
35
  | Embedding dimensions | 384 |
36
  | Text context | 512 tokens, including special tokens |
37
- | Audio input | Mono, 16 kHz, up to 30 seconds |
38
  | Output | L2-normalized vectors; cosine similarity |
39
  | License | Apache 2.0; see component attribution |
40
 
41
- The smaller inference package preserves the encoding computation and all reported scores. [Equivalence and storage details](benchmarks/inference-equivalence.json).
42
 
43
  ## Evaluation
44
 
@@ -50,8 +50,8 @@ Scores are 0–100; higher is better. The comparison uses the same held-out eval
50
  | MASSIVE English · Accuracy | 65.95 | **81.97** |
51
  | COCO · Image → text · R@1 | 40.83 | **60.87** |
52
  | COCO · Text → image · R@1 | 30.18 | **55.82** |
53
- | LibriSpeech · Audio → text · R@1 | 4.21 | **17.16** |
54
- | LibriSpeech · Text → audio · R@1 | 9.58 | **25.44** |
55
 
56
  The common protocol uses labeled TRAIN prototypes for text classification and all matching positives for retrieval; it is separate from official MTEB classification. [All 14 metrics, Macro-F1, exact counts and uncertainty](benchmarks/EVALUATION.md#routing-and-cross-modal-retrieval).
57
 
@@ -61,18 +61,23 @@ The primary metric, **Mean(TaskType)**, weights each task type equally. Mean(Tas
61
 
62
  | Benchmark | Mean(TaskType) | Global rank | Rank at ≤ size | Mean(Task) |
63
  | --- | ---: | ---: | ---: | ---: |
64
- | MTEB English v2 · 41 tasks | 60.78 | 65/188 | 3/66 | 64.88 |
65
- | MAEB audio-only · 19 tasks | 46.97 | 32/64 | 8/22 | 39.77 |
66
 
67
- Nano ranks **3/66** on the complete English panel among models with no more than its 135.4M total parameters, **0.61 points** behind the highest-scoring model in that group. These are snapshot-relative comparisons across reported protocols, including single-modality specialists. The original small has not been evaluated on these complete panels. [Full rankings and both aggregate metrics](benchmarks/complete-panel-ranks.md) · [All task scores and methods](benchmarks/EVALUATION.md).
 
 
68
 
69
  ### Quality and model size
70
 
71
  The plots include every complete model with a known size. The dashed line shows the observed Pareto frontier; both Vela models remain plotted even when below it. Known-size peers below the frontier provide additional reference points.
72
 
73
- ![Complete English benchmark: size–quality comparison and ranking](assets/complete-panel-english41-nano.png)
74
-
75
- ![Complete audio benchmark: size–quality comparison and ranking](assets/complete-panel-audio19-mini.png)
 
 
 
76
 
77
  ### Selected task strengths
78
 
@@ -90,7 +95,7 @@ These task-level comparisons highlight specific strengths; they do not establish
90
  ## Usage
91
 
92
 
93
- Use PyTorch, Transformers 4.57.6, Hugging Face Hub, safetensors, NumPy, and Pillow:
94
 
95
  ```python
96
  import sys
@@ -107,12 +112,14 @@ print(vectors.shape) # (2, 384)
107
 
108
  Text inputs support up to 512 tokens, including special tokens. Longer inputs raise `ValueError`; shorten or explicitly chunk them.
109
 
110
- Pass a list of Pillow images to `model.encode_image(images)`. Pass a list of mono NumPy waveforms to `model.encode_audio(waveforms, sampling_rate=16000)`; each waveform must be at most 30 seconds. Compare normalized vectors with their dot product. To route media, embed each destination’s name and description, then select the closest vector. To discover routing categories, cluster media embeddings and inspect each group. Similarity scores are rankings, not calibrated probabilities.
111
 
112
  ## Training and license
113
 
114
  Cross-modal alignment uses COCO image–caption pairs from the CC-BY 2.0 image subset and [LibriSpeech](https://www.openslr.org/12/) speech–transcript pairs (CC-BY 4.0). COCO annotations are CC-BY 4.0. The frozen [GIST text backbone](https://huggingface.co/avsolatorio/GIST-small-Embedding-v0) is MIT-licensed; its pinned model card and component notices are included.
115
 
 
 
116
  The repository includes the native model code, component configurations, and tokenizer and processor files. See [NOTICE](./NOTICE) and [LICENSE](./LICENSE) for component attribution and license terms.
117
 
118
  [Explore the Vela model collection](https://huggingface.co/collections/llm-semantic-router/vela-10-6aa555ba70cc6997d6d67798)
 
22
 
23
  # Vela Omni Nano
24
 
25
+ Text, images, speech and environmental audio in one embedding space. Vela Omni Nano supports multimodal search, routing and clustering with normalized vectors that can be compared directly.
26
 
27
  [Try Vela Studio](https://huggingface.co/spaces/llm-semantic-router/vela-studio) · [Vela collection](https://huggingface.co/collections/llm-semantic-router/vela-10-6aa555ba70cc6997d6d67798) · [Detailed evaluation](benchmarks/EVALUATION.md)
28
 
 
30
 
31
  | Feature | Value |
32
  | --- | --- |
33
+ | Modalities | Text, images, speech and environmental audio |
34
+ | Total parameters | 163.8M (163,771,288) |
35
  | Embedding dimensions | 384 |
36
  | Text context | 512 tokens, including special tokens |
37
+ | Audio input | Original PCM at 16, 44.1 or 48 kHz; up to 30 seconds |
38
  | Output | L2-normalized vectors; cosine similarity |
39
  | License | Apache 2.0; see component attribution |
40
 
41
+ A frozen CLAP audio branch adds environmental-sound information to the existing speech representation. Text and image computations are retained; audio embeddings are newly trained and evaluated. [Architecture and measured identity](benchmarks/component-equivalence.json).
42
 
43
  ## Evaluation
44
 
 
50
  | MASSIVE English · Accuracy | 65.95 | **81.97** |
51
  | COCO · Image → text · R@1 | 40.83 | **60.87** |
52
  | COCO · Text → image · R@1 | 30.18 | **55.82** |
53
+ | LibriSpeech · Audio → text · R@1 | 4.21 | **16.12** |
54
+ | LibriSpeech · Text → audio · R@1 | 9.58 | **20.34** |
55
 
56
  The common protocol uses labeled TRAIN prototypes for text classification and all matching positives for retrieval; it is separate from official MTEB classification. [All 14 metrics, Macro-F1, exact counts and uncertainty](benchmarks/EVALUATION.md#routing-and-cross-modal-retrieval).
57
 
 
61
 
62
  | Benchmark | Mean(TaskType) | Global rank | Rank at ≤ size | Mean(Task) |
63
  | --- | ---: | ---: | ---: | ---: |
64
+ | MTEB English v2 · 41 tasks | 60.78 | 65/188 | 5/75 | 64.88 |
65
+ | MAEB audio-only · 19 tasks | 52.34 | 19/64 | 6/27 | 43.59 |
66
 
67
+ Nano ranks **5/75** on the English panel and **6/27** on the audio panel among models with no more than its 163.8M total parameters. These are snapshot-relative comparisons across reported protocols, including single-modality specialists. The original small has not been evaluated on these complete panels. [Full rankings and both aggregate metrics](benchmarks/complete-panel-ranks.md) · [All task scores and methods](benchmarks/EVALUATION.md).
68
+
69
+ Audio Mean(TaskType) improves from 46.97 to 52.34 over the previous Nano. Parameters rise 20.97%; speech–text retrieval and VehicleSoundClustering regress. [Complete gains and trade-offs](benchmarks/release-history.md).
70
 
71
  ### Quality and model size
72
 
73
  The plots include every complete model with a known size. The dashed line shows the observed Pareto frontier; both Vela models remain plotted even when below it. Known-size peers below the frontier provide additional reference points.
74
 
75
+ <table>
76
+ <tr>
77
+ <td width="50%"><a href="assets/complete-panel-english41-nano.png"><img src="assets/complete-panel-english41-nano.png" alt="Complete English41: size, quality and ranking; click for full resolution" /></a></td>
78
+ <td width="50%"><a href="assets/complete-panel-audio19-mini.png"><img src="assets/complete-panel-audio19-mini.png" alt="Complete audio19: size, quality and ranking; click for full resolution" /></a></td>
79
+ </tr>
80
+ </table>
81
 
82
  ### Selected task strengths
83
 
 
95
  ## Usage
96
 
97
 
98
+ Use PyTorch and matching torchaudio, Transformers 4.57.6, Hugging Face Hub, safetensors, NumPy, and Pillow:
99
 
100
  ```python
101
  import sys
 
112
 
113
  Text inputs support up to 512 tokens, including special tokens. Longer inputs raise `ValueError`; shorten or explicitly chunk them.
114
 
115
+ Pass a list of Pillow images to `model.encode_image(images)`. Pass original NumPy waveforms to `model.encode_audio(waveforms, sampling_rate=48000)` using their actual 16,000, 44,100 or 48,000 Hz rate; each waveform must be at most 30 seconds. Mono or channels-first arrays are supported. Keep the original waveform: the speech and CLAP branches independently derive their 16 kHz and 48 kHz inputs. Do not downsample to 16 kHz before calling the API when higher-rate PCM is available. Compare normalized vectors with their dot product. To route media, embed each destination’s name and description, then select the closest vector. To discover routing categories, cluster media embeddings and inspect each group. Similarity scores are rankings, not calibrated probabilities.
116
 
117
  ## Training and license
118
 
119
  Cross-modal alignment uses COCO image–caption pairs from the CC-BY 2.0 image subset and [LibriSpeech](https://www.openslr.org/12/) speech–transcript pairs (CC-BY 4.0). COCO annotations are CC-BY 4.0. The frozen [GIST text backbone](https://huggingface.co/avsolatorio/GIST-small-Embedding-v0) is MIT-licensed; its pinned model card and component notices are included.
120
 
121
+ Audio residual alignment uses 3,299 FSD50K TRAIN recordings under CC0 or CC-BY 3.0; [per-recording attribution](training/fsd50k-attributions.jsonl) is included. The frozen [CLAP unfused audio component](https://huggingface.co/laion/clap-htsat-unfused) is Apache 2.0; its original license is retained as [LICENSE-CLAP](LICENSE-CLAP).
122
+
123
  The repository includes the native model code, component configurations, and tokenizer and processor files. See [NOTICE](./NOTICE) and [LICENSE](./LICENSE) for component attribution and license terms.
124
 
125
  [Explore the Vela model collection](https://huggingface.co/collections/llm-semantic-router/vela-10-6aa555ba70cc6997d6d67798)
assets/complete-panel-audio19-mini.png CHANGED

Git LFS Details

  • SHA256: 69b0f8362667abaa5579dc5f78bdf49e918428f40e350f393dae451807ef9dd3
  • Pointer size: 131 Bytes
  • Size of remote file: 344 kB

Git LFS Details

  • SHA256: 131a400c69f3e7b6fc1ff1a1de0f285fd7e5938a2fffbfe98552d8981bfa1b27
  • Pointer size: 131 Bytes
  • Size of remote file: 341 kB
assets/complete-panel-audio19-mini.svg CHANGED
assets/complete-panel-english41-nano.png CHANGED

Git LFS Details

  • SHA256: ea7e14d09f749c22b0d4d32414df5008f9d68c746f8614483b25108ff52c88cc
  • Pointer size: 131 Bytes
  • Size of remote file: 346 kB

Git LFS Details

  • SHA256: 49e1d1a57289291c261e50070ebe2ed5f7fe313ff0d1688c7e67ca2f0c4b27ff
  • Pointer size: 131 Bytes
  • Size of remote file: 345 kB
assets/complete-panel-english41-nano.svg CHANGED
assets/pareto-imdb.png CHANGED

Git LFS Details

  • SHA256: d9259e7654b451d7b71c0b2b4a5c97f0f429475e7c557d3493bfc67763f491b7
  • Pointer size: 131 Bytes
  • Size of remote file: 274 kB

Git LFS Details

  • SHA256: f37151296f70b9436d93a2f51f89e2d401920bf9f83dda65505d09d20cee680f
  • Pointer size: 131 Bytes
  • Size of remote file: 279 kB
assets/pareto-imdb.svg CHANGED
assets/pareto-nmsqa.png CHANGED

Git LFS Details

  • SHA256: e36b9a6383be987ac7e3c9e562fe0931b210b0cf0fe889f993f5317e06b4c6da
  • Pointer size: 131 Bytes
  • Size of remote file: 262 kB

Git LFS Details

  • SHA256: 04597d0e96d04fb5a5b977993d4ce66d6ad3ec353f0965372ce25045bdd385a3
  • Pointer size: 131 Bytes
  • Size of remote file: 262 kB
assets/pareto-nmsqa.svg CHANGED
benchmarks/EVALUATION.md CHANGED
@@ -1,12 +1,10 @@
1
- # Evaluation details
2
 
3
- The primary comparator is [llm-semantic-router/multi-modal-embed-small](https://huggingface.co/llm-semantic-router/multi-modal-embed-small/tree/fdf8e01b7b0f3a69ac1ac8e2a64dcb1ede177ba4), pinned to `fdf8e01b7b0f3a69ac1ac8e2a64dcb1ede177ba4`. It is the original small model, not a previous Vela release. Current Nano native identity is `50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa`. Scores are 0–100; bold applies only to an unrounded improvement over this original comparator. [Original checkpoint and comparison evidence](original-baseline.json).
4
 
5
  ## Routing and cross-modal retrieval
6
 
7
- Both models use the same held-out examples, complete candidate pools and 128-token text limit. These known evaluation pools have been reused across releases, so this is a progress comparison rather than a new blind test. Banking77 and MASSIVE use nearest normalized class prototypes from labeled TRAIN examples; this task-adapted protocol is distinct from official MTEB frozen-embedding classification. Retrieval uses cosine scores and all matching positives.
8
-
9
- | Metric | [Original small](https://huggingface.co/llm-semantic-router/multi-modal-embed-small/tree/fdf8e01b7b0f3a69ac1ac8e2a64dcb1ede177ba4) | Current Nano |
10
  | --- | ---: | ---: |
11
  | Banking77 · Accuracy | 70.4221 | **87.9870** |
12
  | MASSIVE English · Accuracy | 65.9489 | **81.9650** |
@@ -16,108 +14,58 @@ Both models use the same held-out examples, complete candidate pools and 128-tok
16
  | COCO · Text → image · R@1 | 30.1823 | **55.8202** |
17
  | COCO · Text → image · R@5 | 59.4168 | **85.2248** |
18
  | COCO · Text → image · R@10 | 74.2892 | **94.2649** |
19
- | LibriSpeech · Audio → text · R@1 | 4.2129 | **17.1582** |
20
- | LibriSpeech · Audio → text · R@5 | 11.9877 | **36.6526** |
21
- | LibriSpeech · Audio → text · R@10 | 19.0349 | **46.9935** |
22
- | LibriSpeech · Text → audio · R@1 | 9.5785 | **25.4406** |
23
- | LibriSpeech · Text → audio · R@5 | 22.5287 | **46.1303** |
24
- | LibriSpeech · Text → audio · R@10 | 30.6897 | **55.9004** |
25
  | Banking77 · Macro-F1 | 70.2078 | **87.9168** |
26
  | MASSIVE English · Macro-F1 | 61.7178 | **79.2322** |
27
 
28
- The text pools contain 3,080 Banking77 and 2,972 MASSIVE queries, with 77 and 60 prototype classes. COCO uses 823 images and 4,115 captions; LibriSpeech uses 2,611 clips and 2,610 unique transcripts. Macro-F1 equally averages all declared classes (undefined F1 is zero). [scores.json](../scores.json) retains exact numerators, denominators, class definitions and paired uncertainty. Historical paired intervals keep their actual comparator and are not relabeled.
29
-
30
- ## Complete standard panels
31
-
32
- | Benchmark / metric | Current Nano |
33
- | --- | ---: |
34
- | English v2 · 41/41 tasks · mean task | 64.88 |
35
- | English v2 · mean task type | 60.78 |
36
- | MAEB audio-only · 19/19 tasks · mean task | 39.77 |
37
- | MAEB audio-only · 134 slots · mean task type | 46.97 |
38
-
39
- English contains 41 tasks and seven task types. The MAEB audio-only panel contains 19 tasks, five task types and 134 subset/split slots, excluding joint-modality tasks. Aggregation uses official within-task results, then equal task or task-type means. The original small was not evaluated in this matched complete-panel protocol; no score is inferred from common-protocol results or a previous Vela release.
40
-
41
- Nano uses the 135,383,808-parameter inference package. [Inference equivalence](inference-equivalence.json) preserves the original evaluated 96d artifact and all actual text/audio component identities. The participating weights and computation are unchanged; this is not a benchmark rerun.
42
 
43
- ### Every English task
44
 
45
- | Task | Main metric | Current Nano |
46
  | --- | ---: | ---: |
47
- | ArguAna | ndcg_at_10 | 59.2460 |
48
- | ArXivHierarchicalClusteringP2P | v_measure | 64.8885 |
49
- | ArXivHierarchicalClusteringS2S | v_measure | 57.4432 |
50
- | AskUbuntuDupQuestions | map_at_1000 | 62.3290 |
51
- | BIOSSES | cosine_spearman | 86.9875 |
52
- | Banking77Classification | accuracy | 82.1429 |
53
- | BiorxivClusteringP2P.v2 | v_measure | 41.3138 |
54
- | CQADupstackGamingRetrieval | ndcg_at_10 | 56.9790 |
55
- | CQADupstackUnixRetrieval | ndcg_at_10 | 39.5860 |
56
- | ClimateFEVERHardNegatives | ndcg_at_10 | 31.8460 |
57
- | FEVERHardNegatives | ndcg_at_10 | 87.5760 |
58
- | FiQA2018 | ndcg_at_10 | 39.1430 |
59
- | HotpotQAHardNegatives | ndcg_at_10 | 66.3490 |
60
- | ImdbClassification | accuracy | 91.9476 |
61
- | MTOPDomainClassification | accuracy | 94.9179 |
62
- | MassiveIntentClassification | accuracy | 70.9684 |
63
- | MassiveScenarioClassification | accuracy | 76.1130 |
64
- | MedrxivClusteringP2P.v2 | v_measure | 39.8302 |
65
- | MedrxivClusteringS2S.v2 | v_measure | 37.8751 |
66
- | MindSmallReranking | max_over_subqueries_map_at_1000 | 32.3660 |
67
- | SCIDOCS | ndcg_at_10 | 21.8890 |
68
- | SICK-R | cosine_spearman | 80.5317 |
69
- | STS12 | cosine_spearman | 75.5659 |
70
- | STS13 | cosine_spearman | 86.2635 |
71
- | STS14 | cosine_spearman | 82.2988 |
72
- | STS15 | cosine_spearman | 88.7365 |
73
- | STSBenchmark | cosine_spearman | 87.0782 |
74
- | SprintDuplicateQuestions | max_ap | 95.7996 |
75
- | StackExchangeClustering.v2 | v_measure | 58.4244 |
76
- | StackExchangeClusteringP2P.v2 | v_measure | 41.0530 |
77
- | TRECCOVID | ndcg_at_10 | 69.1210 |
78
- | Touche2020Retrieval.v3 | ndcg_at_10 | 48.3630 |
79
- | ToxicConversationsClassification | accuracy | 71.9043 |
80
- | TweetSentimentExtractionClassification | accuracy | 62.7929 |
81
- | TwentyNewsgroupsClustering.v2 | v_measure | 50.8726 |
82
- | TwitterSemEval2015 | max_ap | 72.9515 |
83
- | TwitterURLCorpus | max_ap | 85.3034 |
84
- | SummEvalSummarization.v2 | cosine_spearman | 31.8344 |
85
- | AmazonCounterfactualClassification | accuracy | 71.8507 |
86
- | STS17 | cosine_spearman | 89.0210 |
87
- | STS22.v2 | cosine_spearman | 68.7533 |
88
 
89
  ### Every audio task
90
 
91
- | Task | Main metric | Current Nano |
92
- | --- | ---: | ---: |
93
- | JamAltArtistA2ARetrieval | ndcg_at_10 | 68.9005 |
94
- | BeijingOpera | accuracy | 68.6791 |
95
- | BirdCLEF | accuracy | 12.9000 |
96
- | CREMA_D | accuracy | 29.7637 |
97
- | CommonLanguageAgeDetection | accuracy | 16.0400 |
98
- | GTZANGenre | accuracy | 50.3000 |
99
- | IEMOCAPGender | accuracy | 58.4519 |
100
- | MInDS14 | accuracy | 23.3212 |
101
- | MridinghamTonic | accuracy | 32.3632 |
102
- | SIBFLEURS | accuracy | 18.3047 |
103
- | VoxCelebSA | accuracy | 33.1691 |
104
- | VoxPopuliLanguageID | accuracy | 91.8000 |
105
- | CREMA_DClustering | v_measure | 2.3602 |
106
- | VehicleSoundClustering | v_measure | 13.8007 |
107
- | VoxPopuliGenderClustering | v_measure | 0.1045 |
108
- | CREMADPairClassification | max_ap | 54.2012 |
109
- | NMSQAPairClassification | max_ap | 63.2834 |
110
- | VoxPopuliAccentPairClassification | max_ap | 54.0220 |
111
- | GTZANAudioReranking | map_at_1000 | 63.7940 |
112
-
113
- [English JSON](mteb-eng-v2.json) and [audio JSON](maeb-audio-only.json) retain all raw metrics, dataset revisions, repeated probes, subset/split coverage and actual evaluated identities. Undefined auxiliary statistics remain null; primary scores are finite.
114
-
115
- ## Selected standard tasks
116
-
117
- MTEB 2.21.0, FP32, batch size 8, seed 20260918. Eleven tasks and twelve split rows; MTOP validation is descriptive. Native context limits, official preprocessing, sampling, probes and clustering are preserved; no task instruction is added. This is not a complete multilingual or image benchmark.
118
 
119
  | Task / split | Main metric | Original small | Current Nano |
120
- | --- | ---: | ---: | ---: |
121
  | SICK-R / test | cosine_spearman | 71.3902 | **80.5317** |
122
  | STSBenchmark / test | cosine_spearman | 77.2987 | **87.0782** |
123
  | ArguAna / test | ndcg_at_10 | 32.8890 | **59.2460** |
@@ -127,16 +75,20 @@ MTEB 2.21.0, FP32, batch size 8, seed 20260918. Eleven tasks and twelve split ro
127
  | OxfordPets / test | accuracy | 41.8479 | **91.2674** |
128
  | TinyImageNetClustering / valid | nmi | 46.2849 | **61.6919** |
129
  | OxfordPetsZeroShot / test | accuracy | 16.4077 | 8.0131 |
130
- | CREMA_D / train | accuracy | 29.9520 | 29.7637 |
131
- | CREMA_DClustering / train | v_measure | 2.4058 | 2.3602 |
132
- | SpeechCommandsZeroshotv0.02 / test | accuracy | 5.5228 | **14.1139** |
 
 
 
 
133
 
134
- Original small uses its native 128-token limit; current Nano uses 512. All rows use the same pinned task revisions, official probes, sampling and metrics. Bold marks an unrounded improvement over the original comparator; lower values remain visible. MTOP validation is descriptive. [Full original results and identities](original-standard.json), including every repeated-probe statistic, accompany [current results](../scores.json).
135
 
136
- ## Release history and limitations
137
 
138
- [Release history](release-history.md) keeps all measured improvements and regressions against previous Vela versions, including retrieval and zero-shot regressions, with their original identities. Those historical comparisons do not replace the original small comparator above. [Product progress JSON](net-progress.json) retains its explicit weights; it is not an official aggregate or an overall SOTA claim.
139
 
140
- Cross-modal alignment uses licensed COCO and LibriSpeech TRAIN examples. The text foundation's pretraining overlap with downstream benchmarks is not ruled out. Benchmark encoders are frozen; official probes use their declared labeled data. No current-artifact latency or peak-memory improvement is claimed. Native operation and context limits are described in the model card.
141
 
142
- [Every plotted observation](pareto-data.json) and [Pareto methodology](pareto-methodology.md) preserve reported registry snapshots and separately measured peers. Known peers remain in the plots; individual task frontiers do not establish overall size SOTA.
 
1
+ # Evaluation
2
 
3
+ Current native artifact: `474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2`. Scores are percentages or official scores multiplied by 100. **Bold means an improvement over the original small**, measured on the same named protocol. Known test pools are reused across releases.
4
 
5
  ## Routing and cross-modal retrieval
6
 
7
+ | Metric | Original small | Current Nano |
 
 
8
  | --- | ---: | ---: |
9
  | Banking77 · Accuracy | 70.4221 | **87.9870** |
10
  | MASSIVE English · Accuracy | 65.9489 | **81.9650** |
 
14
  | COCO · Text → image · R@1 | 30.1823 | **55.8202** |
15
  | COCO · Text → image · R@5 | 59.4168 | **85.2248** |
16
  | COCO · Text → image · R@10 | 74.2892 | **94.2649** |
17
+ | LibriSpeech · Audio → text · R@1 | 4.2129 | **16.1241** |
18
+ | LibriSpeech · Audio → text · R@5 | 11.9877 | **34.8142** |
19
+ | LibriSpeech · Audio → text · R@10 | 19.0349 | **45.6147** |
20
+ | LibriSpeech · Text → audio · R@1 | 9.5785 | **20.3448** |
21
+ | LibriSpeech · Text → audio · R@5 | 22.5287 | **39.9617** |
22
+ | LibriSpeech · Text → audio · R@10 | 30.6897 | **50.6130** |
23
  | Banking77 · Macro-F1 | 70.2078 | **87.9168** |
24
  | MASSIVE English · Macro-F1 | 61.7178 | **79.2322** |
25
 
26
+ The text pools use TRAIN class prototypes; they are distinct from official MTEB classifier probes. All image/audio retrieval pools and all annotated positives are retained. [Counts, class F1 and paired intervals](../scores.json). Fresh paired intervals compare the previous Vela Nano; no new original-small interval is claimed.
 
 
 
 
 
 
 
 
 
 
 
 
 
27
 
28
+ ## Complete panels
29
 
30
+ | Panel | Mean(Task) | Mean(TaskType) |
31
  | --- | ---: | ---: |
32
+ | English · 41 tasks / 7 types | 64.8843 | 60.7818 |
33
+ | Audio · 19 tasks / 134 slots / 5 types | 43.5938 | 52.3354 |
34
+
35
+ English retains its actual evaluated artifact through exact text-component applicability. Every audio task is newly evaluated on this full native artifact. The original small has not been measured on these complete panels. [Full rankings](complete-panel-ranks.md).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
 
37
  ### Every audio task
38
 
39
+ | Task | Primary metric | Previous Nano | Current Nano | Delta pp |
40
+ | --- | --- | ---: | ---: | ---: |
41
+ | JamAltArtistA2ARetrieval | ndcg_at_10 | 68.9005 | 85.8015 | +16.9010 |
42
+ | BeijingOpera | accuracy | 68.6791 | 79.2287 | +10.5496 |
43
+ | BirdCLEF | accuracy | 12.9000 | 15.2000 | +2.3000 |
44
+ | CREMA_D | accuracy | 29.7637 | 32.0481 | +2.2845 |
45
+ | CommonLanguageAgeDetection | accuracy | 16.0400 | 16.6550 | +0.6150 |
46
+ | GTZANGenre | accuracy | 50.3000 | 61.7000 | +11.4000 |
47
+ | IEMOCAPGender | accuracy | 58.4519 | 61.6897 | +3.2379 |
48
+ | MInDS14 | accuracy | 23.3212 | 21.8520 | -1.4691 |
49
+ | MridinghamTonic | accuracy | 32.3632 | 58.4198 | +26.0566 |
50
+ | SIBFLEURS | accuracy | 18.3047 | 18.0725 | -0.2321 |
51
+ | VoxCelebSA | accuracy | 33.1691 | 33.1685 | -0.0006 |
52
+ | VoxPopuliLanguageID | accuracy | 91.8000 | 91.2000 | -0.6000 |
53
+ | CREMA_DClustering | v_measure | 2.3602 | 5.2143 | +2.8542 |
54
+ | VehicleSoundClustering | v_measure | 13.8007 | 5.2754 | -8.5253 |
55
+ | VoxPopuliGenderClustering | v_measure | 0.1045 | 0.0233 | -0.0813 |
56
+ | CREMADPairClassification | max_ap | 54.2012 | 55.2073 | +1.0060 |
57
+ | NMSQAPairClassification | max_ap | 63.2834 | 62.8993 | -0.3841 |
58
+ | VoxPopuliAccentPairClassification | max_ap | 54.0220 | 54.1514 | +0.1294 |
59
+ | GTZANAudioReranking | map_at_1000 | 63.7940 | 70.4760 | +6.6820 |
60
+
61
+ [English raw metrics](mteb-eng-v2.json) · [Audio raw metrics](maeb-audio-only.json) · [Exact component applicability](component-equivalence.json).
62
+
63
+ ## Selected standard protocols
64
+
65
+ These 11 selected tasks are not a complete MTEB, MMTEB, MIEB or MAEB benchmark. The original small retains its native 128-token context; Nano uses 512. Official FP32 batch-eight evaluation uses the same pinned datasets, splits and seed. MTOP validation and test remain separate.
66
 
67
  | Task / split | Main metric | Original small | Current Nano |
68
+ | --- | --- | ---: | ---: |
69
  | SICK-R / test | cosine_spearman | 71.3902 | **80.5317** |
70
  | STSBenchmark / test | cosine_spearman | 77.2987 | **87.0782** |
71
  | ArguAna / test | ndcg_at_10 | 32.8890 | **59.2460** |
 
75
  | OxfordPets / test | accuracy | 41.8479 | **91.2674** |
76
  | TinyImageNetClustering / valid | nmi | 46.2849 | **61.6919** |
77
  | OxfordPetsZeroShot / test | accuracy | 16.4077 | 8.0131 |
78
+ | CREMA_D / train | accuracy | 29.9520 | **32.0481** |
79
+ | CREMA_DClustering / train | v_measure | 2.4058 | **5.2143** |
80
+ | SpeechCommandsZeroshotv0.02 / test | accuracy | 5.5228 | **15.7094** |
81
+
82
+ Eight text/image tasks retain the actual previous artifact measurement through exact component identity; three audio tasks are fresh. Every named auxiliary metric remains in scores.json.
83
+
84
+ ## Event routing diagnostic
85
 
86
+ On the fixed 840-clip FSD50K evaluation pool, positive recall@5 is 66.2157% versus 27.3296% previously. The 3,299 TRAIN clips define 149 class prototypes; all 153 known labels retain their annotated-positive denominator, including labels without a TRAIN prototype. This diagnostic is outside the weighted product score. [Protocol and R@1/5/10](event-audio-diagnostic.json).
87
 
88
+ ## Resources and data
89
 
90
+ Total parameters: 163,771,288; FP32 vector storage: 1,536 bytes, excluding index metadata. Native weights: 655,583,800 bytes. No isolated latency or deployment-memory claim is made for this update. The additional CLAP branch costs inference computation as well as parameters.
91
 
92
+ The existing Whisper Tiny encoder and affine head stay frozen. A frozen CLAP audio tower and a 384×512 residual projection are added; only that projection is trained for one fixed 600-step run. Speech alignment and a pointwise parent anchor use 28,535 LibriSpeech TRAIN clips; known-positive event labels and CLAP relational geometry use 3,299 FSD50K TRAIN clips. Center/scale and the geometry coefficient use TRAIN only.
93
 
94
+ The frozen GIST backbone discloses MTEB classification training in its model card; independent pretraining-data separation is not claimed. Original per-record licenses and complete attribution are retained. [Release progress and all regressions](release-history.md).
benchmarks/complete-panel-data.json CHANGED
@@ -24,14 +24,16 @@
24
  ],
25
  "sources": {
26
  "nano_English41": {
27
- "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/mteb-eng-v2.json",
28
- "sha256": "733f6a58faae3d04e1b5a59de73085b156dc9987ba0e04e28548c17dc5904cec",
29
- "evidence_type": "Vela_published_report"
 
30
  },
31
  "nano_audio19": {
32
- "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/maeb-audio-only.json",
33
- "sha256": "e0259c59e04649f42cd2f004d57b525c79442297d28498a7434bddb4b000bdf9",
34
- "evidence_type": "Vela_published_report"
 
35
  },
36
  "mini_English41": {
37
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/mteb-eng-v2.json",
@@ -64,14 +66,20 @@
64
  "sha256": "8bbbd3a362d84a6e8f030803259db69e932e5997b1c7af1a276a0d17862eb184",
65
  "evidence_type": "exact_component_identity_report",
66
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731"
 
 
 
 
 
 
67
  }
68
  },
69
  "models": {
70
  "nano": {
71
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Nano",
72
- "release_commit": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806",
73
- "parameters_exact": 135383808,
74
- "current_native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
75
  "task_evidence": {
76
  "English41": {
77
  "source_id": "nano_English41",
@@ -79,47 +87,58 @@
79
  "model": "local/Vela-Omni-Nano-GIST-Rotation",
80
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
81
  "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
 
82
  "component_applicability": {
83
  "component": "text",
84
- "applies_to_revision": "self",
85
- "applies_to_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
86
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
87
- "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
88
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
89
- "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
90
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/component-equivalence.json",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
91
  "new_full_panel_inference": false,
92
- "basis": "199 exact text tensors and byte-identical inference/tokenizer/configuration under the same official evaluator protocol."
93
- },
94
- "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
95
- "inference_profile_applicability": {
96
- "revision": "self",
97
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
98
- "parameters": 135383808,
99
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
100
- "new_benchmark_run": false
101
  }
102
  },
103
- "self_reference_interpretation": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806"
104
  },
105
  "audio19": {
106
  "source_id": "nano_audio19",
107
  "identity": {
108
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
109
- "evaluated_revision": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
110
- "native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
111
- "revision_semantics": "Native artifact fingerprint identifying the evaluated weights and inference files; the release commit is the repository revision containing this report.",
112
- "inference_profile_applicability": {
113
- "revision": "self",
114
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
115
- "parameters": 135383808,
116
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
117
- "new_benchmark_run": false
118
- }
119
  },
120
- "self_reference_interpretation": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806"
121
  }
122
- }
 
 
123
  },
124
  "mini": {
125
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Mini",
@@ -15579,10 +15598,10 @@
15579
  "id": "current:nano",
15580
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
15581
  "kind": "nano",
15582
- "parameters_exact": 135383808,
15583
- "parameters_billion": 0.135383808,
15584
  "parameter_basis": "exact full released native tensor count, all modalities",
15585
- "current_native_fingerprint": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
15586
  "mean_task": 0.6488431514310018,
15587
  "mean_task_type": 0.6078184975654924,
15588
  "by_task_type": {
@@ -22560,23 +22579,25 @@
22560
  "global_rank": 65,
22561
  "global_population": 188,
22562
  "tied_other_models": 0,
22563
- "same_or_smaller_rank": 5,
22564
- "same_or_smaller_population": 66,
22565
  "best_same_or_smaller_other_model": {
22566
- "id": "scores_MTEB_eng_v2:row:61",
22567
- "model": "avsolatorio/GIST-Embedding-v0",
22568
  "kind": "registry_reported",
22569
- "parameters_billion": 0.109,
22570
- "mean_task": 0.6550466210091898
22571
  },
22572
- "gap_pp_above_best_same_or_smaller": -0.6203469578187959,
22573
  "observed_frontier": false,
22574
- "dominator_count": 4,
22575
  "all_dominator_ids": [
22576
  "scores_MTEB_eng_v2:row:61",
22577
  "scores_MTEB_eng_v2:row:65",
 
22578
  "scores_MTEB_eng_v2:row:71",
22579
- "scores_MTEB_eng_v2:row:72"
 
22580
  ]
22581
  },
22582
  "mean_task_type": {
@@ -22585,8 +22606,8 @@
22585
  "global_rank": 65,
22586
  "global_population": 188,
22587
  "tied_other_models": 0,
22588
- "same_or_smaller_rank": 3,
22589
- "same_or_smaller_population": 66,
22590
  "best_same_or_smaller_other_model": {
22591
  "id": "scores_MTEB_eng_v2:row:61",
22592
  "model": "avsolatorio/GIST-Embedding-v0",
@@ -22596,10 +22617,12 @@
22596
  },
22597
  "gap_pp_above_best_same_or_smaller": -0.6140745488965593,
22598
  "observed_frontier": false,
22599
- "dominator_count": 2,
22600
  "all_dominator_ids": [
22601
  "scores_MTEB_eng_v2:row:61",
22602
- "scores_MTEB_eng_v2:row:65"
 
 
22603
  ]
22604
  }
22605
  },
@@ -23193,7 +23216,7 @@
23193
  "id": "current:nano",
23194
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
23195
  "score": 0.6488431514310018,
23196
- "parameters_billion": 0.135383808
23197
  },
23198
  {
23199
  "rank": 66,
@@ -24511,7 +24534,7 @@
24511
  "id": "current:nano",
24512
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
24513
  "score": 0.6078184975654924,
24514
- "parameters_billion": 0.135383808
24515
  },
24516
  {
24517
  "rank": 66,
@@ -29320,39 +29343,39 @@
29320
  "id": "current:nano",
29321
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
29322
  "kind": "nano",
29323
- "parameters_exact": 135383808,
29324
- "parameters_billion": 0.135383808,
29325
  "parameter_basis": "exact full released native tensor count, all modalities",
29326
- "current_native_fingerprint": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
29327
- "mean_task": 0.3976628038845883,
29328
- "mean_task_type": 0.4696781408655678,
29329
  "by_task_type": {
29330
- "Any2AnyRetrieval": 0.689005,
29331
- "AudioClassification": 0.39553889510295764,
29332
- "AudioClustering": 0.054218008138417706,
29333
- "AudioPairClassification": 0.5716888010864635,
29334
- "AudioReranking": 0.63794
29335
  },
29336
  "scores_by_task": {
29337
- "JamAltArtistA2ARetrieval": 0.689005,
29338
- "BeijingOpera": 0.686790780141844,
29339
- "BirdCLEF": 0.129,
29340
- "CREMA_D": 0.29763670140167686,
29341
- "CommonLanguageAgeDetection": 0.1604,
29342
- "GTZANGenre": 0.503,
29343
- "IEMOCAPGender": 0.5845186270364481,
29344
- "MInDS14": 0.23321168947951418,
29345
- "MridinghamTonic": 0.32363208758254514,
29346
- "SIBFLEURS": 0.1830465783151205,
29347
- "VoxCelebSA": 0.33169138217538546,
29348
- "VoxPopuliLanguageID": 0.9179999999999999,
29349
- "CREMA_DClustering": 0.023601715918443712,
29350
- "VehicleSoundClustering": 0.13800701178563807,
29351
- "VoxPopuliGenderClustering": 0.0010452967111713475,
29352
- "CREMADPairClassification": 0.5420123629581632,
29353
- "NMSQAPairClassification": 0.6328340215485784,
29354
- "VoxPopuliAccentPairClassification": 0.5402200187526489,
29355
- "GTZANAudioReranking": 0.63794
29356
  },
29357
  "coverage": {
29358
  "complete": true,
@@ -29658,13 +29681,13 @@
29658
  "current_ranks": {
29659
  "nano": {
29660
  "mean_task": {
29661
- "score": 0.3976628038845883,
29662
- "display_score_100": 39.76628038845883,
29663
- "global_rank": 34,
29664
  "global_population": 64,
29665
  "tied_other_models": 0,
29666
- "same_or_smaller_rank": 6,
29667
- "same_or_smaller_population": 22,
29668
  "best_same_or_smaller_other_model": {
29669
  "id": "scores_MAEB_beta_audio-only:row:9",
29670
  "model": "openai/whisper-base",
@@ -29672,43 +29695,40 @@
29672
  "parameters_billion": 0.074,
29673
  "mean_task": 0.44856370433436527
29674
  },
29675
- "gap_pp_above_best_same_or_smaller": -5.090090044977696,
29676
  "observed_frontier": false,
29677
- "dominator_count": 5,
29678
  "all_dominator_ids": [
29679
  "scores_MAEB_beta_audio-only:row:9",
29680
- "scores_MAEB_beta_audio-only:row:12",
29681
  "scores_MAEB_beta_audio-only:row:18",
29682
- "scores_MAEB_beta_audio-only:row:20",
29683
- "scores_MAEB_beta_audio-only:row:25"
29684
  ]
29685
  },
29686
  "mean_task_type": {
29687
- "score": 0.4696781408655678,
29688
- "display_score_100": 46.96781408655678,
29689
- "global_rank": 32,
29690
  "global_population": 64,
29691
  "tied_other_models": 0,
29692
- "same_or_smaller_rank": 8,
29693
- "same_or_smaller_population": 22,
29694
  "best_same_or_smaller_other_model": {
29695
- "id": "scores_MAEB_beta_audio-only:row:20",
29696
- "model": "MIT/ast-finetuned-audioset-10-10-0.4593",
29697
  "kind": "registry_reported",
29698
- "parameters_billion": 0.087,
29699
- "mean_task_type": 0.55212392228164
29700
  },
29701
- "gap_pp_above_best_same_or_smaller": -8.244578141607217,
29702
  "observed_frontier": false,
29703
- "dominator_count": 7,
29704
  "all_dominator_ids": [
29705
- "scores_MAEB_beta_audio-only:row:9",
29706
- "scores_MAEB_beta_audio-only:row:12",
29707
- "scores_MAEB_beta_audio-only:row:18",
29708
  "scores_MAEB_beta_audio-only:row:20",
29709
- "scores_MAEB_beta_audio-only:row:25",
29710
- "scores_MAEB_beta_audio-only:row:27",
29711
- "scores_MAEB_beta_audio-only:row:28"
29712
  ]
29713
  }
29714
  },
@@ -29914,95 +29934,95 @@
29914
  },
29915
  {
29916
  "rank": 22,
 
 
 
 
 
 
 
29917
  "id": "scores_MAEB_beta_audio-only:row:19",
29918
  "model": "laion/clap-htsat-fused",
29919
  "score": 0.43475542389060884,
29920
  "parameters_billion": 0.154
29921
  },
29922
  {
29923
- "rank": 23,
29924
  "id": "scores_MAEB_beta_audio-only:row:14",
29925
  "model": "laion/clap-htsat-unfused",
29926
  "score": 0.42935003353973167,
29927
  "parameters_billion": 0.153
29928
  },
29929
  {
29930
- "rank": 24,
29931
  "id": "scores_MAEB_beta_audio-only:row:41",
29932
  "model": "facebook/mms-1b-fl102",
29933
  "score": 0.41367276031991745,
29934
  "parameters_billion": 1.0
29935
  },
29936
  {
29937
- "rank": 25,
29938
  "id": "scores_MAEB_beta_audio-only:row:42",
29939
  "model": "facebook/seamless-m4t-v2-large",
29940
  "score": 0.4131003795149639,
29941
  "parameters_billion": 2.3
29942
  },
29943
  {
29944
- "rank": 26,
29945
  "id": "scores_MAEB_beta_audio-only:row:25",
29946
  "model": "google/vggish",
29947
  "score": 0.4109744040247678,
29948
  "parameters_billion": 0.072
29949
  },
29950
  {
29951
- "rank": 27,
29952
  "id": "scores_MAEB_beta_audio-only:row:12",
29953
  "model": "matthewagi/HeAR-s1.1",
29954
  "score": 0.41088769685242515,
29955
  "parameters_billion": 0.022
29956
  },
29957
  {
29958
- "rank": 28,
29959
  "id": "scores_MAEB_beta_audio-only:row:30",
29960
  "model": "facebook/wav2vec2-xls-r-2b",
29961
  "score": 0.4095841483488132,
29962
  "parameters_billion": 2.0
29963
  },
29964
  {
29965
- "rank": 29,
29966
  "id": "scores_MAEB_beta_audio-only:row:22",
29967
  "model": "OpenMuQ/MuQ-MuLan-large",
29968
  "score": 0.4089651437048504,
29969
  "parameters_billion": 0.63
29970
  },
29971
  {
29972
- "rank": 30,
29973
  "id": "scores_MAEB_beta_audio-only:row:35",
29974
  "model": "facebook/mms-1b-all",
29975
  "score": 0.4070234414344685,
29976
  "parameters_billion": 1.0
29977
  },
29978
  {
29979
- "rank": 31,
29980
  "id": "scores_MAEB_beta_audio-only:row:33",
29981
  "model": "microsoft/msclap-2022",
29982
  "score": 0.40099675438596494,
29983
  "parameters_billion": 0.196
29984
  },
29985
  {
29986
- "rank": 32,
29987
  "id": "scores_MAEB_beta_audio-only:row:31",
29988
  "model": "facebook/mms-1b-l1107",
29989
  "score": 0.3998511984004128,
29990
  "parameters_billion": 1.0
29991
  },
29992
  {
29993
- "rank": 33,
29994
  "id": "scores_MAEB_beta_audio-only:row:29",
29995
  "model": "facebook/wav2vec2-lv-60-espeak-cv-ft",
29996
  "score": 0.3983267115583075,
29997
  "parameters_billion": 0.317
29998
  },
29999
- {
30000
- "rank": 34,
30001
- "id": "current:nano",
30002
- "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
30003
- "score": 0.3976628038845883,
30004
- "parameters_billion": 0.135383808
30005
- },
30006
  {
30007
  "rank": 35,
30008
  "id": "scores_MAEB_beta_audio-only:row:26",
@@ -30343,102 +30363,102 @@
30343
  },
30344
  {
30345
  "rank": 19,
 
 
 
 
 
 
 
30346
  "id": "scores_MAEB_beta_audio-only:row:33",
30347
  "model": "microsoft/msclap-2022",
30348
  "score": 0.5156324,
30349
  "parameters_billion": 0.196
30350
  },
30351
  {
30352
- "rank": 20,
30353
  "id": "scores_MAEB_beta_audio-only:row:23",
30354
  "model": "EximiusLabs/fusion-embedding-2-2b-preview",
30355
  "score": 0.5116159368092692,
30356
  "parameters_billion": 2.8
30357
  },
30358
  {
30359
- "rank": 21,
30360
  "id": "scores_MAEB_beta_audio-only:row:28",
30361
  "model": "google/yamnet",
30362
  "score": 0.510955464795009,
30363
  "parameters_billion": 0.004
30364
  },
30365
  {
30366
- "rank": 22,
30367
  "id": "scores_MAEB_beta_audio-only:row:1",
30368
  "model": "openai/whisper-medium",
30369
  "score": 0.5071337459893048,
30370
  "parameters_billion": 0.769
30371
  },
30372
  {
30373
- "rank": 23,
30374
  "id": "scores_MAEB_beta_audio-only:row:12",
30375
  "model": "matthewagi/HeAR-s1.1",
30376
  "score": 0.5022611316399287,
30377
  "parameters_billion": 0.022
30378
  },
30379
  {
30380
- "rank": 24,
30381
  "id": "scores_MAEB_beta_audio-only:row:0",
30382
  "model": "Qwen/Qwen2-Audio-7B",
30383
  "score": 0.49261104875222816,
30384
  "parameters_billion": 7.0
30385
  },
30386
  {
30387
- "rank": 25,
30388
  "id": "scores_MAEB_beta_audio-only:row:24",
30389
  "model": "lyrebird/wav2clip",
30390
  "score": 0.48892121889483064,
30391
  "parameters_billion": 0.163
30392
  },
30393
  {
30394
- "rank": 26,
30395
  "id": "scores_MAEB_beta_audio-only:row:26",
30396
  "model": "microsoft/wavlm-large",
30397
  "score": 0.4848395147058824,
30398
  "parameters_billion": 0.317
30399
  },
30400
  {
30401
- "rank": 27,
30402
  "id": "scores_MAEB_beta_audio-only:row:9",
30403
  "model": "openai/whisper-base",
30404
  "score": 0.48299134028520496,
30405
  "parameters_billion": 0.074
30406
  },
30407
  {
30408
- "rank": 28,
30409
  "id": "scores_MAEB_beta_audio-only:row:13",
30410
  "model": "openai/whisper-large-v3",
30411
  "score": 0.4818755459001783,
30412
  "parameters_billion": 1.55
30413
  },
30414
  {
30415
- "rank": 29,
30416
  "id": "scores_MAEB_beta_audio-only:row:11",
30417
  "model": "openai/whisper-small",
30418
  "score": 0.48186107352941177,
30419
  "parameters_billion": 0.244
30420
  },
30421
  {
30422
- "rank": 30,
30423
  "id": "scores_MAEB_beta_audio-only:row:18",
30424
  "model": "openai/whisper-tiny",
30425
  "score": 0.4774227983957219,
30426
  "parameters_billion": 0.039
30427
  },
30428
  {
30429
- "rank": 31,
30430
  "id": "scores_MAEB_beta_audio-only:row:27",
30431
  "model": "facebook/hubert-base-ls960",
30432
  "score": 0.4762437683600713,
30433
  "parameters_billion": 0.095
30434
  },
30435
- {
30436
- "rank": 32,
30437
- "id": "current:nano",
30438
- "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
30439
- "score": 0.4696781408655678,
30440
- "parameters_billion": 0.135383808
30441
- },
30442
  {
30443
  "rank": 33,
30444
  "id": "scores_MAEB_beta_audio-only:row:45",
@@ -30798,15 +30818,15 @@
30798
  "focus": "nano",
30799
  "same_size_top10_ids": [
30800
  "scores_MTEB_eng_v2:row:61",
 
30801
  "scores_MTEB_eng_v2:row:65",
 
30802
  "current:nano",
30803
  "scores_MTEB_eng_v2:row:71",
30804
  "scores_MTEB_eng_v2:row:66",
30805
  "scores_MTEB_eng_v2:row:72",
30806
  "scores_MTEB_eng_v2:row:79",
30807
- "scores_MTEB_eng_v2:row:80",
30808
- "scores_MTEB_eng_v2:row:83",
30809
- "scores_MTEB_eng_v2:row:86"
30810
  ],
30811
  "all_known_size_point_ids": [
30812
  "scores_MTEB_eng_v2:row:0",
 
24
  ],
25
  "sources": {
26
  "nano_English41": {
27
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/mteb-eng-v2.json",
28
+ "evidence_type": "Vela_release_report",
29
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
30
+ "sha256": "f3daecad1cf431630598efb04f3667431e3420239240f28a51568a74cef0c64f"
31
  },
32
  "nano_audio19": {
33
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
34
+ "evidence_type": "Vela_release_report",
35
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
36
+ "sha256": "4fc7f217745eb0aa998bf0f8e044bd1fa5de2f73f1a7dead7b829a62c48669ff"
37
  },
38
  "mini_English41": {
39
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/mteb-eng-v2.json",
 
66
  "sha256": "8bbbd3a362d84a6e8f030803259db69e932e5997b1c7af1a276a0d17862eb184",
67
  "evidence_type": "exact_component_identity_report",
68
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731"
69
+ },
70
+ "nano_component_equivalence": {
71
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/component-equivalence.json",
72
+ "evidence_type": "Vela_release_report",
73
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
74
+ "sha256": "ebd40f9a887791d1736132c37f510a69e18a03e247422853f54a42b417b0db9b"
75
  }
76
  },
77
  "models": {
78
  "nano": {
79
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Nano",
80
+ "release_commit": null,
81
+ "parameters_exact": 163771288,
82
+ "current_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
83
  "task_evidence": {
84
  "English41": {
85
  "source_id": "nano_English41",
 
87
  "model": "local/Vela-Omni-Nano-GIST-Rotation",
88
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
89
  "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
90
+ "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
91
  "component_applicability": {
92
  "component": "text",
 
 
93
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
 
94
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
95
+ "applies_to_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
96
+ "prior_component_applicability": {
97
+ "model": "local/Vela-Omni-Nano-GIST-Rotation",
98
+ "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
99
+ "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
100
+ "component_applicability": {
101
+ "component": "text",
102
+ "applies_to_revision": "self",
103
+ "applies_to_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
104
+ "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
105
+ "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
106
+ "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
107
+ "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
108
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/component-equivalence.json",
109
+ "new_full_panel_inference": false,
110
+ "basis": "199 exact text tensors and byte-identical inference/tokenizer/configuration under the same official evaluator protocol."
111
+ },
112
+ "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
113
+ "inference_profile_applicability": {
114
+ "revision": "self",
115
+ "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
116
+ "parameters": 135383808,
117
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
118
+ "new_benchmark_run": false
119
+ }
120
+ },
121
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/component-equivalence.json",
122
  "new_full_panel_inference": false,
123
+ "basis": "The frozen text tensors, tokenizer, configuration and encoding computation are unchanged. The original measured text identity is retained."
 
 
 
 
 
 
 
 
124
  }
125
  },
126
+ "self_reference_interpretation": "Nano native fingerprint 474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2"
127
  },
128
  "audio19": {
129
  "source_id": "nano_audio19",
130
  "identity": {
131
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
132
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
133
+ "release_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
134
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
135
+ "new_full_panel_inference": true
 
 
 
 
 
 
136
  },
137
+ "self_reference_interpretation": "Nano native fingerprint 474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2"
138
  }
139
+ },
140
+ "release_reference": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano",
141
+ "release_revision_semantics": "Exact native fingerprint below; Hub commit is not encoded in this file."
142
  },
143
  "mini": {
144
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Mini",
 
15598
  "id": "current:nano",
15599
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
15600
  "kind": "nano",
15601
+ "parameters_exact": 163771288,
15602
+ "parameters_billion": 0.163771288,
15603
  "parameter_basis": "exact full released native tensor count, all modalities",
15604
+ "current_native_fingerprint": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
15605
  "mean_task": 0.6488431514310018,
15606
  "mean_task_type": 0.6078184975654924,
15607
  "by_task_type": {
 
22579
  "global_rank": 65,
22580
  "global_population": 188,
22581
  "tied_other_models": 0,
22582
+ "same_or_smaller_rank": 7,
22583
+ "same_or_smaller_population": 75,
22584
  "best_same_or_smaller_other_model": {
22585
+ "id": "scores_MTEB_eng_v2:row:68",
22586
+ "model": "Tarka-AIR/Tarka-Embedding-150M-V1",
22587
  "kind": "registry_reported",
22588
+ "parameters_billion": 0.156,
22589
+ "mean_task": 0.6639085365853659
22590
  },
22591
+ "gap_pp_above_best_same_or_smaller": -1.5065385154364064,
22592
  "observed_frontier": false,
22593
+ "dominator_count": 6,
22594
  "all_dominator_ids": [
22595
  "scores_MTEB_eng_v2:row:61",
22596
  "scores_MTEB_eng_v2:row:65",
22597
+ "scores_MTEB_eng_v2:row:68",
22598
  "scores_MTEB_eng_v2:row:71",
22599
+ "scores_MTEB_eng_v2:row:72",
22600
+ "scores_MTEB_eng_v2:row:74"
22601
  ]
22602
  },
22603
  "mean_task_type": {
 
22606
  "global_rank": 65,
22607
  "global_population": 188,
22608
  "tied_other_models": 0,
22609
+ "same_or_smaller_rank": 5,
22610
+ "same_or_smaller_population": 75,
22611
  "best_same_or_smaller_other_model": {
22612
  "id": "scores_MTEB_eng_v2:row:61",
22613
  "model": "avsolatorio/GIST-Embedding-v0",
 
22617
  },
22618
  "gap_pp_above_best_same_or_smaller": -0.6140745488965593,
22619
  "observed_frontier": false,
22620
+ "dominator_count": 4,
22621
  "all_dominator_ids": [
22622
  "scores_MTEB_eng_v2:row:61",
22623
+ "scores_MTEB_eng_v2:row:65",
22624
+ "scores_MTEB_eng_v2:row:68",
22625
+ "scores_MTEB_eng_v2:row:74"
22626
  ]
22627
  }
22628
  },
 
23216
  "id": "current:nano",
23217
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
23218
  "score": 0.6488431514310018,
23219
+ "parameters_billion": 0.163771288
23220
  },
23221
  {
23222
  "rank": 66,
 
24534
  "id": "current:nano",
24535
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
24536
  "score": 0.6078184975654924,
24537
+ "parameters_billion": 0.163771288
24538
  },
24539
  {
24540
  "rank": 66,
 
29343
  "id": "current:nano",
29344
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
29345
  "kind": "nano",
29346
+ "parameters_exact": 163771288,
29347
+ "parameters_billion": 0.163771288,
29348
  "parameter_basis": "exact full released native tensor count, all modalities",
29349
+ "current_native_fingerprint": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
29350
+ "mean_task": 0.43593843064595883,
29351
+ "mean_task_type": 0.5233540613685161,
29352
  "by_task_type": {
29353
+ "Any2AnyRetrieval": 0.858015,
29354
+ "AudioClassification": 0.44475865771818446,
29355
+ "AudioClustering": 0.03504342925419188,
29356
+ "AudioPairClassification": 0.5741932198702043,
29357
+ "AudioReranking": 0.70476
29358
  },
29359
  "scores_by_task": {
29360
+ "JamAltArtistA2ARetrieval": 0.858015,
29361
+ "BeijingOpera": 0.7922872340425531,
29362
+ "BirdCLEF": 0.15199999999999997,
29363
+ "CREMA_D": 0.32048146984697823,
29364
+ "CommonLanguageAgeDetection": 0.16655,
29365
+ "GTZANGenre": 0.617,
29366
+ "IEMOCAPGender": 0.6168973830636596,
29367
+ "MInDS14": 0.21852038094350856,
29368
+ "MridinghamTonic": 0.5841978617863635,
29369
+ "SIBFLEURS": 0.18072541269472217,
29370
+ "VoxCelebSA": 0.331685492522244,
29371
+ "VoxPopuliLanguageID": 0.9120000000000001,
29372
+ "CREMA_DClustering": 0.05214346580804754,
29373
+ "VehicleSoundClustering": 0.05275425308186281,
29374
+ "VoxPopuliGenderClustering": 0.00023256887266527633,
29375
+ "CREMADPairClassification": 0.5520726152267248,
29376
+ "NMSQAPairClassification": 0.6289933955933497,
29377
+ "VoxPopuliAccentPairClassification": 0.5415136487905384,
29378
+ "GTZANAudioReranking": 0.70476
29379
  },
29380
  "coverage": {
29381
  "complete": true,
 
29681
  "current_ranks": {
29682
  "nano": {
29683
  "mean_task": {
29684
+ "score": 0.43593843064595883,
29685
+ "display_score_100": 43.59384306459588,
29686
+ "global_rank": 22,
29687
  "global_population": 64,
29688
  "tied_other_models": 0,
29689
+ "same_or_smaller_rank": 5,
29690
+ "same_or_smaller_population": 27,
29691
  "best_same_or_smaller_other_model": {
29692
  "id": "scores_MAEB_beta_audio-only:row:9",
29693
  "model": "openai/whisper-base",
 
29695
  "parameters_billion": 0.074,
29696
  "mean_task": 0.44856370433436527
29697
  },
29698
+ "gap_pp_above_best_same_or_smaller": -1.2625273688406435,
29699
  "observed_frontier": false,
29700
+ "dominator_count": 4,
29701
  "all_dominator_ids": [
29702
  "scores_MAEB_beta_audio-only:row:9",
29703
+ "scores_MAEB_beta_audio-only:row:17",
29704
  "scores_MAEB_beta_audio-only:row:18",
29705
+ "scores_MAEB_beta_audio-only:row:20"
 
29706
  ]
29707
  },
29708
  "mean_task_type": {
29709
+ "score": 0.5233540613685161,
29710
+ "display_score_100": 52.33540613685162,
29711
+ "global_rank": 19,
29712
  "global_population": 64,
29713
  "tied_other_models": 0,
29714
+ "same_or_smaller_rank": 6,
29715
+ "same_or_smaller_population": 27,
29716
  "best_same_or_smaller_other_model": {
29717
+ "id": "scores_MAEB_beta_audio-only:row:17",
29718
+ "model": "microsoft/msclap-2023",
29719
  "kind": "registry_reported",
29720
+ "parameters_billion": 0.16,
29721
+ "mean_task_type": 0.5584816090017826
29722
  },
29723
+ "gap_pp_above_best_same_or_smaller": -3.512754763326642,
29724
  "observed_frontier": false,
29725
+ "dominator_count": 5,
29726
  "all_dominator_ids": [
29727
+ "scores_MAEB_beta_audio-only:row:14",
29728
+ "scores_MAEB_beta_audio-only:row:17",
29729
+ "scores_MAEB_beta_audio-only:row:19",
29730
  "scores_MAEB_beta_audio-only:row:20",
29731
+ "scores_MAEB_beta_audio-only:row:25"
 
 
29732
  ]
29733
  }
29734
  },
 
29934
  },
29935
  {
29936
  "rank": 22,
29937
+ "id": "current:nano",
29938
+ "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
29939
+ "score": 0.43593843064595883,
29940
+ "parameters_billion": 0.163771288
29941
+ },
29942
+ {
29943
+ "rank": 23,
29944
  "id": "scores_MAEB_beta_audio-only:row:19",
29945
  "model": "laion/clap-htsat-fused",
29946
  "score": 0.43475542389060884,
29947
  "parameters_billion": 0.154
29948
  },
29949
  {
29950
+ "rank": 24,
29951
  "id": "scores_MAEB_beta_audio-only:row:14",
29952
  "model": "laion/clap-htsat-unfused",
29953
  "score": 0.42935003353973167,
29954
  "parameters_billion": 0.153
29955
  },
29956
  {
29957
+ "rank": 25,
29958
  "id": "scores_MAEB_beta_audio-only:row:41",
29959
  "model": "facebook/mms-1b-fl102",
29960
  "score": 0.41367276031991745,
29961
  "parameters_billion": 1.0
29962
  },
29963
  {
29964
+ "rank": 26,
29965
  "id": "scores_MAEB_beta_audio-only:row:42",
29966
  "model": "facebook/seamless-m4t-v2-large",
29967
  "score": 0.4131003795149639,
29968
  "parameters_billion": 2.3
29969
  },
29970
  {
29971
+ "rank": 27,
29972
  "id": "scores_MAEB_beta_audio-only:row:25",
29973
  "model": "google/vggish",
29974
  "score": 0.4109744040247678,
29975
  "parameters_billion": 0.072
29976
  },
29977
  {
29978
+ "rank": 28,
29979
  "id": "scores_MAEB_beta_audio-only:row:12",
29980
  "model": "matthewagi/HeAR-s1.1",
29981
  "score": 0.41088769685242515,
29982
  "parameters_billion": 0.022
29983
  },
29984
  {
29985
+ "rank": 29,
29986
  "id": "scores_MAEB_beta_audio-only:row:30",
29987
  "model": "facebook/wav2vec2-xls-r-2b",
29988
  "score": 0.4095841483488132,
29989
  "parameters_billion": 2.0
29990
  },
29991
  {
29992
+ "rank": 30,
29993
  "id": "scores_MAEB_beta_audio-only:row:22",
29994
  "model": "OpenMuQ/MuQ-MuLan-large",
29995
  "score": 0.4089651437048504,
29996
  "parameters_billion": 0.63
29997
  },
29998
  {
29999
+ "rank": 31,
30000
  "id": "scores_MAEB_beta_audio-only:row:35",
30001
  "model": "facebook/mms-1b-all",
30002
  "score": 0.4070234414344685,
30003
  "parameters_billion": 1.0
30004
  },
30005
  {
30006
+ "rank": 32,
30007
  "id": "scores_MAEB_beta_audio-only:row:33",
30008
  "model": "microsoft/msclap-2022",
30009
  "score": 0.40099675438596494,
30010
  "parameters_billion": 0.196
30011
  },
30012
  {
30013
+ "rank": 33,
30014
  "id": "scores_MAEB_beta_audio-only:row:31",
30015
  "model": "facebook/mms-1b-l1107",
30016
  "score": 0.3998511984004128,
30017
  "parameters_billion": 1.0
30018
  },
30019
  {
30020
+ "rank": 34,
30021
  "id": "scores_MAEB_beta_audio-only:row:29",
30022
  "model": "facebook/wav2vec2-lv-60-espeak-cv-ft",
30023
  "score": 0.3983267115583075,
30024
  "parameters_billion": 0.317
30025
  },
 
 
 
 
 
 
 
30026
  {
30027
  "rank": 35,
30028
  "id": "scores_MAEB_beta_audio-only:row:26",
 
30363
  },
30364
  {
30365
  "rank": 19,
30366
+ "id": "current:nano",
30367
+ "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
30368
+ "score": 0.5233540613685161,
30369
+ "parameters_billion": 0.163771288
30370
+ },
30371
+ {
30372
+ "rank": 20,
30373
  "id": "scores_MAEB_beta_audio-only:row:33",
30374
  "model": "microsoft/msclap-2022",
30375
  "score": 0.5156324,
30376
  "parameters_billion": 0.196
30377
  },
30378
  {
30379
+ "rank": 21,
30380
  "id": "scores_MAEB_beta_audio-only:row:23",
30381
  "model": "EximiusLabs/fusion-embedding-2-2b-preview",
30382
  "score": 0.5116159368092692,
30383
  "parameters_billion": 2.8
30384
  },
30385
  {
30386
+ "rank": 22,
30387
  "id": "scores_MAEB_beta_audio-only:row:28",
30388
  "model": "google/yamnet",
30389
  "score": 0.510955464795009,
30390
  "parameters_billion": 0.004
30391
  },
30392
  {
30393
+ "rank": 23,
30394
  "id": "scores_MAEB_beta_audio-only:row:1",
30395
  "model": "openai/whisper-medium",
30396
  "score": 0.5071337459893048,
30397
  "parameters_billion": 0.769
30398
  },
30399
  {
30400
+ "rank": 24,
30401
  "id": "scores_MAEB_beta_audio-only:row:12",
30402
  "model": "matthewagi/HeAR-s1.1",
30403
  "score": 0.5022611316399287,
30404
  "parameters_billion": 0.022
30405
  },
30406
  {
30407
+ "rank": 25,
30408
  "id": "scores_MAEB_beta_audio-only:row:0",
30409
  "model": "Qwen/Qwen2-Audio-7B",
30410
  "score": 0.49261104875222816,
30411
  "parameters_billion": 7.0
30412
  },
30413
  {
30414
+ "rank": 26,
30415
  "id": "scores_MAEB_beta_audio-only:row:24",
30416
  "model": "lyrebird/wav2clip",
30417
  "score": 0.48892121889483064,
30418
  "parameters_billion": 0.163
30419
  },
30420
  {
30421
+ "rank": 27,
30422
  "id": "scores_MAEB_beta_audio-only:row:26",
30423
  "model": "microsoft/wavlm-large",
30424
  "score": 0.4848395147058824,
30425
  "parameters_billion": 0.317
30426
  },
30427
  {
30428
+ "rank": 28,
30429
  "id": "scores_MAEB_beta_audio-only:row:9",
30430
  "model": "openai/whisper-base",
30431
  "score": 0.48299134028520496,
30432
  "parameters_billion": 0.074
30433
  },
30434
  {
30435
+ "rank": 29,
30436
  "id": "scores_MAEB_beta_audio-only:row:13",
30437
  "model": "openai/whisper-large-v3",
30438
  "score": 0.4818755459001783,
30439
  "parameters_billion": 1.55
30440
  },
30441
  {
30442
+ "rank": 30,
30443
  "id": "scores_MAEB_beta_audio-only:row:11",
30444
  "model": "openai/whisper-small",
30445
  "score": 0.48186107352941177,
30446
  "parameters_billion": 0.244
30447
  },
30448
  {
30449
+ "rank": 31,
30450
  "id": "scores_MAEB_beta_audio-only:row:18",
30451
  "model": "openai/whisper-tiny",
30452
  "score": 0.4774227983957219,
30453
  "parameters_billion": 0.039
30454
  },
30455
  {
30456
+ "rank": 32,
30457
  "id": "scores_MAEB_beta_audio-only:row:27",
30458
  "model": "facebook/hubert-base-ls960",
30459
  "score": 0.4762437683600713,
30460
  "parameters_billion": 0.095
30461
  },
 
 
 
 
 
 
 
30462
  {
30463
  "rank": 33,
30464
  "id": "scores_MAEB_beta_audio-only:row:45",
 
30818
  "focus": "nano",
30819
  "same_size_top10_ids": [
30820
  "scores_MTEB_eng_v2:row:61",
30821
+ "scores_MTEB_eng_v2:row:68",
30822
  "scores_MTEB_eng_v2:row:65",
30823
+ "scores_MTEB_eng_v2:row:74",
30824
  "current:nano",
30825
  "scores_MTEB_eng_v2:row:71",
30826
  "scores_MTEB_eng_v2:row:66",
30827
  "scores_MTEB_eng_v2:row:72",
30828
  "scores_MTEB_eng_v2:row:79",
30829
+ "scores_MTEB_eng_v2:row:80"
 
 
30830
  ],
30831
  "all_known_size_point_ids": [
30832
  "scores_MTEB_eng_v2:row:0",
benchmarks/complete-panel-metadata.json CHANGED
@@ -4,7 +4,7 @@
4
  "primary_aggregate": "Mean(TaskType)",
5
  "supplementary_aggregate": "Mean(Task)",
6
  "data_file": "complete-panel-data.json",
7
- "data_sha256": "fbe0dcb2b929bbc3f283f2fca30af768c9009feff9d2b288a97d995a928727f6",
8
  "score_display": "100 \u00d7 native score",
9
  "snapshot_date": "2026-09-17",
10
  "known_size_point_counts": {
@@ -198,15 +198,15 @@
198
  "unknown_size_count": 13,
199
  "same_size_top10_ids": [
200
  "scores_MTEB_eng_v2:row:61",
 
201
  "scores_MTEB_eng_v2:row:65",
 
202
  "current:nano",
203
  "scores_MTEB_eng_v2:row:71",
204
  "scores_MTEB_eng_v2:row:66",
205
  "scores_MTEB_eng_v2:row:72",
206
  "scores_MTEB_eng_v2:row:79",
207
- "scores_MTEB_eng_v2:row:80",
208
- "scores_MTEB_eng_v2:row:83",
209
- "scores_MTEB_eng_v2:row:86"
210
  ],
211
  "focus_frontier": false,
212
  "label_overlaps": [],
@@ -215,8 +215,8 @@
215
  "all_points_inside_axes": true,
216
  "all_other_panel_ranks_visible": true,
217
  "exports_sha256": {
218
- "complete-panel-english41-nano.png": "ea7e14d09f749c22b0d4d32414df5008f9d68c746f8614483b25108ff52c88cc",
219
- "complete-panel-english41-nano.svg": "79f18961e977139f72b402e83316b6dc523268817d241e61a0ef9efb54a20697"
220
  },
221
  "visible_reference_labels": [
222
  {
@@ -351,8 +351,8 @@
351
  "all_points_inside_axes": true,
352
  "all_other_panel_ranks_visible": true,
353
  "exports_sha256": {
354
- "complete-panel-audio19-mini.png": "69b0f8362667abaa5579dc5f78bdf49e918428f40e350f393dae451807ef9dd3",
355
- "complete-panel-audio19-mini.svg": "66c5e212b37de8896b38d16c03118dd2dd7a8f50fdd4b650b85c016f1aafe96c"
356
  },
357
  "visible_reference_labels": [
358
  {
@@ -406,7 +406,7 @@
406
  ]
407
  }
408
  },
409
- "visual_inspection": "All five PNGs inspected: white background, full eligible point coverage, preserved named references, no clipped or overlapping labels. Mini audio rank and task scores updated; SIB-FLEURS correctly describes topic classification.",
410
  "scope": "Complete dedicated snapshot rows; no per-task best stitching, no partial-panel ranking, no overall SOTA claim. Unknown-size complete peers count globally but cannot be plotted.",
411
  "ranking_population": "Dedicated registry rows plus both current Vela models, unique names asserted.",
412
  "source_urls_and_actual_measured_identity": "See sources and models in complete-panel-data.json."
 
4
  "primary_aggregate": "Mean(TaskType)",
5
  "supplementary_aggregate": "Mean(Task)",
6
  "data_file": "complete-panel-data.json",
7
+ "data_sha256": "ffaedbbc06e8db52a0071323c1b15c503a97871b03129710d7e1bdfcb9fd81f6",
8
  "score_display": "100 \u00d7 native score",
9
  "snapshot_date": "2026-09-17",
10
  "known_size_point_counts": {
 
198
  "unknown_size_count": 13,
199
  "same_size_top10_ids": [
200
  "scores_MTEB_eng_v2:row:61",
201
+ "scores_MTEB_eng_v2:row:68",
202
  "scores_MTEB_eng_v2:row:65",
203
+ "scores_MTEB_eng_v2:row:74",
204
  "current:nano",
205
  "scores_MTEB_eng_v2:row:71",
206
  "scores_MTEB_eng_v2:row:66",
207
  "scores_MTEB_eng_v2:row:72",
208
  "scores_MTEB_eng_v2:row:79",
209
+ "scores_MTEB_eng_v2:row:80"
 
 
210
  ],
211
  "focus_frontier": false,
212
  "label_overlaps": [],
 
215
  "all_points_inside_axes": true,
216
  "all_other_panel_ranks_visible": true,
217
  "exports_sha256": {
218
+ "complete-panel-english41-nano.png": "49e1d1a57289291c261e50070ebe2ed5f7fe313ff0d1688c7e67ca2f0c4b27ff",
219
+ "complete-panel-english41-nano.svg": "b0b8f04d2a408bc98e8e36a6c7024e9e1226ade870a8e526eb9970c150b2d2bd"
220
  },
221
  "visible_reference_labels": [
222
  {
 
351
  "all_points_inside_axes": true,
352
  "all_other_panel_ranks_visible": true,
353
  "exports_sha256": {
354
+ "complete-panel-audio19-mini.png": "131a400c69f3e7b6fc1ff1a1de0f285fd7e5938a2fffbfe98552d8981bfa1b27",
355
+ "complete-panel-audio19-mini.svg": "09955e43a78a5d0eb3b385b725d9db771116a56da47e55ef8cad6903274300c0"
356
  },
357
  "visible_reference_labels": [
358
  {
 
406
  ]
407
  }
408
  },
409
+ "visual_inspection": "All six PNGs inspected: white background, full eligible point coverage, preserved named references, no clipped or overlapping labels. Nano size/audio ranks and task scores updated; Mini scores unchanged; SIB-FLEURS is topic classification.",
410
  "scope": "Complete dedicated snapshot rows; no per-task best stitching, no partial-panel ranking, no overall SOTA claim. Unknown-size complete peers count globally but cannot be plotted.",
411
  "ranking_population": "Dedicated registry rows plus both current Vela models, unique names asserted.",
412
  "source_urls_and_actual_measured_identity": "See sources and models in complete-panel-data.json."
benchmarks/complete-panel-ranks.json CHANGED
@@ -50,23 +50,25 @@
50
  "global_rank": 65,
51
  "global_population": 188,
52
  "tied_other_models": 0,
53
- "same_or_smaller_rank": 5,
54
- "same_or_smaller_population": 66,
55
  "best_same_or_smaller_other_model": {
56
- "id": "scores_MTEB_eng_v2:row:61",
57
- "model": "avsolatorio/GIST-Embedding-v0",
58
  "kind": "registry_reported",
59
- "parameters_billion": 0.109,
60
- "mean_task": 0.6550466210091898
61
  },
62
- "gap_pp_above_best_same_or_smaller": -0.6203469578187959,
63
  "observed_frontier": false,
64
- "dominator_count": 4,
65
  "all_dominator_ids": [
66
  "scores_MTEB_eng_v2:row:61",
67
  "scores_MTEB_eng_v2:row:65",
 
68
  "scores_MTEB_eng_v2:row:71",
69
- "scores_MTEB_eng_v2:row:72"
 
70
  ]
71
  },
72
  "mean_task_type": {
@@ -75,8 +77,8 @@
75
  "global_rank": 65,
76
  "global_population": 188,
77
  "tied_other_models": 0,
78
- "same_or_smaller_rank": 3,
79
- "same_or_smaller_population": 66,
80
  "best_same_or_smaller_other_model": {
81
  "id": "scores_MTEB_eng_v2:row:61",
82
  "model": "avsolatorio/GIST-Embedding-v0",
@@ -86,10 +88,12 @@
86
  },
87
  "gap_pp_above_best_same_or_smaller": -0.6140745488965593,
88
  "observed_frontier": false,
89
- "dominator_count": 2,
90
  "all_dominator_ids": [
91
  "scores_MTEB_eng_v2:row:61",
92
- "scores_MTEB_eng_v2:row:65"
 
 
93
  ]
94
  }
95
  },
@@ -683,7 +687,7 @@
683
  "id": "current:nano",
684
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
685
  "score": 0.6488431514310018,
686
- "parameters_billion": 0.135383808
687
  },
688
  {
689
  "rank": 66,
@@ -2001,7 +2005,7 @@
2001
  "id": "current:nano",
2002
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
2003
  "score": 0.6078184975654924,
2004
- "parameters_billion": 0.135383808
2005
  },
2006
  {
2007
  "rank": 66,
@@ -3105,13 +3109,13 @@
3105
  "ranks": {
3106
  "nano": {
3107
  "mean_task": {
3108
- "score": 0.3976628038845883,
3109
- "display_score_100": 39.76628038845883,
3110
- "global_rank": 34,
3111
  "global_population": 64,
3112
  "tied_other_models": 0,
3113
- "same_or_smaller_rank": 6,
3114
- "same_or_smaller_population": 22,
3115
  "best_same_or_smaller_other_model": {
3116
  "id": "scores_MAEB_beta_audio-only:row:9",
3117
  "model": "openai/whisper-base",
@@ -3119,43 +3123,40 @@
3119
  "parameters_billion": 0.074,
3120
  "mean_task": 0.44856370433436527
3121
  },
3122
- "gap_pp_above_best_same_or_smaller": -5.090090044977696,
3123
  "observed_frontier": false,
3124
- "dominator_count": 5,
3125
  "all_dominator_ids": [
3126
  "scores_MAEB_beta_audio-only:row:9",
3127
- "scores_MAEB_beta_audio-only:row:12",
3128
  "scores_MAEB_beta_audio-only:row:18",
3129
- "scores_MAEB_beta_audio-only:row:20",
3130
- "scores_MAEB_beta_audio-only:row:25"
3131
  ]
3132
  },
3133
  "mean_task_type": {
3134
- "score": 0.4696781408655678,
3135
- "display_score_100": 46.96781408655678,
3136
- "global_rank": 32,
3137
  "global_population": 64,
3138
  "tied_other_models": 0,
3139
- "same_or_smaller_rank": 8,
3140
- "same_or_smaller_population": 22,
3141
  "best_same_or_smaller_other_model": {
3142
- "id": "scores_MAEB_beta_audio-only:row:20",
3143
- "model": "MIT/ast-finetuned-audioset-10-10-0.4593",
3144
  "kind": "registry_reported",
3145
- "parameters_billion": 0.087,
3146
- "mean_task_type": 0.55212392228164
3147
  },
3148
- "gap_pp_above_best_same_or_smaller": -8.244578141607217,
3149
  "observed_frontier": false,
3150
- "dominator_count": 7,
3151
  "all_dominator_ids": [
3152
- "scores_MAEB_beta_audio-only:row:9",
3153
- "scores_MAEB_beta_audio-only:row:12",
3154
- "scores_MAEB_beta_audio-only:row:18",
3155
  "scores_MAEB_beta_audio-only:row:20",
3156
- "scores_MAEB_beta_audio-only:row:25",
3157
- "scores_MAEB_beta_audio-only:row:27",
3158
- "scores_MAEB_beta_audio-only:row:28"
3159
  ]
3160
  }
3161
  },
@@ -3361,95 +3362,95 @@
3361
  },
3362
  {
3363
  "rank": 22,
 
 
 
 
 
 
 
3364
  "id": "scores_MAEB_beta_audio-only:row:19",
3365
  "model": "laion/clap-htsat-fused",
3366
  "score": 0.43475542389060884,
3367
  "parameters_billion": 0.154
3368
  },
3369
  {
3370
- "rank": 23,
3371
  "id": "scores_MAEB_beta_audio-only:row:14",
3372
  "model": "laion/clap-htsat-unfused",
3373
  "score": 0.42935003353973167,
3374
  "parameters_billion": 0.153
3375
  },
3376
  {
3377
- "rank": 24,
3378
  "id": "scores_MAEB_beta_audio-only:row:41",
3379
  "model": "facebook/mms-1b-fl102",
3380
  "score": 0.41367276031991745,
3381
  "parameters_billion": 1.0
3382
  },
3383
  {
3384
- "rank": 25,
3385
  "id": "scores_MAEB_beta_audio-only:row:42",
3386
  "model": "facebook/seamless-m4t-v2-large",
3387
  "score": 0.4131003795149639,
3388
  "parameters_billion": 2.3
3389
  },
3390
  {
3391
- "rank": 26,
3392
  "id": "scores_MAEB_beta_audio-only:row:25",
3393
  "model": "google/vggish",
3394
  "score": 0.4109744040247678,
3395
  "parameters_billion": 0.072
3396
  },
3397
  {
3398
- "rank": 27,
3399
  "id": "scores_MAEB_beta_audio-only:row:12",
3400
  "model": "matthewagi/HeAR-s1.1",
3401
  "score": 0.41088769685242515,
3402
  "parameters_billion": 0.022
3403
  },
3404
  {
3405
- "rank": 28,
3406
  "id": "scores_MAEB_beta_audio-only:row:30",
3407
  "model": "facebook/wav2vec2-xls-r-2b",
3408
  "score": 0.4095841483488132,
3409
  "parameters_billion": 2.0
3410
  },
3411
  {
3412
- "rank": 29,
3413
  "id": "scores_MAEB_beta_audio-only:row:22",
3414
  "model": "OpenMuQ/MuQ-MuLan-large",
3415
  "score": 0.4089651437048504,
3416
  "parameters_billion": 0.63
3417
  },
3418
  {
3419
- "rank": 30,
3420
  "id": "scores_MAEB_beta_audio-only:row:35",
3421
  "model": "facebook/mms-1b-all",
3422
  "score": 0.4070234414344685,
3423
  "parameters_billion": 1.0
3424
  },
3425
  {
3426
- "rank": 31,
3427
  "id": "scores_MAEB_beta_audio-only:row:33",
3428
  "model": "microsoft/msclap-2022",
3429
  "score": 0.40099675438596494,
3430
  "parameters_billion": 0.196
3431
  },
3432
  {
3433
- "rank": 32,
3434
  "id": "scores_MAEB_beta_audio-only:row:31",
3435
  "model": "facebook/mms-1b-l1107",
3436
  "score": 0.3998511984004128,
3437
  "parameters_billion": 1.0
3438
  },
3439
  {
3440
- "rank": 33,
3441
  "id": "scores_MAEB_beta_audio-only:row:29",
3442
  "model": "facebook/wav2vec2-lv-60-espeak-cv-ft",
3443
  "score": 0.3983267115583075,
3444
  "parameters_billion": 0.317
3445
  },
3446
- {
3447
- "rank": 34,
3448
- "id": "current:nano",
3449
- "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
3450
- "score": 0.3976628038845883,
3451
- "parameters_billion": 0.135383808
3452
- },
3453
  {
3454
  "rank": 35,
3455
  "id": "scores_MAEB_beta_audio-only:row:26",
@@ -3790,102 +3791,102 @@
3790
  },
3791
  {
3792
  "rank": 19,
 
 
 
 
 
 
 
3793
  "id": "scores_MAEB_beta_audio-only:row:33",
3794
  "model": "microsoft/msclap-2022",
3795
  "score": 0.5156324,
3796
  "parameters_billion": 0.196
3797
  },
3798
  {
3799
- "rank": 20,
3800
  "id": "scores_MAEB_beta_audio-only:row:23",
3801
  "model": "EximiusLabs/fusion-embedding-2-2b-preview",
3802
  "score": 0.5116159368092692,
3803
  "parameters_billion": 2.8
3804
  },
3805
  {
3806
- "rank": 21,
3807
  "id": "scores_MAEB_beta_audio-only:row:28",
3808
  "model": "google/yamnet",
3809
  "score": 0.510955464795009,
3810
  "parameters_billion": 0.004
3811
  },
3812
  {
3813
- "rank": 22,
3814
  "id": "scores_MAEB_beta_audio-only:row:1",
3815
  "model": "openai/whisper-medium",
3816
  "score": 0.5071337459893048,
3817
  "parameters_billion": 0.769
3818
  },
3819
  {
3820
- "rank": 23,
3821
  "id": "scores_MAEB_beta_audio-only:row:12",
3822
  "model": "matthewagi/HeAR-s1.1",
3823
  "score": 0.5022611316399287,
3824
  "parameters_billion": 0.022
3825
  },
3826
  {
3827
- "rank": 24,
3828
  "id": "scores_MAEB_beta_audio-only:row:0",
3829
  "model": "Qwen/Qwen2-Audio-7B",
3830
  "score": 0.49261104875222816,
3831
  "parameters_billion": 7.0
3832
  },
3833
  {
3834
- "rank": 25,
3835
  "id": "scores_MAEB_beta_audio-only:row:24",
3836
  "model": "lyrebird/wav2clip",
3837
  "score": 0.48892121889483064,
3838
  "parameters_billion": 0.163
3839
  },
3840
  {
3841
- "rank": 26,
3842
  "id": "scores_MAEB_beta_audio-only:row:26",
3843
  "model": "microsoft/wavlm-large",
3844
  "score": 0.4848395147058824,
3845
  "parameters_billion": 0.317
3846
  },
3847
  {
3848
- "rank": 27,
3849
  "id": "scores_MAEB_beta_audio-only:row:9",
3850
  "model": "openai/whisper-base",
3851
  "score": 0.48299134028520496,
3852
  "parameters_billion": 0.074
3853
  },
3854
  {
3855
- "rank": 28,
3856
  "id": "scores_MAEB_beta_audio-only:row:13",
3857
  "model": "openai/whisper-large-v3",
3858
  "score": 0.4818755459001783,
3859
  "parameters_billion": 1.55
3860
  },
3861
  {
3862
- "rank": 29,
3863
  "id": "scores_MAEB_beta_audio-only:row:11",
3864
  "model": "openai/whisper-small",
3865
  "score": 0.48186107352941177,
3866
  "parameters_billion": 0.244
3867
  },
3868
  {
3869
- "rank": 30,
3870
  "id": "scores_MAEB_beta_audio-only:row:18",
3871
  "model": "openai/whisper-tiny",
3872
  "score": 0.4774227983957219,
3873
  "parameters_billion": 0.039
3874
  },
3875
  {
3876
- "rank": 31,
3877
  "id": "scores_MAEB_beta_audio-only:row:27",
3878
  "model": "facebook/hubert-base-ls960",
3879
  "score": 0.4762437683600713,
3880
  "parameters_billion": 0.095
3881
  },
3882
- {
3883
- "rank": 32,
3884
- "id": "current:nano",
3885
- "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
3886
- "score": 0.4696781408655678,
3887
- "parameters_billion": 0.135383808
3888
- },
3889
  {
3890
  "rank": 33,
3891
  "id": "scores_MAEB_beta_audio-only:row:45",
 
50
  "global_rank": 65,
51
  "global_population": 188,
52
  "tied_other_models": 0,
53
+ "same_or_smaller_rank": 7,
54
+ "same_or_smaller_population": 75,
55
  "best_same_or_smaller_other_model": {
56
+ "id": "scores_MTEB_eng_v2:row:68",
57
+ "model": "Tarka-AIR/Tarka-Embedding-150M-V1",
58
  "kind": "registry_reported",
59
+ "parameters_billion": 0.156,
60
+ "mean_task": 0.6639085365853659
61
  },
62
+ "gap_pp_above_best_same_or_smaller": -1.5065385154364064,
63
  "observed_frontier": false,
64
+ "dominator_count": 6,
65
  "all_dominator_ids": [
66
  "scores_MTEB_eng_v2:row:61",
67
  "scores_MTEB_eng_v2:row:65",
68
+ "scores_MTEB_eng_v2:row:68",
69
  "scores_MTEB_eng_v2:row:71",
70
+ "scores_MTEB_eng_v2:row:72",
71
+ "scores_MTEB_eng_v2:row:74"
72
  ]
73
  },
74
  "mean_task_type": {
 
77
  "global_rank": 65,
78
  "global_population": 188,
79
  "tied_other_models": 0,
80
+ "same_or_smaller_rank": 5,
81
+ "same_or_smaller_population": 75,
82
  "best_same_or_smaller_other_model": {
83
  "id": "scores_MTEB_eng_v2:row:61",
84
  "model": "avsolatorio/GIST-Embedding-v0",
 
88
  },
89
  "gap_pp_above_best_same_or_smaller": -0.6140745488965593,
90
  "observed_frontier": false,
91
+ "dominator_count": 4,
92
  "all_dominator_ids": [
93
  "scores_MTEB_eng_v2:row:61",
94
+ "scores_MTEB_eng_v2:row:65",
95
+ "scores_MTEB_eng_v2:row:68",
96
+ "scores_MTEB_eng_v2:row:74"
97
  ]
98
  }
99
  },
 
687
  "id": "current:nano",
688
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
689
  "score": 0.6488431514310018,
690
+ "parameters_billion": 0.163771288
691
  },
692
  {
693
  "rank": 66,
 
2005
  "id": "current:nano",
2006
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
2007
  "score": 0.6078184975654924,
2008
+ "parameters_billion": 0.163771288
2009
  },
2010
  {
2011
  "rank": 66,
 
3109
  "ranks": {
3110
  "nano": {
3111
  "mean_task": {
3112
+ "score": 0.43593843064595883,
3113
+ "display_score_100": 43.59384306459588,
3114
+ "global_rank": 22,
3115
  "global_population": 64,
3116
  "tied_other_models": 0,
3117
+ "same_or_smaller_rank": 5,
3118
+ "same_or_smaller_population": 27,
3119
  "best_same_or_smaller_other_model": {
3120
  "id": "scores_MAEB_beta_audio-only:row:9",
3121
  "model": "openai/whisper-base",
 
3123
  "parameters_billion": 0.074,
3124
  "mean_task": 0.44856370433436527
3125
  },
3126
+ "gap_pp_above_best_same_or_smaller": -1.2625273688406435,
3127
  "observed_frontier": false,
3128
+ "dominator_count": 4,
3129
  "all_dominator_ids": [
3130
  "scores_MAEB_beta_audio-only:row:9",
3131
+ "scores_MAEB_beta_audio-only:row:17",
3132
  "scores_MAEB_beta_audio-only:row:18",
3133
+ "scores_MAEB_beta_audio-only:row:20"
 
3134
  ]
3135
  },
3136
  "mean_task_type": {
3137
+ "score": 0.5233540613685161,
3138
+ "display_score_100": 52.33540613685162,
3139
+ "global_rank": 19,
3140
  "global_population": 64,
3141
  "tied_other_models": 0,
3142
+ "same_or_smaller_rank": 6,
3143
+ "same_or_smaller_population": 27,
3144
  "best_same_or_smaller_other_model": {
3145
+ "id": "scores_MAEB_beta_audio-only:row:17",
3146
+ "model": "microsoft/msclap-2023",
3147
  "kind": "registry_reported",
3148
+ "parameters_billion": 0.16,
3149
+ "mean_task_type": 0.5584816090017826
3150
  },
3151
+ "gap_pp_above_best_same_or_smaller": -3.512754763326642,
3152
  "observed_frontier": false,
3153
+ "dominator_count": 5,
3154
  "all_dominator_ids": [
3155
+ "scores_MAEB_beta_audio-only:row:14",
3156
+ "scores_MAEB_beta_audio-only:row:17",
3157
+ "scores_MAEB_beta_audio-only:row:19",
3158
  "scores_MAEB_beta_audio-only:row:20",
3159
+ "scores_MAEB_beta_audio-only:row:25"
 
 
3160
  ]
3161
  }
3162
  },
 
3362
  },
3363
  {
3364
  "rank": 22,
3365
+ "id": "current:nano",
3366
+ "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
3367
+ "score": 0.43593843064595883,
3368
+ "parameters_billion": 0.163771288
3369
+ },
3370
+ {
3371
+ "rank": 23,
3372
  "id": "scores_MAEB_beta_audio-only:row:19",
3373
  "model": "laion/clap-htsat-fused",
3374
  "score": 0.43475542389060884,
3375
  "parameters_billion": 0.154
3376
  },
3377
  {
3378
+ "rank": 24,
3379
  "id": "scores_MAEB_beta_audio-only:row:14",
3380
  "model": "laion/clap-htsat-unfused",
3381
  "score": 0.42935003353973167,
3382
  "parameters_billion": 0.153
3383
  },
3384
  {
3385
+ "rank": 25,
3386
  "id": "scores_MAEB_beta_audio-only:row:41",
3387
  "model": "facebook/mms-1b-fl102",
3388
  "score": 0.41367276031991745,
3389
  "parameters_billion": 1.0
3390
  },
3391
  {
3392
+ "rank": 26,
3393
  "id": "scores_MAEB_beta_audio-only:row:42",
3394
  "model": "facebook/seamless-m4t-v2-large",
3395
  "score": 0.4131003795149639,
3396
  "parameters_billion": 2.3
3397
  },
3398
  {
3399
+ "rank": 27,
3400
  "id": "scores_MAEB_beta_audio-only:row:25",
3401
  "model": "google/vggish",
3402
  "score": 0.4109744040247678,
3403
  "parameters_billion": 0.072
3404
  },
3405
  {
3406
+ "rank": 28,
3407
  "id": "scores_MAEB_beta_audio-only:row:12",
3408
  "model": "matthewagi/HeAR-s1.1",
3409
  "score": 0.41088769685242515,
3410
  "parameters_billion": 0.022
3411
  },
3412
  {
3413
+ "rank": 29,
3414
  "id": "scores_MAEB_beta_audio-only:row:30",
3415
  "model": "facebook/wav2vec2-xls-r-2b",
3416
  "score": 0.4095841483488132,
3417
  "parameters_billion": 2.0
3418
  },
3419
  {
3420
+ "rank": 30,
3421
  "id": "scores_MAEB_beta_audio-only:row:22",
3422
  "model": "OpenMuQ/MuQ-MuLan-large",
3423
  "score": 0.4089651437048504,
3424
  "parameters_billion": 0.63
3425
  },
3426
  {
3427
+ "rank": 31,
3428
  "id": "scores_MAEB_beta_audio-only:row:35",
3429
  "model": "facebook/mms-1b-all",
3430
  "score": 0.4070234414344685,
3431
  "parameters_billion": 1.0
3432
  },
3433
  {
3434
+ "rank": 32,
3435
  "id": "scores_MAEB_beta_audio-only:row:33",
3436
  "model": "microsoft/msclap-2022",
3437
  "score": 0.40099675438596494,
3438
  "parameters_billion": 0.196
3439
  },
3440
  {
3441
+ "rank": 33,
3442
  "id": "scores_MAEB_beta_audio-only:row:31",
3443
  "model": "facebook/mms-1b-l1107",
3444
  "score": 0.3998511984004128,
3445
  "parameters_billion": 1.0
3446
  },
3447
  {
3448
+ "rank": 34,
3449
  "id": "scores_MAEB_beta_audio-only:row:29",
3450
  "model": "facebook/wav2vec2-lv-60-espeak-cv-ft",
3451
  "score": 0.3983267115583075,
3452
  "parameters_billion": 0.317
3453
  },
 
 
 
 
 
 
 
3454
  {
3455
  "rank": 35,
3456
  "id": "scores_MAEB_beta_audio-only:row:26",
 
3791
  },
3792
  {
3793
  "rank": 19,
3794
+ "id": "current:nano",
3795
+ "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
3796
+ "score": 0.5233540613685161,
3797
+ "parameters_billion": 0.163771288
3798
+ },
3799
+ {
3800
+ "rank": 20,
3801
  "id": "scores_MAEB_beta_audio-only:row:33",
3802
  "model": "microsoft/msclap-2022",
3803
  "score": 0.5156324,
3804
  "parameters_billion": 0.196
3805
  },
3806
  {
3807
+ "rank": 21,
3808
  "id": "scores_MAEB_beta_audio-only:row:23",
3809
  "model": "EximiusLabs/fusion-embedding-2-2b-preview",
3810
  "score": 0.5116159368092692,
3811
  "parameters_billion": 2.8
3812
  },
3813
  {
3814
+ "rank": 22,
3815
  "id": "scores_MAEB_beta_audio-only:row:28",
3816
  "model": "google/yamnet",
3817
  "score": 0.510955464795009,
3818
  "parameters_billion": 0.004
3819
  },
3820
  {
3821
+ "rank": 23,
3822
  "id": "scores_MAEB_beta_audio-only:row:1",
3823
  "model": "openai/whisper-medium",
3824
  "score": 0.5071337459893048,
3825
  "parameters_billion": 0.769
3826
  },
3827
  {
3828
+ "rank": 24,
3829
  "id": "scores_MAEB_beta_audio-only:row:12",
3830
  "model": "matthewagi/HeAR-s1.1",
3831
  "score": 0.5022611316399287,
3832
  "parameters_billion": 0.022
3833
  },
3834
  {
3835
+ "rank": 25,
3836
  "id": "scores_MAEB_beta_audio-only:row:0",
3837
  "model": "Qwen/Qwen2-Audio-7B",
3838
  "score": 0.49261104875222816,
3839
  "parameters_billion": 7.0
3840
  },
3841
  {
3842
+ "rank": 26,
3843
  "id": "scores_MAEB_beta_audio-only:row:24",
3844
  "model": "lyrebird/wav2clip",
3845
  "score": 0.48892121889483064,
3846
  "parameters_billion": 0.163
3847
  },
3848
  {
3849
+ "rank": 27,
3850
  "id": "scores_MAEB_beta_audio-only:row:26",
3851
  "model": "microsoft/wavlm-large",
3852
  "score": 0.4848395147058824,
3853
  "parameters_billion": 0.317
3854
  },
3855
  {
3856
+ "rank": 28,
3857
  "id": "scores_MAEB_beta_audio-only:row:9",
3858
  "model": "openai/whisper-base",
3859
  "score": 0.48299134028520496,
3860
  "parameters_billion": 0.074
3861
  },
3862
  {
3863
+ "rank": 29,
3864
  "id": "scores_MAEB_beta_audio-only:row:13",
3865
  "model": "openai/whisper-large-v3",
3866
  "score": 0.4818755459001783,
3867
  "parameters_billion": 1.55
3868
  },
3869
  {
3870
+ "rank": 30,
3871
  "id": "scores_MAEB_beta_audio-only:row:11",
3872
  "model": "openai/whisper-small",
3873
  "score": 0.48186107352941177,
3874
  "parameters_billion": 0.244
3875
  },
3876
  {
3877
+ "rank": 31,
3878
  "id": "scores_MAEB_beta_audio-only:row:18",
3879
  "model": "openai/whisper-tiny",
3880
  "score": 0.4774227983957219,
3881
  "parameters_billion": 0.039
3882
  },
3883
  {
3884
+ "rank": 32,
3885
  "id": "scores_MAEB_beta_audio-only:row:27",
3886
  "model": "facebook/hubert-base-ls960",
3887
  "score": 0.4762437683600713,
3888
  "parameters_billion": 0.095
3889
  },
 
 
 
 
 
 
 
3890
  {
3891
  "rank": 33,
3892
  "id": "scores_MAEB_beta_audio-only:row:45",
benchmarks/complete-panel-ranks.md CHANGED
@@ -1,12 +1,12 @@
1
  # Complete-panel ranks
2
 
3
- Primary: **Mean(TaskType)**. Supplementary: Mean(Task). Scores are ×100. Same complete September 17, 2026 dedicated registry cohorts; all Vela modality parameters counted. Different upstream protocols remain explicit. No mean is imputed for an incomplete model.
4
 
5
  | Model | Panel | Mean(TaskType) | Global | ≤ current size | Mean(Task) | Global | ≤ current size |
6
- |---|---|---:|---:|---:|---:|---:|---:|
7
- | Vela Omni Nano | English41 | 60.7818 | 65/188 | 3/66 | 64.8843 | 65/188 | 5/66 |
8
- | Vela Omni Nano | audio19 | 46.9678 | 32/64 | 8/22 | 39.7663 | 34/64 | 6/22 |
9
  | Vela Omni Mini | English41 | 58.7867 | 92/188 | 50/134 | 63.4923 | 82/188 | 42/134 |
10
  | Vela Omni Mini | audio19 | 54.8679 | 12/64 | 5/50 | 47.7720 | 9/64 | 3/50 |
11
 
12
- These are complete-panel snapshot comparisons, not official leaderboard submissions or a general multimodal SOTA claim. English applicability retains the original measured text identity. Mini audio is newly evaluated on all 19 tasks and 134 subset/split slots. Every peer and exclusion is retained in complete-panel-data.json.
 
1
  # Complete-panel ranks
2
 
3
+ Primary: **Mean(TaskType)**. Supplementary: Mean(Task). Scores are ×100. Complete September 17, 2026 registry cohorts; all Vela modality parameters counted. Upstream protocols can differ. No aggregate is imputed for an incomplete model.
4
 
5
  | Model | Panel | Mean(TaskType) | Global | ≤ current size | Mean(Task) | Global | ≤ current size |
6
+ |---|---|---:|---|---|---:|---|---|
7
+ | Vela Omni Nano | English41 | 60.7818 | 65/188 | 5/75 | 64.8843 | 65/188 | 7/75 |
8
+ | Vela Omni Nano | audio19 | 52.3354 | 19/64 | 6/27 | 43.5938 | 22/64 | 5/27 |
9
  | Vela Omni Mini | English41 | 58.7867 | 92/188 | 50/134 | 63.4923 | 82/188 | 42/134 |
10
  | Vela Omni Mini | audio19 | 54.8679 | 12/64 | 5/50 | 47.7720 | 9/64 | 3/50 |
11
 
12
+ Snapshot comparisons include single-modality specialists and are not official leaderboard submissions or an overall multimodal SOTA claim. Nano audio is newly evaluated on all 19 tasks and 134 subset/split slots; unchanged English text computation retains its original measured identity. Mini scores and size are unchanged. Full task values, exclusions and provenance are retained in complete-panel-data.json.
benchmarks/component-equivalence.json CHANGED
@@ -1,37 +1,36 @@
1
  {
2
  "schema_version": 1,
3
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
4
- "parameters": 144554112,
5
- "evaluated_model": {
6
- "model": "local/Vela-Omni-Nano-GIST-Rotation",
7
- "revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
8
- "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
9
- "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646"
 
 
 
10
  },
11
- "applies_to": {
12
- "revision": "self",
13
- "native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4"
 
 
 
 
 
 
14
  },
15
- "unchanged_components": [
16
- "text"
17
- ],
18
- "tensor_identity": {
19
- "text_tensors": 199,
20
- "text_parameters": 33360000,
21
- "stored_values_dtypes_shapes_exact": true,
22
- "effective_float32_tensors_exact": true
23
  },
24
- "inference_files_exact": true,
25
- "method": "Both complete native packages were rehashed; all text tensors and inference source, configuration and tokenizer bytes match. Official evaluator, runtime, seed, batching, CLS pooling and 512-token limit are identical.",
26
- "reports": [
27
- "benchmarks/mteb-eng-v2.json"
28
- ],
29
- "scope": "The full English report keeps the artifact that was actually evaluated. Only text-only scores apply through this bridge. Image, audio and cross-modal results are freshly evaluated on the current artifact.",
30
- "inference_profile_applicability": {
31
- "revision": "self",
32
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
33
- "parameters": 135383808,
34
- "report": "benchmarks/inference-equivalence.json",
35
- "new_benchmark_run": false
36
- }
37
  }
 
1
  {
2
  "schema_version": 1,
3
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
4
+ "release_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
5
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
6
+ "parameters": 163771288,
7
+ "release_native_files": "native-manifest.json",
8
+ "audio": {
9
+ "fresh_full_artifact_evaluation": true,
10
+ "task_count": 19,
11
+ "slot_count": 134,
12
+ "score_inheritance": false
13
  },
14
+ "text": {
15
+ "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
16
+ "applies_to_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
17
+ "retained_text_tensor_count": 197,
18
+ "parameters": 33212160,
19
+ "stored_and_effective_FP32_values_exact": true,
20
+ "tokenizer_configuration_and_inference_operations_exact": true,
21
+ "unused_pooler_tensors_removed_previously": 2,
22
+ "new_English_panel_forward": false
23
  },
24
+ "image": {
25
+ "evaluated_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
26
+ "applies_to_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
27
+ "retained_image_tensor_count": 210,
28
+ "parameters": 93815424,
29
+ "stored_and_effective_FP32_values_exact": true,
30
+ "inference_operations_exact": true,
31
+ "new_standard_image_forward": false
32
  },
33
+ "previous_native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
34
+ "proof_sha256": "42d802c7279a1c20728926c079f5a4f952b16629fc49a8ed121da4c51f3d77aa",
35
+ "scope": "Equivalence applies only to the named unchanged text and image computations; it makes no claim of unchanged audio or latency."
 
 
 
 
 
 
 
 
 
 
36
  }
benchmarks/event-audio-diagnostic.json ADDED
@@ -0,0 +1,24 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "scope": "Separate weak-label positive-recall diagnostic, outside the product net and official benchmark panels.",
3
+ "evaluated_native_fingerprint": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
4
+ "previous_native_fingerprint": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
5
+ "train_recordings": 3299,
6
+ "evaluation_recordings": 840,
7
+ "known_label_universe": 153,
8
+ "TRAIN_supported_prototypes": 149,
9
+ "protocol": "Normalize the mean TRAIN audio embedding for each known class. Per clip, every annotated positive remains in the recall denominator, including labels without a TRAIN prototype. Average clip recall at 1/5/10.",
10
+ "scores": {
11
+ "candidate": {
12
+ "positive_recall@1": 0.2538690476190476,
13
+ "positive_recall@10": 0.7918622448979592,
14
+ "positive_recall@5": 0.662157029478458
15
+ },
16
+ "parent": {
17
+ "positive_recall@1": 0.0901388888888889,
18
+ "positive_recall@10": 0.3862528344671202,
19
+ "positive_recall@5": 0.273296485260771
20
+ }
21
+ },
22
+ "missing_prototype_positives_count_as_misses": true,
23
+ "new_native_evaluation": true
24
+ }
benchmarks/maeb-audio-only.json CHANGED
The diff for this file is too large to render. See raw diff
 
benchmarks/mteb-eng-v2.json CHANGED
@@ -4672,351 +4672,64 @@
4672
  "component_applicability": {
4673
  "component": "text",
4674
  "applies_to_revision": "self",
4675
- "applies_to_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
4676
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
4677
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
4678
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
4679
  "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
4680
  "report": "benchmarks/component-equivalence.json",
4681
  "new_full_panel_inference": false,
4682
- "basis": "199 exact text tensors and byte-identical inference/tokenizer/configuration under the same official evaluator protocol."
 
4683
  },
4684
  "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
4685
- "comparison_to_previous_release": [
4686
- {
4687
- "candidate_100": 59.245999999999995,
4688
- "current_100": 37.513999999999996,
4689
- "delta_pp": 21.732,
4690
- "primary_metric": "ndcg_at_10",
4691
- "task": "ArguAna",
4692
- "task_type": "Retrieval"
4693
- },
4694
- {
4695
- "candidate_100": 64.88847259326147,
4696
- "current_100": 65.25395409281045,
4697
- "delta_pp": -0.36548149954897724,
4698
- "primary_metric": "v_measure",
4699
- "task": "ArXivHierarchicalClusteringP2P",
4700
- "task_type": "Clustering"
4701
- },
4702
- {
4703
- "candidate_100": 57.44315313869382,
4704
- "current_100": 55.989532291560415,
4705
- "delta_pp": 1.4536208471334078,
4706
- "primary_metric": "v_measure",
4707
- "task": "ArXivHierarchicalClusteringS2S",
4708
- "task_type": "Clustering"
4709
- },
4710
- {
4711
- "candidate_100": 62.329,
4712
- "current_100": 60.514,
4713
- "delta_pp": 1.8149999999999977,
4714
- "primary_metric": "map_at_1000",
4715
- "task": "AskUbuntuDupQuestions",
4716
- "task_type": "Reranking"
4717
- },
4718
- {
4719
- "candidate_100": 86.98752423842483,
4720
- "current_100": 62.24539420714732,
4721
- "delta_pp": 24.742130031277505,
4722
- "primary_metric": "cosine_spearman",
4723
- "task": "BIOSSES",
4724
- "task_type": "STS"
4725
- },
4726
- {
4727
- "candidate_100": 82.14285714285714,
4728
- "current_100": 69.57467532467533,
4729
- "delta_pp": 12.568181818181813,
4730
- "primary_metric": "accuracy",
4731
- "task": "Banking77Classification",
4732
- "task_type": "Classification"
4733
- },
4734
- {
4735
- "candidate_100": 41.313796273674654,
4736
- "current_100": 39.82923119776807,
4737
- "delta_pp": 1.4845650759065876,
4738
- "primary_metric": "v_measure",
4739
- "task": "BiorxivClusteringP2P.v2",
4740
- "task_type": "Clustering"
4741
- },
4742
- {
4743
- "candidate_100": 56.979,
4744
- "current_100": 42.976,
4745
- "delta_pp": 14.003,
4746
- "primary_metric": "ndcg_at_10",
4747
- "task": "CQADupstackGamingRetrieval",
4748
- "task_type": "Retrieval"
4749
- },
4750
- {
4751
- "candidate_100": 39.586,
4752
- "current_100": 30.362000000000002,
4753
- "delta_pp": 9.223999999999997,
4754
- "primary_metric": "ndcg_at_10",
4755
- "task": "CQADupstackUnixRetrieval",
4756
- "task_type": "Retrieval"
4757
- },
4758
- {
4759
- "candidate_100": 31.846000000000004,
4760
- "current_100": 16.073999999999998,
4761
- "delta_pp": 15.772000000000006,
4762
- "primary_metric": "ndcg_at_10",
4763
- "task": "ClimateFEVERHardNegatives",
4764
- "task_type": "Retrieval"
4765
- },
4766
- {
4767
- "candidate_100": 87.576,
4768
- "current_100": 24.055,
4769
- "delta_pp": 63.520999999999994,
4770
- "primary_metric": "ndcg_at_10",
4771
- "task": "FEVERHardNegatives",
4772
- "task_type": "Retrieval"
4773
- },
4774
- {
4775
- "candidate_100": 39.143,
4776
- "current_100": 22.588,
4777
- "delta_pp": 16.555,
4778
- "primary_metric": "ndcg_at_10",
4779
- "task": "FiQA2018",
4780
- "task_type": "Retrieval"
4781
- },
4782
- {
4783
- "candidate_100": 66.349,
4784
- "current_100": 26.987,
4785
- "delta_pp": 39.36200000000001,
4786
- "primary_metric": "ndcg_at_10",
4787
- "task": "HotpotQAHardNegatives",
4788
- "task_type": "Retrieval"
4789
- },
4790
- {
4791
- "candidate_100": 91.94760000000001,
4792
- "current_100": 62.1364,
4793
- "delta_pp": 29.811200000000007,
4794
- "primary_metric": "accuracy",
4795
- "task": "ImdbClassification",
4796
- "task_type": "Classification"
4797
- },
4798
- {
4799
- "candidate_100": 94.91792065663475,
4800
- "current_100": 86.71910624715002,
4801
- "delta_pp": 8.198814409484726,
4802
- "primary_metric": "accuracy",
4803
- "task": "MTOPDomainClassification",
4804
- "task_type": "Classification"
4805
- },
4806
- {
4807
- "candidate_100": 70.96839273705447,
4808
- "current_100": 60.39340954942838,
4809
- "delta_pp": 10.574983187626096,
4810
- "primary_metric": "accuracy",
4811
- "task": "MassiveIntentClassification",
4812
- "task_type": "Classification"
4813
- },
4814
- {
4815
- "candidate_100": 76.11297915265635,
4816
- "current_100": 73.37256220578345,
4817
- "delta_pp": 2.740416946872898,
4818
- "primary_metric": "accuracy",
4819
- "task": "MassiveScenarioClassification",
4820
- "task_type": "Classification"
4821
- },
4822
- {
4823
- "candidate_100": 39.83023091223438,
4824
- "current_100": 37.76813963073551,
4825
- "delta_pp": 2.0620912814988657,
4826
- "primary_metric": "v_measure",
4827
- "task": "MedrxivClusteringP2P.v2",
4828
- "task_type": "Clustering"
4829
- },
4830
- {
4831
- "candidate_100": 37.875084453707764,
4832
- "current_100": 36.13349306174702,
4833
- "delta_pp": 1.7415913919607462,
4834
- "primary_metric": "v_measure",
4835
- "task": "MedrxivClusteringS2S.v2",
4836
- "task_type": "Clustering"
4837
- },
4838
- {
4839
- "candidate_100": 32.366,
4840
- "current_100": 30.623,
4841
- "delta_pp": 1.7429999999999986,
4842
- "primary_metric": "max_over_subqueries_map_at_1000",
4843
- "task": "MindSmallReranking",
4844
- "task_type": "Reranking"
4845
- },
4846
- {
4847
- "candidate_100": 21.889,
4848
- "current_100": 16.827,
4849
- "delta_pp": 5.061999999999998,
4850
- "primary_metric": "ndcg_at_10",
4851
- "task": "SCIDOCS",
4852
- "task_type": "Retrieval"
4853
- },
4854
- {
4855
- "candidate_100": 80.53174565838819,
4856
- "current_100": 71.69618925026361,
4857
- "delta_pp": 8.835556408124575,
4858
- "primary_metric": "cosine_spearman",
4859
- "task": "SICK-R",
4860
- "task_type": "STS"
4861
- },
4862
- {
4863
- "candidate_100": 75.5658801516535,
4864
- "current_100": 68.16915974310523,
4865
- "delta_pp": 7.396720408548276,
4866
- "primary_metric": "cosine_spearman",
4867
- "task": "STS12",
4868
- "task_type": "STS"
4869
- },
4870
- {
4871
- "candidate_100": 86.26346974038235,
4872
- "current_100": 74.40342742804673,
4873
- "delta_pp": 11.860042312335622,
4874
- "primary_metric": "cosine_spearman",
4875
- "task": "STS13",
4876
- "task_type": "STS"
4877
- },
4878
- {
4879
- "candidate_100": 82.29882793900394,
4880
- "current_100": 68.02653891259911,
4881
- "delta_pp": 14.272289026404835,
4882
- "primary_metric": "cosine_spearman",
4883
- "task": "STS14",
4884
- "task_type": "STS"
4885
- },
4886
- {
4887
- "candidate_100": 88.73648284326485,
4888
- "current_100": 77.86144933320074,
4889
- "delta_pp": 10.875033510064114,
4890
- "primary_metric": "cosine_spearman",
4891
- "task": "STS15",
4892
- "task_type": "STS"
4893
- },
4894
- {
4895
- "candidate_100": 87.0781967302773,
4896
- "current_100": 77.92400226361768,
4897
- "delta_pp": 9.154194466659618,
4898
- "primary_metric": "cosine_spearman",
4899
- "task": "STSBenchmark",
4900
- "task_type": "STS"
4901
- },
4902
- {
4903
- "candidate_100": 95.79964673462752,
4904
- "current_100": 89.79096125036985,
4905
- "delta_pp": 6.0086854842576685,
4906
- "primary_metric": "max_ap",
4907
- "task": "SprintDuplicateQuestions",
4908
- "task_type": "PairClassification"
4909
- },
4910
- {
4911
- "candidate_100": 58.424418738430106,
4912
- "current_100": 57.34464990059502,
4913
- "delta_pp": 1.0797688378350827,
4914
- "primary_metric": "v_measure",
4915
- "task": "StackExchangeClustering.v2",
4916
- "task_type": "Clustering"
4917
- },
4918
- {
4919
- "candidate_100": 41.053017253245386,
4920
- "current_100": 39.37668627014545,
4921
- "delta_pp": 1.6763309830999376,
4922
- "primary_metric": "v_measure",
4923
- "task": "StackExchangeClusteringP2P.v2",
4924
- "task_type": "Clustering"
4925
- },
4926
- {
4927
- "candidate_100": 69.121,
4928
- "current_100": 41.292,
4929
- "delta_pp": 27.828999999999994,
4930
- "primary_metric": "ndcg_at_10",
4931
- "task": "TRECCOVID",
4932
- "task_type": "Retrieval"
4933
- },
4934
- {
4935
- "candidate_100": 48.363,
4936
- "current_100": 31.477,
4937
- "delta_pp": 16.886,
4938
- "primary_metric": "ndcg_at_10",
4939
- "task": "Touche2020Retrieval.v3",
4940
- "task_type": "Retrieval"
4941
- },
4942
- {
4943
- "candidate_100": 71.904296875,
4944
- "current_100": 64.4091796875,
4945
- "delta_pp": 7.4951171875,
4946
- "primary_metric": "accuracy",
4947
- "task": "ToxicConversationsClassification",
4948
- "task_type": "Classification"
4949
- },
4950
- {
4951
- "candidate_100": 62.79286926994907,
4952
- "current_100": 50.77815506508207,
4953
- "delta_pp": 12.014714204867005,
4954
- "primary_metric": "accuracy",
4955
- "task": "TweetSentimentExtractionClassification",
4956
- "task_type": "Classification"
4957
- },
4958
- {
4959
- "candidate_100": 50.87260708286914,
4960
- "current_100": 44.88471745749435,
4961
- "delta_pp": 5.987889625374791,
4962
- "primary_metric": "v_measure",
4963
- "task": "TwentyNewsgroupsClustering.v2",
4964
- "task_type": "Clustering"
4965
- },
4966
- {
4967
- "candidate_100": 72.95151691526065,
4968
- "current_100": 56.34475441213181,
4969
- "delta_pp": 16.60676250312884,
4970
- "primary_metric": "max_ap",
4971
- "task": "TwitterSemEval2015",
4972
- "task_type": "PairClassification"
4973
- },
4974
- {
4975
- "candidate_100": 85.30340320492272,
4976
- "current_100": 78.70209561308617,
4977
- "delta_pp": 6.6013075918365445,
4978
- "primary_metric": "max_ap",
4979
- "task": "TwitterURLCorpus",
4980
- "task_type": "PairClassification"
4981
- },
4982
- {
4983
- "candidate_100": 31.834434345907127,
4984
- "current_100": 29.249640152019634,
4985
- "delta_pp": 2.584794193887493,
4986
- "primary_metric": "cosine_spearman",
4987
- "task": "SummEvalSummarization.v2",
4988
- "task_type": "Summarization"
4989
- },
4990
- {
4991
- "candidate_100": 71.8507462686567,
4992
- "current_100": 56.82089552238806,
4993
- "delta_pp": 15.02985074626865,
4994
- "primary_metric": "accuracy",
4995
- "task": "AmazonCounterfactualClassification",
4996
- "task_type": "Classification"
4997
- },
4998
- {
4999
- "candidate_100": 89.02101343562433,
5000
- "current_100": 84.16589279728525,
5001
- "delta_pp": 4.8551206383390735,
5002
- "primary_metric": "cosine_spearman",
5003
- "task": "STS17",
5004
- "task_type": "STS"
5005
- },
5006
- {
5007
- "candidate_100": 68.75333638044488,
5008
- "current_100": 67.57005849656171,
5009
- "delta_pp": 1.1832778838831643,
5010
- "primary_metric": "cosine_spearman",
5011
- "task": "STS22.v2",
5012
- "task_type": "STS"
5013
- }
5014
- ],
5015
- "inference_profile_applicability": {
5016
- "revision": "self",
5017
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
5018
- "parameters": 135383808,
5019
- "report": "benchmarks/inference-equivalence.json",
5020
- "new_benchmark_run": false
5021
  }
5022
  }
 
4672
  "component_applicability": {
4673
  "component": "text",
4674
  "applies_to_revision": "self",
4675
+ "applies_to_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
4676
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
4677
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
4678
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
4679
  "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
4680
  "report": "benchmarks/component-equivalence.json",
4681
  "new_full_panel_inference": false,
4682
+ "basis": "The evaluated text computation is retained through the published Nano and its inference-only trim. All 197 retained text tensors, effective FP32 values, tokenizer/config and text operation path match; the two removed pooler tensors are unused.",
4683
+ "bridge_sha256": "42d802c7279a1c20728926c079f5a4f952b16629fc49a8ed121da4c51f3d77aa"
4684
  },
4685
  "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
4686
+ "comparison_to_previous_release": {
4687
+ "ArguAna": 0.0,
4688
+ "ArXivHierarchicalClusteringP2P": 0.0,
4689
+ "ArXivHierarchicalClusteringS2S": 0.0,
4690
+ "AskUbuntuDupQuestions": 0.0,
4691
+ "BIOSSES": 0.0,
4692
+ "Banking77Classification": 0.0,
4693
+ "BiorxivClusteringP2P.v2": 0.0,
4694
+ "CQADupstackGamingRetrieval": 0.0,
4695
+ "CQADupstackUnixRetrieval": 0.0,
4696
+ "ClimateFEVERHardNegatives": 0.0,
4697
+ "FEVERHardNegatives": 0.0,
4698
+ "FiQA2018": 0.0,
4699
+ "HotpotQAHardNegatives": 0.0,
4700
+ "ImdbClassification": 0.0,
4701
+ "MTOPDomainClassification": 0.0,
4702
+ "MassiveIntentClassification": 0.0,
4703
+ "MassiveScenarioClassification": 0.0,
4704
+ "MedrxivClusteringP2P.v2": 0.0,
4705
+ "MedrxivClusteringS2S.v2": 0.0,
4706
+ "MindSmallReranking": 0.0,
4707
+ "SCIDOCS": 0.0,
4708
+ "SICK-R": 0.0,
4709
+ "STS12": 0.0,
4710
+ "STS13": 0.0,
4711
+ "STS14": 0.0,
4712
+ "STS15": 0.0,
4713
+ "STSBenchmark": 0.0,
4714
+ "SprintDuplicateQuestions": 0.0,
4715
+ "StackExchangeClustering.v2": 0.0,
4716
+ "StackExchangeClusteringP2P.v2": 0.0,
4717
+ "TRECCOVID": 0.0,
4718
+ "Touche2020Retrieval.v3": 0.0,
4719
+ "ToxicConversationsClassification": 0.0,
4720
+ "TweetSentimentExtractionClassification": 0.0,
4721
+ "TwentyNewsgroupsClustering.v2": 0.0,
4722
+ "TwitterSemEval2015": 0.0,
4723
+ "TwitterURLCorpus": 0.0,
4724
+ "SummEvalSummarization.v2": 0.0,
4725
+ "AmazonCounterfactualClassification": 0.0,
4726
+ "STS17": 0.0,
4727
+ "STS22.v2": 0.0
4728
+ },
4729
+ "applicable_parameters": 163771288,
4730
+ "parameter_semantics": "parameters describes the actual evaluated artifact; applicable_parameters describes this unchanged-text release.",
4731
+ "comparison_baseline": {
4732
+ "revision": "9e747d2e1d4b97d3081ff4ac2ac79986c7a8da95",
4733
+ "native_fingerprint": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa"
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4734
  }
4735
  }
benchmarks/native-manifest.json ADDED
@@ -0,0 +1,149 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
3
+ "revision": "self",
4
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
5
+ "parameters": 163771288,
6
+ "files": {
7
+ "LICENSE": {
8
+ "bytes": 11357,
9
+ "sha256": "c71d239df91726fc519c6eb72d318ec65820627232b2f796219e87dcf35d0ab4"
10
+ },
11
+ "LICENSE-CLAP": {
12
+ "bytes": 7048,
13
+ "sha256": "a2010f343487d3f7618affe54f789f5487602331c0a8d03f49e9a7c547cf0499"
14
+ },
15
+ "NOTICE": {
16
+ "bytes": 1803,
17
+ "sha256": "a05fd6236b1daa54fb9306cf685c15bf3b4315a1007e1c4d56a59d996891a94d"
18
+ },
19
+ "components/audio/config.json": {
20
+ "bytes": 1941,
21
+ "sha256": "372c430053183035fb9d2d7079482f10053b6ff7d7367525ada9fdeead8adeac"
22
+ },
23
+ "components/audio/preprocessor_config.json": {
24
+ "bytes": 184990,
25
+ "sha256": "9b5cd03a36fbb8a627c64d98a5b5b126ead95a77720723944487311f0110b666"
26
+ },
27
+ "components/audio_clap/config.json": {
28
+ "bytes": 5390,
29
+ "sha256": "9efb9557bc804f2ca6e394486af2e45dfed0b18554909735a99c6220b84e4288"
30
+ },
31
+ "components/audio_clap/preprocessor_config.json": {
32
+ "bytes": 541,
33
+ "sha256": "9739f58296aa6f9ac18008fd0150fb2649bc554985fbde86d0a4041c882ac753"
34
+ },
35
+ "components/image/config.json": {
36
+ "bytes": 322,
37
+ "sha256": "c2b09e8c0a60f3405d3e53897358a8a373bc4286234f30fcc9324d065fc13728"
38
+ },
39
+ "components/image/preprocessor_config.json": {
40
+ "bytes": 368,
41
+ "sha256": "5a0e5062d04603dc0e5eff01440f6bebfc66d396f9ca16e0128b31c5730cafc5"
42
+ },
43
+ "components/text/1_Pooling/config.json": {
44
+ "bytes": 190,
45
+ "sha256": "d1caf60c96f5fba2157c0c26b76d80818fad6cf0b8eb5e73ec372ff9818eba5c"
46
+ },
47
+ "components/text/config.json": {
48
+ "bytes": 719,
49
+ "sha256": "419ef27bf56a60adc670c50610f8099c64b23e4a1912f4ee7e321139518569b2"
50
+ },
51
+ "components/text/config_sentence_transformers.json": {
52
+ "bytes": 124,
53
+ "sha256": "940d5f50db195fa6e5e6a4f122c095f77880de259d74b14a65779ed48bdd7c56"
54
+ },
55
+ "components/text/modules.json": {
56
+ "bytes": 349,
57
+ "sha256": "84e40c8e006c9b1d6c122e02cba9b02458120b5fb0c87b746c41e0207cf642cf"
58
+ },
59
+ "components/text/sentence_bert_config.json": {
60
+ "bytes": 52,
61
+ "sha256": "84e39fda68ccbff05bfa723ae9c0e70e23e2ec373b76e0f8c6e71af72a693cbf"
62
+ },
63
+ "components/text/special_tokens_map.json": {
64
+ "bytes": 695,
65
+ "sha256": "5d5b662e421ea9fac075174bb0688ee0d9431699900b90662acd44b2a350503a"
66
+ },
67
+ "components/text/tokenizer.json": {
68
+ "bytes": 711396,
69
+ "sha256": "d241a60d5e8f04cc1b2b3e9ef7a4921b27bf526d9f6050ab90f9267a1f9e5c66"
70
+ },
71
+ "components/text/tokenizer_config.json": {
72
+ "bytes": 1242,
73
+ "sha256": "0b29c7bfc889e53b36d9dd3e686dd4300f6525110eaa98c76a5dafceb2029f53"
74
+ },
75
+ "components/text/vocab.txt": {
76
+ "bytes": 231508,
77
+ "sha256": "07eced375cec144d27c900241f3e339478dec958f92fddbc551f295c992038a3"
78
+ },
79
+ "config.json": {
80
+ "bytes": 561,
81
+ "sha256": "ebd558ac77fa99b721b2f2c5d07e6d19c6aa1cc516c12df2900077ce937e78a3"
82
+ },
83
+ "licenses/GIST-MODEL-CARD.md": {
84
+ "bytes": 67964,
85
+ "sha256": "e2905b6dc66b39a697cd82f3211eea4a7ce45b17cc720eceae7fdd337b601c3c"
86
+ },
87
+ "model.safetensors": {
88
+ "bytes": 655583800,
89
+ "sha256": "d5aa7f00e217bd4fd5314140a3603b9a873636320c8d09732a14f79f4cb0667e"
90
+ },
91
+ "omni_components/__init__.py": {
92
+ "bytes": 41,
93
+ "sha256": "543f72d073d77a71e7f5ad857f26d53abbecbbba1ed094336f5bf730af0bd641"
94
+ },
95
+ "omni_components/audio_encoder.py": {
96
+ "bytes": 8492,
97
+ "sha256": "2c1e8e347c3e49a438084cc61f486dda9fde00d9f4b4d7caac828aa9b2f77488"
98
+ },
99
+ "omni_components/audio_io.py": {
100
+ "bytes": 426,
101
+ "sha256": "23392eb0784bd92e4b0176004455025faa0d9d47f6216e12da34079c49765546"
102
+ },
103
+ "omni_components/embedder.py": {
104
+ "bytes": 14049,
105
+ "sha256": "ad709fd80bc1850f609443777ddf3327765f3ea330659f5c154f90fa8ab5e8fd"
106
+ },
107
+ "omni_components/fusion.py": {
108
+ "bytes": 11098,
109
+ "sha256": "a3f6132f190f1533887a77570f4457a0da4eba698574b90888da477919d5c279"
110
+ },
111
+ "omni_components/image_encoder.py": {
112
+ "bytes": 8195,
113
+ "sha256": "40e64a4fde0cb1a1416b764a1a8db47105985e3b979bac7b0aa3bc9d789d0f7a"
114
+ },
115
+ "omni_components/mini.py": {
116
+ "bytes": 10075,
117
+ "sha256": "c62439716bcf06a0051a8d9ec7ea41f917816133f425a7fa004561279daac4fe"
118
+ },
119
+ "omni_components/records.py": {
120
+ "bytes": 118,
121
+ "sha256": "acb3a3cda9d260e39007b6aafd74efd002de974700e14a5535f66590b8b801d6"
122
+ },
123
+ "omni_components/residual_projection.py": {
124
+ "bytes": 1668,
125
+ "sha256": "30028685698c9b4779c44f8ebe28bf80c1689b99594ca4b287db16d133227ad4"
126
+ },
127
+ "omni_components/single_modality.py": {
128
+ "bytes": 4218,
129
+ "sha256": "a4493f2f53bec989f9eff7fbb4b63b9a47aafd2004fd74560911be62ed3237ea"
130
+ },
131
+ "omni_components/text_backbone.py": {
132
+ "bytes": 3244,
133
+ "sha256": "c11a92203f3010ef857d2c809257781150a790acf1877ed83e1a0021449268c3"
134
+ },
135
+ "omni_components/text_encoder.py": {
136
+ "bytes": 9409,
137
+ "sha256": "578329b6bf656423e0b9543783f30a6455449a7402fd3a1a9f601a8afc9705af"
138
+ },
139
+ "omni_components/tiny_clap_residual.py": {
140
+ "bytes": 4315,
141
+ "sha256": "bfbd151f52f91da907affbc33f7a61cea5b11c11cc5fa41d234277a99f112f6c"
142
+ },
143
+ "vela_omni.py": {
144
+ "bytes": 10436,
145
+ "sha256": "216c267591e0af3b4beb583e3be239b1af1c23300ccb74c5a79d58e9f79e52d3"
146
+ }
147
+ },
148
+ "same_as_actual_evaluated_native": true
149
+ }
benchmarks/net-progress.json CHANGED
@@ -1,547 +1,91 @@
1
  {
2
- "scope": "Product-weighted progress comparison, not an official benchmark aggregate or SOTA claim.",
3
- "baseline_model": "llm-semantic-router/Vela-1.0-Omni-Nano",
4
- "baseline_revision": "3a1efc48995e2b71622d65c6cf71cfc00bbf3aa3",
5
- "baseline_native_artifact_sha256": "fb9cd39c9980a085b75456c715c2295b997784fe733314be68c13635f6384edf",
6
- "candidate_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
7
- "score_scale": "0-100; changes in percentage points",
8
  "components": [
9
  {
10
- "baseline_100": 59.37287180704907,
11
- "candidate_100": 67.57407300861536,
12
  "component": "common14",
13
- "delta_pp": 8.201201201566288,
 
14
  "weight": 0.3
15
  },
16
  {
17
- "baseline_100": 51.97553164396411,
18
  "candidate_100": 60.781849756549235,
19
  "component": "English41_MeanTaskType",
20
- "delta_pp": 8.806318112585124,
 
21
  "weight": 0.3
22
  },
23
  {
24
- "baseline_100": 54.865785925824184,
25
  "candidate_100": 53.65746394890041,
26
  "component": "image3",
27
- "delta_pp": -1.208321976923776,
 
28
  "weight": 0.2
29
  },
30
  {
31
- "baseline_100": 46.9743843608,
32
- "candidate_100": 46.96781408655678,
33
  "component": "audio19_MeanTaskType",
34
- "delta_pp": -0.00657027424322,
 
35
  "weight": 0.2
36
  }
37
  ],
38
- "weighted_net_delta_pp": 4.859277344012025,
39
- "common_definition": "Equal mean of three families: two text accuracies, all six image retrieval recalls, all six audio retrieval recalls.",
40
- "image3_definition": "Equal mean of OxfordPets classification accuracy, TinyImageNet clustering NMI and OxfordPets zero-shot accuracy.",
41
- "text_audio_definition": "Full English41 and full audio19 mean task-type scores. English uses the explicit exact text-component bridge; audio is freshly evaluated.",
42
- "coverage": {
43
- "common_metrics": 14,
44
- "standard_tasks": 11,
45
- "standard_split_rows": 12,
46
- "English_tasks": 41,
47
- "audio_tasks": 19,
48
- "audio_subset_split_slots": 134
49
- },
50
  "common_metric_delta_pp": {
51
- "audio.audio_to_text.recall@1": 8.157793948678668,
52
- "audio.audio_to_text.recall@10": 15.51129835312141,
53
- "audio.audio_to_text.recall@5": 14.362313289927231,
54
- "audio.text_to_audio.recall@1": 11.417624521072797,
55
- "audio.text_to_audio.recall@10": 15.478927203065133,
56
- "audio.text_to_audio.recall@5": 14.444444444444445,
57
- "image.image_to_text.recall@1": -4.738760631834751,
58
- "image.image_to_text.recall@10": -1.2150668286755772,
59
- "image.image_to_text.recall@5": -1.701093560145808,
60
- "image.text_to_image.recall@1": 2.7460510328068044,
61
- "image.text_to_image.recall@10": 0.7776427703523694,
62
- "image.text_to_image.recall@5": 1.0206561360874848,
63
- "text.banking77.accuracy": 11.168831168831169,
64
- "text.massive-en.accuracy": 12.617765814266487
65
  },
66
- "standard_task_deltas": [
67
- {
68
- "baseline": 0.7169618925026362,
69
- "candidate": 0.8053174565838819,
70
- "hf_subset": "default",
71
- "main_metric": "cosine_spearman",
72
- "split": "test",
73
- "status": "complete",
74
- "task": "SICK-R",
75
- "delta_pp": 8.835556408124567
76
- },
77
- {
78
- "baseline": 0.7792400226361769,
79
- "candidate": 0.870781967302773,
80
- "hf_subset": "default",
81
- "main_metric": "cosine_spearman",
82
- "split": "test",
83
- "status": "complete",
84
- "task": "STSBenchmark",
85
- "delta_pp": 9.154194466659614
86
- },
87
- {
88
- "baseline": 0.37514,
89
- "candidate": 0.59246,
90
- "hf_subset": "default",
91
- "main_metric": "ndcg_at_10",
92
- "split": "test",
93
- "status": "complete",
94
- "task": "ArguAna",
95
- "delta_pp": 21.732000000000003
96
- },
97
- {
98
- "baseline": 0.44884717457494344,
99
- "candidate": 0.5087260708286914,
100
- "hf_subset": "default",
101
- "main_metric": "v_measure",
102
- "split": "test",
103
- "status": "complete",
104
- "task": "TwentyNewsgroupsClustering.v2",
105
- "delta_pp": 5.987889625374793
106
- },
107
- {
108
- "baseline": 0.8758836689038031,
109
- "candidate": 0.9476510067114093,
110
- "hf_subset": "en",
111
- "main_metric": "accuracy",
112
- "split": "validation",
113
- "status": "complete",
114
- "task": "MTOPDomainClassification",
115
- "delta_pp": 7.176733780760625
116
- },
117
- {
118
- "baseline": 0.8671910624715002,
119
- "candidate": 0.9491792065663475,
120
- "hf_subset": "en",
121
- "main_metric": "accuracy",
122
- "split": "test",
123
- "status": "complete",
124
- "task": "MTOPDomainClassification",
125
- "delta_pp": 8.198814409484722
126
- },
127
- {
128
- "baseline": 0.9126737530662306,
129
- "candidate": 0.9126737530662306,
130
- "hf_subset": "default",
131
- "main_metric": "accuracy",
132
- "split": "test",
133
- "status": "complete",
134
- "task": "OxfordPets",
135
- "delta_pp": 0.0
136
- },
137
- {
138
- "baseline": 0.6169193395626786,
139
- "candidate": 0.6169193395626786,
140
- "hf_subset": "default",
141
- "main_metric": "nmi",
142
- "split": "valid",
143
- "status": "complete",
144
- "task": "TinyImageNetClustering",
145
- "delta_pp": 0.0
146
- },
147
- {
148
- "baseline": 0.1163804851458163,
149
- "candidate": 0.08013082583810302,
150
- "hf_subset": "default",
151
- "main_metric": "accuracy",
152
- "split": "test",
153
- "status": "complete",
154
- "task": "OxfordPetsZeroShot",
155
- "delta_pp": -3.6249659307713276
156
- },
157
- {
158
- "baseline": 0.29763670140167686,
159
- "candidate": 0.29763670140167686,
160
- "hf_subset": "default",
161
- "main_metric": "accuracy",
162
- "split": "train",
163
- "status": "complete",
164
- "task": "CREMA_D",
165
- "delta_pp": 0.0
166
- },
167
- {
168
- "baseline": 0.023601715918443712,
169
- "candidate": 0.023601715918443712,
170
- "hf_subset": "default",
171
- "main_metric": "v_measure",
172
- "split": "train",
173
- "status": "complete",
174
- "task": "CREMA_DClustering",
175
- "delta_pp": 0.0
176
- },
177
- {
178
- "baseline": 0.14972999509081983,
179
- "candidate": 0.1411389297987236,
180
- "hf_subset": "default",
181
- "main_metric": "accuracy",
182
- "split": "test",
183
- "status": "complete",
184
- "task": "SpeechCommandsZeroshotv0.02",
185
- "delta_pp": -0.8591065292096217
186
- }
187
- ],
188
- "English_task_deltas": [
189
- {
190
- "candidate_100": 59.245999999999995,
191
- "current_100": 37.513999999999996,
192
- "delta_pp": 21.732,
193
- "primary_metric": "ndcg_at_10",
194
- "task": "ArguAna",
195
- "task_type": "Retrieval"
196
- },
197
- {
198
- "candidate_100": 64.88847259326147,
199
- "current_100": 65.25395409281045,
200
- "delta_pp": -0.36548149954897724,
201
- "primary_metric": "v_measure",
202
- "task": "ArXivHierarchicalClusteringP2P",
203
- "task_type": "Clustering"
204
- },
205
- {
206
- "candidate_100": 57.44315313869382,
207
- "current_100": 55.989532291560415,
208
- "delta_pp": 1.4536208471334078,
209
- "primary_metric": "v_measure",
210
- "task": "ArXivHierarchicalClusteringS2S",
211
- "task_type": "Clustering"
212
- },
213
- {
214
- "candidate_100": 62.329,
215
- "current_100": 60.514,
216
- "delta_pp": 1.8149999999999977,
217
- "primary_metric": "map_at_1000",
218
- "task": "AskUbuntuDupQuestions",
219
- "task_type": "Reranking"
220
- },
221
- {
222
- "candidate_100": 86.98752423842483,
223
- "current_100": 62.24539420714732,
224
- "delta_pp": 24.742130031277505,
225
- "primary_metric": "cosine_spearman",
226
- "task": "BIOSSES",
227
- "task_type": "STS"
228
- },
229
- {
230
- "candidate_100": 82.14285714285714,
231
- "current_100": 69.57467532467533,
232
- "delta_pp": 12.568181818181813,
233
- "primary_metric": "accuracy",
234
- "task": "Banking77Classification",
235
- "task_type": "Classification"
236
- },
237
- {
238
- "candidate_100": 41.313796273674654,
239
- "current_100": 39.82923119776807,
240
- "delta_pp": 1.4845650759065876,
241
- "primary_metric": "v_measure",
242
- "task": "BiorxivClusteringP2P.v2",
243
- "task_type": "Clustering"
244
- },
245
- {
246
- "candidate_100": 56.979,
247
- "current_100": 42.976,
248
- "delta_pp": 14.003,
249
- "primary_metric": "ndcg_at_10",
250
- "task": "CQADupstackGamingRetrieval",
251
- "task_type": "Retrieval"
252
- },
253
- {
254
- "candidate_100": 39.586,
255
- "current_100": 30.362000000000002,
256
- "delta_pp": 9.223999999999997,
257
- "primary_metric": "ndcg_at_10",
258
- "task": "CQADupstackUnixRetrieval",
259
- "task_type": "Retrieval"
260
- },
261
- {
262
- "candidate_100": 31.846000000000004,
263
- "current_100": 16.073999999999998,
264
- "delta_pp": 15.772000000000006,
265
- "primary_metric": "ndcg_at_10",
266
- "task": "ClimateFEVERHardNegatives",
267
- "task_type": "Retrieval"
268
- },
269
- {
270
- "candidate_100": 87.576,
271
- "current_100": 24.055,
272
- "delta_pp": 63.520999999999994,
273
- "primary_metric": "ndcg_at_10",
274
- "task": "FEVERHardNegatives",
275
- "task_type": "Retrieval"
276
- },
277
- {
278
- "candidate_100": 39.143,
279
- "current_100": 22.588,
280
- "delta_pp": 16.555,
281
- "primary_metric": "ndcg_at_10",
282
- "task": "FiQA2018",
283
- "task_type": "Retrieval"
284
- },
285
- {
286
- "candidate_100": 66.349,
287
- "current_100": 26.987,
288
- "delta_pp": 39.36200000000001,
289
- "primary_metric": "ndcg_at_10",
290
- "task": "HotpotQAHardNegatives",
291
- "task_type": "Retrieval"
292
- },
293
- {
294
- "candidate_100": 91.94760000000001,
295
- "current_100": 62.1364,
296
- "delta_pp": 29.811200000000007,
297
- "primary_metric": "accuracy",
298
- "task": "ImdbClassification",
299
- "task_type": "Classification"
300
- },
301
- {
302
- "candidate_100": 94.91792065663475,
303
- "current_100": 86.71910624715002,
304
- "delta_pp": 8.198814409484726,
305
- "primary_metric": "accuracy",
306
- "task": "MTOPDomainClassification",
307
- "task_type": "Classification"
308
- },
309
- {
310
- "candidate_100": 70.96839273705447,
311
- "current_100": 60.39340954942838,
312
- "delta_pp": 10.574983187626096,
313
- "primary_metric": "accuracy",
314
- "task": "MassiveIntentClassification",
315
- "task_type": "Classification"
316
- },
317
- {
318
- "candidate_100": 76.11297915265635,
319
- "current_100": 73.37256220578345,
320
- "delta_pp": 2.740416946872898,
321
- "primary_metric": "accuracy",
322
- "task": "MassiveScenarioClassification",
323
- "task_type": "Classification"
324
- },
325
- {
326
- "candidate_100": 39.83023091223438,
327
- "current_100": 37.76813963073551,
328
- "delta_pp": 2.0620912814988657,
329
- "primary_metric": "v_measure",
330
- "task": "MedrxivClusteringP2P.v2",
331
- "task_type": "Clustering"
332
- },
333
- {
334
- "candidate_100": 37.875084453707764,
335
- "current_100": 36.13349306174702,
336
- "delta_pp": 1.7415913919607462,
337
- "primary_metric": "v_measure",
338
- "task": "MedrxivClusteringS2S.v2",
339
- "task_type": "Clustering"
340
- },
341
- {
342
- "candidate_100": 32.366,
343
- "current_100": 30.623,
344
- "delta_pp": 1.7429999999999986,
345
- "primary_metric": "max_over_subqueries_map_at_1000",
346
- "task": "MindSmallReranking",
347
- "task_type": "Reranking"
348
- },
349
- {
350
- "candidate_100": 21.889,
351
- "current_100": 16.827,
352
- "delta_pp": 5.061999999999998,
353
- "primary_metric": "ndcg_at_10",
354
- "task": "SCIDOCS",
355
- "task_type": "Retrieval"
356
- },
357
- {
358
- "candidate_100": 80.53174565838819,
359
- "current_100": 71.69618925026361,
360
- "delta_pp": 8.835556408124575,
361
- "primary_metric": "cosine_spearman",
362
- "task": "SICK-R",
363
- "task_type": "STS"
364
- },
365
- {
366
- "candidate_100": 75.5658801516535,
367
- "current_100": 68.16915974310523,
368
- "delta_pp": 7.396720408548276,
369
- "primary_metric": "cosine_spearman",
370
- "task": "STS12",
371
- "task_type": "STS"
372
- },
373
- {
374
- "candidate_100": 86.26346974038235,
375
- "current_100": 74.40342742804673,
376
- "delta_pp": 11.860042312335622,
377
- "primary_metric": "cosine_spearman",
378
- "task": "STS13",
379
- "task_type": "STS"
380
- },
381
- {
382
- "candidate_100": 82.29882793900394,
383
- "current_100": 68.02653891259911,
384
- "delta_pp": 14.272289026404835,
385
- "primary_metric": "cosine_spearman",
386
- "task": "STS14",
387
- "task_type": "STS"
388
- },
389
- {
390
- "candidate_100": 88.73648284326485,
391
- "current_100": 77.86144933320074,
392
- "delta_pp": 10.875033510064114,
393
- "primary_metric": "cosine_spearman",
394
- "task": "STS15",
395
- "task_type": "STS"
396
- },
397
- {
398
- "candidate_100": 87.0781967302773,
399
- "current_100": 77.92400226361768,
400
- "delta_pp": 9.154194466659618,
401
- "primary_metric": "cosine_spearman",
402
- "task": "STSBenchmark",
403
- "task_type": "STS"
404
- },
405
- {
406
- "candidate_100": 95.79964673462752,
407
- "current_100": 89.79096125036985,
408
- "delta_pp": 6.0086854842576685,
409
- "primary_metric": "max_ap",
410
- "task": "SprintDuplicateQuestions",
411
- "task_type": "PairClassification"
412
- },
413
- {
414
- "candidate_100": 58.424418738430106,
415
- "current_100": 57.34464990059502,
416
- "delta_pp": 1.0797688378350827,
417
- "primary_metric": "v_measure",
418
- "task": "StackExchangeClustering.v2",
419
- "task_type": "Clustering"
420
- },
421
- {
422
- "candidate_100": 41.053017253245386,
423
- "current_100": 39.37668627014545,
424
- "delta_pp": 1.6763309830999376,
425
- "primary_metric": "v_measure",
426
- "task": "StackExchangeClusteringP2P.v2",
427
- "task_type": "Clustering"
428
- },
429
- {
430
- "candidate_100": 69.121,
431
- "current_100": 41.292,
432
- "delta_pp": 27.828999999999994,
433
- "primary_metric": "ndcg_at_10",
434
- "task": "TRECCOVID",
435
- "task_type": "Retrieval"
436
- },
437
- {
438
- "candidate_100": 48.363,
439
- "current_100": 31.477,
440
- "delta_pp": 16.886,
441
- "primary_metric": "ndcg_at_10",
442
- "task": "Touche2020Retrieval.v3",
443
- "task_type": "Retrieval"
444
- },
445
- {
446
- "candidate_100": 71.904296875,
447
- "current_100": 64.4091796875,
448
- "delta_pp": 7.4951171875,
449
- "primary_metric": "accuracy",
450
- "task": "ToxicConversationsClassification",
451
- "task_type": "Classification"
452
- },
453
- {
454
- "candidate_100": 62.79286926994907,
455
- "current_100": 50.77815506508207,
456
- "delta_pp": 12.014714204867005,
457
- "primary_metric": "accuracy",
458
- "task": "TweetSentimentExtractionClassification",
459
- "task_type": "Classification"
460
- },
461
- {
462
- "candidate_100": 50.87260708286914,
463
- "current_100": 44.88471745749435,
464
- "delta_pp": 5.987889625374791,
465
- "primary_metric": "v_measure",
466
- "task": "TwentyNewsgroupsClustering.v2",
467
- "task_type": "Clustering"
468
- },
469
- {
470
- "candidate_100": 72.95151691526065,
471
- "current_100": 56.34475441213181,
472
- "delta_pp": 16.60676250312884,
473
- "primary_metric": "max_ap",
474
- "task": "TwitterSemEval2015",
475
- "task_type": "PairClassification"
476
- },
477
- {
478
- "candidate_100": 85.30340320492272,
479
- "current_100": 78.70209561308617,
480
- "delta_pp": 6.6013075918365445,
481
- "primary_metric": "max_ap",
482
- "task": "TwitterURLCorpus",
483
- "task_type": "PairClassification"
484
- },
485
- {
486
- "candidate_100": 31.834434345907127,
487
- "current_100": 29.249640152019634,
488
- "delta_pp": 2.584794193887493,
489
- "primary_metric": "cosine_spearman",
490
- "task": "SummEvalSummarization.v2",
491
- "task_type": "Summarization"
492
- },
493
- {
494
- "candidate_100": 71.8507462686567,
495
- "current_100": 56.82089552238806,
496
- "delta_pp": 15.02985074626865,
497
- "primary_metric": "accuracy",
498
- "task": "AmazonCounterfactualClassification",
499
- "task_type": "Classification"
500
- },
501
- {
502
- "candidate_100": 89.02101343562433,
503
- "current_100": 84.16589279728525,
504
- "delta_pp": 4.8551206383390735,
505
- "primary_metric": "cosine_spearman",
506
- "task": "STS17",
507
- "task_type": "STS"
508
- },
509
- {
510
- "candidate_100": 68.75333638044488,
511
- "current_100": 67.57005849656171,
512
- "delta_pp": 1.1832778838831643,
513
- "primary_metric": "cosine_spearman",
514
- "task": "STS22.v2",
515
- "task_type": "STS"
516
- }
517
- ],
518
  "audio_task_delta_pp": {
519
- "BeijingOpera": 0.0,
520
- "BirdCLEF": 0.0,
521
- "CREMADPairClassification": 0.12104996267769952,
522
- "CREMA_D": 0.0,
523
- "CREMA_DClustering": 0.0,
524
- "CommonLanguageAgeDetection": 0.0,
525
- "GTZANAudioReranking": 0.0,
526
- "GTZANGenre": 0.0,
527
- "IEMOCAPGender": 0.0,
528
- "JamAltArtistA2ARetrieval": 0.0,
529
- "MInDS14": 0.0,
530
- "MridinghamTonic": -0.14336917562723928,
531
- "NMSQAPairClassification": -0.1520590751812989,
532
- "SIBFLEURS": 0.0,
533
- "VehicleSoundClustering": 0.0,
534
- "VoxCelebSA": 0.0,
535
- "VoxPopuliAccentPairClassification": -0.02844431688273641,
536
- "VoxPopuliGenderClustering": 0.0,
537
- "VoxPopuliLanguageID": 0.0
 
 
 
 
 
 
 
 
 
538
  },
539
- "paired_interval_scope": "Paired intervals cover the original common metrics, family and two-primary-per-family utility. They are not confidence intervals for the new all-14 common mean or the four-component product aggregate.",
540
- "inference_profile_applicability": {
541
- "revision": "self",
542
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
543
- "parameters": 135383808,
544
- "report": "benchmarks/inference-equivalence.json",
545
- "new_benchmark_run": false
546
- }
547
  }
 
1
  {
2
+ "scope": "Fixed product progress metric; not an official benchmark aggregate, significance claim, or proof of overall SOTA.",
3
+ "baseline_revision": "9e747d2e1d4b97d3081ff4ac2ac79986c7a8da95",
4
+ "baseline_native_fingerprint": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
5
+ "baseline_evaluated_fingerprint": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
6
+ "candidate_native_fingerprint": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
 
7
  "components": [
8
  {
9
+ "candidate_100": 66.41835251574038,
 
10
  "component": "common14",
11
+ "delta_pp": -1.1557204928749676,
12
+ "parent_100": 67.57407300861536,
13
  "weight": 0.3
14
  },
15
  {
 
16
  "candidate_100": 60.781849756549235,
17
  "component": "English41_MeanTaskType",
18
+ "delta_pp": 0.0,
19
+ "parent_100": 60.781849756549235,
20
  "weight": 0.3
21
  },
22
  {
 
23
  "candidate_100": 53.65746394890041,
24
  "component": "image3",
25
+ "delta_pp": 0.0,
26
+ "parent_100": 53.65746394890041,
27
  "weight": 0.2
28
  },
29
  {
30
+ "candidate_100": 52.33540613685162,
 
31
  "component": "audio19_MeanTaskType",
32
+ "delta_pp": 5.367592050294833,
33
+ "parent_100": 46.96781408655678,
34
  "weight": 0.2
35
  }
36
  ],
37
+ "weighted_net_delta_pp": 0.7268022621964765,
38
+ "score_scale": "0–100; deltas in percentage points",
39
+ "parameter_increase_percent": 20.968150046422096,
40
+ "common_definition": "Mean of text accuracy (2), image recall (6), and audio recall (6), with equal one-third family weights.",
41
+ "image3_definition": "Mean of OxfordPets accuracy, TinyImageNetClustering NMI, and OxfordPetsZeroShot accuracy.",
 
 
 
 
 
 
 
42
  "common_metric_delta_pp": {
43
+ "audio.audio_to_text.recall@1": -1.0340865568747608,
44
+ "audio.audio_to_text.recall@10": -1.3787820758330127,
45
+ "audio.audio_to_text.recall@5": -1.8383761011106892,
46
+ "audio.text_to_audio.recall@1": -5.095785440613029,
47
+ "audio.text_to_audio.recall@10": -5.287356321839088,
48
+ "audio.text_to_audio.recall@5": -6.168582375478926,
49
+ "image.image_to_text.recall@1": 0.0,
50
+ "image.image_to_text.recall@10": 0.0,
51
+ "image.image_to_text.recall@5": 0.0,
52
+ "image.text_to_image.recall@1": 0.0,
53
+ "image.text_to_image.recall@10": 0.0,
54
+ "image.text_to_image.recall@5": 0.0,
55
+ "text.banking77.accuracy": 0.0,
56
+ "text.massive-en.accuracy": 0.0
57
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
58
  "audio_task_delta_pp": {
59
+ "BeijingOpera": 10.549645390070905,
60
+ "BirdCLEF": 2.2999999999999963,
61
+ "CREMADPairClassification": 1.0060252268561554,
62
+ "CREMA_D": 2.2844768445301367,
63
+ "CREMA_DClustering": 2.8541749889603825,
64
+ "CommonLanguageAgeDetection": 0.6150000000000017,
65
+ "GTZANAudioReranking": 6.68200000000001,
66
+ "GTZANGenre": 11.399999999999999,
67
+ "IEMOCAPGender": 3.237875602721152,
68
+ "JamAltArtistA2ARetrieval": 16.90099999999998,
69
+ "MInDS14": -1.4691308536005643,
70
+ "MridinghamTonic": 26.05657742038183,
71
+ "NMSQAPairClassification": -0.3840625955228716,
72
+ "SIBFLEURS": -0.23211656203982745,
73
+ "VehicleSoundClustering": -8.525275870377525,
74
+ "VoxCelebSA": -0.00058896531414665,
75
+ "VoxPopuliAccentPairClassification": 0.1293630037889515,
76
+ "VoxPopuliGenderClustering": -0.08127278385060711,
77
+ "VoxPopuliLanguageID": -0.5999999999999783
78
+ },
79
+ "audio_aggregate_delta_pp": {
80
+ "Any2AnyRetrieval": 16.90099999999998,
81
+ "AudioClassification": 4.921976261522682,
82
+ "AudioClustering": -1.9174578884225837,
83
+ "AudioPairClassification": 0.2504418783740858,
84
+ "AudioReranking": 6.68200000000001,
85
+ "Mean(Task)": 3.8275626761370307,
86
+ "Mean(TaskType)": 5.367592050294833
87
  },
88
+ "English_and_image_unchanged": true,
89
+ "event840_in_net": false,
90
+ "paired_interval_scope": "Fresh paired intervals compare this candidate with the preceding Vela Nano computation. They are not an interval for the all-14 aggregate or the weighted net. An original-small interval was not recomputed for this update."
 
 
 
 
 
91
  }
benchmarks/original-baseline.json CHANGED
@@ -113,8 +113,11 @@
113
  "score": null
114
  },
115
  "selected_standard_11": {
116
- "status": "not_evaluated_in_this_matched_protocol",
117
- "score": null
 
 
 
118
  }
119
  },
120
  "not_previous_Vela": true,
 
113
  "score": null
114
  },
115
  "selected_standard_11": {
116
+ "status": "evaluated_all_11_selected_tasks",
117
+ "task_count": 11,
118
+ "split_rows": 12,
119
+ "report": "original-standard.json",
120
+ "report_sha256": "f07e4624749868498923cf3c7d521ca9547276dbdbfb002cbe97bd074195b55a"
121
  }
122
  },
123
  "not_previous_Vela": true,
benchmarks/pareto-data.json CHANGED
@@ -2569,8 +2569,8 @@
2569
  },
2570
  {
2571
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
2572
- "parameters_billion": 0.135383808,
2573
- "parameters_exact": 135383808,
2574
  "score": 64.88847259326147,
2575
  "open_weights": true,
2576
  "modalities": [
@@ -2582,18 +2582,18 @@
2582
  "kind": "nano",
2583
  "evaluated_task_revision": "0bbdb47bcbe3a90093699aefeed338a0f28a7ee8",
2584
  "on_displayed_frontier": false,
2585
- "evidence_type": "exact_public_inference_applicability",
2586
  "observation_id": "vela_release_evaluation:llm-semantic-router/Vela-1.0-Omni-Nano",
2587
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
2588
- "revision": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
2589
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
2590
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
2591
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
2592
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
2593
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/mteb-eng-v2.json",
2594
  "prior_applicable_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
2595
- "inference_equivalence_report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
2596
- "revision_kind": "native_artifact_fingerprint"
2597
  },
2598
  {
2599
  "model": "ibm-granite/granite-embedding-english-r2",
@@ -9675,9 +9675,9 @@
9675
  },
9676
  {
9677
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
9678
- "parameters_billion": 0.135383808,
9679
- "parameters_exact": 135383808,
9680
- "score": 63.28340215485784,
9681
  "open_weights": true,
9682
  "modalities": [
9683
  "text",
@@ -9688,17 +9688,16 @@
9688
  "kind": "nano",
9689
  "evaluated_task_revision": "d33eea302d8a64c0a2ee094cf39d77c0814fe87a",
9690
  "on_displayed_frontier": true,
9691
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
9692
- "revision": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
9693
- "evaluated_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
9694
- "evaluated_revision": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
9695
  "evaluated_model": "llm-semantic-router/Vela-1.0-Omni-Nano",
9696
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
9697
- "evidence_type": "exact_public_inference_applicability",
9698
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/maeb-audio-only.json",
9699
- "prior_applicable_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
9700
- "inference_equivalence_report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
9701
- "revision_kind": "native_artifact_fingerprint"
9702
  },
9703
  {
9704
  "model": "laion/clap-htsat-unfused",
@@ -10829,9 +10828,9 @@
10829
  },
10830
  {
10831
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
10832
- "parameters_billion": 0.135383808,
10833
- "score": 63.28340215485784,
10834
- "evidence_type": "exact_public_inference_applicability"
10835
  },
10836
  {
10837
  "model": "microsoft/speecht5_multimodal",
@@ -10873,7 +10872,7 @@
10873
  "measured_peer_observations": 0,
10874
  "highlight_variant": "nano",
10875
  "highlight_dominators": [],
10876
- "highlight_gap_above_best_equal_or_smaller_peer_pp": 5.003102154857842
10877
  },
10878
  "sibfleurs": {
10879
  "task": "SIBFLEURS",
@@ -11356,9 +11355,9 @@
11356
  },
11357
  {
11358
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
11359
- "parameters_billion": 0.135383808,
11360
- "parameters_exact": 135383808,
11361
- "score": 18.304657831512046,
11362
  "open_weights": true,
11363
  "modalities": [
11364
  "text",
@@ -11369,17 +11368,16 @@
11369
  "kind": "nano",
11370
  "evaluated_task_revision": "186b61175fbd77059f769b6bf1110d449a2ff311",
11371
  "on_displayed_frontier": false,
11372
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
11373
- "revision": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
11374
- "evaluated_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
11375
- "evaluated_revision": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
11376
  "evaluated_model": "llm-semantic-router/Vela-1.0-Omni-Nano",
11377
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
11378
- "evidence_type": "exact_public_inference_applicability",
11379
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/maeb-audio-only.json",
11380
- "prior_applicable_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
11381
- "inference_equivalence_report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
11382
- "revision_kind": "native_artifact_fingerprint"
11383
  },
11384
  {
11385
  "model": "laion/clap-htsat-unfused",
@@ -13083,9 +13081,9 @@
13083
  },
13084
  {
13085
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
13086
- "parameters_billion": 0.135383808,
13087
- "parameters_exact": 135383808,
13088
- "score": 13.800701178563807,
13089
  "open_weights": true,
13090
  "modalities": [
13091
  "text",
@@ -13096,17 +13094,16 @@
13096
  "kind": "nano",
13097
  "evaluated_task_revision": "5f905f9544d0c78b607639662c88c2ddd1216384",
13098
  "on_displayed_frontier": false,
13099
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
13100
- "revision": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
13101
- "evaluated_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
13102
- "evaluated_revision": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
13103
  "evaluated_model": "llm-semantic-router/Vela-1.0-Omni-Nano",
13104
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
13105
- "evidence_type": "exact_public_inference_applicability",
13106
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/maeb-audio-only.json",
13107
- "prior_applicable_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
13108
- "inference_equivalence_report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
13109
- "revision_kind": "native_artifact_fingerprint"
13110
  },
13111
  {
13112
  "model": "laion/clap-htsat-unfused",
@@ -14297,6 +14294,25 @@
14297
  "measured_peer_observations": 1,
14298
  "highlight_variant": "nano",
14299
  "highlight_dominators": [
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14300
  {
14301
  "model": "MIT/ast-finetuned-audioset-10-10-0.4593",
14302
  "parameters_billion": 0.086594063,
@@ -14319,17 +14335,38 @@
14319
  "report": "benchmarks/peer-evaluations/ast-vehicle.json",
14320
  "license": "bsd-3-clause",
14321
  "on_displayed_frontier": true
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14322
  }
14323
  ],
14324
- "highlight_gap_above_best_equal_or_smaller_peer_pp": -0.5421709345231776
14325
  }
14326
  },
14327
  "models": {
14328
  "nano": {
14329
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Nano",
14330
- "release_commit": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806",
14331
- "parameters_exact": 135383808,
14332
- "current_native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
14333
  "task_evidence": {
14334
  "English41": {
14335
  "source_id": "nano_English41",
@@ -14337,47 +14374,58 @@
14337
  "model": "local/Vela-Omni-Nano-GIST-Rotation",
14338
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14339
  "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
 
14340
  "component_applicability": {
14341
  "component": "text",
14342
- "applies_to_revision": "self",
14343
- "applies_to_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
14344
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
14345
- "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14346
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14347
- "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
14348
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/component-equivalence.json",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14349
  "new_full_panel_inference": false,
14350
- "basis": "199 exact text tensors and byte-identical inference/tokenizer/configuration under the same official evaluator protocol."
14351
- },
14352
- "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
14353
- "inference_profile_applicability": {
14354
- "revision": "self",
14355
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
14356
- "parameters": 135383808,
14357
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
14358
- "new_benchmark_run": false
14359
  }
14360
  },
14361
- "self_reference_interpretation": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806"
14362
  },
14363
  "audio19": {
14364
  "source_id": "nano_audio19",
14365
  "identity": {
14366
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
14367
- "evaluated_revision": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
14368
- "native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
14369
- "revision_semantics": "Native artifact fingerprint identifying the evaluated weights and inference files; the release commit is the repository revision containing this report.",
14370
- "inference_profile_applicability": {
14371
- "revision": "self",
14372
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
14373
- "parameters": 135383808,
14374
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
14375
- "new_benchmark_run": false
14376
- }
14377
  },
14378
- "self_reference_interpretation": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806"
14379
  }
14380
- }
 
 
14381
  },
14382
  "mini": {
14383
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Mini",
@@ -14438,14 +14486,16 @@
14438
  },
14439
  "sources": {
14440
  "nano_English41": {
14441
- "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/mteb-eng-v2.json",
14442
- "sha256": "733f6a58faae3d04e1b5a59de73085b156dc9987ba0e04e28548c17dc5904cec",
14443
- "evidence_type": "Vela_published_report"
 
14444
  },
14445
  "nano_audio19": {
14446
- "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/maeb-audio-only.json",
14447
- "sha256": "e0259c59e04649f42cd2f004d57b525c79442297d28498a7434bddb4b000bdf9",
14448
- "evidence_type": "Vela_published_report"
 
14449
  },
14450
  "mini_English41": {
14451
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/mteb-eng-v2.json",
@@ -14478,6 +14528,12 @@
14478
  "sha256": "8bbbd3a362d84a6e8f030803259db69e932e5997b1c7af1a276a0d17862eb184",
14479
  "evidence_type": "exact_component_identity_report",
14480
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731"
 
 
 
 
 
 
14481
  }
14482
  }
14483
  }
 
2569
  },
2570
  {
2571
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
2572
+ "parameters_billion": 0.163771288,
2573
+ "parameters_exact": 163771288,
2574
  "score": 64.88847259326147,
2575
  "open_weights": true,
2576
  "modalities": [
 
2582
  "kind": "nano",
2583
  "evaluated_task_revision": "0bbdb47bcbe3a90093699aefeed338a0f28a7ee8",
2584
  "on_displayed_frontier": false,
2585
+ "evidence_type": "exact_text_component_applicability",
2586
  "observation_id": "vela_release_evaluation:llm-semantic-router/Vela-1.0-Omni-Nano",
2587
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
2588
+ "revision": null,
2589
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
2590
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
2591
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
2592
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
2593
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/mteb-eng-v2.json",
2594
  "prior_applicable_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
2595
+ "revision_kind": "native_artifact_fingerprint",
2596
+ "identity_ref": "models.nano.task_evidence.English41"
2597
  },
2598
  {
2599
  "model": "ibm-granite/granite-embedding-english-r2",
 
9675
  },
9676
  {
9677
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
9678
+ "parameters_billion": 0.163771288,
9679
+ "parameters_exact": 163771288,
9680
+ "score": 62.89933955933497,
9681
  "open_weights": true,
9682
  "modalities": [
9683
  "text",
 
9688
  "kind": "nano",
9689
  "evaluated_task_revision": "d33eea302d8a64c0a2ee094cf39d77c0814fe87a",
9690
  "on_displayed_frontier": true,
9691
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
9692
+ "revision": null,
9693
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
9694
+ "evaluated_revision": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
9695
  "evaluated_model": "llm-semantic-router/Vela-1.0-Omni-Nano",
9696
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
9697
+ "evidence_type": "measured_release_evaluation",
9698
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
9699
+ "revision_kind": "native_artifact_fingerprint",
9700
+ "identity_ref": "models.nano.task_evidence.audio19"
 
9701
  },
9702
  {
9703
  "model": "laion/clap-htsat-unfused",
 
10828
  },
10829
  {
10830
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
10831
+ "parameters_billion": 0.163771288,
10832
+ "score": 62.89933955933497,
10833
+ "evidence_type": "measured_release_evaluation"
10834
  },
10835
  {
10836
  "model": "microsoft/speecht5_multimodal",
 
10872
  "measured_peer_observations": 0,
10873
  "highlight_variant": "nano",
10874
  "highlight_dominators": [],
10875
+ "highlight_gap_above_best_equal_or_smaller_peer_pp": 4.619039559334972
10876
  },
10877
  "sibfleurs": {
10878
  "task": "SIBFLEURS",
 
11355
  },
11356
  {
11357
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
11358
+ "parameters_billion": 0.163771288,
11359
+ "parameters_exact": 163771288,
11360
+ "score": 18.072541269472218,
11361
  "open_weights": true,
11362
  "modalities": [
11363
  "text",
 
11368
  "kind": "nano",
11369
  "evaluated_task_revision": "186b61175fbd77059f769b6bf1110d449a2ff311",
11370
  "on_displayed_frontier": false,
11371
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
11372
+ "revision": null,
11373
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
11374
+ "evaluated_revision": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
11375
  "evaluated_model": "llm-semantic-router/Vela-1.0-Omni-Nano",
11376
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
11377
+ "evidence_type": "measured_release_evaluation",
11378
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
11379
+ "revision_kind": "native_artifact_fingerprint",
11380
+ "identity_ref": "models.nano.task_evidence.audio19"
 
11381
  },
11382
  {
11383
  "model": "laion/clap-htsat-unfused",
 
13081
  },
13082
  {
13083
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
13084
+ "parameters_billion": 0.163771288,
13085
+ "parameters_exact": 163771288,
13086
+ "score": 5.275425308186281,
13087
  "open_weights": true,
13088
  "modalities": [
13089
  "text",
 
13094
  "kind": "nano",
13095
  "evaluated_task_revision": "5f905f9544d0c78b607639662c88c2ddd1216384",
13096
  "on_displayed_frontier": false,
13097
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
13098
+ "revision": null,
13099
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
13100
+ "evaluated_revision": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
13101
  "evaluated_model": "llm-semantic-router/Vela-1.0-Omni-Nano",
13102
  "revision_semantics": "Evaluated revisions are native artifact fingerprints; applies-to identity is the current whole artifact.",
13103
+ "evidence_type": "measured_release_evaluation",
13104
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
13105
+ "revision_kind": "native_artifact_fingerprint",
13106
+ "identity_ref": "models.nano.task_evidence.audio19"
 
13107
  },
13108
  {
13109
  "model": "laion/clap-htsat-unfused",
 
14294
  "measured_peer_observations": 1,
14295
  "highlight_variant": "nano",
14296
  "highlight_dominators": [
14297
+ {
14298
+ "model": "openai/whisper-tiny",
14299
+ "parameters_billion": 0.039,
14300
+ "score": 12.3883,
14301
+ "open_weights": true,
14302
+ "modalities": [
14303
+ "audio"
14304
+ ],
14305
+ "training_tasks": null,
14306
+ "source_urls": [
14307
+ "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MAEB%28beta%2C%20audio-only%29/scores",
14308
+ "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MAEB%28beta%29/scores"
14309
+ ],
14310
+ "size_basis": "reported_rounded_total",
14311
+ "kind": "peer",
14312
+ "on_displayed_frontier": true,
14313
+ "evidence_type": "registry_reported",
14314
+ "observation_id": "registry_reported:openai/whisper-tiny"
14315
+ },
14316
  {
14317
  "model": "MIT/ast-finetuned-audioset-10-10-0.4593",
14318
  "parameters_billion": 0.086594063,
 
14335
  "report": "benchmarks/peer-evaluations/ast-vehicle.json",
14336
  "license": "bsd-3-clause",
14337
  "on_displayed_frontier": true
14338
+ },
14339
+ {
14340
+ "model": "MIT/ast-finetuned-audioset-10-10-0.4593",
14341
+ "parameters_billion": 0.087,
14342
+ "score": 13.366700000000002,
14343
+ "open_weights": true,
14344
+ "modalities": [
14345
+ "audio"
14346
+ ],
14347
+ "training_tasks": [
14348
+ "AudioSetMini"
14349
+ ],
14350
+ "source_urls": [
14351
+ "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MAEB%28beta%2C%20audio-only%29/scores",
14352
+ "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MAEB%28beta%29/scores"
14353
+ ],
14354
+ "size_basis": "reported_rounded_total",
14355
+ "kind": "peer",
14356
+ "on_displayed_frontier": false,
14357
+ "evidence_type": "registry_reported",
14358
+ "observation_id": "registry_reported:MIT/ast-finetuned-audioset-10-10-0.4593"
14359
  }
14360
  ],
14361
+ "highlight_gap_above_best_equal_or_smaller_peer_pp": -9.067446804900705
14362
  }
14363
  },
14364
  "models": {
14365
  "nano": {
14366
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Nano",
14367
+ "release_commit": null,
14368
+ "parameters_exact": 163771288,
14369
+ "current_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
14370
  "task_evidence": {
14371
  "English41": {
14372
  "source_id": "nano_English41",
 
14374
  "model": "local/Vela-Omni-Nano-GIST-Rotation",
14375
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14376
  "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14377
+ "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
14378
  "component_applicability": {
14379
  "component": "text",
 
 
14380
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
 
14381
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14382
+ "applies_to_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
14383
+ "prior_component_applicability": {
14384
+ "model": "local/Vela-Omni-Nano-GIST-Rotation",
14385
+ "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14386
+ "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14387
+ "component_applicability": {
14388
+ "component": "text",
14389
+ "applies_to_revision": "self",
14390
+ "applies_to_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
14391
+ "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
14392
+ "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14393
+ "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
14394
+ "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
14395
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/component-equivalence.json",
14396
+ "new_full_panel_inference": false,
14397
+ "basis": "199 exact text tensors and byte-identical inference/tokenizer/configuration under the same official evaluator protocol."
14398
+ },
14399
+ "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
14400
+ "inference_profile_applicability": {
14401
+ "revision": "self",
14402
+ "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
14403
+ "parameters": 135383808,
14404
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
14405
+ "new_benchmark_run": false
14406
+ }
14407
+ },
14408
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/component-equivalence.json",
14409
  "new_full_panel_inference": false,
14410
+ "basis": "The frozen text tensors, tokenizer, configuration and encoding computation are unchanged. The original measured text identity is retained."
 
 
 
 
 
 
 
 
14411
  }
14412
  },
14413
+ "self_reference_interpretation": "Nano native fingerprint 474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2"
14414
  },
14415
  "audio19": {
14416
  "source_id": "nano_audio19",
14417
  "identity": {
14418
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
14419
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
14420
+ "release_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
14421
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
14422
+ "new_full_panel_inference": true
 
 
 
 
 
 
14423
  },
14424
+ "self_reference_interpretation": "Nano native fingerprint 474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2"
14425
  }
14426
+ },
14427
+ "release_reference": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano",
14428
+ "release_revision_semantics": "Exact native fingerprint below; Hub commit is not encoded in this file."
14429
  },
14430
  "mini": {
14431
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Mini",
 
14486
  },
14487
  "sources": {
14488
  "nano_English41": {
14489
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/mteb-eng-v2.json",
14490
+ "evidence_type": "Vela_release_report",
14491
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
14492
+ "sha256": "f3daecad1cf431630598efb04f3667431e3420239240f28a51568a74cef0c64f"
14493
  },
14494
  "nano_audio19": {
14495
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
14496
+ "evidence_type": "Vela_release_report",
14497
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
14498
+ "sha256": "4fc7f217745eb0aa998bf0f8e044bd1fa5de2f73f1a7dead7b829a62c48669ff"
14499
  },
14500
  "mini_English41": {
14501
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/mteb-eng-v2.json",
 
14528
  "sha256": "8bbbd3a362d84a6e8f030803259db69e932e5997b1c7af1a276a0d17862eb184",
14529
  "evidence_type": "exact_component_identity_report",
14530
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731"
14531
+ },
14532
+ "nano_component_equivalence": {
14533
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/component-equivalence.json",
14534
+ "evidence_type": "Vela_release_report",
14535
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
14536
+ "sha256": "ebd40f9a887791d1736132c37f510a69e18a03e247422853f54a42b417b0db9b"
14537
  }
14538
  }
14539
  }
benchmarks/pareto-gallery.md CHANGED
@@ -1,18 +1,16 @@
1
- # Vela Nano: quality, size and rank
2
 
3
- Complete-panel comparisons show overall position; selected tasks show specific strengths.
4
 
5
- ![Complete English benchmark: size–quality comparison and ranking](../assets/complete-panel-english41-nano.png)
6
 
7
- ![Complete audio benchmark: size–quality comparison and ranking](../assets/complete-panel-audio19-mini.png)
8
 
9
- ## Selected tasks
10
 
11
- <table>
12
- <tr>
13
- <td width="50%"><img src="../assets/pareto-imdb.png" alt="Vela Nano: IMDb · Text classification" /></td>
14
- <td width="50%"><img src="../assets/pareto-nmsqa.png" alt="Vela Nano: NMSQA · Audio pair classification" /></td>
15
- </tr>
16
- </table>
17
 
18
- [Data and methodology](pareto-methodology.md) · [Complete evaluation](EVALUATION.md)
 
1
+ # Figure gallery
2
 
3
+ Complete-panel comparisons are primary. Scores are measured or registry-reported under the protocols recorded in the data; no incomplete panel contributes an aggregate.
4
 
5
+ ![Complete English41](../assets/complete-panel-english41-nano.png)
6
 
7
+ ![Complete audio19](../assets/complete-panel-audio19-mini.png)
8
 
9
+ Task comparisons are secondary and do not establish overall SOTA. Highlighting does not imply frontier membership.
10
 
11
+ | Nano | Mini |
12
+ | --- | --- |
13
+ | IMDb sentiment classification | Mridingham musical tonic classification |
14
+ | NMSQA audio pair classification | SIB-FLEURS spoken-topic classification |
 
 
15
 
16
+ [Full ranks](complete-panel-ranks.md) · [Methods and provenance](pareto-methodology.md)
benchmarks/pareto-methodology.md CHANGED
@@ -8,18 +8,20 @@ Rankings use the September 17, 2026 dedicated MTEB registry snapshots plus both
8
 
9
  The plots show every complete model with a known positive total parameter count. The dashed Pareto line connects nondominated observations: no other eligible model has both no more parameters and no lower score, with one strict improvement. Vela points remain visible even when they are below this line. The adjacent Top 10 restricts the complete population to models no larger than the highlighted Vela model. All comparisons use unrounded scores.
10
 
11
- Vela sizes count all stored model parameters across text, image and audio: Nano 135,383,808 and Mini 1,361,475,288. Peer sizes use registry-reported totals, which can be rounded. Peers include both single-modality specialists and multimodal encoders. Their scores were reported under differing prompts, model revisions and evaluation protocols; these are snapshot-relative comparisons, not matched reproductions or official leaderboard submissions. There is no combined text–image–audio rank. Neither current Vela model lies on these complete-panel frontiers.
12
 
13
- [All ranks](complete-panel-ranks.md) · [Complete vectors, source URLs, exclusions and plotted points](complete-panel-data.json) · [Figure metadata](complete-panel-metadata.json) · [Vela evaluation protocol](../EVALUATION.md)
14
 
15
  ## Selected task strengths
16
 
17
- The task gallery highlights Nano on IMDb classification and NMSQA audio pair classification, and Mini on Mridingham tonic classification and SIB-FLEURS spoken-topic classification. Each highlighted model is nondominated among the documented references for that particular task. A task-level position does not establish overall SOTA, and neither a label nor a displayed highlight removes other eligible points from the frontier calculation.
18
 
19
  The task plots retain finite scores with known positive total size, including closed-weight models. Registry observations and separately measured peers retain their individual provenance. A stronger new observation can change the frontier without any change in Vela's own score. Complete benchmark results, including regressions, remain in the evaluation report.
20
 
21
  [IMDb and Mridingham points and sources](pareto-new-frontiers.json) · [NMSQA and SIB-FLEURS points and sources](pareto-data.json) · [Other retained task observations (CSV)](pareto-points.csv)
22
 
23
- Mini audio uses the newly completed 19-task evaluation of this release; its unchanged English text computation retains the original evaluated identity through an explicit component bridge. Per-task regressions remain in the complete report. No reported peer coordinate was replaced by a partial-panel rerun.
24
 
25
  SIB-FLEURS measures topic classification of multilingual audio, not language identification. This description follows pinned MTEB 2.21.0 task metadata (`mteb/tasks/classification/multilingual/sibfleurs.py`, SHA-256 `f88cc34acc7414dc84ccf8cbbd2f79456076df126f9acc696bd7f45c9481ad55`), with [dataset revision 186b6117](https://huggingface.co/datasets/mteb/sib-fleurs-multilingual-mini/tree/186b61175fbd77059f769b6bf1110d449a2ff311).
 
 
 
8
 
9
  The plots show every complete model with a known positive total parameter count. The dashed Pareto line connects nondominated observations: no other eligible model has both no more parameters and no lower score, with one strict improvement. Vela points remain visible even when they are below this line. The adjacent Top 10 restricts the complete population to models no larger than the highlighted Vela model. All comparisons use unrounded scores.
10
 
11
+ Vela sizes count all stored model parameters across text, image and audio: Nano 163,771,288 and Mini 1,361,475,288. Peer sizes use registry-reported totals, which can be rounded. Peers include both single-modality specialists and multimodal encoders. Their scores were reported under differing prompts, model revisions and evaluation protocols; these are snapshot-relative comparisons, not matched reproductions or official leaderboard submissions. There is no combined text–image–audio rank. Neither current Vela model lies on these complete-panel frontiers.
12
 
13
+ [All ranks](complete-panel-ranks.md) · [Complete vectors, source URLs, exclusions and plotted points](complete-panel-data.json) · [Figure metadata](complete-panel-metadata.json) · [Vela evaluation protocol](https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/blob/main/benchmarks/EVALUATION.md)
14
 
15
  ## Selected task strengths
16
 
17
+ The task gallery highlights Nano on IMDb classification and NMSQA audio pair classification, and Mini on Mridingham tonic classification and SIB-FLEURS spoken-topic classification. Highlighted models are shown whether or not they lie on the observed frontier; frontier membership is recomputed from all documented eligible points. A task-level position does not establish overall SOTA, and neither a label nor a displayed highlight removes other eligible points from the frontier calculation.
18
 
19
  The task plots retain finite scores with known positive total size, including closed-weight models. Registry observations and separately measured peers retain their individual provenance. A stronger new observation can change the frontier without any change in Vela's own score. Complete benchmark results, including regressions, remain in the evaluation report.
20
 
21
  [IMDb and Mridingham points and sources](pareto-new-frontiers.json) · [NMSQA and SIB-FLEURS points and sources](pareto-data.json) · [Other retained task observations (CSV)](pareto-points.csv)
22
 
23
+ Mini audio retains its existing complete 19-task evaluation; its unchanged English text computation retains the original evaluated identity through an explicit component bridge. Per-task regressions remain in the complete report. No reported peer coordinate was replaced by a partial-panel rerun.
24
 
25
  SIB-FLEURS measures topic classification of multilingual audio, not language identification. This description follows pinned MTEB 2.21.0 task metadata (`mteb/tasks/classification/multilingual/sibfleurs.py`, SHA-256 `f88cc34acc7414dc84ccf8cbbd2f79456076df126f9acc696bd7f45c9481ad55`), with [dataset revision 186b6117](https://huggingface.co/datasets/mteb/sib-fleurs-multilingual-mini/tree/186b61175fbd77059f769b6bf1110d449a2ff311).
26
+
27
+ Nano now includes the complete 19-task audio evaluation at 163,771,288 total parameters. Its English computation is unchanged; Mini scores and size are unchanged. The new Nano audio point improves the complete-panel aggregate but individual regressions remain visible, including NMSQA and VehicleSoundClustering. The NMSQA highlight is a task comparison, not a claim of frontier membership. The general complete-panel rankings take priority over selected task plots.
benchmarks/pareto-new-frontiers.json CHANGED
@@ -6,39 +6,17 @@
6
  "observation_policy": "Only exactly equal model/name/size/score duplicates are coalesced with every source retained. Differing registry scores remain separate observations, with scored revision unknown. F2/AST independent reruns remain separate pinned observations and do not replace reported coordinates. Current other Vela variant is included.",
7
  "dominance": "A peer has no more total parameters and no lower score, with at least one strict inequality. Scores use original available floating values, not rounded display labels.",
8
  "sources": {
9
- "scores_MTEB_eng_v2": {
10
- "url": "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MTEB%28eng%2C%20v2%29/scores",
11
- "snapshot_sha256": "e574f14dc778d72c08255450091191b103930c487fa172a2dfabe41a0b56569d",
12
- "retrieved_at_utc": "2026-09-17T20:28:47.419876+00:00",
13
- "bytes": 1071836,
14
- "benchmark": "MTEB(eng, v2)",
15
- "scored_checkpoint_revisions_exposed": false
16
- },
17
- "scores_MAEB_beta": {
18
- "url": "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MAEB%28beta%29/scores",
19
- "snapshot_sha256": "601467f576706074bb5724aa0c10a50e948b0e8be8158742b01baf2fd7f7f14c",
20
- "retrieved_at_utc": "2026-09-17T20:28:48.161556+00:00",
21
- "bytes": 222458,
22
- "benchmark": "MAEB(beta)",
23
- "scored_checkpoint_revisions_exposed": false
24
- },
25
- "scores_MAEB_beta_audio-only": {
26
- "url": "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MAEB%28beta%2C%20audio-only%29/scores",
27
- "snapshot_sha256": "6153942c26a31577cc1e3a2fbb6eb10b09dc2f1fe08749db598d261e653f6bca",
28
- "retrieved_at_utc": "2026-09-17T20:28:48.670374+00:00",
29
- "bytes": 194971,
30
- "benchmark": "MAEB(beta, audio-only)",
31
- "scored_checkpoint_revisions_exposed": false
32
- },
33
  "nano_English41": {
34
- "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/mteb-eng-v2.json",
35
- "sha256": "733f6a58faae3d04e1b5a59de73085b156dc9987ba0e04e28548c17dc5904cec",
36
- "evidence_type": "Vela_published_report"
 
37
  },
38
  "nano_audio19": {
39
- "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/maeb-audio-only.json",
40
- "sha256": "e0259c59e04649f42cd2f004d57b525c79442297d28498a7434bddb4b000bdf9",
41
- "evidence_type": "Vela_published_report"
 
42
  },
43
  "mini_English41": {
44
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/mteb-eng-v2.json",
@@ -52,19 +30,39 @@
52
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731",
53
  "sha256": "b9adc831e1de9caed37d38292dc95d65f4f933178fa7abbf7c1211f23617efbf"
54
  },
 
 
 
 
 
 
 
 
 
 
 
 
 
 
55
  "mini_component_equivalence": {
56
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/component-equivalence.json",
57
  "sha256": "8bbbd3a362d84a6e8f030803259db69e932e5997b1c7af1a276a0d17862eb184",
58
  "evidence_type": "exact_component_identity_report",
59
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731"
 
 
 
 
 
 
60
  }
61
  },
62
  "models": {
63
  "nano": {
64
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Nano",
65
- "release_commit": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806",
66
- "parameters_exact": 135383808,
67
- "current_native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
68
  "task_evidence": {
69
  "English41": {
70
  "source_id": "nano_English41",
@@ -72,47 +70,58 @@
72
  "model": "local/Vela-Omni-Nano-GIST-Rotation",
73
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
74
  "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
 
75
  "component_applicability": {
76
  "component": "text",
77
- "applies_to_revision": "self",
78
- "applies_to_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
79
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
80
- "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
81
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
82
- "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
83
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/component-equivalence.json",
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
84
  "new_full_panel_inference": false,
85
- "basis": "199 exact text tensors and byte-identical inference/tokenizer/configuration under the same official evaluator protocol."
86
- },
87
- "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
88
- "inference_profile_applicability": {
89
- "revision": "self",
90
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
91
- "parameters": 135383808,
92
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
93
- "new_benchmark_run": false
94
  }
95
  },
96
- "self_reference_interpretation": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806"
97
  },
98
  "audio19": {
99
  "source_id": "nano_audio19",
100
  "identity": {
101
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
102
- "evaluated_revision": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
103
- "native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
104
- "revision_semantics": "Native artifact fingerprint identifying the evaluated weights and inference files; the release commit is the repository revision containing this report.",
105
- "inference_profile_applicability": {
106
- "revision": "self",
107
- "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
108
- "parameters": 135383808,
109
- "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
110
- "new_benchmark_run": false
111
- }
112
  },
113
- "self_reference_interpretation": "d8ac5b5ac2274a501fc61aeb5be70cec1855a806"
114
  }
115
- }
 
 
116
  },
117
  "mini": {
118
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Mini",
@@ -1735,7 +1744,7 @@
1735
  }
1736
  ],
1737
  "kind": "peer",
1738
- "on_displayed_frontier": false
1739
  },
1740
  {
1741
  "model": "Omartificial-Intelligence-Space/Marbert-all-nli-triplet-Matryoshka",
@@ -4103,14 +4112,18 @@
4103
  {
4104
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
4105
  "kind": "nano",
4106
- "parameters_billion": 0.135383808,
4107
- "parameters_exact": 135383808,
4108
  "score": 91.94760000000001,
4109
- "evidence_type": "measured_with_explicit_applicability",
4110
  "identity_ref": "models.nano.task_evidence.English41",
4111
  "size_basis": "current_whole_model_exact_tensor_count",
4112
  "open_weights": true,
4113
- "on_displayed_frontier": true
 
 
 
 
4114
  },
4115
  {
4116
  "model": "llm-semantic-router/Vela-1.0-Omni-Mini",
@@ -4135,6 +4148,7 @@
4135
  "Mihaiii/Squirtle",
4136
  "MongoDB/mdbr-leaf-mt",
4137
  "BAAI/bge-small-en",
 
4138
  "llm-semantic-router/Vela-1.0-Omni-Nano",
4139
  "jinaai/jina-embeddings-v5-text-nano",
4140
  "jxm/cde-small-v2",
@@ -4373,7 +4387,7 @@
4373
  }
4374
  ],
4375
  "unknown_size_observation_count": 28,
4376
- "highlight_gap_above_best_equal_or_smaller_peer_pp": 0.43240000000001544,
4377
  "highlight_dominators": []
4378
  },
4379
  "mridingham": {
@@ -5770,14 +5784,20 @@
5770
  {
5771
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
5772
  "kind": "nano",
5773
- "parameters_billion": 0.135383808,
5774
- "parameters_exact": 135383808,
5775
- "score": 32.363208758254515,
5776
- "evidence_type": "measured_with_explicit_applicability",
5777
  "identity_ref": "models.nano.task_evidence.audio19",
5778
  "size_basis": "current_whole_model_exact_tensor_count",
5779
  "open_weights": true,
5780
- "on_displayed_frontier": false
 
 
 
 
 
 
5781
  },
5782
  {
5783
  "model": "llm-semantic-router/Vela-1.0-Omni-Mini",
@@ -5800,12 +5820,13 @@
5800
  "google/yamnet",
5801
  "matthewagi/HeAR-s1.1",
5802
  "MIT/ast-finetuned-audioset-10-10-0.4593",
 
5803
  "llm-semantic-router/Vela-1.0-Omni-Mini"
5804
  ],
5805
  "reported_peer_observations": 65,
5806
  "unknown_size_observations_excluded": [],
5807
  "unknown_size_observation_count": 0,
5808
- "highlight_gap_above_best_equal_or_smaller_peer_pp": 13.974058526666042,
5809
  "highlight_dominators": []
5810
  }
5811
  },
 
6
  "observation_policy": "Only exactly equal model/name/size/score duplicates are coalesced with every source retained. Differing registry scores remain separate observations, with scored revision unknown. F2/AST independent reruns remain separate pinned observations and do not replace reported coordinates. Current other Vela variant is included.",
7
  "dominance": "A peer has no more total parameters and no lower score, with at least one strict inequality. Scores use original available floating values, not rounded display labels.",
8
  "sources": {
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9
  "nano_English41": {
10
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/mteb-eng-v2.json",
11
+ "evidence_type": "Vela_release_report",
12
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
13
+ "sha256": "f3daecad1cf431630598efb04f3667431e3420239240f28a51568a74cef0c64f"
14
  },
15
  "nano_audio19": {
16
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
17
+ "evidence_type": "Vela_release_report",
18
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
19
+ "sha256": "4fc7f217745eb0aa998bf0f8e044bd1fa5de2f73f1a7dead7b829a62c48669ff"
20
  },
21
  "mini_English41": {
22
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/mteb-eng-v2.json",
 
30
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731",
31
  "sha256": "b9adc831e1de9caed37d38292dc95d65f4f933178fa7abbf7c1211f23617efbf"
32
  },
33
+ "scores_MTEB_eng_v2": {
34
+ "url": "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MTEB%28eng%2C%20v2%29/scores",
35
+ "sha256": "e574f14dc778d72c08255450091191b103930c487fa172a2dfabe41a0b56569d",
36
+ "retrieved_at_utc": "2026-09-17T20:28:47.419876+00:00",
37
+ "bytes": 1071836,
38
+ "benchmark": "MTEB(eng, v2)"
39
+ },
40
+ "scores_MAEB_beta_audio-only": {
41
+ "url": "https://mteb-leaderboard-backend.hf.space/v1/benchmarks/MAEB%28beta%2C%20audio-only%29/scores",
42
+ "sha256": "6153942c26a31577cc1e3a2fbb6eb10b09dc2f1fe08749db598d261e653f6bca",
43
+ "retrieved_at_utc": "2026-09-17T20:28:48.670374+00:00",
44
+ "bytes": 194971,
45
+ "benchmark": "MAEB(beta, audio-only)"
46
+ },
47
  "mini_component_equivalence": {
48
  "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Mini/resolve/main/benchmarks/component-equivalence.json",
49
  "sha256": "8bbbd3a362d84a6e8f030803259db69e932e5997b1c7af1a276a0d17862eb184",
50
  "evidence_type": "exact_component_identity_report",
51
  "native_artifact_sha256": "d3a742fbe0d833b0a72e5104e0381209b7d3da78f948d3439a5da7d986232731"
52
+ },
53
+ "nano_component_equivalence": {
54
+ "url": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/component-equivalence.json",
55
+ "evidence_type": "Vela_release_report",
56
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
57
+ "sha256": "ebd40f9a887791d1736132c37f510a69e18a03e247422853f54a42b417b0db9b"
58
  }
59
  },
60
  "models": {
61
  "nano": {
62
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Nano",
63
+ "release_commit": null,
64
+ "parameters_exact": 163771288,
65
+ "current_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
66
  "task_evidence": {
67
  "English41": {
68
  "source_id": "nano_English41",
 
70
  "model": "local/Vela-Omni-Nano-GIST-Rotation",
71
  "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
72
  "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
73
+ "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
74
  "component_applicability": {
75
  "component": "text",
 
 
76
  "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
 
77
  "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
78
+ "applies_to_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
79
+ "prior_component_applicability": {
80
+ "model": "local/Vela-Omni-Nano-GIST-Rotation",
81
+ "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
82
+ "native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
83
+ "component_applicability": {
84
+ "component": "text",
85
+ "applies_to_revision": "self",
86
+ "applies_to_native_artifact_sha256": "96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4",
87
+ "evaluated_model": "local/Vela-Omni-Nano-GIST-Rotation",
88
+ "evaluated_revision": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
89
+ "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646",
90
+ "revision_kind": "local native artifact fingerprint, not a Hugging Face commit",
91
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/component-equivalence.json",
92
+ "new_full_panel_inference": false,
93
+ "basis": "199 exact text tensors and byte-identical inference/tokenizer/configuration under the same official evaluator protocol."
94
+ },
95
+ "revision_semantics": "local native artifact fingerprint, not a Hugging Face commit",
96
+ "inference_profile_applicability": {
97
+ "revision": "self",
98
+ "native_artifact_sha256": "50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa",
99
+ "parameters": 135383808,
100
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/d8ac5b5ac2274a501fc61aeb5be70cec1855a806/benchmarks/inference-equivalence.json",
101
+ "new_benchmark_run": false
102
+ }
103
+ },
104
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/component-equivalence.json",
105
  "new_full_panel_inference": false,
106
+ "basis": "The frozen text tensors, tokenizer, configuration and encoding computation are unchanged. The original measured text identity is retained."
 
 
 
 
 
 
 
 
107
  }
108
  },
109
+ "self_reference_interpretation": "Nano native fingerprint 474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2"
110
  },
111
  "audio19": {
112
  "source_id": "nano_audio19",
113
  "identity": {
114
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
115
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
116
+ "release_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
117
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
118
+ "new_full_panel_inference": true
 
 
 
 
 
 
119
  },
120
+ "self_reference_interpretation": "Nano native fingerprint 474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2"
121
  }
122
+ },
123
+ "release_reference": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano",
124
+ "release_revision_semantics": "Exact native fingerprint below; Hub commit is not encoded in this file."
125
  },
126
  "mini": {
127
  "repo_id": "llm-semantic-router/Vela-1.0-Omni-Mini",
 
1744
  }
1745
  ],
1746
  "kind": "peer",
1747
+ "on_displayed_frontier": true
1748
  },
1749
  {
1750
  "model": "Omartificial-Intelligence-Space/Marbert-all-nli-triplet-Matryoshka",
 
4112
  {
4113
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
4114
  "kind": "nano",
4115
+ "parameters_billion": 0.163771288,
4116
+ "parameters_exact": 163771288,
4117
  "score": 91.94760000000001,
4118
+ "evidence_type": "exact_text_component_applicability",
4119
  "identity_ref": "models.nano.task_evidence.English41",
4120
  "size_basis": "current_whole_model_exact_tensor_count",
4121
  "open_weights": true,
4122
+ "on_displayed_frontier": true,
4123
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
4124
+ "revision": null,
4125
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/mteb-eng-v2.json",
4126
+ "evaluated_native_artifact_sha256": "29994142677559ca657c27b657eda0dddbf583bd4cf42a3a31095ab3f99d1646"
4127
  },
4128
  {
4129
  "model": "llm-semantic-router/Vela-1.0-Omni-Mini",
 
4148
  "Mihaiii/Squirtle",
4149
  "MongoDB/mdbr-leaf-mt",
4150
  "BAAI/bge-small-en",
4151
+ "codefuse-ai/F2LLM-v2-160M",
4152
  "llm-semantic-router/Vela-1.0-Omni-Nano",
4153
  "jinaai/jina-embeddings-v5-text-nano",
4154
  "jxm/cde-small-v2",
 
4387
  }
4388
  ],
4389
  "unknown_size_observation_count": 28,
4390
+ "highlight_gap_above_best_equal_or_smaller_peer_pp": 0.23280000000001166,
4391
  "highlight_dominators": []
4392
  },
4393
  "mridingham": {
 
5784
  {
5785
  "model": "llm-semantic-router/Vela-1.0-Omni-Nano",
5786
  "kind": "nano",
5787
+ "parameters_billion": 0.163771288,
5788
+ "parameters_exact": 163771288,
5789
+ "score": 58.41978617863634,
5790
+ "evidence_type": "measured_release_evaluation",
5791
  "identity_ref": "models.nano.task_evidence.audio19",
5792
  "size_basis": "current_whole_model_exact_tensor_count",
5793
  "open_weights": true,
5794
+ "on_displayed_frontier": true,
5795
+ "native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
5796
+ "revision": null,
5797
+ "report": "https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/resolve/main/benchmarks/maeb-audio-only.json",
5798
+ "evaluated_native_artifact_sha256": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
5799
+ "evaluated_revision": "474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2",
5800
+ "evaluated_model": "llm-semantic-router/Vela-1.0-Omni-Nano"
5801
  },
5802
  {
5803
  "model": "llm-semantic-router/Vela-1.0-Omni-Mini",
 
5820
  "google/yamnet",
5821
  "matthewagi/HeAR-s1.1",
5822
  "MIT/ast-finetuned-audioset-10-10-0.4593",
5823
+ "llm-semantic-router/Vela-1.0-Omni-Nano",
5824
  "llm-semantic-router/Vela-1.0-Omni-Mini"
5825
  ],
5826
  "reported_peer_observations": 65,
5827
  "unknown_size_observations_excluded": [],
5828
  "unknown_size_observation_count": 0,
5829
+ "highlight_gap_above_best_equal_or_smaller_peer_pp": 9.717472348029702,
5830
  "highlight_dominators": []
5831
  }
5832
  },
benchmarks/pareto-points.csv CHANGED
@@ -81,7 +81,7 @@ ArXivHierarchicalClusteringP2P,v_measure,nomic-ai/nomic-embed-text-v1,peer,0.137
81
  ArXivHierarchicalClusteringP2P,v_measure,nomic-ai/nomic-embed-text-v1-ablated,peer,0.137,57.903400000000005,True,False,registry_reported,registry_reported:nomic-ai/nomic-embed-text-v1-ablated,,
82
  ArXivHierarchicalClusteringP2P,v_measure,nomic-ai/nomic-embed-text-v1-unsupervised,peer,0.137,56.7453,True,False,registry_reported,registry_reported:nomic-ai/nomic-embed-text-v1-unsupervised,,
83
  ArXivHierarchicalClusteringP2P,v_measure,nomic-ai/nomic-embed-text-v1.5,peer,0.137,60.08051520875409,True,False,registry_reported,registry_reported:nomic-ai/nomic-embed-text-v1.5,,
84
- ArXivHierarchicalClusteringP2P,v_measure,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.135383808,64.88847259326147,True,False,exact_public_inference_applicability,vela_release_evaluation:llm-semantic-router/Vela-1.0-Omni-Nano,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa
85
  ArXivHierarchicalClusteringP2P,v_measure,ibm-granite/granite-embedding-english-r2,peer,0.149,59.0566,True,False,registry_reported,registry_reported:ibm-granite/granite-embedding-english-r2,,
86
  ArXivHierarchicalClusteringP2P,v_measure,Tarka-AIR/Tarka-Embedding-150M-V1,peer,0.156,63.397099999999995,True,False,registry_reported,registry_reported:Tarka-AIR/Tarka-Embedding-150M-V1,,
87
  ArXivHierarchicalClusteringP2P,v_measure,codefuse-ai/F2LLM-v2-160M,peer,0.159,62.4399,True,False,registry_reported,registry_reported:codefuse-ai/F2LLM-v2-160M,,
@@ -234,7 +234,7 @@ NMSQAPairClassification,max_ap,microsoft/wavlm-base-plus-sv,peer,0.095,48.4168,T
234
  NMSQAPairClassification,max_ap,microsoft/wavlm-base-sd,peer,0.095,52.9964,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sd,,
235
  NMSQAPairClassification,max_ap,microsoft/wavlm-base-sv,peer,0.095,52.9964,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sv,,
236
  NMSQAPairClassification,max_ap,asapp/sew-d-mid-400k-ft-ls100h,peer,0.139,48.4208,True,False,registry_reported,registry_reported:asapp/sew-d-mid-400k-ft-ls100h,,
237
- NMSQAPairClassification,max_ap,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.135383808,63.28340215485784,True,True,exact_public_inference_applicability,,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa
238
  NMSQAPairClassification,max_ap,laion/clap-htsat-unfused,peer,0.153,46.9791,True,False,registry_reported,registry_reported:laion/clap-htsat-unfused,,
239
  NMSQAPairClassification,max_ap,laion/clap-htsat-fused,peer,0.154,47.8507,True,False,registry_reported,registry_reported:laion/clap-htsat-fused,,
240
  NMSQAPairClassification,max_ap,microsoft/msclap-2023,peer,0.16,51.0934,True,False,registry_reported,registry_reported:microsoft/msclap-2023,,
@@ -302,7 +302,7 @@ SIBFLEURS,accuracy,microsoft/wavlm-base-plus-sv,peer,0.095,11.857322549019608,Tr
302
  SIBFLEURS,accuracy,microsoft/wavlm-base-sd,peer,0.095,13.31318431372549,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sd,,
303
  SIBFLEURS,accuracy,microsoft/wavlm-base-sv,peer,0.095,13.31318431372549,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sv,,
304
  SIBFLEURS,accuracy,asapp/sew-d-mid-400k-ft-ls100h,peer,0.139,12.634273529411763,True,False,registry_reported,registry_reported:asapp/sew-d-mid-400k-ft-ls100h,,
305
- SIBFLEURS,accuracy,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.135383808,18.304657831512046,True,False,exact_public_inference_applicability,,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa
306
  SIBFLEURS,accuracy,laion/clap-htsat-unfused,peer,0.153,8.674730392156862,True,False,registry_reported,registry_reported:laion/clap-htsat-unfused,,
307
  SIBFLEURS,accuracy,laion/clap-htsat-fused,peer,0.154,9.069597058823529,True,False,registry_reported,registry_reported:laion/clap-htsat-fused,,
308
  SIBFLEURS,accuracy,microsoft/msclap-2023,peer,0.16,14.073474509803921,True,False,registry_reported,registry_reported:microsoft/msclap-2023,,
@@ -372,7 +372,7 @@ VehicleSoundClustering,v_measure,microsoft/wavlm-base-plus-sv,peer,0.095,1.7932,
372
  VehicleSoundClustering,v_measure,microsoft/wavlm-base-sd,peer,0.095,0.40940000000000004,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sd,,
373
  VehicleSoundClustering,v_measure,microsoft/wavlm-base-sv,peer,0.095,0.40940000000000004,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sv,,
374
  VehicleSoundClustering,v_measure,asapp/sew-d-mid-400k-ft-ls100h,peer,0.139,0.6276,True,False,registry_reported,registry_reported:asapp/sew-d-mid-400k-ft-ls100h,,
375
- VehicleSoundClustering,v_measure,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.135383808,13.800701178563807,True,False,exact_public_inference_applicability,,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa,50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa
376
  VehicleSoundClustering,v_measure,laion/clap-htsat-unfused,peer,0.153,2.6642,True,False,registry_reported,registry_reported:laion/clap-htsat-unfused,,
377
  VehicleSoundClustering,v_measure,laion/clap-htsat-fused,peer,0.154,4.7132,True,False,registry_reported,registry_reported:laion/clap-htsat-fused,,
378
  VehicleSoundClustering,v_measure,microsoft/msclap-2023,peer,0.16,2.86,True,False,registry_reported,registry_reported:microsoft/msclap-2023,,
 
81
  ArXivHierarchicalClusteringP2P,v_measure,nomic-ai/nomic-embed-text-v1-ablated,peer,0.137,57.903400000000005,True,False,registry_reported,registry_reported:nomic-ai/nomic-embed-text-v1-ablated,,
82
  ArXivHierarchicalClusteringP2P,v_measure,nomic-ai/nomic-embed-text-v1-unsupervised,peer,0.137,56.7453,True,False,registry_reported,registry_reported:nomic-ai/nomic-embed-text-v1-unsupervised,,
83
  ArXivHierarchicalClusteringP2P,v_measure,nomic-ai/nomic-embed-text-v1.5,peer,0.137,60.08051520875409,True,False,registry_reported,registry_reported:nomic-ai/nomic-embed-text-v1.5,,
84
+ ArXivHierarchicalClusteringP2P,v_measure,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.163771288,64.88847259326147,True,False,exact_text_component_applicability,vela_release_evaluation:llm-semantic-router/Vela-1.0-Omni-Nano,,474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2
85
  ArXivHierarchicalClusteringP2P,v_measure,ibm-granite/granite-embedding-english-r2,peer,0.149,59.0566,True,False,registry_reported,registry_reported:ibm-granite/granite-embedding-english-r2,,
86
  ArXivHierarchicalClusteringP2P,v_measure,Tarka-AIR/Tarka-Embedding-150M-V1,peer,0.156,63.397099999999995,True,False,registry_reported,registry_reported:Tarka-AIR/Tarka-Embedding-150M-V1,,
87
  ArXivHierarchicalClusteringP2P,v_measure,codefuse-ai/F2LLM-v2-160M,peer,0.159,62.4399,True,False,registry_reported,registry_reported:codefuse-ai/F2LLM-v2-160M,,
 
234
  NMSQAPairClassification,max_ap,microsoft/wavlm-base-sd,peer,0.095,52.9964,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sd,,
235
  NMSQAPairClassification,max_ap,microsoft/wavlm-base-sv,peer,0.095,52.9964,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sv,,
236
  NMSQAPairClassification,max_ap,asapp/sew-d-mid-400k-ft-ls100h,peer,0.139,48.4208,True,False,registry_reported,registry_reported:asapp/sew-d-mid-400k-ft-ls100h,,
237
+ NMSQAPairClassification,max_ap,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.163771288,62.89933955933497,True,True,measured_release_evaluation,,,474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2
238
  NMSQAPairClassification,max_ap,laion/clap-htsat-unfused,peer,0.153,46.9791,True,False,registry_reported,registry_reported:laion/clap-htsat-unfused,,
239
  NMSQAPairClassification,max_ap,laion/clap-htsat-fused,peer,0.154,47.8507,True,False,registry_reported,registry_reported:laion/clap-htsat-fused,,
240
  NMSQAPairClassification,max_ap,microsoft/msclap-2023,peer,0.16,51.0934,True,False,registry_reported,registry_reported:microsoft/msclap-2023,,
 
302
  SIBFLEURS,accuracy,microsoft/wavlm-base-sd,peer,0.095,13.31318431372549,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sd,,
303
  SIBFLEURS,accuracy,microsoft/wavlm-base-sv,peer,0.095,13.31318431372549,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sv,,
304
  SIBFLEURS,accuracy,asapp/sew-d-mid-400k-ft-ls100h,peer,0.139,12.634273529411763,True,False,registry_reported,registry_reported:asapp/sew-d-mid-400k-ft-ls100h,,
305
+ SIBFLEURS,accuracy,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.163771288,18.072541269472218,True,False,measured_release_evaluation,,,474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2
306
  SIBFLEURS,accuracy,laion/clap-htsat-unfused,peer,0.153,8.674730392156862,True,False,registry_reported,registry_reported:laion/clap-htsat-unfused,,
307
  SIBFLEURS,accuracy,laion/clap-htsat-fused,peer,0.154,9.069597058823529,True,False,registry_reported,registry_reported:laion/clap-htsat-fused,,
308
  SIBFLEURS,accuracy,microsoft/msclap-2023,peer,0.16,14.073474509803921,True,False,registry_reported,registry_reported:microsoft/msclap-2023,,
 
372
  VehicleSoundClustering,v_measure,microsoft/wavlm-base-sd,peer,0.095,0.40940000000000004,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sd,,
373
  VehicleSoundClustering,v_measure,microsoft/wavlm-base-sv,peer,0.095,0.40940000000000004,True,False,registry_reported,registry_reported:microsoft/wavlm-base-sv,,
374
  VehicleSoundClustering,v_measure,asapp/sew-d-mid-400k-ft-ls100h,peer,0.139,0.6276,True,False,registry_reported,registry_reported:asapp/sew-d-mid-400k-ft-ls100h,,
375
+ VehicleSoundClustering,v_measure,llm-semantic-router/Vela-1.0-Omni-Nano,nano,0.163771288,5.275425308186281,True,False,measured_release_evaluation,,,474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2
376
  VehicleSoundClustering,v_measure,laion/clap-htsat-unfused,peer,0.153,2.6642,True,False,registry_reported,registry_reported:laion/clap-htsat-unfused,,
377
  VehicleSoundClustering,v_measure,laion/clap-htsat-fused,peer,0.154,4.7132,True,False,registry_reported,registry_reported:laion/clap-htsat-fused,,
378
  VehicleSoundClustering,v_measure,microsoft/msclap-2023,peer,0.16,2.86,True,False,registry_reported,registry_reported:microsoft/msclap-2023,,
benchmarks/previous-release-scores.json ADDED
The diff for this file is too large to render. See raw diff
 
benchmarks/release-history.md CHANGED
@@ -1,94 +1,74 @@
1
- # Release-to-release history
2
 
3
- These are historical Vela-to-Vela comparisons, not the original small baseline used by the current scorecard. All values and comparators remain as originally measured.
4
 
5
- ## Historical identities
6
 
7
- This report retains measurements from Nano artifact `96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4` (144,554,112 parameters). The previous Nano is [3a1efc48](https://huggingface.co/llm-semantic-router/Vela-1.0-Omni-Nano/tree/3a1efc48995e2b71622d65c6cf71cfc00bbf3aa3) (`fb9cd39c9980a085b75456c715c2295b997784fe733314be68c13635f6384edf`). Scores are shown on a 0–100 scale. Bold means an unrounded improvement over the explicitly named comparator.
8
-
9
-
10
- ## Every English change
11
-
12
- | Task | Main metric | Previous Vela | Measured Vela | Change (pp) |
13
  | --- | ---: | ---: | ---: | ---: |
14
- | ArguAna | ndcg_at_10 | 37.5140 | 59.2460 | +21.7320 |
15
- | ArXivHierarchicalClusteringP2P | v_measure | 65.2540 | 64.8885 | -0.3655 |
16
- | ArXivHierarchicalClusteringS2S | v_measure | 55.9895 | 57.4432 | +1.4536 |
17
- | AskUbuntuDupQuestions | map_at_1000 | 60.5140 | 62.3290 | +1.8150 |
18
- | BIOSSES | cosine_spearman | 62.2454 | 86.9875 | +24.7421 |
19
- | Banking77Classification | accuracy | 69.5747 | 82.1429 | +12.5682 |
20
- | BiorxivClusteringP2P.v2 | v_measure | 39.8292 | 41.3138 | +1.4846 |
21
- | CQADupstackGamingRetrieval | ndcg_at_10 | 42.9760 | 56.9790 | +14.0030 |
22
- | CQADupstackUnixRetrieval | ndcg_at_10 | 30.3620 | 39.5860 | +9.2240 |
23
- | ClimateFEVERHardNegatives | ndcg_at_10 | 16.0740 | 31.8460 | +15.7720 |
24
- | FEVERHardNegatives | ndcg_at_10 | 24.0550 | 87.5760 | +63.5210 |
25
- | FiQA2018 | ndcg_at_10 | 22.5880 | 39.1430 | +16.5550 |
26
- | HotpotQAHardNegatives | ndcg_at_10 | 26.9870 | 66.3490 | +39.3620 |
27
- | ImdbClassification | accuracy | 62.1364 | 91.9476 | +29.8112 |
28
- | MTOPDomainClassification | accuracy | 86.7191 | 94.9179 | +8.1988 |
29
- | MassiveIntentClassification | accuracy | 60.3934 | 70.9684 | +10.5750 |
30
- | MassiveScenarioClassification | accuracy | 73.3726 | 76.1130 | +2.7404 |
31
- | MedrxivClusteringP2P.v2 | v_measure | 37.7681 | 39.8302 | +2.0621 |
32
- | MedrxivClusteringS2S.v2 | v_measure | 36.1335 | 37.8751 | +1.7416 |
33
- | MindSmallReranking | max_over_subqueries_map_at_1000 | 30.6230 | 32.3660 | +1.7430 |
34
- | SCIDOCS | ndcg_at_10 | 16.8270 | 21.8890 | +5.0620 |
35
- | SICK-R | cosine_spearman | 71.6962 | 80.5317 | +8.8356 |
36
- | STS12 | cosine_spearman | 68.1692 | 75.5659 | +7.3967 |
37
- | STS13 | cosine_spearman | 74.4034 | 86.2635 | +11.8600 |
38
- | STS14 | cosine_spearman | 68.0265 | 82.2988 | +14.2723 |
39
- | STS15 | cosine_spearman | 77.8614 | 88.7365 | +10.8750 |
40
- | STSBenchmark | cosine_spearman | 77.9240 | 87.0782 | +9.1542 |
41
- | SprintDuplicateQuestions | max_ap | 89.7910 | 95.7996 | +6.0087 |
42
- | StackExchangeClustering.v2 | v_measure | 57.3446 | 58.4244 | +1.0798 |
43
- | StackExchangeClusteringP2P.v2 | v_measure | 39.3767 | 41.0530 | +1.6763 |
44
- | TRECCOVID | ndcg_at_10 | 41.2920 | 69.1210 | +27.8290 |
45
- | Touche2020Retrieval.v3 | ndcg_at_10 | 31.4770 | 48.3630 | +16.8860 |
46
- | ToxicConversationsClassification | accuracy | 64.4092 | 71.9043 | +7.4951 |
47
- | TweetSentimentExtractionClassification | accuracy | 50.7782 | 62.7929 | +12.0147 |
48
- | TwentyNewsgroupsClustering.v2 | v_measure | 44.8847 | 50.8726 | +5.9879 |
49
- | TwitterSemEval2015 | max_ap | 56.3448 | 72.9515 | +16.6068 |
50
- | TwitterURLCorpus | max_ap | 78.7021 | 85.3034 | +6.6013 |
51
- | SummEvalSummarization.v2 | cosine_spearman | 29.2496 | 31.8344 | +2.5848 |
52
- | AmazonCounterfactualClassification | accuracy | 56.8209 | 71.8507 | +15.0299 |
53
- | STS17 | cosine_spearman | 84.1659 | 89.0210 | +4.8551 |
54
- | STS22.v2 | cosine_spearman | 67.5701 | 68.7533 | +1.1833 |
55
 
56
- ## Every audio change
57
 
58
- | Task | Main metric | Previous Vela | Measured Vela | Change (pp) |
59
- | --- | ---: | ---: | ---: | ---: |
60
- | JamAltArtistA2ARetrieval | ndcg_at_10 | 68.9005 | 68.9005 | +0.0000 |
61
- | BeijingOpera | accuracy | 68.6791 | 68.6791 | +0.0000 |
62
- | BirdCLEF | accuracy | 12.9000 | 12.9000 | +0.0000 |
63
- | CREMA_D | accuracy | 29.7637 | 29.7637 | +0.0000 |
64
- | CommonLanguageAgeDetection | accuracy | 16.0400 | 16.0400 | +0.0000 |
65
- | GTZANGenre | accuracy | 50.3000 | 50.3000 | +0.0000 |
66
- | IEMOCAPGender | accuracy | 58.4519 | 58.4519 | +0.0000 |
67
- | MInDS14 | accuracy | 23.3212 | 23.3212 | +0.0000 |
68
- | MridinghamTonic | accuracy | 32.5066 | 32.3632 | -0.1434 |
69
- | SIBFLEURS | accuracy | 18.3047 | 18.3047 | +0.0000 |
70
- | VoxCelebSA | accuracy | 33.1691 | 33.1691 | +0.0000 |
71
- | VoxPopuliLanguageID | accuracy | 91.8000 | 91.8000 | +0.0000 |
72
- | CREMA_DClustering | v_measure | 2.3602 | 2.3602 | +0.0000 |
73
- | VehicleSoundClustering | v_measure | 13.8007 | 13.8007 | +0.0000 |
74
- | VoxPopuliGenderClustering | v_measure | 0.1045 | 0.1045 | +0.0000 |
75
- | CREMADPairClassification | max_ap | 54.0802 | 54.2012 | +0.1210 |
76
- | NMSQAPairClassification | max_ap | 63.4355 | 63.2834 | -0.1521 |
77
- | VoxPopuliAccentPairClassification | max_ap | 54.0504 | 54.0220 | -0.0284 |
78
- | GTZANAudioReranking | map_at_1000 | 63.7940 | 63.7940 | +0.0000 |
79
 
 
80
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
81
 
82
- | Component | Weight | Previous | Current | Change (pp) |
83
- | --- | ---: | ---: | ---: | ---: |
84
- | common14 | 30% | 59.3729 | 67.5741 | +8.2012 |
85
- | English41_MeanTaskType | 30% | 51.9755 | 60.7818 | +8.8063 |
86
- | image3 | 20% | 54.8658 | 53.6575 | -1.2083 |
87
- | audio19_MeanTaskType | 20% | 46.9744 | 46.9678 | -0.0066 |
88
-
89
- Weighted product progress is +4.8593 points. This is a product comparison, not an official leaderboard aggregate or overall SOTA claim. The common component equally weights text, image and audio families and uses all 14 metrics. The image component equally averages Pets classification, TinyImageNet NMI and Pets zero-shot accuracy.
90
-
91
- Regressions remain visible: all three image-to-text recall ranks decline; Pets zero-shot decreases 3.625 points; SpeechCommands zero-shot decreases 0.859; ArXiv P2P decreases 0.3655. Three audio-panel tasks decline slightly (MridinghamTonic, NMSQA and VoxPopuliAccent); all other changes are listed above. Classification/clustering preservation does not imply preservation of cross-modal or short-label zero-shot behavior.
92
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
93
 
94
- The exact original common metrics, full paired intervals, and numerical history remain in [scores.json](../scores.json) and [net-progress.json](net-progress.json). This file is historical evidence, not an additional product comparator.
 
1
+ # Release progress and trade-offs
2
 
3
+ This update compares native `474ad3d52f6ee7f818a3dbb6112e3d66c2423542d08228e7996a8cd600bd99a2` against the previous Nano `50ce808197fb43a1b8913a66d4a47339ec66ecd10d52cbd9f9e752d02f8c43fa` (measured as `96d76f97478e114b816df3ab1b3c6947e4c5ef0babfb40ab865d7e3c838d34c4`). The primary product scorecard continues to compare against the original small.
4
 
5
+ Fixed product-weighted change: **+0.726802 pp**. Parameters increase **20.9682%**. This custom weighted result is not an official benchmark aggregate, statistical significance claim or full-size SOTA result.
6
 
7
+ | Component | Previous | Current | Delta pp | Weight |
 
 
 
 
 
8
  | --- | ---: | ---: | ---: | ---: |
9
+ | common14 | 67.574073 | 66.418353 | -1.155720 | 0.3 |
10
+ | English41_MeanTaskType | 60.781850 | 60.781850 | +0.000000 | 0.3 |
11
+ | image3 | 53.657464 | 53.657464 | +0.000000 | 0.2 |
12
+ | audio19_MeanTaskType | 46.967814 | 52.335406 | +5.367592 | 0.2 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
13
 
14
+ ## Common metrics
15
 
16
+ | Metric | Delta pp |
17
+ | --- | ---: |
18
+ | audio.audio_to_text.recall@1 | -1.034087 |
19
+ | audio.audio_to_text.recall@10 | -1.378782 |
20
+ | audio.audio_to_text.recall@5 | -1.838376 |
21
+ | audio.text_to_audio.recall@1 | -5.095785 |
22
+ | audio.text_to_audio.recall@10 | -5.287356 |
23
+ | audio.text_to_audio.recall@5 | -6.168582 |
24
+ | image.image_to_text.recall@1 | +0.000000 |
25
+ | image.image_to_text.recall@10 | +0.000000 |
26
+ | image.image_to_text.recall@5 | +0.000000 |
27
+ | image.text_to_image.recall@1 | +0.000000 |
28
+ | image.text_to_image.recall@10 | +0.000000 |
29
+ | image.text_to_image.recall@5 | +0.000000 |
30
+ | text.banking77.accuracy | +0.000000 |
31
+ | text.massive-en.accuracy | +0.000000 |
 
 
 
 
 
32
 
33
+ ## Every audio task
34
 
35
+ | Task | Delta pp |
36
+ | --- | ---: |
37
+ | BeijingOpera | +10.549645 |
38
+ | BirdCLEF | +2.300000 |
39
+ | CREMADPairClassification | +1.006025 |
40
+ | CREMA_D | +2.284477 |
41
+ | CREMA_DClustering | +2.854175 |
42
+ | CommonLanguageAgeDetection | +0.615000 |
43
+ | GTZANAudioReranking | +6.682000 |
44
+ | GTZANGenre | +11.400000 |
45
+ | IEMOCAPGender | +3.237876 |
46
+ | JamAltArtistA2ARetrieval | +16.901000 |
47
+ | MInDS14 | -1.469131 |
48
+ | MridinghamTonic | +26.056577 |
49
+ | NMSQAPairClassification | -0.384063 |
50
+ | SIBFLEURS | -0.232117 |
51
+ | VehicleSoundClustering | -8.525276 |
52
+ | VoxCelebSA | -0.000589 |
53
+ | VoxPopuliAccentPairClassification | +0.129363 |
54
+ | VoxPopuliGenderClustering | -0.081273 |
55
+ | VoxPopuliLanguageID | -0.600000 |
56
 
57
+ ## Selected standard protocols
 
 
 
 
 
 
 
 
 
58
 
59
+ | Task / split | Delta pp |
60
+ | --- | ---: |
61
+ | SICK-R / test | +0.000000 |
62
+ | STSBenchmark / test | +0.000000 |
63
+ | ArguAna / test | +0.000000 |
64
+ | TwentyNewsgroupsClustering.v2 / test | +0.000000 |
65
+ | MTOPDomainClassification / validation | +0.000000 |
66
+ | MTOPDomainClassification / test | +0.000000 |
67
+ | OxfordPets / test | +0.000000 |
68
+ | TinyImageNetClustering / valid | +0.000000 |
69
+ | OxfordPetsZeroShot / test | +0.000000 |
70
+ | CREMA_D / train | +2.284477 |
71
+ | CREMA_DClustering / train | +2.854175 |
72
+ | SpeechCommandsZeroshotv0.02 / test | +1.595484 |
73
 
74
+ The 840-clip event diagnostic is reported separately and does not enter the weighted value. English and image computations are unchanged; their original measurement identities are preserved. Prior results remain in [previous-release-scores.json](previous-release-scores.json).
components/audio_clap/config.json ADDED
@@ -0,0 +1,206 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_commit_hash": null,
3
+ "architectures": [
4
+ "ClapModel"
5
+ ],
6
+ "audio_config": {
7
+ "_name_or_path": "",
8
+ "add_cross_attention": false,
9
+ "aff_block_r": 4,
10
+ "architectures": null,
11
+ "attention_probs_dropout_prob": 0.0,
12
+ "bad_words_ids": null,
13
+ "begin_suppress_tokens": null,
14
+ "bos_token_id": null,
15
+ "chunk_size_feed_forward": 0,
16
+ "cross_attention_hidden_size": null,
17
+ "decoder_start_token_id": null,
18
+ "depths": [
19
+ 2,
20
+ 2,
21
+ 6,
22
+ 2
23
+ ],
24
+ "diversity_penalty": 0.0,
25
+ "do_sample": false,
26
+ "drop_path_rate": 0.0,
27
+ "early_stopping": false,
28
+ "enable_fusion": false,
29
+ "enable_patch_layer_norm": true,
30
+ "encoder_no_repeat_ngram_size": 0,
31
+ "eos_token_id": null,
32
+ "exponential_decay_length_penalty": null,
33
+ "finetuning_task": null,
34
+ "flatten_patch_embeds": true,
35
+ "forced_bos_token_id": null,
36
+ "forced_eos_token_id": null,
37
+ "fusion_num_hidden_layers": 2,
38
+ "fusion_type": null,
39
+ "hidden_act": "gelu",
40
+ "hidden_dropout_prob": 0.1,
41
+ "hidden_size": 768,
42
+ "id2label": {
43
+ "0": "LABEL_0",
44
+ "1": "LABEL_1"
45
+ },
46
+ "initializer_factor": 1.0,
47
+ "is_decoder": false,
48
+ "is_encoder_decoder": false,
49
+ "label2id": {
50
+ "LABEL_0": 0,
51
+ "LABEL_1": 1
52
+ },
53
+ "layer_norm_eps": 1e-05,
54
+ "length_penalty": 1.0,
55
+ "max_length": 20,
56
+ "min_length": 0,
57
+ "mlp_ratio": 4.0,
58
+ "model_type": "clap_audio_model",
59
+ "no_repeat_ngram_size": 0,
60
+ "num_attention_heads": [
61
+ 4,
62
+ 8,
63
+ 16,
64
+ 32
65
+ ],
66
+ "num_beam_groups": 1,
67
+ "num_beams": 1,
68
+ "num_classes": 527,
69
+ "num_hidden_layers": 4,
70
+ "num_mel_bins": 64,
71
+ "num_return_sequences": 1,
72
+ "output_attentions": false,
73
+ "output_hidden_states": false,
74
+ "output_scores": false,
75
+ "pad_token_id": null,
76
+ "patch_embed_input_channels": 1,
77
+ "patch_embeds_hidden_size": 96,
78
+ "patch_size": 4,
79
+ "patch_stride": [
80
+ 4,
81
+ 4
82
+ ],
83
+ "prefix": null,
84
+ "problem_type": null,
85
+ "projection_dim": 512,
86
+ "projection_hidden_act": "relu",
87
+ "projection_hidden_size": 768,
88
+ "pruned_heads": {},
89
+ "qkv_bias": true,
90
+ "remove_invalid_values": false,
91
+ "repetition_penalty": 1.0,
92
+ "return_dict": true,
93
+ "return_dict_in_generate": false,
94
+ "sep_token_id": null,
95
+ "spec_size": 256,
96
+ "suppress_tokens": null,
97
+ "task_specific_params": null,
98
+ "temperature": 1.0,
99
+ "tf_legacy_loss": false,
100
+ "tie_encoder_decoder": false,
101
+ "tie_word_embeddings": true,
102
+ "tokenizer_class": null,
103
+ "top_k": 50,
104
+ "top_p": 1.0,
105
+ "torch_dtype": null,
106
+ "torchscript": false,
107
+ "transformers_version": "4.27.0.dev0",
108
+ "typical_p": 1.0,
109
+ "use_bfloat16": false,
110
+ "window_size": 8
111
+ },
112
+ "hidden_size": 768,
113
+ "initializer_factor": 1.0,
114
+ "logit_scale_init_value": 14.285714285714285,
115
+ "model_type": "clap",
116
+ "num_hidden_layers": 16,
117
+ "projection_dim": 512,
118
+ "projection_hidden_act": "relu",
119
+ "text_config": {
120
+ "_name_or_path": "",
121
+ "add_cross_attention": false,
122
+ "architectures": null,
123
+ "attention_probs_dropout_prob": 0.1,
124
+ "bad_words_ids": null,
125
+ "begin_suppress_tokens": null,
126
+ "bos_token_id": 0,
127
+ "chunk_size_feed_forward": 0,
128
+ "classifier_dropout": null,
129
+ "cross_attention_hidden_size": null,
130
+ "decoder_start_token_id": null,
131
+ "diversity_penalty": 0.0,
132
+ "do_sample": false,
133
+ "early_stopping": false,
134
+ "encoder_no_repeat_ngram_size": 0,
135
+ "eos_token_id": 2,
136
+ "exponential_decay_length_penalty": null,
137
+ "finetuning_task": null,
138
+ "forced_bos_token_id": null,
139
+ "forced_eos_token_id": null,
140
+ "fusion_hidden_size": 768,
141
+ "fusion_num_hidden_layers": 2,
142
+ "hidden_act": "gelu",
143
+ "hidden_dropout_prob": 0.1,
144
+ "hidden_size": 768,
145
+ "id2label": {
146
+ "0": "LABEL_0",
147
+ "1": "LABEL_1"
148
+ },
149
+ "initializer_factor": 1.0,
150
+ "initializer_range": 0.02,
151
+ "intermediate_size": 3072,
152
+ "is_decoder": false,
153
+ "is_encoder_decoder": false,
154
+ "label2id": {
155
+ "LABEL_0": 0,
156
+ "LABEL_1": 1
157
+ },
158
+ "layer_norm_eps": 1e-12,
159
+ "length_penalty": 1.0,
160
+ "max_length": 20,
161
+ "max_position_embeddings": 514,
162
+ "min_length": 0,
163
+ "model_type": "clap_text_model",
164
+ "no_repeat_ngram_size": 0,
165
+ "num_attention_heads": 12,
166
+ "num_beam_groups": 1,
167
+ "num_beams": 1,
168
+ "num_hidden_layers": 12,
169
+ "num_return_sequences": 1,
170
+ "output_attentions": false,
171
+ "output_hidden_states": false,
172
+ "output_scores": false,
173
+ "pad_token_id": 1,
174
+ "position_embedding_type": "absolute",
175
+ "prefix": null,
176
+ "problem_type": null,
177
+ "projection_dim": 512,
178
+ "projection_hidden_act": "relu",
179
+ "projection_hidden_size": 768,
180
+ "pruned_heads": {},
181
+ "remove_invalid_values": false,
182
+ "repetition_penalty": 1.0,
183
+ "return_dict": true,
184
+ "return_dict_in_generate": false,
185
+ "sep_token_id": null,
186
+ "suppress_tokens": null,
187
+ "task_specific_params": null,
188
+ "temperature": 1.0,
189
+ "tf_legacy_loss": false,
190
+ "tie_encoder_decoder": false,
191
+ "tie_word_embeddings": true,
192
+ "tokenizer_class": null,
193
+ "top_k": 50,
194
+ "top_p": 1.0,
195
+ "torch_dtype": null,
196
+ "torchscript": false,
197
+ "transformers_version": "4.27.0.dev0",
198
+ "type_vocab_size": 1,
199
+ "typical_p": 1.0,
200
+ "use_bfloat16": false,
201
+ "use_cache": true,
202
+ "vocab_size": 50265
203
+ },
204
+ "torch_dtype": "float32",
205
+ "transformers_version": null
206
+ }
components/audio_clap/preprocessor_config.json ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "chunk_length_s": 10,
3
+ "feature_extractor_type": "ClapFeatureExtractor",
4
+ "feature_size": 64,
5
+ "fft_window_size": 1024,
6
+ "frequency_max": 14000,
7
+ "frequency_min": 50,
8
+ "hop_length": 480,
9
+ "max_length_s": 10,
10
+ "n_fft": 1024,
11
+ "nb_frequency_bins": 513,
12
+ "nb_max_frames": 1000,
13
+ "nb_max_samples": 480000,
14
+ "padding": "repeatpad",
15
+ "padding_side": "right",
16
+ "padding_value": 0.0,
17
+ "processor_class": "ClapProcessor",
18
+ "return_attention_mask": false,
19
+ "sampling_rate": 48000,
20
+ "top_db": null,
21
+ "truncation": "rand_trunc"
22
+ }
config.json CHANGED
@@ -2,16 +2,18 @@
2
  "architectures": [
3
  "VelaOmni"
4
  ],
 
5
  "audio_max_seconds": 30,
6
  "audio_pooling": "mean",
7
  "audio_projection": "linear",
8
  "audio_sampling_rate": 16000,
9
  "embedding_dim": 384,
10
- "format_version": 5,
11
  "inference_profile": "single_modality",
12
  "library_name": "pytorch",
13
  "max_text_length": 512,
14
- "parameter_count": 135383808,
 
15
  "text_pooling": "cls",
16
  "text_projection": "identity",
17
  "torch_dtype": "float32",
 
2
  "architectures": [
3
  "VelaOmni"
4
  ],
5
+ "audio_backend": "tiny_clap_residual_v1",
6
  "audio_max_seconds": 30,
7
  "audio_pooling": "mean",
8
  "audio_projection": "linear",
9
  "audio_sampling_rate": 16000,
10
  "embedding_dim": 384,
11
+ "format_version": 9,
12
  "inference_profile": "single_modality",
13
  "library_name": "pytorch",
14
  "max_text_length": 512,
15
+ "parameter_count": 163771288,
16
+ "residual_audio_readout": "tiny_full_mean_clap_endpoint_zero_linear",
17
  "text_pooling": "cls",
18
  "text_projection": "identity",
19
  "torch_dtype": "float32",
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:401759e85cec305aca7ff08b31a1a3efa2c410b24b7bb49f9a81de3aabafa52d
3
- size 541598024
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d5aa7f00e217bd4fd5314140a3603b9a873636320c8d09732a14f79f4cb0667e
3
+ size 655583800
omni_components/single_modality.py CHANGED
@@ -1,5 +1,7 @@
1
  """Nano inference runtime for the three public modality encoders."""
2
  from pathlib import Path
 
 
3
  from torch import nn
4
 
5
  from .audio_encoder import AudioEncoder
@@ -11,11 +13,14 @@ from .text_encoder import TextEncoder
11
  class SingleModalityEmbedder(nn.Module):
12
  """Full-depth encoders in a shared space, without optional fusion/exit heads."""
13
 
14
- def __init__(self, component_dir, legacy=False, max_text_length=512, text_pooling="cls"):
15
  super().__init__()
16
  if legacy:
17
  raise ValueError("Inference profile requires native Nano components")
18
  root = Path(component_dir)
 
 
 
19
  self.output_dim = 384
20
  self.normalize = True
21
  self.max_text_length = max_text_length
@@ -34,6 +39,10 @@ class SingleModalityEmbedder(nn.Module):
34
  enable_layer_outputs=False,
35
  )
36
 
 
 
 
 
37
  def encode_text(self, texts):
38
  return MultimodalEmbedder.encode_text(self, texts)
39
 
@@ -41,5 +50,36 @@ class SingleModalityEmbedder(nn.Module):
41
  return MultimodalEmbedder.encode_image(self, images)
42
 
43
  def encode_audio(self, audio, sampling_rate=16000):
 
 
44
  return MultimodalEmbedder.encode_audio(self, audio, sampling_rate=sampling_rate)
45
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  """Nano inference runtime for the three public modality encoders."""
2
  from pathlib import Path
3
+ import numpy as np
4
+ import torch
5
  from torch import nn
6
 
7
  from .audio_encoder import AudioEncoder
 
13
  class SingleModalityEmbedder(nn.Module):
14
  """Full-depth encoders in a shared space, without optional fusion/exit heads."""
15
 
16
+ def __init__(self, component_dir, legacy=False, max_text_length=512, text_pooling="cls", audio_backend="whisper"):
17
  super().__init__()
18
  if legacy:
19
  raise ValueError("Inference profile requires native Nano components")
20
  root = Path(component_dir)
21
+ if audio_backend not in ("whisper", "tiny_clap_residual_v1"):
22
+ raise ValueError("Unsupported Nano audio backend")
23
+ self.audio_backend = audio_backend
24
  self.output_dim = 384
25
  self.normalize = True
26
  self.max_text_length = max_text_length
 
39
  enable_layer_outputs=False,
40
  )
41
 
42
+ if audio_backend == "tiny_clap_residual_v1":
43
+ from .tiny_clap_residual import TinyClapResidual
44
+ self.audio_residual = TinyClapResidual(root)
45
+
46
  def encode_text(self, texts):
47
  return MultimodalEmbedder.encode_text(self, texts)
48
 
 
50
  return MultimodalEmbedder.encode_image(self, images)
51
 
52
  def encode_audio(self, audio, sampling_rate=16000):
53
+ if self.audio_backend == "tiny_clap_residual_v1":
54
+ return self._encode_audio_original(audio, sampling_rate)
55
  return MultimodalEmbedder.encode_audio(self, audio, sampling_rate=sampling_rate)
56
 
57
+
58
+ def _encode_audio_original(self, waveforms, sampling_rate):
59
+ if self.audio_backend != "tiny_clap_residual_v1":
60
+ raise ValueError("Original-rate dual audio requires residual backend")
61
+ from .tiny_clap_residual import validate_original, native_rate
62
+ waves = [np.array(validate_original(w, sampling_rate), dtype=np.float32, copy=True, order="C") for w in waveforms]
63
+ if not waves:
64
+ raise ValueError("Provide at least one original waveform")
65
+ results = []
66
+ for start in range(0, len(waves), 8):
67
+ current = waves[start:start + 8]
68
+ audio16 = [native_rate(w, sampling_rate, 16000) for w in current]
69
+ audio48 = [native_rate(w, sampling_rate, 48000) for w in current]
70
+ results.append(self._encode_audio_prepared(audio16, audio48))
71
+ return torch.cat(results)
72
+
73
+ def _encode_audio_prepared(self, audio16, audio48):
74
+ """Internal paired native-rate path; benchmark resamples before cropping."""
75
+ if self.audio_backend != "tiny_clap_residual_v1" or not 1 <= len(audio16) == len(audio48) <= 8:
76
+ raise ValueError("One through eight paired native-rate streams required")
77
+ for streams, rate in ((audio16, 16000), (audio48, 48000)):
78
+ if any(np.asarray(w).ndim != 1 or not 1 <= len(w) <= 30 * rate or not np.isfinite(w).all() for w in streams):
79
+ raise ValueError("Finite mono native-rate streams of at most30 seconds required")
80
+ if self.audio_encoder.normalize or not self.normalize or self.audio_encoder.pooling_mode != "mean":
81
+ raise ValueError("Pinned unnormalized Tiny affine and one embedding normalization required")
82
+ # Delegate the complete original AudioEncoder call, including processor and full mean.
83
+ original_affine = self.audio_encoder(audio=audio16, sampling_rate=16000)
84
+ clap = self.audio_residual.encode_clap(audio48)
85
+ return self.audio_residual(original_affine, clap)
omni_components/tiny_clap_residual.py ADDED
@@ -0,0 +1,82 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Pinned CLAP residual; retained Tiny branch and affine remain unchanged."""
2
+ from pathlib import Path
3
+ import numpy as np
4
+ import torch
5
+ from torch import nn
6
+ from torch.nn import functional as F
7
+ from transformers import ClapConfig,ClapFeatureExtractor,ClapAudioModelWithProjection
8
+
9
+ def unit(x):
10
+ if not torch.isfinite(x).all() or torch.any(x.norm(dim=-1)<=1e-12):
11
+ raise ValueError('Finite nonzero embedding required')
12
+ return F.normalize(x,dim=-1)
13
+
14
+
15
+ def validate_original(wave,rate):
16
+ if isinstance(rate,bool) or not isinstance(rate,(int,np.integer)) or not 1<=rate<=384000:
17
+ raise ValueError('Positive integer original sampling rate required')
18
+ x=np.asarray(wave,dtype=np.float32)
19
+ if x.ndim not in (1,2) or not np.isfinite(x).all() or x.shape[-1]<1:
20
+ raise ValueError('Finite nonempty mono or channels-first PCM required')
21
+ if x.ndim==2 and not 1<=x.shape[0]<=8:raise ValueError('Require1..8 channels-first PCM')
22
+ if x.shape[-1]>30*rate:raise ValueError('Audio inputs must be at most30 seconds')
23
+ return x
24
+
25
+
26
+ def native_rate(wave,source_rate,target_rate):
27
+ """Same operation order/default sinc kernel as official MTEB2.21 AudioCollator."""
28
+ if target_rate not in (16000,48000):raise ValueError('Unsupported branch rate')
29
+ # Each invocation sees the original waveform. Never cascade16k into48k.
30
+ x=validate_original(wave,source_rate)
31
+ if source_rate!=target_rate:
32
+ import torchaudio
33
+ x=torchaudio.transforms.Resample(orig_freq=source_rate,new_freq=target_rate)(torch.from_numpy(x).float()).numpy()
34
+ else:x=np.asarray(x)
35
+ if x.ndim>1 and x.shape[0]>1:x=np.mean(x,axis=0)
36
+ if x.shape[-1]>target_rate*30:x=x[...,:target_rate*30]
37
+ if x.ndim==2 and x.shape[0]==1:x=x[0]
38
+ if x.ndim!=1 or not np.isfinite(x).all() or not 1<=len(x)<=30*target_rate:raise ValueError('Invalid resampled mono PCM')
39
+ return np.ascontiguousarray(x,dtype=np.float32)
40
+
41
+
42
+ def windows(samples):
43
+ if not 1<=samples<=1440000:raise ValueError('CLAP requires at most30 seconds at48kHz')
44
+ if samples<=480000:return [(0,samples)]
45
+ count=(samples+479999)//480000;last=samples-480000
46
+ return [(i*last//(count-1),i*last//(count-1)+480000) for i in range(count)]
47
+
48
+
49
+
50
+ class TinyClapResidual(nn.Module):
51
+ def __init__(self, component_root):
52
+ super().__init__()
53
+ root=Path(component_root)/'audio_clap'
54
+ config=ClapConfig.from_pretrained(root,local_files_only=True)
55
+ self.clap=ClapAudioModelWithProjection(config.audio_config)
56
+ self.processor=ClapFeatureExtractor.from_pretrained(root,local_files_only=True)
57
+ self.weight=nn.Parameter(torch.zeros(384,512))
58
+ self.register_buffer('mean',torch.zeros(512))
59
+ self.register_buffer('scale',torch.ones(512))
60
+ if sum(p.numel() for p in self.clap.parameters())!=28190872:
61
+ raise ValueError('Pinned CLAP audio architecture mismatch')
62
+ if self.processor.sampling_rate!=48000 or self.processor.truncation!='rand_trunc' or self.processor.padding!='repeatpad':
63
+ raise ValueError('Pinned CLAP processor mismatch')
64
+
65
+ def encode_clap(self, audio48):
66
+ if not 1<=len(audio48)<=8:raise ValueError('One through eight native-rate streams required')
67
+ chunks=[];groups=[]
68
+ for wave in audio48:
69
+ start=len(chunks);chunks.extend(wave[a:b] for a,b in windows(len(wave)));groups.append((start,len(chunks)))
70
+ encoded=[];device=self.weight.device
71
+ for start in range(0,len(chunks),8):
72
+ inputs=self.processor(chunks[start:start+8],sampling_rate=48000,return_tensors='pt')
73
+ encoded.append(unit(self.clap(**{k:v.to(device) for k,v in inputs.items()}).audio_embeds.float()))
74
+ encoded=torch.cat(encoded)
75
+ return torch.stack([encoded[a] if b-a==1 else unit(encoded[a:b].mean(0)) for a,b in groups])
76
+
77
+ def forward(self, original_affine, clap):
78
+ if original_affine.shape!=(len(clap),384) or clap.shape[1:]!=(512,):
79
+ raise ValueError('Paired original affine384 and CLAP512 required')
80
+ if not torch.isfinite(self.mean).all() or not torch.isfinite(self.scale).all() or torch.any(self.scale<=0):
81
+ raise ValueError('Invalid frozen TRAIN statistics')
82
+ return unit(original_affine+F.linear((clap-self.mean)/self.scale,self.weight))
scores.json CHANGED
The diff for this file is too large to render. See raw diff
 
training/fsd50k-attributions.jsonl ADDED
The diff for this file is too large to render. See raw diff
 
vela_omni.py CHANGED
@@ -15,9 +15,14 @@ from omni_components.mini import MultiModalSentenceEmbedder
15
 
16
 
17
  class VelaOmni(nn.Module):
18
- def __init__(self, root, variant, legacy=False, audio_projection="identity", text_projection="identity", audio_pooling="mean", max_text_length=None, text_pooling="mean", inference_profile="full"):
19
  super().__init__()
20
  self.variant = variant
 
 
 
 
 
21
  if inference_profile not in ("full", "single_modality") or (inference_profile == "single_modality" and variant != "nano"):
22
  raise ValueError("Unsupported inference profile")
23
  self.inference_profile = inference_profile
@@ -40,8 +45,9 @@ class VelaOmni(nn.Module):
40
  self.max_text_length = max_text_length
41
  if variant == "nano":
42
  encoder_type = SingleModalityEmbedder if inference_profile == "single_modality" else MultimodalEmbedder
 
43
  self.model = encoder_type(Path(root) / "components", legacy=legacy,
44
- max_text_length=self.max_text_length, text_pooling=self.text_pooling)
45
  self.tokenizer = self.model.text_encoder.tokenizer
46
  if text_projection == "residual_rank64":
47
  self.model.text_encoder.projection = ResidualTextProjection(384)
@@ -59,6 +65,29 @@ class VelaOmni(nn.Module):
59
  else:
60
  raise ValueError("Unknown Vela Omni variant")
61
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
62
  @classmethod
63
  def from_pretrained(cls, repo_id, revision=None, device="cpu", dtype=torch.float32):
64
  root = Path(repo_id)
@@ -67,20 +96,14 @@ class VelaOmni(nn.Module):
67
  root = Path(snapshot_download(repo_id, revision=revision,
68
  allow_patterns=["config.json", "model.safetensors", "components/*"]))
69
  config = json.loads((root / "config.json").read_text())
70
- if config.get("format_version") not in (1, 2, 3, 4, 5):
71
- raise ValueError("Unsupported Vela Omni artifact format")
72
- if config.get("format_version") == 5 and (config.get("variant") != "nano" or config.get("inference_profile") != "single_modality"):
73
- raise ValueError("Format 5 requires Nano single-modality profile")
74
- if config.get("format_version") != 5 and config.get("inference_profile", "full") != "full":
75
- raise ValueError("Inference profile requires format 5")
76
- if config.get("image_pooling", "mean") != "mean":
77
- raise ValueError("Unsupported image pooling")
78
  instance = cls(root, config["variant"], audio_projection=config.get("audio_projection", "identity"),
79
  text_projection=config.get("text_projection", "identity"),
80
  audio_pooling=config.get("audio_pooling", "mean"),
81
  max_text_length=config.get("max_text_length"),
82
  text_pooling=config.get("text_pooling", "mean"),
83
- inference_profile=config.get("inference_profile", "full"))
 
84
  weights = load_file(str(root / "model.safetensors"), device="cpu")
85
  result = instance.model.load_state_dict(weights, strict=True, assign=True)
86
  if result.missing_keys or result.unexpected_keys:
@@ -122,6 +145,11 @@ class VelaOmni(nn.Module):
122
 
123
  @torch.inference_mode()
124
  def encode_audio(self, waveforms, sampling_rate=16000):
 
 
 
 
 
125
  if sampling_rate != 16000:
126
  raise ValueError("Resample audio to 16000 Hz before encoding")
127
  waveforms = [np.asarray(w, dtype=np.float32) for w in waveforms]
 
15
 
16
 
17
  class VelaOmni(nn.Module):
18
+ def __init__(self, root, variant, legacy=False, audio_projection="identity", text_projection="identity", audio_pooling="mean", max_text_length=None, text_pooling="mean", inference_profile="full", audio_backend="whisper"):
19
  super().__init__()
20
  self.variant = variant
21
+ if audio_backend not in ("whisper", "tiny_clap_residual_v1"):
22
+ raise ValueError("Unsupported audio backend")
23
+ if audio_backend != "whisper" and (variant != "nano" or inference_profile != "single_modality" or audio_projection != "linear" or audio_pooling != "mean"):
24
+ raise ValueError("Residual audio requires Nano single-modality mean/linear")
25
+ self.audio_backend = audio_backend
26
  if inference_profile not in ("full", "single_modality") or (inference_profile == "single_modality" and variant != "nano"):
27
  raise ValueError("Unsupported inference profile")
28
  self.inference_profile = inference_profile
 
45
  self.max_text_length = max_text_length
46
  if variant == "nano":
47
  encoder_type = SingleModalityEmbedder if inference_profile == "single_modality" else MultimodalEmbedder
48
+ extra = {"audio_backend": audio_backend} if inference_profile == "single_modality" else {}
49
  self.model = encoder_type(Path(root) / "components", legacy=legacy,
50
+ max_text_length=self.max_text_length, text_pooling=self.text_pooling, **extra)
51
  self.tokenizer = self.model.text_encoder.tokenizer
52
  if text_projection == "residual_rank64":
53
  self.model.text_encoder.projection = ResidualTextProjection(384)
 
65
  else:
66
  raise ValueError("Unknown Vela Omni variant")
67
 
68
+ @staticmethod
69
+ def validate_config(config):
70
+ version = config.get("format_version")
71
+ if version not in (1, 2, 3, 4, 5, 9):
72
+ raise ValueError("Unsupported Vela Omni artifact format")
73
+ backend = config.get("audio_backend", "whisper")
74
+ if version == 9:
75
+ required = {"variant": "nano", "inference_profile": "single_modality",
76
+ "audio_backend": "tiny_clap_residual_v1", "audio_projection": "linear",
77
+ "audio_pooling": "mean", "text_pooling": "cls", "max_text_length": 512,
78
+ "text_projection": "identity", "parameter_count": 163771288,
79
+ "residual_audio_readout": "tiny_full_mean_clap_endpoint_zero_linear"}
80
+ if any(config.get(k) != v for k, v in required.items()):
81
+ raise ValueError("Format 9 requires the pinned Nano residual configuration")
82
+ elif backend != "whisper":
83
+ raise ValueError("Residual audio semantics require format 9")
84
+ if version in (5, 9) and (config.get("variant") != "nano" or config.get("inference_profile") != "single_modality"):
85
+ raise ValueError("Format 5/9 requires Nano single-modality profile")
86
+ if version not in (5, 9) and config.get("inference_profile", "full") != "full":
87
+ raise ValueError("Inference profile requires format 5/9")
88
+ if config.get("image_pooling", "mean") != "mean":
89
+ raise ValueError("Unsupported image pooling")
90
+
91
  @classmethod
92
  def from_pretrained(cls, repo_id, revision=None, device="cpu", dtype=torch.float32):
93
  root = Path(repo_id)
 
96
  root = Path(snapshot_download(repo_id, revision=revision,
97
  allow_patterns=["config.json", "model.safetensors", "components/*"]))
98
  config = json.loads((root / "config.json").read_text())
99
+ cls.validate_config(config)
 
 
 
 
 
 
 
100
  instance = cls(root, config["variant"], audio_projection=config.get("audio_projection", "identity"),
101
  text_projection=config.get("text_projection", "identity"),
102
  audio_pooling=config.get("audio_pooling", "mean"),
103
  max_text_length=config.get("max_text_length"),
104
  text_pooling=config.get("text_pooling", "mean"),
105
+ inference_profile=config.get("inference_profile", "full"),
106
+ audio_backend=config.get("audio_backend", "whisper"))
107
  weights = load_file(str(root / "model.safetensors"), device="cpu")
108
  result = instance.model.load_state_dict(weights, strict=True, assign=True)
109
  if result.missing_keys or result.unexpected_keys:
 
145
 
146
  @torch.inference_mode()
147
  def encode_audio(self, waveforms, sampling_rate=16000):
148
+ if self.audio_backend == "tiny_clap_residual_v1":
149
+ if isinstance(sampling_rate, bool) or sampling_rate not in (16000, 44100, 48000):
150
+ raise ValueError("Supported original sampling rates are16000,44100,48000")
151
+ values = self.model._encode_audio_original(list(waveforms), sampling_rate)
152
+ return F.normalize(values.float(), dim=-1)
153
  if sampling_rate != 16000:
154
  raise ValueError("Resample audio to 16000 Hz before encoding")
155
  waveforms = [np.asarray(w, dtype=np.float32) for w in waveforms]