File size: 10,245 Bytes
5099405
df5639f
 
 
 
 
 
 
a2cbb72
df5639f
 
 
 
5099405
df5639f
a2cbb72
df5639f
a2cbb72
df5639f
a2cbb72
 
df5639f
a2cbb72
df5639f
 
a2cbb72
 
 
eec42a1
 
 
 
 
 
a2cbb72
 
 
 
df5639f
eec42a1
df5639f
a2cbb72
 
eec42a1
 
 
 
 
 
 
 
 
 
df5639f
a2cbb72
df5639f
a2cbb72
df5639f
a2cbb72
 
 
df5639f
a2cbb72
df5639f
a2cbb72
df5639f
a2cbb72
 
 
eec42a1
a2cbb72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
eec42a1
a2cbb72
eec42a1
a2cbb72
 
 
 
 
 
 
eec42a1
 
 
a2cbb72
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
df5639f
 
a2cbb72
df5639f
 
a2cbb72
 
 
 
 
eec42a1
a2cbb72
 
 
 
 
 
 
 
 
eec42a1
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
---
license: cc-by-nc-4.0
base_model:
  - jinaai/jina-embeddings-v5-omni-nano
tags:
  - axera
  - ax650
  - embeddings
  - retrieval
  - multimodal
  - image
  - audio
  - video
---

# jina-embeddings-v5-omni-nano-retrieval on AXERA NPU

This repository packages the retrieval-task AX650 deployment of `jinaai/jina-embeddings-v5-omni-nano`.

The package contains compiled AX650 runtime artifacts and runs with `axllm serve` on AX650 using the expected bidirectional text-embedding behavior.
It was rebuilt with the correct text RoPE base and passes the packaged text, image, audio, and frame-based video precision checks.

## Current Validation Status

- Target platform: AX650 / NPU3
- Current service model id: `AXERA-TECH/jina-embeddings-v5-omni-nano-retrieval-AX650-P128-CTX2047`
- Vision encoder static shape: `256x256`
- Vision soft tokens per frame: `64`
- Audio encoder static profiles:
  - short audio: `8s`, `16kHz`, mono, `800` mel frames, `200` soft tokens
  - long audio: `30s`, `16kHz`, mono, `3000` mel frames, `750` soft tokens
- Audio request sequence length with the packaged `query` prompt:
  - `8s` audio: `200` soft tokens, about `220` total LLM prefill tokens
  - `30s` audio: `750` soft tokens, about `772` total LLM prefill tokens
- Current LLM prefill build:
  - `prefill_len = 128`
  - warm-prefill groups = `128 / 256 / 384 / 512 / 640 / 768 / 896`
  - effective `prefill_max_token_num = 1024`

Board cold-start measurements from this package with both `8s` and `30s` audio profiles loaded:

| Item | Value |
|---|---:|
| `/v1/models` ready time | `25 s` |
| CMM used before startup | `275004 KB` |
| CMM used after startup | `2455548 KB` |
| Cold-start CMM consumption | `2180544 KB` |
| Process VmRSS after startup | `2011016 KB` |
| Process PSS after startup | `2010169 KB` |
| Process anonymous memory | `2006272 KB` |
| Process VmSize after startup | `4191272 KB` |

`CMM` means AXERA contiguous multimedia memory. `RSS` means resident set size in Linux process memory. `PSS` means proportional set size.

## Functional Status

This package can:

- start `./bin/axllm serve .` on AX650
- expose `/v1/models` and `/v1/embeddings`
- run text, image, audio, and frame-based video embedding requests without depending on Hugging Face `safetensors` at runtime

This package is retrieval-only. The upstream model family defines other task adapters, but this AX650 release contains the retrieval route and its cached retrieval reference cases only.

The packaged runtime files are:

- `llama_p128_l0_together.axmodel` ... `llama_p128_l11_together.axmodel`
- `llama_post.axmodel`
- `jina_v5_omni_nano_vision_256x256.axmodel`
- `jina_v5_omni_nano_audio_8s.axmodel`
- `jina_v5_omni_nano_audio_30s.axmodel`
- `model.embed_tokens.weight.bfloat16.bin`
- `jina_v5_omni_tokenizer.txt`
- `jina_v5_omni_tokenizer/`
- `bin/axllm`

## Performance

This package exposes `/v1/embeddings` rather than token streaming, so chat-style `TTFT` is not the right primary metric.
For an embedding model on AX650, the useful execution metric is:

- text: `LLM prefill` only
- image / video: `encoder + LLM prefill`
- audio: `audio encoder + LLM prefill`

There is no decode stage in the normal embedding path, so a `TTFT`-style number mostly collapses into prefill completion plus service overhead.
If you measure only the HTTP response time of `axllm serve`, lightweight text and image requests can look artificially similar because fixed server overhead dominates them.

All values below were re-measured on the validated AX650 / NPU3 board.
Each number is the average of `3` repeated direct-runtime runs after model initialization.
These measurements exclude network overhead and client-side preprocessing outside the packaged runtime path.

| Scenario | Prompt | Input tokens | Encoder avg | LLM prefill avg | Total avg |
|---|---|---:|---:|---:|---:|
| Text document | `document` | `19` text tokens | `-` | `68.88 ms` | `68.88 ms` |
| Text query | `query` | `12` text tokens | `-` | `61.59 ms` | `61.59 ms` |
| Image | `query` | `64` soft tokens, `91` total sequence | `17.24 ms` | `65.66 ms` | `82.91 ms` |
| Video (`3` frames) | `query` | `192` soft tokens, `219` total sequence | `52.49 ms` | `148.40 ms` | `200.89 ms` |

Audio is different on this `3 GB` board:

- the packaged `axllm serve` path is validated and returns correct audio embeddings
- but the current Python direct-benchmark implementation was OOM-killed during this measurement pass
- that failure is a board-side benchmark-method issue in Linux userspace memory, not a proof that the packaged AX650 runtime path cannot run audio embeddings
- because of that, this README does not publish a misleading direct audio split-latency number yet

## Current Board Precision

Current board-side `axllm serve` precision against the packaged Hugging Face reference embeddings is:

| Modality | Case | Output shape | Cosine vs HF |
|---|---|---:|---:|
| Text | `embedding_doc` | `[1, 768]` | `0.999655` |
| Text | `red_planet_query` | `[1, 768]` | `0.999539` |
| Image | `vision_sample` | `[1, 768]` | `0.993549` |
| Audio | `audio_test_chunk0_8s_wav` | `[1, 768]` | `0.989344` |
| Audio | `audio_test_chunk0_30s_wav` | `[1, 768]` | `0.994962` |
| Video | `video_visual_red_panda_openai_mp4` | `[1, 768]` | `0.997008` |

Key observations:

- The previous severe service mismatch is fixed.
- The previous `embedding_doc = 0.9604` result was traced to an LLM build config issue, not to the OpenAI-compatible service wrapper.
- `jina-embeddings-v5-omni-nano` uses `text_config.rope_parameters.rope_theta = 1000000.0`; the earlier `llama` build path silently fell back to `10000.0`.
- The current LLM AXModels were rebuilt after flattening `rope_theta` into `text_config.rope_theta` for `pulsar2 llm_build`.
- With the corrected RoPE base, all non-audio-short packaged precision cases are in the `0.9935 ~ 0.9997` cosine range.
- The `8s` audio profile is usable but lower than the `30s` profile in this P128 build: `0.989344` cosine for the packaged `audio_test_chunk0_8s_wav` case.
  The previous `5s` audio profile produced only about `0.94` cosine after LLM prefill, so it is not used as a default packaged precision case.

## Board Precision Validation Flow

The board-side comparison does not run the original Hugging Face model.
The intended flow is:

- Run the original `jinaai/jina-embeddings-v5-omni-nano` model once on the server for each fixed test case.
- Save the resulting HF embedding as `torch_embedding.npy` under the matching `python/testdata/service_cases/<case>/` directory.
- Package the same case metadata and assets used by the HF run.
- On AX650, start `axllm serve .`, call `/v1/embeddings`, and compare the returned AXERA embedding with the packaged `torch_embedding.npy`.

This guarantees that AX650 validation uses the same inputs as the server-side HF reference, while the board only needs the packaged `.axmodel`, `.bin`, tokenizer, assets, and cached reference `.npy` files.
The board package does not require upstream `.safetensors` files.

Run the packaged comparison on the board with:

```bash
python3 python/compare_openai_api_vs_hf_multimodal.py
```

The script fails if any packaged case is missing its cached `torch_embedding.npy` reference.

## Conversion Note

If you rebuild the LLM AXModels from the original Hugging Face checkpoint, check the RoPE config before running `pulsar2 llm_build`.
The upstream nano config stores the text RoPE base as:

```json
"text_config": {
  "rope_parameters": {
    "rope_theta": 1000000.0
  }
}
```

The current AXERA `llama` `llm_build` route does not read `text_config.rope_parameters.rope_theta`.
It reads the flattened field `text_config.rope_theta`.
Before compilation, make sure the `llm_build_input/config.json` contains:

```json
"text_config": {
  "rope_theta": 1000000.0
}
```

The packaged conversion script handles this automatically.
If you prepare `llm_build_input/config.json` manually, copy the value yourself and confirm the build log prints `rope_theta=1000000.0`.
If the build silently uses the default `10000.0`, the generated LLM AXModels will have noticeably worse embedding precision.

## Runtime Compatibility Notes

The following compatibility details are required to reproduce the validated precision:

- `jina-embeddings-v5-omni-nano` uses a bidirectional EuroBERT / `LlamaModel` text tower, but the service embedding path was still building a decoder-style causal prefill mask.
- The embedding path in `axllm` must use bidirectional masking through `prefill_mask_mode`.
- The service tokenizer path was also missing the final end-of-text token used by the Python reference route.
- For `Qwen3Omni`, the correct end-of-text token id is not the older hard-coded `151643`; it must be resolved from the loaded tokenizer and is `128001` for this Jina nano package.
- LLM precision depends on the `llama` `llm_build` path receiving the flattened `text_config.rope_theta` value.
  The nano text model requires RoPE theta `1000000.0`, and the corrected package was rebuilt with that value.

With these settings, the service-side outputs match the direct board-side `axmodel` route.

## Packaged Test Cases

The package includes cached Hugging Face reference embeddings under:

```text
python/testdata/service_cases/
```

Current packaged cases:

- `embedding_doc`
- `red_planet_query`
- `vision_sample`
- `audio_test_chunk0_8s_wav`
- `audio_test_chunk0_30s_wav`
- `video_visual_red_panda_openai_mp4`

All cached reference embeddings use the upstream retrieval route and output shape `[1, 768]`.
The packaged `meta.json` files define the exact text, prompt name, media asset path, and soft-token count used for both the server-side HF reference and the board-side API request.

## Known Limitation

- Re-measure direct audio split latency if a low-memory benchmark implementation is needed; the packaged service audio path is already validated.
- Audio encoder loading is controlled by `filename_audio_encoder_axmodel_short` and `filename_audio_encoder_axmodel_long` in `config.json`.
  If only one field is configured, only that audio profile is loaded; if both are configured, the runtime selects the short profile for short audio and falls back to the long profile for longer audio.