yongqiang
Add 8s audio profile for Jina nano retrieval
eec42a1
|
Raw
History Blame
10.2 kB
---
license: cc-by-nc-4.0
base_model:
- jinaai/jina-embeddings-v5-omni-nano
tags:
- axera
- ax650
- embeddings
- retrieval
- multimodal
- image
- audio
- video
---
# jina-embeddings-v5-omni-nano-retrieval on AXERA NPU
This repository packages the retrieval-task AX650 deployment of `jinaai/jina-embeddings-v5-omni-nano`.
The package contains compiled AX650 runtime artifacts and runs with `axllm serve` on AX650 using the expected bidirectional text-embedding behavior.
It was rebuilt with the correct text RoPE base and passes the packaged text, image, audio, and frame-based video precision checks.
## Current Validation Status
- Target platform: AX650 / NPU3
- Current service model id: `AXERA-TECH/jina-embeddings-v5-omni-nano-retrieval-AX650-P128-CTX2047`
- Vision encoder static shape: `256x256`
- Vision soft tokens per frame: `64`
- Audio encoder static profiles:
- short audio: `8s`, `16kHz`, mono, `800` mel frames, `200` soft tokens
- long audio: `30s`, `16kHz`, mono, `3000` mel frames, `750` soft tokens
- Audio request sequence length with the packaged `query` prompt:
- `8s` audio: `200` soft tokens, about `220` total LLM prefill tokens
- `30s` audio: `750` soft tokens, about `772` total LLM prefill tokens
- Current LLM prefill build:
- `prefill_len = 128`
- warm-prefill groups = `128 / 256 / 384 / 512 / 640 / 768 / 896`
- effective `prefill_max_token_num = 1024`
Board cold-start measurements from this package with both `8s` and `30s` audio profiles loaded:
| Item | Value |
|---|---:|
| `/v1/models` ready time | `25 s` |
| CMM used before startup | `275004 KB` |
| CMM used after startup | `2455548 KB` |
| Cold-start CMM consumption | `2180544 KB` |
| Process VmRSS after startup | `2011016 KB` |
| Process PSS after startup | `2010169 KB` |
| Process anonymous memory | `2006272 KB` |
| Process VmSize after startup | `4191272 KB` |
`CMM` means AXERA contiguous multimedia memory. `RSS` means resident set size in Linux process memory. `PSS` means proportional set size.
## Functional Status
This package can:
- start `./bin/axllm serve .` on AX650
- expose `/v1/models` and `/v1/embeddings`
- run text, image, audio, and frame-based video embedding requests without depending on Hugging Face `safetensors` at runtime
This package is retrieval-only. The upstream model family defines other task adapters, but this AX650 release contains the retrieval route and its cached retrieval reference cases only.
The packaged runtime files are:
- `llama_p128_l0_together.axmodel` ... `llama_p128_l11_together.axmodel`
- `llama_post.axmodel`
- `jina_v5_omni_nano_vision_256x256.axmodel`
- `jina_v5_omni_nano_audio_8s.axmodel`
- `jina_v5_omni_nano_audio_30s.axmodel`
- `model.embed_tokens.weight.bfloat16.bin`
- `jina_v5_omni_tokenizer.txt`
- `jina_v5_omni_tokenizer/`
- `bin/axllm`
## Performance
This package exposes `/v1/embeddings` rather than token streaming, so chat-style `TTFT` is not the right primary metric.
For an embedding model on AX650, the useful execution metric is:
- text: `LLM prefill` only
- image / video: `encoder + LLM prefill`
- audio: `audio encoder + LLM prefill`
There is no decode stage in the normal embedding path, so a `TTFT`-style number mostly collapses into prefill completion plus service overhead.
If you measure only the HTTP response time of `axllm serve`, lightweight text and image requests can look artificially similar because fixed server overhead dominates them.
All values below were re-measured on the validated AX650 / NPU3 board.
Each number is the average of `3` repeated direct-runtime runs after model initialization.
These measurements exclude network overhead and client-side preprocessing outside the packaged runtime path.
| Scenario | Prompt | Input tokens | Encoder avg | LLM prefill avg | Total avg |
|---|---|---:|---:|---:|---:|
| Text document | `document` | `19` text tokens | `-` | `68.88 ms` | `68.88 ms` |
| Text query | `query` | `12` text tokens | `-` | `61.59 ms` | `61.59 ms` |
| Image | `query` | `64` soft tokens, `91` total sequence | `17.24 ms` | `65.66 ms` | `82.91 ms` |
| Video (`3` frames) | `query` | `192` soft tokens, `219` total sequence | `52.49 ms` | `148.40 ms` | `200.89 ms` |
Audio is different on this `3 GB` board:
- the packaged `axllm serve` path is validated and returns correct audio embeddings
- but the current Python direct-benchmark implementation was OOM-killed during this measurement pass
- that failure is a board-side benchmark-method issue in Linux userspace memory, not a proof that the packaged AX650 runtime path cannot run audio embeddings
- because of that, this README does not publish a misleading direct audio split-latency number yet
## Current Board Precision
Current board-side `axllm serve` precision against the packaged Hugging Face reference embeddings is:
| Modality | Case | Output shape | Cosine vs HF |
|---|---|---:|---:|
| Text | `embedding_doc` | `[1, 768]` | `0.999655` |
| Text | `red_planet_query` | `[1, 768]` | `0.999539` |
| Image | `vision_sample` | `[1, 768]` | `0.993549` |
| Audio | `audio_test_chunk0_8s_wav` | `[1, 768]` | `0.989344` |
| Audio | `audio_test_chunk0_30s_wav` | `[1, 768]` | `0.994962` |
| Video | `video_visual_red_panda_openai_mp4` | `[1, 768]` | `0.997008` |
Key observations:
- The previous severe service mismatch is fixed.
- The previous `embedding_doc = 0.9604` result was traced to an LLM build config issue, not to the OpenAI-compatible service wrapper.
- `jina-embeddings-v5-omni-nano` uses `text_config.rope_parameters.rope_theta = 1000000.0`; the earlier `llama` build path silently fell back to `10000.0`.
- The current LLM AXModels were rebuilt after flattening `rope_theta` into `text_config.rope_theta` for `pulsar2 llm_build`.
- With the corrected RoPE base, all non-audio-short packaged precision cases are in the `0.9935 ~ 0.9997` cosine range.
- The `8s` audio profile is usable but lower than the `30s` profile in this P128 build: `0.989344` cosine for the packaged `audio_test_chunk0_8s_wav` case.
The previous `5s` audio profile produced only about `0.94` cosine after LLM prefill, so it is not used as a default packaged precision case.
## Board Precision Validation Flow
The board-side comparison does not run the original Hugging Face model.
The intended flow is:
- Run the original `jinaai/jina-embeddings-v5-omni-nano` model once on the server for each fixed test case.
- Save the resulting HF embedding as `torch_embedding.npy` under the matching `python/testdata/service_cases/<case>/` directory.
- Package the same case metadata and assets used by the HF run.
- On AX650, start `axllm serve .`, call `/v1/embeddings`, and compare the returned AXERA embedding with the packaged `torch_embedding.npy`.
This guarantees that AX650 validation uses the same inputs as the server-side HF reference, while the board only needs the packaged `.axmodel`, `.bin`, tokenizer, assets, and cached reference `.npy` files.
The board package does not require upstream `.safetensors` files.
Run the packaged comparison on the board with:
```bash
python3 python/compare_openai_api_vs_hf_multimodal.py
```
The script fails if any packaged case is missing its cached `torch_embedding.npy` reference.
## Conversion Note
If you rebuild the LLM AXModels from the original Hugging Face checkpoint, check the RoPE config before running `pulsar2 llm_build`.
The upstream nano config stores the text RoPE base as:
```json
"text_config": {
"rope_parameters": {
"rope_theta": 1000000.0
}
}
```
The current AXERA `llama` `llm_build` route does not read `text_config.rope_parameters.rope_theta`.
It reads the flattened field `text_config.rope_theta`.
Before compilation, make sure the `llm_build_input/config.json` contains:
```json
"text_config": {
"rope_theta": 1000000.0
}
```
The packaged conversion script handles this automatically.
If you prepare `llm_build_input/config.json` manually, copy the value yourself and confirm the build log prints `rope_theta=1000000.0`.
If the build silently uses the default `10000.0`, the generated LLM AXModels will have noticeably worse embedding precision.
## Runtime Compatibility Notes
The following compatibility details are required to reproduce the validated precision:
- `jina-embeddings-v5-omni-nano` uses a bidirectional EuroBERT / `LlamaModel` text tower, but the service embedding path was still building a decoder-style causal prefill mask.
- The embedding path in `axllm` must use bidirectional masking through `prefill_mask_mode`.
- The service tokenizer path was also missing the final end-of-text token used by the Python reference route.
- For `Qwen3Omni`, the correct end-of-text token id is not the older hard-coded `151643`; it must be resolved from the loaded tokenizer and is `128001` for this Jina nano package.
- LLM precision depends on the `llama` `llm_build` path receiving the flattened `text_config.rope_theta` value.
The nano text model requires RoPE theta `1000000.0`, and the corrected package was rebuilt with that value.
With these settings, the service-side outputs match the direct board-side `axmodel` route.
## Packaged Test Cases
The package includes cached Hugging Face reference embeddings under:
```text
python/testdata/service_cases/
```
Current packaged cases:
- `embedding_doc`
- `red_planet_query`
- `vision_sample`
- `audio_test_chunk0_8s_wav`
- `audio_test_chunk0_30s_wav`
- `video_visual_red_panda_openai_mp4`
All cached reference embeddings use the upstream retrieval route and output shape `[1, 768]`.
The packaged `meta.json` files define the exact text, prompt name, media asset path, and soft-token count used for both the server-side HF reference and the board-side API request.
## Known Limitation
- Re-measure direct audio split latency if a low-memory benchmark implementation is needed; the packaged service audio path is already validated.
- Audio encoder loading is controlled by `filename_audio_encoder_axmodel_short` and `filename_audio_encoder_axmodel_long` in `config.json`.
If only one field is configured, only that audio profile is loaded; if both are configured, the runtime selects the short profile for short audio and falls back to the long profile for longer audio.