--- license: cc-by-nc-4.0 base_model: - jinaai/jina-embeddings-v5-omni-nano tags: - axera - ax650 - embeddings - retrieval - multimodal - image - audio - video --- # jina-embeddings-v5-omni-nano-retrieval on AXERA NPU This repository packages the retrieval-task AX650 deployment of `jinaai/jina-embeddings-v5-omni-nano`. The package contains compiled AX650 runtime artifacts and runs with `axllm serve` on AX650 using the expected bidirectional text-embedding behavior. It was rebuilt with the correct text RoPE base and passes the packaged text, image, audio, and frame-based video precision checks. ## Current Validation Status - Target platform: AX650 / NPU3 - Current service model id: `AXERA-TECH/jina-embeddings-v5-omni-nano-retrieval-AX650-P128-CTX2047` - Vision encoder static shape: `256x256` - Vision soft tokens per frame: `64` - Audio encoder static profiles: - short audio: `8s`, `16kHz`, mono, `800` mel frames, `200` soft tokens - long audio: `30s`, `16kHz`, mono, `3000` mel frames, `750` soft tokens - Audio request sequence length with the packaged `query` prompt: - `8s` audio: `200` soft tokens, about `220` total LLM prefill tokens - `30s` audio: `750` soft tokens, about `772` total LLM prefill tokens - Current LLM prefill build: - `prefill_len = 128` - warm-prefill groups = `128 / 256 / 384 / 512 / 640 / 768 / 896` - effective `prefill_max_token_num = 1024` Board cold-start measurements from this package with both `8s` and `30s` audio profiles loaded: | Item | Value | |---|---:| | `/v1/models` ready time | `25 s` | | CMM used before startup | `275004 KB` | | CMM used after startup | `2455548 KB` | | Cold-start CMM consumption | `2180544 KB` | | Process VmRSS after startup | `2011016 KB` | | Process PSS after startup | `2010169 KB` | | Process anonymous memory | `2006272 KB` | | Process VmSize after startup | `4191272 KB` | `CMM` means AXERA contiguous multimedia memory. `RSS` means resident set size in Linux process memory. `PSS` means proportional set size. ## Functional Status This package can: - start `./bin/axllm serve .` on AX650 - expose `/v1/models` and `/v1/embeddings` - run text, image, audio, and frame-based video embedding requests without depending on Hugging Face `safetensors` at runtime This package is retrieval-only. The upstream model family defines other task adapters, but this AX650 release contains the retrieval route and its cached retrieval reference cases only. The packaged runtime files are: - `llama_p128_l0_together.axmodel` ... `llama_p128_l11_together.axmodel` - `llama_post.axmodel` - `jina_v5_omni_nano_vision_256x256.axmodel` - `jina_v5_omni_nano_audio_8s.axmodel` - `jina_v5_omni_nano_audio_30s.axmodel` - `model.embed_tokens.weight.bfloat16.bin` - `jina_v5_omni_tokenizer.txt` - `jina_v5_omni_tokenizer/` - `bin/axllm` ## Performance This package exposes `/v1/embeddings` rather than token streaming, so chat-style `TTFT` is not the right primary metric. For an embedding model on AX650, the useful execution metric is: - text: `LLM prefill` only - image / video: `encoder + LLM prefill` - audio: `audio encoder + LLM prefill` There is no decode stage in the normal embedding path, so a `TTFT`-style number mostly collapses into prefill completion plus service overhead. If you measure only the HTTP response time of `axllm serve`, lightweight text and image requests can look artificially similar because fixed server overhead dominates them. All values below were re-measured on the validated AX650 / NPU3 board. Each number is the average of `3` repeated direct-runtime runs after model initialization. These measurements exclude network overhead and client-side preprocessing outside the packaged runtime path. | Scenario | Prompt | Input tokens | Encoder avg | LLM prefill avg | Total avg | |---|---|---:|---:|---:|---:| | Text document | `document` | `19` text tokens | `-` | `68.88 ms` | `68.88 ms` | | Text query | `query` | `12` text tokens | `-` | `61.59 ms` | `61.59 ms` | | Image | `query` | `64` soft tokens, `91` total sequence | `17.24 ms` | `65.66 ms` | `82.91 ms` | | Video (`3` frames) | `query` | `192` soft tokens, `219` total sequence | `52.49 ms` | `148.40 ms` | `200.89 ms` | Audio is different on this `3 GB` board: - the packaged `axllm serve` path is validated and returns correct audio embeddings - but the current Python direct-benchmark implementation was OOM-killed during this measurement pass - that failure is a board-side benchmark-method issue in Linux userspace memory, not a proof that the packaged AX650 runtime path cannot run audio embeddings - because of that, this README does not publish a misleading direct audio split-latency number yet ## Current Board Precision Current board-side `axllm serve` precision against the packaged Hugging Face reference embeddings is: | Modality | Case | Output shape | Cosine vs HF | |---|---|---:|---:| | Text | `embedding_doc` | `[1, 768]` | `0.999655` | | Text | `red_planet_query` | `[1, 768]` | `0.999539` | | Image | `vision_sample` | `[1, 768]` | `0.993549` | | Audio | `audio_test_chunk0_8s_wav` | `[1, 768]` | `0.989344` | | Audio | `audio_test_chunk0_30s_wav` | `[1, 768]` | `0.994962` | | Video | `video_visual_red_panda_openai_mp4` | `[1, 768]` | `0.997008` | Key observations: - The previous severe service mismatch is fixed. - The previous `embedding_doc = 0.9604` result was traced to an LLM build config issue, not to the OpenAI-compatible service wrapper. - `jina-embeddings-v5-omni-nano` uses `text_config.rope_parameters.rope_theta = 1000000.0`; the earlier `llama` build path silently fell back to `10000.0`. - The current LLM AXModels were rebuilt after flattening `rope_theta` into `text_config.rope_theta` for `pulsar2 llm_build`. - With the corrected RoPE base, all non-audio-short packaged precision cases are in the `0.9935 ~ 0.9997` cosine range. - The `8s` audio profile is usable but lower than the `30s` profile in this P128 build: `0.989344` cosine for the packaged `audio_test_chunk0_8s_wav` case. The previous `5s` audio profile produced only about `0.94` cosine after LLM prefill, so it is not used as a default packaged precision case. ## Board Precision Validation Flow The board-side comparison does not run the original Hugging Face model. The intended flow is: - Run the original `jinaai/jina-embeddings-v5-omni-nano` model once on the server for each fixed test case. - Save the resulting HF embedding as `torch_embedding.npy` under the matching `python/testdata/service_cases//` directory. - Package the same case metadata and assets used by the HF run. - On AX650, start `axllm serve .`, call `/v1/embeddings`, and compare the returned AXERA embedding with the packaged `torch_embedding.npy`. This guarantees that AX650 validation uses the same inputs as the server-side HF reference, while the board only needs the packaged `.axmodel`, `.bin`, tokenizer, assets, and cached reference `.npy` files. The board package does not require upstream `.safetensors` files. Run the packaged comparison on the board with: ```bash python3 python/compare_openai_api_vs_hf_multimodal.py ``` The script fails if any packaged case is missing its cached `torch_embedding.npy` reference. ## Conversion Note If you rebuild the LLM AXModels from the original Hugging Face checkpoint, check the RoPE config before running `pulsar2 llm_build`. The upstream nano config stores the text RoPE base as: ```json "text_config": { "rope_parameters": { "rope_theta": 1000000.0 } } ``` The current AXERA `llama` `llm_build` route does not read `text_config.rope_parameters.rope_theta`. It reads the flattened field `text_config.rope_theta`. Before compilation, make sure the `llm_build_input/config.json` contains: ```json "text_config": { "rope_theta": 1000000.0 } ``` The packaged conversion script handles this automatically. If you prepare `llm_build_input/config.json` manually, copy the value yourself and confirm the build log prints `rope_theta=1000000.0`. If the build silently uses the default `10000.0`, the generated LLM AXModels will have noticeably worse embedding precision. ## Runtime Compatibility Notes The following compatibility details are required to reproduce the validated precision: - `jina-embeddings-v5-omni-nano` uses a bidirectional EuroBERT / `LlamaModel` text tower, but the service embedding path was still building a decoder-style causal prefill mask. - The embedding path in `axllm` must use bidirectional masking through `prefill_mask_mode`. - The service tokenizer path was also missing the final end-of-text token used by the Python reference route. - For `Qwen3Omni`, the correct end-of-text token id is not the older hard-coded `151643`; it must be resolved from the loaded tokenizer and is `128001` for this Jina nano package. - LLM precision depends on the `llama` `llm_build` path receiving the flattened `text_config.rope_theta` value. The nano text model requires RoPE theta `1000000.0`, and the corrected package was rebuilt with that value. With these settings, the service-side outputs match the direct board-side `axmodel` route. ## Packaged Test Cases The package includes cached Hugging Face reference embeddings under: ```text python/testdata/service_cases/ ``` Current packaged cases: - `embedding_doc` - `red_planet_query` - `vision_sample` - `audio_test_chunk0_8s_wav` - `audio_test_chunk0_30s_wav` - `video_visual_red_panda_openai_mp4` All cached reference embeddings use the upstream retrieval route and output shape `[1, 768]`. The packaged `meta.json` files define the exact text, prompt name, media asset path, and soft-token count used for both the server-side HF reference and the board-side API request. ## Known Limitation - Re-measure direct audio split latency if a low-memory benchmark implementation is needed; the packaged service audio path is already validated. - Audio encoder loading is controlled by `filename_audio_encoder_axmodel_short` and `filename_audio_encoder_axmodel_long` in `config.json`. If only one field is configured, only that audio profile is loaded; if both are configured, the runtime selects the short profile for short audio and falls back to the long profile for longer audio.