yongqiang
Add 8s audio profile for Jina nano retrieval
eec42a1
|
Raw
History Blame
10.2 kB
metadata
license: cc-by-nc-4.0
base_model:
  - jinaai/jina-embeddings-v5-omni-nano
tags:
  - axera
  - ax650
  - embeddings
  - retrieval
  - multimodal
  - image
  - audio
  - video

jina-embeddings-v5-omni-nano-retrieval on AXERA NPU

This repository packages the retrieval-task AX650 deployment of jinaai/jina-embeddings-v5-omni-nano.

The package contains compiled AX650 runtime artifacts and runs with axllm serve on AX650 using the expected bidirectional text-embedding behavior. It was rebuilt with the correct text RoPE base and passes the packaged text, image, audio, and frame-based video precision checks.

Current Validation Status

  • Target platform: AX650 / NPU3
  • Current service model id: AXERA-TECH/jina-embeddings-v5-omni-nano-retrieval-AX650-P128-CTX2047
  • Vision encoder static shape: 256x256
  • Vision soft tokens per frame: 64
  • Audio encoder static profiles:
    • short audio: 8s, 16kHz, mono, 800 mel frames, 200 soft tokens
    • long audio: 30s, 16kHz, mono, 3000 mel frames, 750 soft tokens
  • Audio request sequence length with the packaged query prompt:
    • 8s audio: 200 soft tokens, about 220 total LLM prefill tokens
    • 30s audio: 750 soft tokens, about 772 total LLM prefill tokens
  • Current LLM prefill build:
    • prefill_len = 128
    • warm-prefill groups = 128 / 256 / 384 / 512 / 640 / 768 / 896
    • effective prefill_max_token_num = 1024

Board cold-start measurements from this package with both 8s and 30s audio profiles loaded:

Item Value
/v1/models ready time 25 s
CMM used before startup 275004 KB
CMM used after startup 2455548 KB
Cold-start CMM consumption 2180544 KB
Process VmRSS after startup 2011016 KB
Process PSS after startup 2010169 KB
Process anonymous memory 2006272 KB
Process VmSize after startup 4191272 KB

CMM means AXERA contiguous multimedia memory. RSS means resident set size in Linux process memory. PSS means proportional set size.

Functional Status

This package can:

  • start ./bin/axllm serve . on AX650
  • expose /v1/models and /v1/embeddings
  • run text, image, audio, and frame-based video embedding requests without depending on Hugging Face safetensors at runtime

This package is retrieval-only. The upstream model family defines other task adapters, but this AX650 release contains the retrieval route and its cached retrieval reference cases only.

The packaged runtime files are:

  • llama_p128_l0_together.axmodel ... llama_p128_l11_together.axmodel
  • llama_post.axmodel
  • jina_v5_omni_nano_vision_256x256.axmodel
  • jina_v5_omni_nano_audio_8s.axmodel
  • jina_v5_omni_nano_audio_30s.axmodel
  • model.embed_tokens.weight.bfloat16.bin
  • jina_v5_omni_tokenizer.txt
  • jina_v5_omni_tokenizer/
  • bin/axllm

Performance

This package exposes /v1/embeddings rather than token streaming, so chat-style TTFT is not the right primary metric. For an embedding model on AX650, the useful execution metric is:

  • text: LLM prefill only
  • image / video: encoder + LLM prefill
  • audio: audio encoder + LLM prefill

There is no decode stage in the normal embedding path, so a TTFT-style number mostly collapses into prefill completion plus service overhead. If you measure only the HTTP response time of axllm serve, lightweight text and image requests can look artificially similar because fixed server overhead dominates them.

All values below were re-measured on the validated AX650 / NPU3 board. Each number is the average of 3 repeated direct-runtime runs after model initialization. These measurements exclude network overhead and client-side preprocessing outside the packaged runtime path.

Scenario Prompt Input tokens Encoder avg LLM prefill avg Total avg
Text document document 19 text tokens - 68.88 ms 68.88 ms
Text query query 12 text tokens - 61.59 ms 61.59 ms
Image query 64 soft tokens, 91 total sequence 17.24 ms 65.66 ms 82.91 ms
Video (3 frames) query 192 soft tokens, 219 total sequence 52.49 ms 148.40 ms 200.89 ms

Audio is different on this 3 GB board:

  • the packaged axllm serve path is validated and returns correct audio embeddings
  • but the current Python direct-benchmark implementation was OOM-killed during this measurement pass
  • that failure is a board-side benchmark-method issue in Linux userspace memory, not a proof that the packaged AX650 runtime path cannot run audio embeddings
  • because of that, this README does not publish a misleading direct audio split-latency number yet

Current Board Precision

Current board-side axllm serve precision against the packaged Hugging Face reference embeddings is:

Modality Case Output shape Cosine vs HF
Text embedding_doc [1, 768] 0.999655
Text red_planet_query [1, 768] 0.999539
Image vision_sample [1, 768] 0.993549
Audio audio_test_chunk0_8s_wav [1, 768] 0.989344
Audio audio_test_chunk0_30s_wav [1, 768] 0.994962
Video video_visual_red_panda_openai_mp4 [1, 768] 0.997008

Key observations:

  • The previous severe service mismatch is fixed.
  • The previous embedding_doc = 0.9604 result was traced to an LLM build config issue, not to the OpenAI-compatible service wrapper.
  • jina-embeddings-v5-omni-nano uses text_config.rope_parameters.rope_theta = 1000000.0; the earlier llama build path silently fell back to 10000.0.
  • The current LLM AXModels were rebuilt after flattening rope_theta into text_config.rope_theta for pulsar2 llm_build.
  • With the corrected RoPE base, all non-audio-short packaged precision cases are in the 0.9935 ~ 0.9997 cosine range.
  • The 8s audio profile is usable but lower than the 30s profile in this P128 build: 0.989344 cosine for the packaged audio_test_chunk0_8s_wav case. The previous 5s audio profile produced only about 0.94 cosine after LLM prefill, so it is not used as a default packaged precision case.

Board Precision Validation Flow

The board-side comparison does not run the original Hugging Face model. The intended flow is:

  • Run the original jinaai/jina-embeddings-v5-omni-nano model once on the server for each fixed test case.
  • Save the resulting HF embedding as torch_embedding.npy under the matching python/testdata/service_cases/<case>/ directory.
  • Package the same case metadata and assets used by the HF run.
  • On AX650, start axllm serve ., call /v1/embeddings, and compare the returned AXERA embedding with the packaged torch_embedding.npy.

This guarantees that AX650 validation uses the same inputs as the server-side HF reference, while the board only needs the packaged .axmodel, .bin, tokenizer, assets, and cached reference .npy files. The board package does not require upstream .safetensors files.

Run the packaged comparison on the board with:

python3 python/compare_openai_api_vs_hf_multimodal.py

The script fails if any packaged case is missing its cached torch_embedding.npy reference.

Conversion Note

If you rebuild the LLM AXModels from the original Hugging Face checkpoint, check the RoPE config before running pulsar2 llm_build. The upstream nano config stores the text RoPE base as:

"text_config": {
  "rope_parameters": {
    "rope_theta": 1000000.0
  }
}

The current AXERA llama llm_build route does not read text_config.rope_parameters.rope_theta. It reads the flattened field text_config.rope_theta. Before compilation, make sure the llm_build_input/config.json contains:

"text_config": {
  "rope_theta": 1000000.0
}

The packaged conversion script handles this automatically. If you prepare llm_build_input/config.json manually, copy the value yourself and confirm the build log prints rope_theta=1000000.0. If the build silently uses the default 10000.0, the generated LLM AXModels will have noticeably worse embedding precision.

Runtime Compatibility Notes

The following compatibility details are required to reproduce the validated precision:

  • jina-embeddings-v5-omni-nano uses a bidirectional EuroBERT / LlamaModel text tower, but the service embedding path was still building a decoder-style causal prefill mask.
  • The embedding path in axllm must use bidirectional masking through prefill_mask_mode.
  • The service tokenizer path was also missing the final end-of-text token used by the Python reference route.
  • For Qwen3Omni, the correct end-of-text token id is not the older hard-coded 151643; it must be resolved from the loaded tokenizer and is 128001 for this Jina nano package.
  • LLM precision depends on the llama llm_build path receiving the flattened text_config.rope_theta value. The nano text model requires RoPE theta 1000000.0, and the corrected package was rebuilt with that value.

With these settings, the service-side outputs match the direct board-side axmodel route.

Packaged Test Cases

The package includes cached Hugging Face reference embeddings under:

python/testdata/service_cases/

Current packaged cases:

  • embedding_doc
  • red_planet_query
  • vision_sample
  • audio_test_chunk0_8s_wav
  • audio_test_chunk0_30s_wav
  • video_visual_red_panda_openai_mp4

All cached reference embeddings use the upstream retrieval route and output shape [1, 768]. The packaged meta.json files define the exact text, prompt name, media asset path, and soft-token count used for both the server-side HF reference and the board-side API request.

Known Limitation

  • Re-measure direct audio split latency if a low-memory benchmark implementation is needed; the packaged service audio path is already validated.
  • Audio encoder loading is controlled by filename_audio_encoder_axmodel_short and filename_audio_encoder_axmodel_long in config.json. If only one field is configured, only that audio profile is loaded; if both are configured, the runtime selects the short profile for short audio and falls back to the long profile for longer audio.