Instructions to use n8dgr8/ov_intent_analysis_sft-RKLLM-RK3588 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RKLLM
How to use n8dgr8/ov_intent_analysis_sft-RKLLM-RK3588 with RKLLM:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
ov_intent_analysis_sft β RKLLM (RK3588) conversion
Pre-converted W8A8 RKLLM runtime file of guoxuter/ov_intent_analysis_sft (the OpenViking retrieval intent-analysis / query-planner model, fine-tuned from Qwen3.5-0.8B).
Run the same tuned query planner on your RK3588 NPU β no x86 conversion rig required.
Why this exists
OpenViking's docs recommend this model for the query_planner slot: it decides whether a search needs context retrieval, skips chitchat (no queries β no token spend), and emits structured skill / resource / memory queries. The stock model ships as HF safetensors; to run it on the RK3588 NPU you need a .rkllm file, which can only be produced by rkllm-toolkit (x86_64-only). This repo is that conversion, done once, so RK3588 owners can skip the whole rig.
File
| File | Size | Spec |
|---|---|---|
ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm |
1.3 GB | W8A8, RK3588, 3 NPU cores, max_context 4096 |
Runtime requirements: librkllmrt.so 1.3.0 (rkllm-toolkit 1.3.0 generation). Verified on kernel 6.1 vendor with rknpu driver 0.9.8.
How to serve it
Any RKLLM-capable server works. Two options:
Option A β full rkllama server (Ollama API; recommended for OpenViking)
pip install rkllama # python 3.9β3.12
rkllama_server --models /path/to/models
Place the .rkllm in your models dir. This gives an Ollama-compatible API, which is what OpenViking's query_planner speaks.
Option B β minimal OpenAI+Ollama server (no transformers/torch)
The same ctypes wrapper, no heavy deps (see the conversion recipe below for the gist). Serves /v1/* and /api/*.
OpenViking wiring
{
"query_planner": {
"provider": "litellm",
"model": "ollama/guoxuter/ov_intent_analysis_sft:v7_q8",
"api_base": "http://127.0.0.1:8091",
"temperature": 0.0,
"timeout": 60,
"extra_request_body": { "think": false }
}
}
Keep the model string exactly as-is β OpenViking auto-matches the bundled v7 prompt by string (retrieval.ov_intent_analysis_sft_v7 in intent_analyzer.py). Only api_base changes: point it at your RKLLM server instead of Ollama.
Benchmark (RK3588, same prompt)
| Runtime | Wall time | Notes |
|---|---|---|
| Ollama / CPU (GGUF Q8) | 16.5 s | output lands in thinking unless think:false |
| rk-llama.cpp NPU (GGUF Q8) | 10.9 s | needed --reasoning off |
| RKLLM NPU (this file, W8A8) | ~9 s (1.8 s warm) | prefill ~200 t/s, decode ~13 t/s |
Query planning is prefill-dominated, which is exactly where the NPU wins. Decode is memory-bandwidth-bound, so don't expect magic on long generations β this model's outputs are short JSON.
Conversion recipe (for reproducing / other models)
The converter is x86-only, so run it on an x86 box (any Linux, or a serverless cloud like Modal):
# rkllm-toolkit 1.3.0 (wheel from airockchip/rknn-llm release-v1.3.0,
# rkllm-toolkit/packages/rkllm_toolkit-1.3.0-cp311-cp311-linux_x86_64.whl)
# deps pinned from that release's requirements.txt (torch 2.6.0, transformers 5.8.0, ...)
from rkllm.api import RKLLM
llm = RKLLM()
llm.load_huggingface(model="guoxuter/ov_intent_analysis_sft", device="cpu") # or cuda
llm.build(
do_quantization=True,
optimization_level=1,
quantized_dtype="W8A8",
quantized_algorithm="normal",
target_platform="RK3588",
num_npu_core=3,
dataset="data_quant.json", # calibration: input/target pairs
hybrid_rate=0, # REQUIRED arg in 1.3.0
max_context=4096, # NOTE: `max_context`, NOT `max_context_len`
)
llm.export_rkllm("ov_intent_analysis_sft_v7_w8a8_rk3588.rkllm")
Gotchas hit along the way (so you don't):
build()takesmax_context, notmax_context_len(raisesTypeError: unexpected keyword argument).hybrid_rate=0is required in the 1.3.0 signature.- The toolkit only ships x86_64 wheels β there is no aarch64 path; don't fight it on an ARM SBC.
- Keep the toolkit major version aligned with your runtime's
librkllmrt.so(1.3.0 β 1.3.0).
Attribution & license
- Base model:
guoxuter/ov_intent_analysis_sftβ Apache-2.0, which its card states covers the fine-tuned checkpoint (same license as Qwen3.5-0.8B). - Conversion performed with Rockchip's rkllm-toolkit 1.3.0 (airockchip/rknn-llm).
- This conversion is published under Apache-2.0.
Big thanks to guoxuter for the tuned model and the OpenViking team for the recommended workflow.
- Downloads last month
- 7
Model tree for n8dgr8/ov_intent_analysis_sft-RKLLM-RK3588
Base model
Qwen/Qwen3.5-0.8B-Base