--- license: gemma library_name: rkllm language: - en base_model: - google/gemma-3-4b-it pipeline_tag: text-generation tags: - rkllm - rk3588 - rockchip - edge-ai - llm - gemma - text-generation-inference --- # gemma-3-4b-it — RKLLM build for RK3588 boards #### Built with Gemma 3 **Author:** @jamescallander **Source model:** [google/gemma-3-4b-it · Hugging Face](https://huggingface.co/google/gemma-3-4b-it) **Target:** Rockchip RK3588 NPU via **RKNN-LLM Runtime** > This repository hosts a **conversion** of `gemma-3-4b-it` for use on Rockchip RK3588 single-board computers (Orange Pi 5 plus, Radxa Rock 5b+, Banana Pi M7, etc.). Conversion was performed using the [RKNN-LLM toolkit](https://github.com/airockchip/rknn-llm?utm_source=chatgpt.com) #### Conversion details - **RKLLM-Toolkit version:** v1.2.1 - **NPU driver:** v0.9.8 - **Python:** 3.10 - **Quantization:** `w8a8_g128` - **Output:** single-file `.rkllm` artifact - **Modifications:** quantization (w8a8_g128), export to `.rkllm` format for RK3588 SBCs - ****Tokenizer:**** not required at runtime (UI handles prompt I/O) ## Intended use - On-device chat and instruction following on RK3588 SBCs. - gemma-3-4b-it is tuned for general conversational tasks, Q&A, and reasoning, making it suitable for **edge inference deployments** where low power and privacy matter. ## Limitations - Requires 5GB free memory - Quantized build (`w8a8_g128`) may show small quality differences vs. full-precision upstream. - Tested on a Radxa Rock 5B+; other devices may require different drivers/toolkit versions. ## Quick start (RK3588) ### 1) Install runtime The RKNN-LLM toolkit and instructions can be found on the specific development board's manufacturer website or from [airockchip's github page](https://github.com/airockchip). Download and install the required packages as per the toolkit's instructions. ### 2) Simple Flask server deployment The simplest way the deploy the `.rkllm` converted model is using an example script provided in the toolkit in this directory: `rknn-llm/examples/rkllm_server_demo` ```bash python3 /rknn-llm/examples/rkllm_server_demo/flask_server.py \ --rkllm_model_path /gemma-3-4b-it_w8a8_g128_rk3588.rkllm \ --target_platform rk3588 ``` ### 3) Sending a request A basic format for message request is: ```json { "model":"gemma-3-4b-it", "messages":[{ "role":"user", "content":""}], "stream":false } ``` Example request using `curl`: ```bash curl -s -X POST :8080/rkllm_chat \ -H 'Content-Type: application/json' \ -d '{"model":"gemma-3-4b-it","messages":[{"role":"user","content":"Explain who Napoleon Bonaparte is in two or three sentences."}],"stream":false}' ``` The response is formated in the following way: ```json { "choices":[{ "finish_reason":"stop", "index":0, "logprobs":null, "message":{ "content":", "role":"assistant"}}], "created":null, "id":"rkllm_chat", "object":"rkllm_chat", "usage":{ "completion_tokens":null, "prompt_tokens":null, "total_tokens":null} } ``` Example response: ```json {"choices":[{"finish_reason":"stop","index":0,"logprobs":null,"message":{"content":"Napoleon Bonaparte was a brilliant military leader and the Emperor of France, rising to power during the late 1790s and dominating Europe throughout much of the early 1800s. He rose from humble beginnings to become one of history's most iconic figures, known for his strategic genius and ambition, though also infamous for his role in the Napoleonic Wars which ultimately led to his downfall and exile.","role":"assistant"}}],"created":null,"id":"rkllm_chat","object":"rkllm_chat","usage":{"completion_tokens":null,"prompt_tokens":null,"total_tokens":null}} ``` ### 4) UI compatibility This server exposes an **OpenAI-compatible Chat Completions API**. You can connect it to any OpenAI-compatible client or UI (for example: [Open WebUI](https://github.com/open-webui/open-webui?utm_source=chatgpt.com)) - Configure your client with the API base: `http://:8080` and use the endpoint: `/rkllm_chat` - Make sure the `model` field matches the converted model’s name, for example: ```json { "model": "gemma-3-4b-it", "messages": [{"role":"user","content":"Hello!"}], "stream": false } ``` ## ⚠️ Safety disclaimer - 🛑 Not a substitute for professional advice, diagnosis, or treatment. - Intended for **research and educational purposes** only. - Do not rely on outputs for decisions related to health, safety, or legal/financial matters. - Always consult a qualified professional for real-world guidance. - Follow Google’s [Prohibited Use Policy](https://ai.google.dev/gemma/prohibited_use_policy) # License This conversion follows the license of the source model: [Gemma Terms of Use](https://ai.google.dev/gemma/terms) - -**Required notice:** see [`NOTICE`](NOTICE)