Instructions to use Ruiruiz30/OmniJev-MLX-5bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Ruiruiz30/OmniJev-MLX-5bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("Ruiruiz30/OmniJev-MLX-5bit") config = load_config("Ruiruiz30/OmniJev-MLX-5bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Ruiruiz30/OmniJev-MLX-5bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Ruiruiz30/OmniJev-MLX-5bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Ruiruiz30/OmniJev-MLX-5bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use Ruiruiz30/OmniJev-MLX-5bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Ruiruiz30/OmniJev-MLX-5bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Ruiruiz30/OmniJev-MLX-5bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Ruiruiz30/OmniJev-MLX-5bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Ruiruiz30/OmniJev-MLX-5bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Ruiruiz30/OmniJev-MLX-5bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
OmniJev-MLX-5bit
An independent Apple Silicon MLX conversion of the official OmniJev v1.1 4B checkpoint.
The model uses the official Qwen/Qwen3.5-4B backbone and OmniJev decision head. The official LoRA adapter was merged into the backbone in bfloat16, and the fused backbone was then converted to 5-bit affine quantization with group size 64. The decision head remains a separate FP32 file and is loaded at inference time.
This repository is a conversion and packaging release. It is not a new fine-tune and does not contain additional training of the 4B model or decision head. It is not an official release of the upstream OmniJev or Qwen projects.
What is included
model.safetensors: fused Qwen3.5/OmniJev backbone in MLX 5-bit format.decision_head/head.npz: the official OmniJev 4B decision head exported from the releasedhead.pt.decision_head/head_meta.json: upstream calibration metadata, including the choice temperature.omni_mlx/classifier.py: a small MLX inference wrapper for typed visual decisions.conversion.json: source, quantization and validation provenance.head_provenance.json: checksum and export provenance for the decision head.benchmarks/: the local fixed-subset ScienceQA measurements used for this conversion.
The head is separate from the quantized backbone. The runtime first obtains option/question representations from the backbone and then applies the decision head. This is different from merging the head into the Qwen weight tensors.
Conversion provenance
Qwen/Qwen3.5-4B
+ official OmniJev v1.1 LoRA adapter
└─ merge in bfloat16
fused OmniJev backbone
└─ MLX affine quantization, 5 bits, group size 64
model.safetensors
official OmniJev head.pt
└─ format export only
decision_head/head.npz (float32)
The source OmniJev release reports a rank-32 LoRA adapter and an independently calibrated decision head. This repository preserves those learned weights; it does not retrain them.
Local validation
These measurements use one fixed 50-image ScienceQA visual subset, the same subset for both resolutions. They are local conversion checks, not the official OmniJev benchmark and not a universal accuracy claim.
Hardware: Mac mini M4, 16 GB unified memory. Batch size 1.
| Input setting | Correct | Accuracy | Warm median | P95 |
|---|---|---|---|---|
max_pixels=87808 (112-level input) |
44/50 | 88.0% | ~737 ms | ~1,013 ms |
max_pixels=175616 (224-level input) |
44/50 | 88.0% | ~1,050 ms | ~1,354 ms |
The complete per-item outputs are in benchmarks/scienceqa-50-112.json and benchmarks/scienceqa-50-224.json. The text-only JevBench and the upstream OmniJev score tables should not be inferred from this 50-image check.
Installation
This package targets Apple Silicon and MLX.
hf download Ruiruiz30/OmniJev-MLX-5bit \
--local-dir OmniJev-MLX-5bit
cd OmniJev-MLX-5bit
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
The model is about 3.3 GB. Keep the repository's directory structure intact so that decision_head/head.npz and decision_head/head_meta.json remain next to each other.
Image decision
The included wrapper returns probabilities for the supplied options and selects the highest-probability option. It does not generate a free-form explanation.
python -m omni_mlx.classifier \
--model . \
--image /path/to/frame.png \
--state "A kart is approaching a right turn." \
--question "Which steering action is best?" \
--options "Turn left" "Hold center" "Turn right" \
--max-pixels 87808
The output contains prediction, option probabilities, abstain_probability, elapsed time and the quantization setting. The supplied options must be distinct; the head supports typed choice, yes/no and score-style candidate sets through the same candidate-scoring path.
The default max_pixels=87808 is the setting used for the faster local validation above. Increase it when small visual details matter, and validate the resulting accuracy and latency on your own frames.
Upstream sources and license
The OmniJev project and this conversion wrapper are released under Apache-2.0. The Qwen backbone retains its own upstream terms. See LICENSE and NOTICE.md.
- Downloads last month
- 55
5-bit