OmniJev-MLX-5bit

An independent Apple Silicon MLX conversion of the official OmniJev v1.1 4B checkpoint.

The model uses the official Qwen/Qwen3.5-4B backbone and OmniJev decision head. The official LoRA adapter was merged into the backbone in bfloat16, and the fused backbone was then converted to 5-bit affine quantization with group size 64. The decision head remains a separate FP32 file and is loaded at inference time.

This repository is a conversion and packaging release. It is not a new fine-tune and does not contain additional training of the 4B model or decision head. It is not an official release of the upstream OmniJev or Qwen projects.

What is included

  • model.safetensors: fused Qwen3.5/OmniJev backbone in MLX 5-bit format.
  • decision_head/head.npz: the official OmniJev 4B decision head exported from the released head.pt.
  • decision_head/head_meta.json: upstream calibration metadata, including the choice temperature.
  • omni_mlx/classifier.py: a small MLX inference wrapper for typed visual decisions.
  • conversion.json: source, quantization and validation provenance.
  • head_provenance.json: checksum and export provenance for the decision head.
  • benchmarks/: the local fixed-subset ScienceQA measurements used for this conversion.

The head is separate from the quantized backbone. The runtime first obtains option/question representations from the backbone and then applies the decision head. This is different from merging the head into the Qwen weight tensors.

Conversion provenance

Qwen/Qwen3.5-4B
    + official OmniJev v1.1 LoRA adapter
    └─ merge in bfloat16
       fused OmniJev backbone
    └─ MLX affine quantization, 5 bits, group size 64
       model.safetensors

official OmniJev head.pt
    └─ format export only
       decision_head/head.npz (float32)

The source OmniJev release reports a rank-32 LoRA adapter and an independently calibrated decision head. This repository preserves those learned weights; it does not retrain them.

Local validation

These measurements use one fixed 50-image ScienceQA visual subset, the same subset for both resolutions. They are local conversion checks, not the official OmniJev benchmark and not a universal accuracy claim.

Hardware: Mac mini M4, 16 GB unified memory. Batch size 1.

Input setting Correct Accuracy Warm median P95
max_pixels=87808 (112-level input) 44/50 88.0% ~737 ms ~1,013 ms
max_pixels=175616 (224-level input) 44/50 88.0% ~1,050 ms ~1,354 ms

The complete per-item outputs are in benchmarks/scienceqa-50-112.json and benchmarks/scienceqa-50-224.json. The text-only JevBench and the upstream OmniJev score tables should not be inferred from this 50-image check.

Installation

This package targets Apple Silicon and MLX.

hf download Ruiruiz30/OmniJev-MLX-5bit \
  --local-dir OmniJev-MLX-5bit

cd OmniJev-MLX-5bit
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt

The model is about 3.3 GB. Keep the repository's directory structure intact so that decision_head/head.npz and decision_head/head_meta.json remain next to each other.

Image decision

The included wrapper returns probabilities for the supplied options and selects the highest-probability option. It does not generate a free-form explanation.

python -m omni_mlx.classifier \
  --model . \
  --image /path/to/frame.png \
  --state "A kart is approaching a right turn." \
  --question "Which steering action is best?" \
  --options "Turn left" "Hold center" "Turn right" \
  --max-pixels 87808

The output contains prediction, option probabilities, abstain_probability, elapsed time and the quantization setting. The supplied options must be distinct; the head supports typed choice, yes/no and score-style candidate sets through the same candidate-scoring path.

The default max_pixels=87808 is the setting used for the faster local validation above. Increase it when small visual details matter, and validate the resulting accuracy and latency on your own frames.

Upstream sources and license

The OmniJev project and this conversion wrapper are released under Apache-2.0. The Qwen backbone retains its own upstream terms. See LICENSE and NOTICE.md.

Downloads last month
55
Safetensors
Model size
5B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

5-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Ruiruiz30/OmniJev-MLX-5bit

Finetuned
Qwen/Qwen3.5-4B
Quantized
(512)
this model