--- model_name: Mizar-3B base_model: KE-Team/Ke-Omni-R-3B library_name: peft license: bsd-3-clause-clear tags: - audio-language-model - audio-question-answering - lora - mizar - rt-opd - knowledge-distillation language: - en --- # Mizar-3B: Audio Understanding with RT-OPD **[Paper](https://arxiv.org/abs/2609.28778)** · **[Mizar family](https://huggingface.co/collections/KaiyangLi/mizar-audio-language-model-family-6aa96a97630d4868ab979b4d)** · **[GitHub code](https://github.com/KaiyangLi1992/RT-OPD)** · **[Mizar-159M](https://huggingface.co/KaiyangLi/Mizar-159M)** · **[Mizar-3B](https://huggingface.co/KaiyangLi/Mizar-3B)** **Mizar-3B** is the Ke-based 3B member of the Mizar audio-language model family. This repository contains its **LoRA adapter** from the main experiment in [*Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models*](https://arxiv.org/abs/2609.28778). It needs the pinned **KE-Team/Ke-Omni-R-3B** base. It is not a standalone base model. The model accepts audio and a multiple-choice question and generates text. ## Mizar model family | Model | Model weights | Code | Foundation / training | |---|---|---|---| | **Mizar-159M** | [KaiyangLi/Mizar-159M](https://huggingface.co/KaiyangLi/Mizar-159M) | [Mizar_159M](https://github.com/KaiyangLi1992/Mizar_159M) | CED-Small + SmolLM2-135M; three-stage audio-language training | | **Mizar-3B** | [KaiyangLi/Mizar-3B](https://huggingface.co/KaiyangLi/Mizar-3B) | [RT-OPD](https://github.com/KaiyangLi1992/RT-OPD) | Ke-Omni-R-3B + RT-OPD; released as a LoRA adapter | These models share the **Mizar** family name and audio-understanding focus. They use different foundations and training recipes. Mizar-3B is the name of the released Ke-based model; RT-OPD is its distillation method. The Qwen profile in this repository remains a separate RT-OPD experiment. ## Which checkpoint is this? Seed **85**, fixed final step **626**, selected **post hoc by the highest Macro-3 among the five main-experiment seeds (82–86)**. This release choice does not replace the paper's five-seed mean and is not an independent test-set estimate. | Result | MMAU full (9,000) | MMAR (1,000) | ADQA-cl (1,577) | Macro-3 | |---|---:|---:|---:|---:| | This seed 85 adapter | 72.7778 | 61.1000 | 57.0704 | 63.6494 | | Paper: five-seed mean | 72.7222 | 60.0800 | 56.4490 | 63.0837 ± 0.3592 | All numbers are percentages; ± is sample standard deviation of per-seed Macro-3. Seed 86 has the highest MMAU alone (73.1222%), but seed 85 has the highest Macro-3. MMAU uses the official hidden-label scoring service; mini is a separate split. `main_results.json` records all ten Ke/Qwen runs and evidence hashes. ## Evaluation inputs and scoring code - Prompt construction and answer parsing: [`ke/source/portable/ke_opd_v2/modeling.py`](https://github.com/KaiyangLi1992/RT-OPD/blob/main/ke/source/portable/ke_opd_v2/modeling.py) - Evaluation entry point: [`ke/scripts/evaluate.py`](https://github.com/KaiyangLi1992/RT-OPD/blob/main/ke/scripts/evaluate.py) - Protocol: [`docs/EVALUATION.md`](https://github.com/KaiyangLi1992/RT-OPD/blob/main/docs/EVALUATION.md) ## Loading and running Use the matching code and pinned Python 3.10 environment from [the RT-OPD release](https://github.com/KaiyangLi1992/RT-OPD). ```bash git clone https://github.com/KaiyangLi1992/RT-OPD.git cd RT-OPD bash ke/environment/setup.sh ke/.venv/bin/hf auth login ke/.venv/bin/python tools/infer.py \ --adapter KaiyangLi/Mizar-3B \ --audio /absolute/path/example.wav \ --question "Which sound is audible?" \ --choices "A dog barking" "A piano playing" ``` The helper downloads the pinned base, loads its **Thinker** component with Transformers **4.52.4**, attaches this adapter with PEFT **0.19.1**, and runs greedy text generation. The teacher is unnecessary at inference. On GPUs without native BF16 support, `--dtype float16` is a convenience option; this is not the paper's frozen vLLM benchmark pipeline. Follow `docs/EVALUATION.md` for paper evaluation. For immutable runs, pass the HF commit ID to `--revision`. ## Training recipe Frozen Ke-Omni-R 7B teacher; fresh Ke-Omni-R-3B student; 10,000 fixed training rows; two epochs / 626 steps; global batch 32 (4 GPUs × microbatch 4 × accumulation 2); learning rate 7.5e-5; LoRA rank 64, alpha 128, dropout 0.05; CE + 0.25 × gated reverse KL; target alpha 1.0; student rollouts at temperature 1.0, top-p 0.95, top-k 64, maximum 96 new tokens. The teacher scores the same student prefix with and without audio; there is no separate reference model or donor audio in this main method. ## Files and provenance - `adapter_model.safetensors`: original, unmodified trained adapter bytes. - `adapter_config.json`: original PEFT settings with the base path replaced by its public model ID and pinned revision for portability. - `base_model.json`: exact base revision and per-file SHA-256 values. - `CHECKPOINT_PROVENANCE.json`: original and published identities, selection policy, and original configuration hash. Optimizer/RNG state is not needed for inference and is not part of this inference release. - `SHA256SUMS.json`: checksums of published payload files. The model is for research on audio understanding. Dataset and base-model terms remain applicable; see the source repositories linked in the code's data guide. This package does not include hidden MMAU gold labels. ## Frozen data manifests `reproducibility/frozen_data.tar.gz` contains the exact 10,000-row training manifest, frozen teacher gate, valid vocabulary, benchmark manifests and audio hashes shared by Ke and Qwen. The code automatically downloads this archive at a pinned revision before preparing models/audio. The archive contains no audio or model weights and no hidden MMAU test labels. Upstream dataset terms apply. ## License The adapter weights and the files authored for this release are licensed under the [BSD 3-Clause Clear License](LICENSE). The adapter requires its base model, KE-Team/Ke-Omni-R-3B, which is fine-tuned from Qwen/Qwen2.5-Omni-3B; using the adapter with that base remains subject to the base models' terms, including the Qwen Research License. Dataset manifests are not relicensed and keep their upstream terms. See [NOTICE](NOTICE).