Instructions to use KaiyangLi/Mizar-3B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use KaiyangLi/Mizar-3B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("KE-Team/Ke-Omni-R-3B") model = PeftModel.from_pretrained(base_model, "KaiyangLi/Mizar-3B") - Notebooks
- Google Colab
- Kaggle
Mizar-3B: Audio Understanding with RT-OPD
Paper · Mizar family · GitHub code · Mizar-159M · Mizar-3B
Mizar-3B is the Ke-based 3B member of the Mizar audio-language model family. This repository contains its LoRA adapter from the main experiment in Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models. It needs the pinned KE-Team/Ke-Omni-R-3B base. It is not a standalone base model. The model accepts audio and a multiple-choice question and generates text.
Mizar model family
| Model | Model weights | Code | Foundation / training |
|---|---|---|---|
| Mizar-159M | KaiyangLi/Mizar-159M | Mizar_159M | CED-Small + SmolLM2-135M; three-stage audio-language training |
| Mizar-3B | KaiyangLi/Mizar-3B | RT-OPD | Ke-Omni-R-3B + RT-OPD; released as a LoRA adapter |
These models share the Mizar family name and audio-understanding focus. They use different foundations and training recipes. Mizar-3B is the name of the released Ke-based model; RT-OPD is its distillation method. The Qwen profile in this repository remains a separate RT-OPD experiment.
Which checkpoint is this?
Seed 85, fixed final step 626, selected post hoc by the highest Macro-3 among the five main-experiment seeds (82–86). This release choice does not replace the paper's five-seed mean and is not an independent test-set estimate.
| Result | MMAU full (9,000) | MMAR (1,000) | ADQA-cl (1,577) | Macro-3 |
|---|---|---|---|---|
| This seed 85 adapter | 72.7778 | 61.1000 | 57.0704 | 63.6494 |
| Paper: five-seed mean | 72.7222 | 60.0800 | 56.4490 | 63.0837 ± 0.3592 |
All numbers are percentages; ± is sample standard deviation of per-seed Macro-3.
Seed 86 has the highest MMAU alone (73.1222%), but seed 85 has the highest Macro-3.
MMAU uses the official hidden-label scoring service; mini is a separate split.
main_results.json records all ten Ke/Qwen runs and evidence hashes.
Evaluation inputs and scoring code
- Prompt construction and answer parsing:
ke/source/portable/ke_opd_v2/modeling.py - Evaluation entry point:
ke/scripts/evaluate.py - Protocol:
docs/EVALUATION.md
Loading and running
Use the matching code and pinned Python 3.10 environment from the RT-OPD release.
git clone https://github.com/KaiyangLi1992/RT-OPD.git
cd RT-OPD
bash ke/environment/setup.sh
ke/.venv/bin/hf auth login
ke/.venv/bin/python tools/infer.py \
--adapter KaiyangLi/Mizar-3B \
--audio /absolute/path/example.wav \
--question "Which sound is audible?" \
--choices "A dog barking" "A piano playing"
The helper downloads the pinned base, loads its Thinker component with
Transformers 4.52.4, attaches this adapter with PEFT 0.19.1, and runs greedy
text generation. The teacher is unnecessary at inference. On GPUs without native
BF16 support, --dtype float16 is a convenience option; this is not the paper's
frozen vLLM benchmark pipeline. Follow docs/EVALUATION.md for paper evaluation.
For immutable runs, pass the HF commit ID to --revision.
Training recipe
Frozen Ke-Omni-R 7B teacher; fresh Ke-Omni-R-3B student; 10,000 fixed training rows; two epochs / 626 steps; global batch 32 (4 GPUs × microbatch 4 × accumulation 2); learning rate 7.5e-5; LoRA rank 64, alpha 128, dropout 0.05; CE + 0.25 × gated reverse KL; target alpha 1.0; student rollouts at temperature 1.0, top-p 0.95, top-k 64, maximum 96 new tokens. The teacher scores the same student prefix with and without audio; there is no separate reference model or donor audio in this main method.
Files and provenance
adapter_model.safetensors: original, unmodified trained adapter bytes.adapter_config.json: original PEFT settings with the base path replaced by its public model ID and pinned revision for portability.base_model.json: exact base revision and per-file SHA-256 values.CHECKPOINT_PROVENANCE.json: original and published identities, selection policy, and original configuration hash. Optimizer/RNG state is not needed for inference and is not part of this inference release.SHA256SUMS.json: checksums of published payload files.
The model is for research on audio understanding. Dataset and base-model terms remain applicable; see the source repositories linked in the code's data guide. This package does not include hidden MMAU gold labels.
Frozen data manifests
reproducibility/frozen_data.tar.gz contains the exact 10,000-row training
manifest, frozen teacher gate, valid vocabulary, benchmark manifests and audio
hashes shared by Ke and Qwen. The code automatically downloads this archive at
a pinned revision before preparing models/audio. The archive contains no audio
or model weights and no hidden MMAU test labels. Upstream dataset terms apply.
License
The adapter weights and the files authored for this release are licensed under the BSD 3-Clause Clear License. The adapter requires its base model, KE-Team/Ke-Omni-R-3B, which is fine-tuned from Qwen/Qwen2.5-Omni-3B; using the adapter with that base remains subject to the base models' terms, including the Qwen Research License. Dataset manifests are not relicensed and keep their upstream terms. See NOTICE.
- Downloads last month
- 31