--- library_name: mlx license: apache-2.0 base_model: apus-ailab/APUS-OpenJev-v1-35B-A3B base_model_relation: quantized pipeline_tag: text-generation language: - en - zh tags: - apus-openjev - decision-model - mlx - apple-silicon --- # APUS-OpenJev-v1-35B-A3B-MLX-4bit [English](README.md) | [中文](README.zh-CN.md) · [Source model](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-35B-A3B) · [Collection](https://huggingface.co/collections/apus-ailab/apus-openjev-v1-6ab1ee888eb002fcdd3a2825) · [GGUF collection](https://huggingface.co/collections/apus-ailab/apus-openjev-v1-gguf-6ab39d5e724c4d8a1021198f) · [MLX collection](https://huggingface.co/collections/apus-ailab/apus-openjev-v1-mlx-6ab39d5fc988a1b1cb89dfc8) · [MLX-8bit](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-35B-A3B-MLX-8bit) · [GGUF / Ollama](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-35B-A3B-GGUF) MLX weights (mixed 4/6-bit affine (mlx-lm `mixed_4_6`), group size 64) of [APUS-OpenJev-v1-35B-A3B](https://huggingface.co/apus-ailab/APUS-OpenJev-v1-35B-A3B) for **Apple Silicon Macs** (mlx-lm, LM Studio). OpenJev is a **decision model**: each request supplies a state, an instruction and 2–16 candidates, and the model scores candidate labels A–P. It is not a chat model. ## Quick start ```bash pip install mlx-lm hf download apus-ailab/APUS-OpenJev-v1-35B-A3B-MLX-4bit --local-dir ./openjev python ./openjev/examples/openjev_mlx.py --model ./openjev ``` `examples/openjev_mlx.py` renders prompts with [openjev_contracts.py](openjev_contracts.py) (the training contract) and returns the exact candidate distribution. ## Parity Frozen80 with identical prompt tokens, compared with the HF BF16 release (full depth, **71/80 · 88.75%**): | Run / 运行 | Backend / 后端 | Frozen80 | = HF BF16 | Max Δp | |---|---|---:|---:|---:| | NVIDIA RTX PRO 6000 (CUDA) | mlx 0.32.2 on Linux x86_64 (Device(gpu, 0)) | 70/80 · 87.50% | 78/80 | 0.8700 | Converted and scored with MLX on Linux (CUDA); the files are platform-independent and load unchanged on Apple Silicon (for 4B, the same kind of file gave identical decisions on CUDA and Metal). Peak memory on Frozen80 was 23.17 GB; plan for a Mac with at least 32 GB of unified memory. Frozen80 is a reused development panel, not a blind benchmark. Per-question rows (candidate probabilities, choice, correctness; join with Frozen80 by `panel_index`): [evaluation/per-question/](evaluation/per-question/). ## Conversion - mlx-lm `0.31.3` / mlx `0.32.2`; mixed 4/6-bit affine (mlx-lm `mixed_4_6`), group size 64. - 6-bit lm_head and v_proj/down_proj in sensitive layers, 4-bit elsewhere (the MLX analogue of Q4_K_M); MoE router gates stay 8-bit ([recipe](conversion.json)). Calibrated DWQ/GPTQ were tried and were not practical for this hybrid-attention model on the conversion hardware. - GDN `A_log` and `linear_attn.norm.weight` keep their source precision (FP32 in the 35B release). - Full depth only, text only, probabilities **not calibrated**. ## License Apache-2.0, inherited from the source model; see [LICENSE](LICENSE). Base model: [Qwen/Qwen3.5-35B-A3B](https://huggingface.co/Qwen/Qwen3.5-35B-A3B). **Authors:** gumpcheng ([xDAN2099](https://huggingface.co/xDAN2099)), zhangxu, [APUS AI-LAB](https://github.com/APUS-AI-Lab).