AliceAI-Foundation-80B-A3B-Base — MLX mixed 2-bit/4-bit quantization

A mixed-precision MLX quantization of yandex/AliceAI-Foundation-80B-A3B-Base, an 80B-parameter (79.6B, ~3B active) decoder-only MoE model with 48 layers, hybrid linear/full attention, hidden size 2048 and vocabulary 129024. The custom architecture (alice_ai) is fully supported: model.py and inference.py are bundled with the weights.

Built directly from the original BF16 release yandex/AliceAI-Foundation-80B-A3B-Base (~151 GiB): every quantized tensor is quantized exactly once from BF16 — MoE experts to affine 2-bit g32, all precision-sensitive tensors to affine 4-bit g64. No intermediate 4-bit pass, no double quantization. Result: 28.29 GiB (9 shards, 2803 tensors) — small enough to run at full speed in 48 GB unified memory, unlike the 4-bit release which causes heavy memory swapping.

Mixed-precision quantization scheme

Tensor group Bits Group size Source
MoE experts (model.layers.*.mlp.experts.gate_up_proj / down_proj, 96 triples) 2 32 Quantized once from the original BF16 weights to affine 2-bit g32 (single quantization)
Attention q/k/v/o_proj, linear-attention projections, shared_expert, embed_tokens, lm_head (662 triples) 4 64 Quantized once from the original BF16 weights to affine 4-bit g64 (single quantization)
Norms, routers, convs, dt_bias BF16 — Unchanged

config.json declares global 4-bit/g64 quantization with per-layer top-level overrides in config["quantization"] for the 96 expert paths ({"bits": 2, "group_size": 32}). mlx_lm's load_model class_predicate honours this map natively — no patched loader is required.

Why mixed precision?

The MoE experts account for 96.9% of all parameters. A uniform 2-bit quantization of the whole model (23.26 GiB) collapses quality: it answers only 1/5 of a simple QA battery correctly and produces garbage ("capital of France" → "закон") or empty outputs. Keeping the remaining precision-sensitive tensors (attention, embeddings, head, shared expert) at 4-bit costs only ~5 GiB extra but restores the model to 5/5 correct — matching the 4-bit baseline — while still fitting comfortably in 48 GB unified memory.

Validation (greedy, temperature 0, bundled inference.py)

Variant Size QA battery Notes
Uniform 2-bit g64 23.26 GiB 1/5 "2+2=4" ok; "capital of France" → "закон"; 3 prompts → empty
This mixed 2/4-bit 28.29 GiB 5/5 Fits and runs at full speed on M4 Pro 48 GB
4-bit baseline 44.87 GB 5/5 Heavy swap on 48 GB machines

Sample outputs of this variant (greedy): "2+2=4."; "Как называется столица России?" → "Москва."; "Как называется столица Франции?" → "Париж."; "Какого цвета небо?" → "Небо голубое."; "Translate hello world to French" → "'bonjour le monde'".

Usage

With mlx_lm

from mlx_lm.utils import load_model
model, tokenizer = load_model("Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit")

With the bundled runner

# Russian few-shot preset
python inference.py --model . --preset ru --prompt "Как называется столица России?"

# English few-shot preset
python inference.py --model . --preset en --prompt "Translate hello world to French"

# Raw completion (no preset)
python inference.py --model . --preset raw --prompt "2+2="

Presets: ru (preset.txt, Russian few-shot), en (preset_en.txt, English few-shot), raw. The original upstream repo shipped a Japanese preset; Russian and English presets are provided here instead.

Conversion method

The conversion script is included as convert_f80b_direct.py. It:

  1. Loads the original BF16 safetensors shards (~151 GiB, 49 shards) directly.
  2. Quantizes the 96 expert triples (3D per-expert tensors) once, to affine 2-bit g32.
  3. Quantizes all other quantized tensors once, to affine 4-bit g64.
  4. Keeps plain tensors (norms, routers, convs) at BF16 and drops the MTP module (as in the MLX release).
  5. Writes the per-layer quantization map into config.json (honoured natively by mlx_lm).

Conversion took 166 s on an M4 Pro. Compared to the earlier double-quantized build (BF16 → 4-bit → 2-bit experts), the single-pass expert weights are measurably closer to the originals (mean relative L2 error 0.363 vs 0.379 on sampled layers).

Conversion history

The model went through three iterations before reaching the current build. They are documented here because each step was driven by a measurable quality or correctness finding.

Step 1 — Uniform 2-bit (failed)

The first attempt quantized every linear weight of the model uniformly to affine 2-bit g64 (convert_f80b_2bit.py), starting from the community 4-bit MLX release Yamada114514/AliceAI-Foundation-80B-A3B-Base-MLX-4bit (45 GB). Result: 23.26 GiB, but quality collapsed — 1/5 on the QA battery, garbage answers ("capital of France" → "закон") and empty outputs. Root cause: the MoE experts tolerate 2 bits, but attention projections, embeddings and the LM head do not.

Step 2 — Mixed 2/4-bit via requantization (first working build)

Parameter analysis showed the 512-experts-per-layer MoE tensors account for 96.9% of all parameters, while attention, embeddings and the head are precision-sensitive. The fix was a mixed scheme (convert_f80b_mixed.py): experts dequantized from the 4-bit release and requantized to 2-bit g32; everything else kept at 4-bit g64; norms, routers and convs left in BF16. This restored 5/5 on the QA battery at 28.29 GiB and became the first published release.

Two limitations remained:

  • Double quantization: the experts passed through 4-bit before reaching 2-bit, adding avoidable rounding error.
  • Per-layer quantization support: mlx_lm needed a per-tensor override mechanism, solved with top-level per-path keys in config["quantization"] (honoured natively by load_model's class_predicate — no patched loader).

Step 3 — Direct conversion from the original BF16 (current build)

The original BF16 release yandex/AliceAI-Foundation-80B-A3B-Base (~151 GiB, 49 shards) was downloaded and a new converter (convert_f80b_direct.py) was written to quantize each tensor exactly once from BF16, deriving the exact target tensor set and per-tensor bit widths from the reference MLX-2bit build so the output stays drop-in compatible. Two source-format quirks had to be handled:

  • Expert tensors in the original release carry no .weight suffix (model.layers.0.mlp.experts.gate_up_proj) and are 3D (512 experts, rows, cols); every other tensor uses the .weight naming.
  • After mx.quantize on a flattened 3D expert tensor, the packed weights, scales and biases must be reshaped back to 3D (512, rows, packed_cols) — SwitchLinear rejects 2D scales.

Validation of the direct build:

  • QA battery: 5/5 (identical answers to the double-quantized build).
  • Weight fidelity: mean relative L2 error of dequantized expert weights vs the original BF16, sampled over layers 0/12/24/47: 0.363 (direct) vs 0.379 (double-quantized) — the single-pass build is strictly closer to the original weights at identical size and speed.
  • Long-form generation: a 10-chapter (948-paragraph) Chinese→Russian literary translation run completed with 99.6% of paragraphs resolved and zero CJK leakage. (A repetition artefact seen in early runs turned out to be confident pattern-continuation of the base model — p≈1.0 per token — identical in both builds, and is handled by cutting the output at the first CJK re-onset; see Limitations.)

The current release replaces the Step-2 weights wholesale; the tensor layout, config.json, model.py, inference.py and presets are unchanged, so existing pipelines keep working without modification.

Limitations

  • The base model is a base (pretrained, non-instruct) model; expect raw-completion behaviour rather than instruction following. It handles Russian and English well (validated on both), and the bundled few-shot presets (ru, en) elicit reliable short-form answers. As a base model it may continue the prompt's structure (e.g. fabricate new "source — translation" pairs) instead of stopping; downstream pipelines should cut the output at the first CJK re-onset after the translation.
  • This is an experimental community quantization; the QA battery above is a smoke test, not a rigorous benchmark.

Hardware requirements

  • Apple Silicon Mac with ≥36 GB unified memory recommended.
  • Tested on a MacBook Pro with Apple M4 Pro (48 GB) — the model runs at full speed with no memory swapping.
  • The 4-bit release (44.87 GB) does not fit comfortably in 48 GB and swaps heavily.

License and credits

Downloads last month
564
Safetensors
Model size
80B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit

Quantized
(6)
this model