Instructions to use Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit" --prompt "Once upon a time"
- Atomic Chat
AliceAI-Foundation-80B-A3B-Base — MLX mixed 2-bit/4-bit quantization
A mixed-precision MLX quantization of yandex/AliceAI-Foundation-80B-A3B-Base, an 80B-parameter (79.6B, ~3B active) decoder-only MoE model with 48 layers, hybrid linear/full attention, hidden size 2048 and vocabulary 129024. The custom architecture (alice_ai) is fully supported: model.py and inference.py are bundled with the weights.
Built directly from the original BF16 release yandex/AliceAI-Foundation-80B-A3B-Base (~151 GiB): every quantized tensor is quantized exactly once from BF16 — MoE experts to affine 2-bit g32, all precision-sensitive tensors to affine 4-bit g64. No intermediate 4-bit pass, no double quantization. Result: 28.29 GiB (9 shards, 2803 tensors) — small enough to run at full speed in 48 GB unified memory, unlike the 4-bit release which causes heavy memory swapping.
Mixed-precision quantization scheme
| Tensor group | Bits | Group size | Source |
|---|---|---|---|
MoE experts (model.layers.*.mlp.experts.gate_up_proj / down_proj, 96 triples) |
2 | 32 | Quantized once from the original BF16 weights to affine 2-bit g32 (single quantization) |
Attention q/k/v/o_proj, linear-attention projections, shared_expert, embed_tokens, lm_head (662 triples) |
4 | 64 | Quantized once from the original BF16 weights to affine 4-bit g64 (single quantization) |
Norms, routers, convs, dt_bias |
BF16 | — | Unchanged |
config.json declares global 4-bit/g64 quantization with per-layer top-level overrides in config["quantization"] for the 96 expert paths ({"bits": 2, "group_size": 32}). mlx_lm's load_model class_predicate honours this map natively — no patched loader is required.
Why mixed precision?
The MoE experts account for 96.9% of all parameters. A uniform 2-bit quantization of the whole model (23.26 GiB) collapses quality: it answers only 1/5 of a simple QA battery correctly and produces garbage ("capital of France" → "закон") or empty outputs. Keeping the remaining precision-sensitive tensors (attention, embeddings, head, shared expert) at 4-bit costs only ~5 GiB extra but restores the model to 5/5 correct — matching the 4-bit baseline — while still fitting comfortably in 48 GB unified memory.
Validation (greedy, temperature 0, bundled inference.py)
| Variant | Size | QA battery | Notes |
|---|---|---|---|
| Uniform 2-bit g64 | 23.26 GiB | 1/5 | "2+2=4" ok; "capital of France" → "закон"; 3 prompts → empty |
| This mixed 2/4-bit | 28.29 GiB | 5/5 | Fits and runs at full speed on M4 Pro 48 GB |
| 4-bit baseline | 44.87 GB | 5/5 | Heavy swap on 48 GB machines |
Sample outputs of this variant (greedy): "2+2=4."; "Как называется столица России?" → "Москва."; "Как называется столица Франции?" → "Париж."; "Какого цвета небо?" → "Небо голубое."; "Translate hello world to French" → "'bonjour le monde'".
Usage
With mlx_lm
from mlx_lm.utils import load_model
model, tokenizer = load_model("Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit")
With the bundled runner
# Russian few-shot preset
python inference.py --model . --preset ru --prompt "Как называется столица России?"
# English few-shot preset
python inference.py --model . --preset en --prompt "Translate hello world to French"
# Raw completion (no preset)
python inference.py --model . --preset raw --prompt "2+2="
Presets: ru (preset.txt, Russian few-shot), en (preset_en.txt, English few-shot), raw. The original upstream repo shipped a Japanese preset; Russian and English presets are provided here instead.
Conversion method
The conversion script is included as convert_f80b_direct.py. It:
- Loads the original BF16 safetensors shards (~151 GiB, 49 shards) directly.
- Quantizes the 96 expert triples (3D per-expert tensors) once, to affine 2-bit g32.
- Quantizes all other quantized tensors once, to affine 4-bit g64.
- Keeps plain tensors (norms, routers, convs) at BF16 and drops the MTP module (as in the MLX release).
- Writes the per-layer quantization map into
config.json(honoured natively bymlx_lm).
Conversion took 166 s on an M4 Pro. Compared to the earlier double-quantized build (BF16 → 4-bit → 2-bit experts), the single-pass expert weights are measurably closer to the originals (mean relative L2 error 0.363 vs 0.379 on sampled layers).
Conversion history
The model went through three iterations before reaching the current build. They are documented here because each step was driven by a measurable quality or correctness finding.
Step 1 — Uniform 2-bit (failed)
The first attempt quantized every linear weight of the model uniformly to affine 2-bit g64 (convert_f80b_2bit.py), starting from the community 4-bit MLX release Yamada114514/AliceAI-Foundation-80B-A3B-Base-MLX-4bit (45 GB). Result: 23.26 GiB, but quality collapsed — 1/5 on the QA battery, garbage answers ("capital of France" → "закон") and empty outputs. Root cause: the MoE experts tolerate 2 bits, but attention projections, embeddings and the LM head do not.
Step 2 — Mixed 2/4-bit via requantization (first working build)
Parameter analysis showed the 512-experts-per-layer MoE tensors account for 96.9% of all parameters, while attention, embeddings and the head are precision-sensitive. The fix was a mixed scheme (convert_f80b_mixed.py): experts dequantized from the 4-bit release and requantized to 2-bit g32; everything else kept at 4-bit g64; norms, routers and convs left in BF16. This restored 5/5 on the QA battery at 28.29 GiB and became the first published release.
Two limitations remained:
- Double quantization: the experts passed through 4-bit before reaching 2-bit, adding avoidable rounding error.
- Per-layer quantization support:
mlx_lmneeded a per-tensor override mechanism, solved with top-level per-path keys inconfig["quantization"](honoured natively byload_model'sclass_predicate— no patched loader).
Step 3 — Direct conversion from the original BF16 (current build)
The original BF16 release yandex/AliceAI-Foundation-80B-A3B-Base (~151 GiB, 49 shards) was downloaded and a new converter (convert_f80b_direct.py) was written to quantize each tensor exactly once from BF16, deriving the exact target tensor set and per-tensor bit widths from the reference MLX-2bit build so the output stays drop-in compatible. Two source-format quirks had to be handled:
- Expert tensors in the original release carry no
.weightsuffix (model.layers.0.mlp.experts.gate_up_proj) and are 3D(512 experts, rows, cols); every other tensor uses the.weightnaming. - After
mx.quantizeon a flattened 3D expert tensor, the packed weights, scales and biases must be reshaped back to 3D(512, rows, packed_cols)—SwitchLinearrejects 2D scales.
Validation of the direct build:
- QA battery: 5/5 (identical answers to the double-quantized build).
- Weight fidelity: mean relative L2 error of dequantized expert weights vs the original BF16, sampled over layers 0/12/24/47: 0.363 (direct) vs 0.379 (double-quantized) — the single-pass build is strictly closer to the original weights at identical size and speed.
- Long-form generation: a 10-chapter (948-paragraph) Chinese→Russian literary translation run completed with 99.6% of paragraphs resolved and zero CJK leakage. (A repetition artefact seen in early runs turned out to be confident pattern-continuation of the base model — p≈1.0 per token — identical in both builds, and is handled by cutting the output at the first CJK re-onset; see Limitations.)
The current release replaces the Step-2 weights wholesale; the tensor layout, config.json, model.py, inference.py and presets are unchanged, so existing pipelines keep working without modification.
Limitations
- The base model is a base (pretrained, non-instruct) model; expect raw-completion behaviour rather than instruction following. It handles Russian and English well (validated on both), and the bundled few-shot presets (
ru,en) elicit reliable short-form answers. As a base model it may continue the prompt's structure (e.g. fabricate new "source — translation" pairs) instead of stopping; downstream pipelines should cut the output at the first CJK re-onset after the translation. - This is an experimental community quantization; the QA battery above is a smoke test, not a rigorous benchmark.
Hardware requirements
- Apple Silicon Mac with ≥36 GB unified memory recommended.
- Tested on a MacBook Pro with Apple M4 Pro (48 GB) — the model runs at full speed with no memory swapping.
- The 4-bit release (44.87 GB) does not fit comfortably in 48 GB and swaps heavily.
License and credits
- License: Apache-2.0 (see
LICENSE,APACHE-2.0.txt,NOTICES). - Base model: yandex/AliceAI-Foundation-80B-A3B-Base by Yandex — credit for the original model and weights.
- 4-bit MLX release: Yamada114514/AliceAI-Foundation-80B-A3B-Base-MLX-4bit by Yamada114514 — an earlier build of this quantization was based on that release; the current weights are converted directly from the original BF16.
- Downloads last month
- 564
4-bit
Model tree for Hosstia/AliceAI-Foundation-80B-A3B-Base-MLX-2bit
Base model
yandex/AliceAI-Foundation-80B-A3B-Base