Snowball 67B-A2B, 262K context, QK 1.75, skew 4×

BF16 export of step 157,000 from the Snowball long-context comparison. The run extended context to 262,144 tokens from step 156,000 through step 157,000 with qk_mult=1.75. Long-context documents were upsampled 4× during extension.

The model has 67B total parameters, approximately 2B active parameters per token, 26 layers, 256 experts with four selected per token, and a 128,256-token vocabulary. This is a base model. The tokenizer's bundled chat template does not indicate instruction tuning; use text completions.

Serving requires the Marin vLLM fork, which registers GrugMoeForCausalLM / grug_moe. The existing serving configuration uses eight H100 GPUs, tensor parallelism 1, data parallelism 8, and expert parallelism because the model has five KV heads. See the serving and evaluation record for the QK 1.57 and 1.75 baseline exports. This upload validates tensor names, BF16 dtype, config, sizes, and SHA-256 checksums. No new generation or numerical parity test has been run for this upload.

The export manifest records each file’s size and SHA-256 checksum, the exporter revision, and the verified Hub weight revision. The exporter applies pending QB router biases before conversion. Optimizer state is excluded.

Model materials are licensed under OpenMDW 1.1.

Downloads last month
258
Safetensors
Model size
67B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including open-athena/snowball-67b-a2b-base-262k-qk175-skew4