| --- |
| license: openmdw-1.1 |
| language: |
| - en |
| pipeline_tag: text-generation |
| library_name: vllm |
| tags: |
| - moe |
| - marin |
| - grug |
| - base-model |
| --- |
| # Snowball 67B-A2B, 262K context, QK 1.75, skew 4× |
|
|
| BF16 export of step 157,000 from the [Snowball long-context comparison](https://github.com/marin-community/marin/issues/8977). |
| The run extended context to 262,144 tokens from step 156,000 through step 157,000 with `qk_mult=1.75`. |
| Long-context documents were upsampled 4× during extension. |
|
|
| The model has 67B total parameters, approximately 2B active parameters per token, 26 layers, 256 experts with four selected per token, and a 128,256-token vocabulary. |
| This is a base model. The tokenizer's bundled chat template does not indicate instruction tuning; use text completions. |
|
|
| Serving requires the [Marin vLLM fork](https://github.com/marin-community/vllm), which registers `GrugMoeForCausalLM` / `grug_moe`. |
| The existing serving configuration uses eight H100 GPUs, tensor parallelism 1, data parallelism 8, and expert parallelism because the model has five KV heads. |
| See the [serving and evaluation record](https://github.com/marin-community/marin/issues/8702) for the QK 1.57 and 1.75 baseline exports. |
| This upload validates tensor names, BF16 dtype, config, sizes, and SHA-256 checksums. No new generation or numerical parity test has been run for this upload. |
|
|
| The [export manifest](export-manifest.json) records each file’s size and SHA-256 checksum, the exporter revision, and the verified Hub weight revision. |
| The exporter applies pending QB router biases before conversion. Optimizer state is excluded. |
|
|
| Model materials are licensed under [OpenMDW 1.1](LICENSE). |
|
|