WillHeld commited on
Commit
2b1f526
·
verified ·
1 Parent(s): 7580086

Document Snowball long-context checkpoint

Browse files
Files changed (1) hide show
  1. README.md +30 -0
README.md ADDED
@@ -0,0 +1,30 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ pipeline_tag: text-generation
5
+ library_name: vllm
6
+ tags:
7
+ - moe
8
+ - marin
9
+ - grug
10
+ - base-model
11
+ ---
12
+ # Snowball 67B-A2B, 262K context, QK 1.57, skew 1×
13
+
14
+ BF16 export of step 157,000 from the [Snowball long-context comparison](https://github.com/marin-community/marin/issues/8977).
15
+ The run extended context to 262,144 tokens from step 156,000 through step 157,000 with `qk_mult=1.57`. Unchanged long-context document sampling.
16
+
17
+ The model has 67B total parameters, approximately 2B active parameters per token, 26 layers, 256 experts with four selected per token, and a 128,256-token vocabulary.
18
+ This is a base model. The tokenizer's bundled chat template does not indicate instruction tuning; use text completions.
19
+
20
+ Serving requires the [Marin vLLM fork](https://github.com/marin-community/vllm), which registers `GrugMoeForCausalLM` / `grug_moe`.
21
+ The existing serving configuration uses eight H100 GPUs, tensor parallelism 1, data parallelism 8, and expert parallelism because the model has five KV heads.
22
+ See the [serving and evaluation record](https://github.com/marin-community/marin/issues/8702) for the QK 1.57 and 1.75 baseline exports.
23
+ These uploads validate tensor names, BF16 dtype, config, sizes, and SHA-256 checksums. No new generation or numerical parity test has been run for this upload.
24
+
25
+ Source checkpoint: `gs://marin-us-central2/grug/moe_67b_a2b_d2560_ep1_rep1_ctx4_bs256_seq262144_ctxext_step156k_qk157-0997fb/checkpoints/step-157000/`
26
+
27
+ Export source: `gs://marin-us-central2/marin/exports/grug/mrcr-8701/step-157000-qk157/hf-bf16-vllm`
28
+
29
+ The [export manifest](export-manifest.json) records each source object's generation and SHA-256 checksum, the exporter revision, and the verified Hub weight revision.
30
+ The exporter applies pending QB router biases before conversion. Optimizer state is excluded.