duico commited on
Commit
852aaa4
·
verified ·
1 Parent(s): 00e95f0

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +12 -12
README.md CHANGED
@@ -12,7 +12,7 @@ pipeline_tag: text-generation
12
 
13
  # Laguna-XS.2-dense
14
 
15
- A **~3B dense** model distilled from **[poolside/Laguna-XS.2](https://huggingface.co/poolside/Laguna-XS.2)** — a 33B Mixture-of-Experts coding model with a ~3B active path. We replace the MoE feed-forward layers with a single **dense** FFN of the same size as the active path (8 routed + 1 shared expert), turning the ~3B *active* compute into a genuine ~3B *dense* model that keeps XS.2's attention.
16
 
17
  > ⚠️ **Research / hackathon artifact — heavily under-trained.** This checkpoint was produced in a time-boxed hackathon with a tiny distillation budget. It is **not** production-ready: generations are still degenerate/repetitive (see below). It demonstrates the *method* and is a starting point for longer distillation.
18
 
@@ -20,8 +20,8 @@ A **~3B dense** model distilled from **[poolside/Laguna-XS.2](https://huggingfac
20
 
21
  Two stages, both distilling from the frozen FP8 XS.2 teacher:
22
 
23
- 1. **[Stage 1](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1) — per-layer MoE→dense init** (RADLADS-style). Each of the 39 sparse MoE blocks is replaced by a dense SwiGLU FFN (intermediate 4608) and trained *independently, in parallel* to match the teacher MoE block's output (NMSE on the residual contribution), fed the teacher's own hidden states (no error compounding). ~90M tokens. Result: a dense init with held-out perplexity **~25** (vs teacher **~4.4**) — functional but rough, because cross-layer error compounding is left uncorrected by design.
24
- 2. **Stage 2 — synchronous logit-KD** (this model). The stitched ~3B dense student is trained **end-to-end** against the fp8 teacher's full-vocab logits (forward-KL), on a 50/50 code+general corpus. ~14M tokens, one H100, **KL 2.5 → 1.40**.
25
 
26
  Data: 50% [DCLM](https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0) + 50% [StarCoder2/the-stack-v2-train](https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids).
27
 
@@ -35,7 +35,7 @@ Data: 50% [DCLM](https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0
35
 
36
  ## Architecture / loading note
37
 
38
- The dense FFNs are intermediate 4608 (layers 1–39) and 8192 (layer 0). For a uniform config that loads with the **stock** `modeling_laguna.py`, the 4608 FFNs are **zero-padded to 8192** (numerically identical — `silu(0)·0 = 0`). So the exported checkpoint reports ~3.8B params (padded); the true model is ~3.0B.
39
 
40
  ```python
41
  import torch
@@ -57,9 +57,9 @@ print(tok.decode(m.generate(**ids, max_new_tokens=64)[0], skip_special_tokens=Tr
57
 
58
  | Model | Params (resident) | Held-out PPL | HumanEval pass@1 |
59
  |---|---|---|---|
60
- | Teacher (Laguna XS.2, fp8) | 33B (3B active) | ~4.4 | _running_ |
61
- | [Stage-1 dense](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1) | ~3B | ~25 | 0.0% (0/164) |
62
- | **Stage-2 dense (this)** | ~3B | _tbd_ | **0.6%** (1/164) |
63
 
64
  _(HumanEval pass@1 via [evalplus](https://github.com/evalplus/evalplus), greedy. Teacher number being measured; will be filled in.)_
65
 
@@ -76,18 +76,18 @@ The KD moves the model directionally toward coherent code; it just needs far mor
76
 
77
  ## VRAM / footprint
78
 
79
- The MoE keeps all 33B params resident even though only ~3B are active per token; the dense model keeps only the ~3B.
80
 
81
  | Model | Params resident | bf16 weights | fp8 weights |
82
  |---|---|---|---|
83
- | XS.2 (33B MoE) | 33.4 B | ~67 GB | ~34 GB |
84
- | **XS.2-dense (stage_2)** | ~3.0 B | **~6 GB** | ~3 GB |
85
 
86
- → **~11× smaller weight footprint** (~61 GB saved, bf16). Caveats: attention is unchanged, so **KV-cache memory is identical** to the teacher (all savings are in the weights); per-token **FLOPs are ~unchanged** (the MoE was already ~3B-active) — the win is **memory/deployability**, not speed. XS.2 needs an 80 GB-class GPU (or fp8 on 48 GB); the dense model fits a 16 GB consumer GPU. _(The exported checkpoint here is zero-padded to ~3.8B / 7.7 GB for stock-modeling compat; the true model is 3.0B / ~6 GB.)_
87
 
88
  ## Limitations & next steps
89
 
90
- - **Severely under-trained.** ~14M KD tokens is 20–400× below typical recovery budgets (RADLADS used 250–700M; MoE→dense work ~4B). Expect near-zero on real coding tasks.
91
  - **Next:** extend Stage 2 substantially — ideally with **cached teacher top-K logits** to remove the teacher forward from the loop (3–5× throughput), reaching 50–100M+ tokens. Then run the paper-faithful agentic evals (SWE-bench / Terminal-Bench via Harbor).
92
 
93
  Code: https://github.com/postscarcity-inc/laguna-xs.2-dense · [Stage-1 model](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1)
 
12
 
13
  # Laguna-XS.2-dense
14
 
15
+ A **≈3B dense** model distilled from **[poolside/Laguna-XS.2](https://huggingface.co/poolside/Laguna-XS.2)** — a 33B Mixture-of-Experts coding model with a ≈3B active path. We replace the MoE feed-forward layers with a single **dense** FFN of the same size as the active path (8 routed + 1 shared expert), turning the ≈3B *active* compute into a genuine ≈3B *dense* model that keeps XS.2's attention.
16
 
17
  > ⚠️ **Research / hackathon artifact — heavily under-trained.** This checkpoint was produced in a time-boxed hackathon with a tiny distillation budget. It is **not** production-ready: generations are still degenerate/repetitive (see below). It demonstrates the *method* and is a starting point for longer distillation.
18
 
 
20
 
21
  Two stages, both distilling from the frozen FP8 XS.2 teacher:
22
 
23
+ 1. **[Stage 1](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1) — per-layer MoE→dense init** (RADLADS-style). Each of the 39 sparse MoE blocks is replaced by a dense SwiGLU FFN (intermediate 4608) and trained *independently, in parallel* to match the teacher MoE block's output (NMSE on the residual contribution), fed the teacher's own hidden states (no error compounding). ≈90M tokens. Result: a dense init with held-out perplexity **≈25** (vs teacher **≈4.4**) — functional but rough, because cross-layer error compounding is left uncorrected by design.
24
+ 2. **Stage 2 — synchronous logit-KD** (this model). The stitched ≈3B dense student is trained **end-to-end** against the fp8 teacher's full-vocab logits (forward-KL), on a 50/50 code+general corpus. ≈14M tokens, one H100, **KL 2.5 → 1.40**.
25
 
26
  Data: 50% [DCLM](https://huggingface.co/datasets/mlfoundations/dclm-baseline-1.0) + 50% [StarCoder2/the-stack-v2-train](https://huggingface.co/datasets/bigcode/the-stack-v2-train-smol-ids).
27
 
 
35
 
36
  ## Architecture / loading note
37
 
38
+ The dense FFNs are intermediate 4608 (layers 1–39) and 8192 (layer 0). For a uniform config that loads with the **stock** `modeling_laguna.py`, the 4608 FFNs are **zero-padded to 8192** (numerically identical — `silu(0)·0 = 0`). So the exported checkpoint reports ≈3.8B params (padded); the true model is ≈3.0B.
39
 
40
  ```python
41
  import torch
 
57
 
58
  | Model | Params (resident) | Held-out PPL | HumanEval pass@1 |
59
  |---|---|---|---|
60
+ | Teacher (Laguna XS.2, fp8) | 33B (3B active) | ≈4.4 | _running_ |
61
+ | [Stage-1 dense](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1) | ≈3B | ≈25 | 0.0% (0/164) |
62
+ | **Stage-2 dense (this)** | ≈3B | _tbd_ | **0.6%** (1/164) |
63
 
64
  _(HumanEval pass@1 via [evalplus](https://github.com/evalplus/evalplus), greedy. Teacher number being measured; will be filled in.)_
65
 
 
76
 
77
  ## VRAM / footprint
78
 
79
+ The MoE keeps all 33B params resident even though only ≈3B are active per token; the dense model keeps only the ≈3B.
80
 
81
  | Model | Params resident | bf16 weights | fp8 weights |
82
  |---|---|---|---|
83
+ | XS.2 (33B MoE) | 33.4 B | ≈67 GB | ≈34 GB |
84
+ | **XS.2-dense (stage_2)** | ≈3.0 B | **≈6 GB** | ≈3 GB |
85
 
86
+ → **≈11× smaller weight footprint** (≈61 GB saved, bf16). Caveats: attention is unchanged, so **KV-cache memory is identical** to the teacher (all savings are in the weights); per-token **FLOPs are ≈unchanged** (the MoE was already ≈3B-active) — the win is **memory/deployability**, not speed. XS.2 needs an 80 GB-class GPU (or fp8 on 48 GB); the dense model fits a 16 GB consumer GPU. _(The exported checkpoint here is zero-padded to ≈3.8B / 7.7 GB for stock-modeling compat; the true model is 3.0B / ≈6 GB.)_
87
 
88
  ## Limitations & next steps
89
 
90
+ - **Severely under-trained.** ≈14M KD tokens is 20–400× below typical recovery budgets (RADLADS used 250–700M; MoE→dense work ≈4B). Expect near-zero on real coding tasks.
91
  - **Next:** extend Stage 2 substantially — ideally with **cached teacher top-K logits** to remove the teacher forward from the loop (3–5× throughput), reaching 50–100M+ tokens. Then run the paper-faithful agentic evals (SWE-bench / Terminal-Bench via Harbor).
92
 
93
  Code: https://github.com/postscarcity-inc/laguna-xs.2-dense · [Stage-1 model](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1)