duico commited on
Commit
daf8bea
·
verified ·
1 Parent(s): 852aaa4

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +11 -12
README.md CHANGED
@@ -55,24 +55,23 @@ print(tok.decode(m.generate(**ids, max_new_tokens=64)[0], skip_special_tokens=Tr
55
 
56
  ## Results
57
 
58
- | Model | Params (resident) | Held-out PPL | HumanEval pass@1 |
59
- |---|---|---|---|
60
- | Teacher (Laguna XS.2, fp8) | 33B (3B active) | ≈4.4 | _running_ |
61
- | [Stage-1 dense](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1) | ≈3B | ≈25 | 0.0% (0/164) |
62
- | **Stage-2 dense (this)** | ≈3B | _tbd_ | **0.6%** (1/164) |
63
 
64
- _(HumanEval pass@1 via [evalplus](https://github.com/evalplus/evalplus), greedy. Teacher number being measured; will be filled in.)_
 
 
 
 
65
 
66
- Both dense checkpoints are **near-zero on HumanEval** — expected at this tiny distillation budget; the model is **not yet usable** for code generation. What this run demonstrates is the **architecture + footprint** (below) and a **directional** improvement (Stage-1 0.0% → Stage-2 0.6%, KL 2.5 → 1.40). Closing the task-accuracy gap needs much more KD.
67
 
68
- ### Sample generations
69
 
70
- On a simple completion prompt — `def fibonacci(n): """Return the n-th Fibonacci number."""` (greedy) — both models are still degenerate, but Stage-2 is visibly less broken:
71
 
72
- - **Stage-1:** non-Pythonic repetition — `return n;` repeated indefinitely.
73
- - **Stage-2:** *valid* Python with the correct function structure, but stuck in a loop — `return 0` repeated.
74
 
75
- The KD moves the model directionally toward coherent code; it just needs far more than 14M tokens.
76
 
77
  ## VRAM / footprint
78
 
 
55
 
56
  ## Results
57
 
58
+ HumanEval pass@1 (greedy, [evalplus](https://github.com/evalplus/evalplus)), base / plus:
 
 
 
 
59
 
60
+ | Model | Params (resident) | PPL | HumanEval (raw completion) | HumanEval (chat template) |
61
+ |---|---|---|---|---|
62
+ | Teacher (Laguna XS.2, fp8) | 33B (3B active) | ≈4.4 | — | **88.4% / 84.8%** |
63
+ | [Stage-1 dense](https://huggingface.co/poolside-laguna-hackathon/laguna-xs2-dense-stage1) | ≈3B | ≈25 | 0.0% | 0.0% |
64
+ | **Stage-2 dense (this)** | ≈3B | — | 0.6% | 0.0% |
65
 
66
+ ## Diagnosis & next steps (honest)
67
 
68
+ The dense models are **near-zero on HumanEval in *both* eval formats** — i.e. genuinely **not yet usable**, not just an eval artifact. The teacher scores a normal **88.4%** with its chat template, so the harness is sound and the eval format matters a lot (raw completion is off-distribution for this instruct model).
69
 
70
+ **Root cause:** XS.2 is an **instruct/agentic** model (chat template, special tokens, EOS-terminated turns), but we distilled it on **raw concatenated pretraining text** (DCLM + code, packed). That pushed the student off its instruct distribution and it never learned to *stop* — so generations degenerate (Stage-1: `return n;` repeated; Stage-2 in chat mode: control-token spam `</think>…</assistant>`). The ≈14M-token KD budget is also 20–400× below typical recovery budgets.
71
 
72
+ **Fix (the real next step):** distill in the model's **native chat format** — coding instruction→response conversations rendered through `chat_template.jinja`, EOS-terminated, ideally with teacher-generated responses (so we match the teacher's actual operating distribution). Same KD loss/loop; only the data + tokenization change. More *raw* tokens would not fix this.
 
73
 
74
+ What this release **does** demonstrate: the MoE→dense **architecture** works (per-layer init converges, see the NMSE curves) and the **≈11× weight-VRAM reduction** (below) at matched active-compute. Task accuracy awaits the chat-format KD run.
75
 
76
  ## VRAM / footprint
77