EvanOLeary commited on
Commit
b08d20a
·
verified ·
1 Parent(s): 1d287e6

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +53 -0
README.md ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: EvanOLeary/laguna-xs2-dense-k8-cuda-sft
4
+ pipeline_tag: text-generation
5
+ library_name: transformers
6
+ tags: [moe-to-dense, densification, laguna, cuda, kernels, sft, code]
7
+ language: [en]
8
+ ---
9
+
10
+ # Laguna-XS.2 → Dense (K=8) · **CUDA-SFT-extended** (follow-up SFT)
11
+
12
+ A **~3.0 B dense** CUDA-kernel model — a **follow-up SFT** on top of
13
+ [laguna-xs2-dense-k8-cuda-sft](https://huggingface.co/EvanOLeary/laguna-xs2-dense-k8-cuda-sft),
14
+ trained on **more CUDA / C++ kernel code**.
15
+
16
+ ## Lineage
17
+ ```
18
+ poolside/Laguna-XS.2 (33B/3B-active MoE, 256 experts)
19
+ → densify (K=8 dense SwiGLU) → DO-ACP warm-start
20
+ → reconstruction-pretrain (kernel mixture, "V2")
21
+ → SFT (SakanaAI CUDA, level_1+2, 400 steps) = laguna-xs2-dense-k8-cuda-sft
22
+ → SFT-extended (level_1+2+3, +500 steps) = THIS MODEL
23
+ → RFT/GRPO (verifiable reward) = next (laguna-xs2-dense-k8-cuda-rft)
24
+ ```
25
+
26
+ ## Why a follow-up SFT (rationale)
27
+ The first SFT (400 steps, level_1+2) produced a model that emits working CUDA on **simple ops**
28
+ (ReLU/Tanh ~3/4 at pass@k) but showed two gaps:
29
+ - **Thin C++ idiom coverage** — it botches more involved C++/CUDA constructs (e.g. `float4*
30
+ v = float4* ptr;` instead of `reinterpret_cast<float4*>(ptr)`), so vectorized kernels fail to compile.
31
+ - **Limited CUDA breadth** — harder ops (Sigmoid/GeLU/Softmax) compile/verify inconsistently.
32
+
33
+ This follow-up **extends the SFT** with **more CUDA + C++ kernel data** (Sakana `level_1+2+3`, +500
34
+ steps from the previous checkpoint) to broaden C++/CUDA coverage before RL. It is the **mid checkpoint**
35
+ in a 3-way comparison: `SFT` → `SFT-extended` (this) → `SFT-extended-RFT`.
36
+
37
+ ## Training
38
+ | | |
39
+ |---|---|
40
+ | Base | `laguna-xs2-dense-k8-cuda-sft` (continued, not from scratch) |
41
+ | Data | `SakanaAI/AI-CUDA-Engineer-Archive` `level_1,level_2,level_3` (correct kernels), PyTorch→CUDA, chat-formatted, prompt masked |
42
+ | Objective | causal-LM cross-entropy on the CUDA completion only |
43
+ | Trainable | `routed_dense` + `lm_head` + norms (1.19 B) |
44
+ | Optimizer | AdamW 1e-5, grad-clip 1.0, grad-accum 8, seq 2048, **500 steps** |
45
+
46
+ ## Evaluation
47
+ Benchmarked 3-way (SFT / SFT-extended / SFT-extended-RFT) on **KernelBench-Lite L1** (10 elementwise
48
+ ops, **K=4**, subprocess-isolated compile+correctness vs PyTorch eager). Results table:
49
+ [github.com/Tyronita/laguna-dense-cuda-kernels](https://github.com/Tyronita/laguna-dense-cuda-kernels).
50
+
51
+ ## Intended use
52
+ Research base for **RFT** (RL on verified compile+correctness+speedup). Kernels are not verified at
53
+ generation time — compile & check before use, and isolate execution (a bad kernel corrupts the CUDA context).