sfanm commited on
Commit
a9a7a62
·
verified ·
1 Parent(s): 3595dc0

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +31 -0
README.md ADDED
@@ -0,0 +1,31 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ language: en
4
+ library_name: transformers
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - nanochat
8
+ - nemotron
9
+ - from-scratch
10
+ - perlmutter
11
+ - gpt2-tokenizer
12
+ ---
13
+ # d24-sft-v1base-olmo3-2.3B
14
+
15
+ v1-base SFT chat model, OLMo-3 Dolmino-style midtrain.
16
+
17
+ nanochat-style **depth-24** decoder — 24 layers × 1536 hidden × 12 heads, SwiGLU / RoPE / RMSNorm, tied embeddings, GPT-2 BPE vocab (50304), **0.757B params**, 2048-token context.
18
+
19
+ **Lineage.** v1 pretrain (5.84B ClimbMix) → OLMo-3 Dolmino-style midtrain (2.3B corpus, 20 components incl. instruction/QA) → SFT (nanochat mix).
20
+
21
+ **Metrics.** GSM8K (greedy, full 1319): **4.93%** · SFT val lm-loss 0.222 (overfits SFT train via format familiarity).
22
+
23
+ ## Load
24
+ ```python
25
+ from transformers import AutoModelForCausalLM, AutoTokenizer
26
+ m = "sfanm/d24-sft-v1base-olmo3-2.3B"
27
+ tok = AutoTokenizer.from_pretrained(m)
28
+ model = AutoModelForCausalLM.from_pretrained(m, torch_dtype="bfloat16")
29
+ ```
30
+
31
+ *Research checkpoint from a from-scratch nanochat-d24 replication (pretrain → midtrain → SFT → RL) on NERSC Perlmutter. Trained on third-party corpora (ClimbMix, FineMath, OpenMath, MetaMath, OpenThoughts, OLMo-3 Dolmino, SmolTalk, …) — see those datasets' licenses; provided as-is for research.*