ftajwar commited on
Commit
6096ac4
·
verified ·
1 Parent(s): 57e4918

Add d24 0.75B base pretrained on ~100B ClimbMix tokens (Megatron->HF export of iter_0095368)

Browse files
README.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ library_name: transformers
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - pretrained
9
+ - base-model
10
+ - llama
11
+ - climbmix
12
+ datasets:
13
+ - karpathy/climbmix-400b-shuffle
14
+ ---
15
+
16
+ # d24-climbmix-100b
17
+
18
+ A **0.75B-parameter** dense decoder-only language model ("d24", nanochat depth-24 shape),
19
+ **pretrained from scratch on ~100B tokens** of [ClimbMix](https://huggingface.co/datasets/karpathy/climbmix-400b-shuffle).
20
+
21
+ This is a **base / foundation model** — it is *not* instruction-tuned and has no chat template.
22
+ Use it for continued pretraining, mid-training, SFT, or few-shot/raw text-completion experiments.
23
+
24
+ ## Architecture
25
+
26
+ | | |
27
+ |---|---|
28
+ | Class | `LlamaForCausalLM` (SwiGLU / RoPE / RMSNorm) |
29
+ | Parameters | 756,819,456 (~0.75B) |
30
+ | Layers | 24 |
31
+ | Hidden size | 1536 |
32
+ | Attention heads | 12 (head_dim 128, no GQA: 12 KV heads) |
33
+ | FFN hidden | 4096 (gated SwiGLU) |
34
+ | Context length | 2048 |
35
+ | Vocab | 50304 (GPT-2 BPE, 50257 padded to a multiple of 128) |
36
+ | Tied embeddings | yes |
37
+ | dtype | bf16 |
38
+ | Tokenizer | GPT-2 (`<\|endoftext\|>` = id 50256 as bos/eos) |
39
+
40
+ ## Training
41
+
42
+ - **Data:** ClimbMix (`karpathy/climbmix-400b-shuffle`), tokenized to GPT-2 bin/idx.
43
+ - **Tokens:** 95,368 iters × global batch 512 × seq 2048 ≈ **100B tokens**.
44
+ - **Optimizer:** cosine LR 3e-4 → 3e-5, warmup 100, AdamW.
45
+ - **Hardware:** ALCF Polaris, 256× A100 (64 nodes × 4 GPUs), pure data-parallel (TP=PP=1).
46
+ - **Framework:** NVIDIA NeMo / Megatron-Bridge (`nemo:26.04`); exported Megatron → HF with `convert/megatron_to_hf`.
47
+
48
+ ## Usage
49
+
50
+ ```python
51
+ import torch
52
+ from transformers import AutoModelForCausalLM, AutoTokenizer
53
+
54
+ tok = AutoTokenizer.from_pretrained("ftajwar/d24-climbmix-100b")
55
+ model = AutoModelForCausalLM.from_pretrained("ftajwar/d24-climbmix-100b", torch_dtype=torch.bfloat16)
56
+
57
+ prompt = "The capital of France is"
58
+ ids = tok(prompt, return_tensors="pt").input_ids
59
+ out = model.generate(ids, max_new_tokens=32, do_sample=False)
60
+ print(tok.decode(out[0], skip_special_tokens=True))
61
+ ```
config.json ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "LlamaForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 50256,
8
+ "dtype": "float32",
9
+ "eos_token_id": 50256,
10
+ "head_dim": 128,
11
+ "hidden_act": "silu",
12
+ "hidden_size": 1536,
13
+ "initializer_range": 0.02,
14
+ "intermediate_size": 4096,
15
+ "max_position_embeddings": 2048,
16
+ "mlp_bias": false,
17
+ "model_type": "llama",
18
+ "num_attention_heads": 12,
19
+ "num_hidden_layers": 24,
20
+ "num_key_value_heads": 12,
21
+ "pad_token_id": null,
22
+ "pretraining_tp": 1,
23
+ "rms_norm_eps": 1e-05,
24
+ "rope_parameters": {
25
+ "rope_theta": 10000.0,
26
+ "rope_type": "default"
27
+ },
28
+ "tie_word_embeddings": true,
29
+ "transformers_version": "5.3.0",
30
+ "use_cache": true,
31
+ "vocab_size": 50304
32
+ }
generation_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 50256,
4
+ "eos_token_id": 50256,
5
+ "output_attentions": false,
6
+ "output_hidden_states": false,
7
+ "transformers_version": "5.3.0",
8
+ "use_cache": true
9
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:33993caef8b170f1876eaeb6d354c5df7ac88fbf8b204c6e20976ea4a3fe2534
3
+ size 1513663904
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": "<|endoftext|>",
5
+ "cls_token": "<|endoftext|>",
6
+ "eos_token": "<|endoftext|>",
7
+ "errors": "replace",
8
+ "is_local": true,
9
+ "model_max_length": 1024,
10
+ "pad_token": null,
11
+ "sep_token": "<|endoftext|>",
12
+ "tokenizer_class": "GPT2Tokenizer",
13
+ "unk_token": "<|endoftext|>"
14
+ }