ferjorosa commited on
Commit
926314d
·
verified ·
1 Parent(s): 0ddb093

Upload tiny-lm LLaMA 3 Swallow Code model

Browse files
Files changed (1) hide show
  1. README.md +33 -4
README.md CHANGED
@@ -1,23 +1,52 @@
1
  ---
2
- language: en
 
 
3
  tags:
4
  - llama3
5
  - swallow-code
6
  - tiny-lm
 
7
  datasets:
8
- - nlp-waseda/EvalPlus-Merged-Filtered-Slim
9
  ---
10
 
11
  # Swallow Code Ibis-16 (LLaMA 3, 16 layers, 8k vocab)
12
 
13
  This model was trained with the
14
  [tiny-lm](https://github.com/ferjorosa/tiny-lm) repository on the
15
- [EvalPlus-Merged-Filtered-Slim dataset](https://huggingface.co/datasets/nlp-waseda/EvalPlus-Merged-Filtered-Slim)
16
- (see [Swallow Project](https://tokyotech-llm.github.io/swallow-llama/)).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
17
 
18
  ## Architecture
19
 
20
  - LLaMA 3 style decoder-only transformer
 
21
  - Layers: 16
22
  - Vocab size: 8192
23
  - Context length: 1024
 
1
  ---
2
+ language:
3
+ - en
4
+ - ja
5
  tags:
6
  - llama3
7
  - swallow-code
8
  - tiny-lm
9
+ - code
10
  datasets:
11
+ - tokyotech-llm/swallow-code
12
  ---
13
 
14
  # Swallow Code Ibis-16 (LLaMA 3, 16 layers, 8k vocab)
15
 
16
  This model was trained with the
17
  [tiny-lm](https://github.com/ferjorosa/tiny-lm) repository on the
18
+ [SwallowCode dataset](https://huggingface.co/datasets/tokyotech-llm/swallow-code).
19
+
20
+ The goal is educational: a compact pretraining run for studying tokenization,
21
+ data pipelines, and transformer training end to end.
22
+
23
+ ## Data
24
+
25
+ The tokenizer is a custom 8k-token BPE trained on SwallowCode with Karpathy's
26
+ [`rustbpe`](https://github.com/karpathy/rustbpe) approach, then exported as a
27
+ `tiktoken` encoding for inference.
28
+
29
+ The dataset was split into 99% train and 1% validation before tokenization.
30
+ The resulting tokenized files contain 25.95B train tokens and 262M validation
31
+ tokens. Training used 1,024-token contiguous windows over the token stream.
32
+
33
+ ## Training
34
+
35
+ The model was trained with PyTorch Lightning using bf16 mixed precision.
36
+
37
+ - Context window: 1,024 tokens
38
+ - Batch size: 64 sequences
39
+ - Gradient accumulation: 4
40
+ - Effective batch size: 262,144 tokens per optimizer step
41
+ - Training budget: about 25B tokens over 95,350 optimizer steps
42
+ - Optimizer: AdamW with cosine LR decay and 1% warmup
43
+ - Peak LR: 6e-4
44
+ - Weight decay: 0.1
45
 
46
  ## Architecture
47
 
48
  - LLaMA 3 style decoder-only transformer
49
+ - Parameters: 18,886,912 total; 16,789,760 non-embedding
50
  - Layers: 16
51
  - Vocab size: 8192
52
  - Context length: 1024