Upload tiny-lm LLaMA 3 Swallow Code model
Browse files
README.md
CHANGED
|
@@ -1,23 +1,52 @@
|
|
| 1 |
---
|
| 2 |
-
language:
|
|
|
|
|
|
|
| 3 |
tags:
|
| 4 |
- llama3
|
| 5 |
- swallow-code
|
| 6 |
- tiny-lm
|
|
|
|
| 7 |
datasets:
|
| 8 |
-
-
|
| 9 |
---
|
| 10 |
|
| 11 |
# Swallow Code Ibis-16 (LLaMA 3, 16 layers, 8k vocab)
|
| 12 |
|
| 13 |
This model was trained with the
|
| 14 |
[tiny-lm](https://github.com/ferjorosa/tiny-lm) repository on the
|
| 15 |
-
[
|
| 16 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 17 |
|
| 18 |
## Architecture
|
| 19 |
|
| 20 |
- LLaMA 3 style decoder-only transformer
|
|
|
|
| 21 |
- Layers: 16
|
| 22 |
- Vocab size: 8192
|
| 23 |
- Context length: 1024
|
|
|
|
| 1 |
---
|
| 2 |
+
language:
|
| 3 |
+
- en
|
| 4 |
+
- ja
|
| 5 |
tags:
|
| 6 |
- llama3
|
| 7 |
- swallow-code
|
| 8 |
- tiny-lm
|
| 9 |
+
- code
|
| 10 |
datasets:
|
| 11 |
+
- tokyotech-llm/swallow-code
|
| 12 |
---
|
| 13 |
|
| 14 |
# Swallow Code Ibis-16 (LLaMA 3, 16 layers, 8k vocab)
|
| 15 |
|
| 16 |
This model was trained with the
|
| 17 |
[tiny-lm](https://github.com/ferjorosa/tiny-lm) repository on the
|
| 18 |
+
[SwallowCode dataset](https://huggingface.co/datasets/tokyotech-llm/swallow-code).
|
| 19 |
+
|
| 20 |
+
The goal is educational: a compact pretraining run for studying tokenization,
|
| 21 |
+
data pipelines, and transformer training end to end.
|
| 22 |
+
|
| 23 |
+
## Data
|
| 24 |
+
|
| 25 |
+
The tokenizer is a custom 8k-token BPE trained on SwallowCode with Karpathy's
|
| 26 |
+
[`rustbpe`](https://github.com/karpathy/rustbpe) approach, then exported as a
|
| 27 |
+
`tiktoken` encoding for inference.
|
| 28 |
+
|
| 29 |
+
The dataset was split into 99% train and 1% validation before tokenization.
|
| 30 |
+
The resulting tokenized files contain 25.95B train tokens and 262M validation
|
| 31 |
+
tokens. Training used 1,024-token contiguous windows over the token stream.
|
| 32 |
+
|
| 33 |
+
## Training
|
| 34 |
+
|
| 35 |
+
The model was trained with PyTorch Lightning using bf16 mixed precision.
|
| 36 |
+
|
| 37 |
+
- Context window: 1,024 tokens
|
| 38 |
+
- Batch size: 64 sequences
|
| 39 |
+
- Gradient accumulation: 4
|
| 40 |
+
- Effective batch size: 262,144 tokens per optimizer step
|
| 41 |
+
- Training budget: about 25B tokens over 95,350 optimizer steps
|
| 42 |
+
- Optimizer: AdamW with cosine LR decay and 1% warmup
|
| 43 |
+
- Peak LR: 6e-4
|
| 44 |
+
- Weight decay: 0.1
|
| 45 |
|
| 46 |
## Architecture
|
| 47 |
|
| 48 |
- LLaMA 3 style decoder-only transformer
|
| 49 |
+
- Parameters: 18,886,912 total; 16,789,760 non-embedding
|
| 50 |
- Layers: 16
|
| 51 |
- Vocab size: 8192
|
| 52 |
- Context length: 1024
|