superAVTR commited on
Commit
6498813
·
verified ·
1 Parent(s): 4afe8c2

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +139 -0
README.md CHANGED
@@ -1,3 +1,142 @@
1
  ---
 
 
 
 
 
 
 
 
 
 
2
  license: apache-2.0
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
+ library_name: transformers
3
+ pipeline_tag: text-generation
4
+ language:
5
+ - en
6
+ tags:
7
+ - llama
8
+ - causal-lm
9
+ - text-generation
10
+ - experimental
11
+ - pytorch
12
  license: apache-2.0
13
  ---
14
+
15
+ # Tinizong-50M
16
+
17
+ Tinizong-50M is a small experimental decoder-only language model developed from scratch as part of the Tinizong LLM project.
18
+
19
+ The model was created primarily as a research and engineering platform for exploring language-model architecture, tokenizer design, dataset construction, pretraining, scaling, and instruction tuning.
20
+
21
+ This release contains a Hugging Face-compatible conversion of the native Tinizong checkpoint.
22
+
23
+ It is a **base language model**, not an instruction-tuned or chat model.
24
+
25
+ ## Model details
26
+
27
+ | Property | Value |
28
+ |---|---|
29
+ | Architecture | Decoder-only Transformer, Llama-compatible |
30
+ | Parameters | ~46M |
31
+ | Vocabulary | 16,000 tokens |
32
+ | Tokenizer | SentencePiece BPE |
33
+ | Context length | 1,024 tokens |
34
+ | Hidden size | 512 |
35
+ | Transformer layers | 12 |
36
+ | Attention heads | 8 |
37
+ | Attention head dimension | 64 |
38
+ | MLP intermediate size | 1,365 |
39
+ | Positional encoding | RoPE |
40
+ | RoPE theta | 10,000 |
41
+ | Normalization | RMSNorm |
42
+ | MLP | SwiGLU |
43
+ | Weight tying | Input embeddings / LM head |
44
+ | Attention | Multi-head causal self-attention |
45
+
46
+ The native implementation was written directly in PyTorch and later converted to the Hugging Face `LlamaForCausalLM` architecture.
47
+
48
+ This checkpoint corresponds to the **v15 Wikipedia + synthetic-data training experiment**.
49
+
50
+ ## Training
51
+
52
+ Tinizong-50M was trained as a causal language model using next-token prediction.
53
+
54
+ Training data included a mixture of:
55
+
56
+ - Wikipedia-derived text
57
+ - synthetic factual / educational text
58
+ - other experimental pretraining material used during development
59
+
60
+ The training corpus and methodology evolved during the project, so this release should be considered an experimental research checkpoint rather than a fully documented production model.
61
+
62
+ The model uses a custom 16K SentencePiece tokenizer with byte fallback.
63
+
64
+ ## Hugging Face conversion
65
+
66
+ The original Tinizong model uses a compact PyTorch implementation with:
67
+
68
+ - combined Q/K/V projection
69
+ - RMSNorm
70
+ - rotary position embeddings
71
+ - SwiGLU feed-forward layers
72
+ - tied input/output embeddings
73
+
74
+ The trained weights were mapped into Hugging Face's `LlamaForCausalLM` representation.
75
+
76
+ The combined native QKV projection was split into Hugging Face `q_proj`, `k_proj`, and `v_proj` tensors. Other layers were mapped directly to their corresponding Llama components.
77
+
78
+ ### Numerical validation
79
+
80
+ The native model and the converted Hugging Face model were evaluated with identical input token IDs.
81
+
82
+ For the released checkpoint:
83
+
84
+ - Maximum absolute logit difference: approximately `9.54e-6`
85
+ - Mean absolute logit difference: approximately `1.14e-6`
86
+ - Logit cosine similarity: `1.0000000000`
87
+ - Next-token argmax: identical
88
+ - Top-token ordering: identical in the validation test
89
+
90
+ Autoregressive generation was also tested using the same random seed and sampling configuration.
91
+
92
+ The native implementation and Hugging Face implementation produced an **exact token-for-token match** using:
93
+
94
+ - temperature: `0.5`
95
+ - top-k: `30`
96
+ - top-p: `1.0`
97
+
98
+ Hugging Face generation with KV caching was also verified against full-context recomputation.
99
+
100
+ ## Usage
101
+
102
+ ```python
103
+ import torch
104
+ from transformers import AutoTokenizer, AutoModelForCausalLM
105
+
106
+ model_id = "YOUR-HF-USERNAME/Tinizong-50M"
107
+
108
+ tokenizer = AutoTokenizer.from_pretrained(
109
+ model_id,
110
+ use_fast=False
111
+ )
112
+
113
+ model = AutoModelForCausalLM.from_pretrained(
114
+ model_id,
115
+ dtype=torch.float32
116
+ )
117
+
118
+ prompt = "The film was released in"
119
+
120
+ inputs = tokenizer(
121
+ prompt,
122
+ return_tensors="pt",
123
+ add_special_tokens=False
124
+ )
125
+
126
+ with torch.no_grad():
127
+ output = model.generate(
128
+ **inputs,
129
+ max_new_tokens=60,
130
+ do_sample=True,
131
+ temperature=0.5,
132
+ top_k=30,
133
+ top_p=1.0,
134
+ use_cache=True
135
+ )
136
+
137
+ print(
138
+ tokenizer.decode(
139
+ output[0],
140
+ skip_special_tokens=False
141
+ )
142
+ )