coderian commited on
Commit
6d8af17
·
verified ·
1 Parent(s): 7bb4a30

QraXAi modeli eklendi

Browse files
LICENSE ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ MIT License
2
+
3
+ Copyright (c) 2026 coderian
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
README.md ADDED
@@ -0,0 +1,163 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ library_name: transformers
3
+ pipeline_tag: text-generation
4
+ language:
5
+ - en
6
+ datasets:
7
+ - roneneldan/TinyStories
8
+ tags:
9
+ - qraxai
10
+ - gpt2-tokenizer
11
+ - causal-lm
12
+ - tiny-stories
13
+ - from-scratch
14
+ license: mit
15
+ ---
16
+
17
+ # QraXAi-Basic-45M
18
+
19
+ A small, decoder-only Transformer (~44.75M parameters) trained **from scratch** on
20
+ [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) with the GPT-2 BPE
21
+ tokenizer (50,257 tokens). QraXAi is a hand-written PyTorch GPT implementation
22
+ (`model.py` / `configuration_qraxai.py`, shipped with this repo), not a fine-tune of GPT-2 —
23
+ only the tokenizer is shared with GPT-2.
24
+
25
+ This is an experimental research model, intended for learning/demo purposes.
26
+
27
+ ## Model details
28
+
29
+ | | |
30
+ |---|---|
31
+ | Architecture | Decoder-only Transformer (GPT-style), pre-norm |
32
+ | Parameters | **44,751,872** (~44.75M), all trainable |
33
+ | Layers | 24 |
34
+ | Hidden size | 256 |
35
+ | Attention heads | 8 (head dim 32) |
36
+ | Feed-forward | 4× hidden, GELU |
37
+ | Context length | 256 tokens (hard limit) |
38
+ | Vocabulary | 50,257 (GPT-2 BPE) |
39
+ | Position encoding | Learned absolute embeddings |
40
+ | Normalization | LayerNorm |
41
+ | Weight tying | No (`lm_head` is separate) |
42
+ | KV cache | No — generation recomputes the full context at every step |
43
+ | Weights | fp32, 179 MB (`model.safetensors`) |
44
+ | Special tokens | `bos = eos = <\|endoftext\|>` (id 50256) |
45
+ | Custom code | Yes — requires `trust_remote_code=True` |
46
+
47
+ Parameter breakdown: token embeddings 12.87M + position embeddings 0.07M +
48
+ 24 × 0.79M transformer blocks (18.95M) + final norm + untied `lm_head` 12.87M.
49
+
50
+ ## Quick start
51
+
52
+ ```python
53
+ import torch
54
+ from transformers import AutoModelForCausalLM, AutoTokenizer
55
+
56
+ repo_id = "coderian/QraXAi-Basic-45M"
57
+
58
+ tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
59
+ model = AutoModelForCausalLM.from_pretrained(
60
+ repo_id,
61
+ trust_remote_code=True,
62
+ dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
63
+ ).to("cuda" if torch.cuda.is_available() else "cpu").eval()
64
+
65
+ prompt = "Once upon a time, there was a little girl named Lily"
66
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
67
+
68
+ with torch.inference_mode():
69
+ output = model.generate(
70
+ **inputs,
71
+ max_new_tokens=244, # prompt + new tokens must stay <= 256
72
+ do_sample=True,
73
+ temperature=0.8,
74
+ top_k=50,
75
+ top_p=0.95,
76
+ repetition_penalty=1.1,
77
+ pad_token_id=tokenizer.eos_token_id,
78
+ eos_token_id=tokenizer.eos_token_id,
79
+ )
80
+
81
+ print(tokenizer.decode(output[0], skip_special_tokens=True))
82
+ ```
83
+
84
+ Notes for generation:
85
+ - **Context is a hard limit of 256 tokens.** The custom `forward` raises `ValueError`
86
+ if the sequence gets longer, so keep `len(prompt) + max_new_tokens <= 256`.
87
+ - The model has **no KV cache**: each new token re-runs the full context, so generation
88
+ cost grows quickly with sequence length.
89
+ - The custom `forward` does not use an attention mask for padding. Generate one prompt
90
+ at a time instead of batching.
91
+ - Generation usually stops at `<|endoftext|>`, but because the training data is a
92
+ continuous stream of stories, the model sometimes starts a new story instead.
93
+
94
+ ## Example output
95
+
96
+ Prompt: `Once upon a time, there was a little girl named Lily`
97
+ (`temperature=0.8, top_k=50, top_p=0.95, repetition_penalty=1.1, max_new_tokens=244`):
98
+
99
+ > Once upon a time, there was a little girl named Lily who loved to play in the big, green field. One day, she found a shiny stone on top of her backyard. She picked it up and showed it to her mom.
100
+ >
101
+ > "Look mommy, I found a pretty mineral!" said Lily excitedly. "It's very pretty!"
102
+ >
103
+ > Her mom smiled and said, "That's right, sweetie. It'll make sure you touch it. But remember, be careful with it because you might find something else inside."
104
+ >
105
+ > Lily nodded her head and kept playing with the jewel until she noticed that the box had fallen into a hole. She felt sad for her mom, but then remembered what her mom said about when something is hurt.
106
+ >
107
+ > The next day, Lily went back to the park and saw that the unknown stone was broken. She asked her mom if they could try and fix it. Her mom told her that it's okay to ask for help and that sometimes you can't use it without asking permission. So, Lily listened to her mom and never touched the stone again.
108
+ > Once upon a time, there was a boy named Timmy. He loved to play with his toy car. One day, he went to
109
+
110
+ (The last line is a new story the model began, cut off by the token limit.)
111
+
112
+ ## Training
113
+
114
+ | | |
115
+ |---|---|
116
+ | Dataset | `roneneldan/TinyStories` (train split, streaming), first 130,000 stories |
117
+ | Data size | 115.7M characters → 28.76M tokens → ~112,350 training blocks |
118
+ | Objective | Next-token prediction (causal LM), cross-entropy |
119
+ | Epochs | 1 (~7,000 optimizer steps) |
120
+ | Batch size | 16 |
121
+ | Block size | 256 |
122
+ | Optimizer | AdamW, lr 3e-4 |
123
+ | Gradient clipping | 1.0 |
124
+ | Mixed precision | bf16 (on CUDA) |
125
+ | Seed | 42 |
126
+
127
+ ## Files in this repository
128
+
129
+ | File | Description |
130
+ |---|---|
131
+ | `config.json` | Model config (`GPTConfig` + `auto_map` for remote code) |
132
+ | `model.safetensors` | fp32 weights (179 MB) |
133
+ | `model.py` | `QraXAiForCausalLM` — custom modeling code |
134
+ | `configuration_qraxai.py` | `GPTConfig` — custom configuration |
135
+ | `tokenizer.json`, `tokenizer_config.json` | GPT-2 BPE tokenizer |
136
+ | `generation_config.json` | Default generation settings |
137
+ | `LICENSE` | MIT License |
138
+
139
+ ## Limitations
140
+
141
+ - English only; trained exclusively on synthetic children's stories (TinyStories), so it
142
+ knows little about the real world.
143
+ - Trained for a single epoch: grammar is mostly coherent, but content can be repetitive,
144
+ inconsistent or nonsensical.
145
+ - Hard 256-token context; no sliding window, so long inputs must be truncated.
146
+ - Not instruction-tuned — it does not follow instructions and is not a chat model.
147
+ - No safety filtering or alignment of any kind. Do not use in production or for
148
+ user-facing applications.
149
+ - TinyStories is synthetic data; nothing prevents the model from producing odd or
150
+ inappropriate continuations.
151
+
152
+ ## Acknowledgements
153
+
154
+ - Dataset and idea: **TinyStories: How Small Can Language Models Be and Still Speak
155
+ Coherent English?** — Ronen Eldan and Yuanzhi Li, 2023 ([arXiv:2305.07759](https://arxiv.org/abs/2305.07759)).
156
+ - Tokenizer: GPT-2 BPE (`gpt2`).
157
+ - Built with PyTorch and Hugging Face Transformers.
158
+
159
+ ## License
160
+
161
+ Released under the [MIT License](LICENSE). Note that the training dataset
162
+ ([TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)) has its own terms;
163
+ check them if you plan to redistribute the data or use the model commercially.
config.json ADDED
@@ -0,0 +1,20 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "QraXAiForCausalLM"
4
+ ],
5
+ "auto_map": {
6
+ "AutoConfig": "configuration_qraxai.GPTConfig",
7
+ "AutoModelForCausalLM": "model.QraXAiForCausalLM"
8
+ },
9
+ "bos_token_id": 50256,
10
+ "dtype": "float32",
11
+ "embed_dim": 256,
12
+ "eos_token_id": 50256,
13
+ "max_seq_len": 256,
14
+ "model_type": "qrax_ai",
15
+ "n_layers": 24,
16
+ "tie_word_embeddings": false,
17
+ "transformers_version": "5.17.0",
18
+ "use_cache": false,
19
+ "vocab_size": 50257
20
+ }
configuration_qraxai.py ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from transformers import PretrainedConfig
2
+
3
+ class GPTConfig(PretrainedConfig):
4
+
5
+ model_type = "qrax_ai"
6
+
7
+ def __init__(
8
+ self,
9
+ vocab_size=10000,
10
+ n_layers=6,
11
+ max_seq_len=512,
12
+ embed_dim=256,
13
+ use_cache=False,
14
+ **kwargs
15
+ ):
16
+ super().__init__(**kwargs)
17
+
18
+ self.use_cache = use_cache
19
+ self.vocab_size = vocab_size
20
+ self.n_layers = n_layers
21
+ self.max_seq_len = max_seq_len
22
+ self.embed_dim = embed_dim
generation_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 50256,
4
+ "eos_token_id": 50256,
5
+ "output_attentions": false,
6
+ "output_hidden_states": false,
7
+ "transformers_version": "5.17.0",
8
+ "use_cache": false
9
+ }
model.py ADDED
@@ -0,0 +1,243 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ from transformers import GenerationMixin, PreTrainedModel
2
+ from transformers.modeling_outputs import CausalLMOutput
3
+ import torch.nn as nn
4
+ import torch
5
+ import math
6
+
7
+ try:
8
+ from .configuration_qraxai import GPTConfig
9
+ except ImportError:
10
+ from configuration_qraxai import GPTConfig
11
+
12
+ class CausalSelfAttention(nn.Module):
13
+
14
+ def __init__(self, embed_dim, num_heads):
15
+ super().__init__()
16
+
17
+ assert embed_dim % num_heads == 0, \
18
+ "embed_dim must be divisible by num_heads"
19
+
20
+ self.embed_dim = embed_dim
21
+ self.num_heads = num_heads
22
+ self.head_dim = embed_dim // num_heads
23
+
24
+ # Q, K, V projections
25
+ self.q_proj = nn.Linear(embed_dim, embed_dim)
26
+ self.k_proj = nn.Linear(embed_dim, embed_dim)
27
+ self.v_proj = nn.Linear(embed_dim, embed_dim)
28
+
29
+ self.o_proj = nn.Linear(embed_dim, embed_dim)
30
+
31
+ def forward(self, x):
32
+ batch_size, seq_len, embed_dim = x.shape
33
+
34
+ Q = self.q_proj(x)
35
+ V = self.v_proj(x)
36
+ K = self.k_proj(x)
37
+
38
+ Q = Q.view(
39
+ batch_size,
40
+ seq_len,
41
+ self.num_heads,
42
+ self.head_dim
43
+ )
44
+
45
+ K = K.view(
46
+ batch_size,
47
+ seq_len,
48
+ self.num_heads,
49
+ self.head_dim
50
+ )
51
+
52
+ V = V.view(
53
+ batch_size,
54
+ seq_len,
55
+ self.num_heads,
56
+ self.head_dim
57
+ )
58
+
59
+ Q = Q.transpose(1, 2)
60
+ K = K.transpose(1, 2)
61
+ V = V.transpose(1, 2)
62
+
63
+
64
+ scores = Q @ K.transpose(-2, -1)
65
+
66
+ scores = scores / math.sqrt(self.head_dim)
67
+
68
+ mask = torch.triu(
69
+ torch.ones(
70
+ seq_len,
71
+ seq_len,
72
+ device=x.device
73
+ ),
74
+ diagonal=1
75
+ ).bool()
76
+
77
+ scores = scores.masked_fill(
78
+ mask,
79
+ torch.finfo(scores.dtype).min
80
+ )
81
+
82
+ attention_w = torch.softmax(
83
+ scores,
84
+ dim=-1
85
+ )
86
+
87
+ output = attention_w @ V
88
+
89
+ output = output.transpose(1, 2)
90
+
91
+ output = output.contiguous().view(
92
+ batch_size,
93
+ seq_len,
94
+ embed_dim
95
+ )
96
+
97
+ # Final projection
98
+ output = self.o_proj(output)
99
+
100
+ return output
101
+
102
+
103
+ class TransformerBlock(nn.Module):
104
+
105
+ def __init__(
106
+ self,
107
+ embed_dim
108
+ ):
109
+ super().__init__()
110
+
111
+ self.ln1 = nn.LayerNorm(embed_dim)
112
+
113
+ self.attention = CausalSelfAttention(embed_dim, 8)
114
+
115
+ self.ln2 = nn.LayerNorm(embed_dim)
116
+
117
+ # feed forward network
118
+ self.ffn = nn.Sequential(
119
+ nn.Linear(
120
+ in_features=embed_dim,
121
+ out_features=4*embed_dim
122
+ ),
123
+
124
+ nn.GELU(),
125
+
126
+ nn.Linear(
127
+ in_features=4*embed_dim,
128
+ out_features=embed_dim
129
+ )
130
+ )
131
+
132
+ def forward(self, x):
133
+ x = x + self.attention(
134
+ self.ln1(x)
135
+ )
136
+
137
+ x = x + self.ffn(
138
+ self.ln2(x)
139
+ )
140
+
141
+ return x
142
+
143
+ class QraXAiForCausalLM(PreTrainedModel, GenerationMixin):
144
+
145
+ config_class = GPTConfig
146
+
147
+ def __init__(
148
+ self,
149
+ config
150
+ ):
151
+
152
+ super().__init__(config)
153
+
154
+ self.token_embedding = nn.Embedding(
155
+ config.vocab_size,
156
+ config.embed_dim
157
+ )
158
+
159
+ self.position_embedding = nn.Embedding(
160
+ config.max_seq_len,
161
+ config.embed_dim
162
+ )
163
+
164
+ self.transformer_blocks = nn.ModuleList([
165
+ TransformerBlock(config.embed_dim)
166
+ for _ in range(config.n_layers)
167
+ ])
168
+
169
+ self.ln_f = nn.LayerNorm(
170
+ config.embed_dim
171
+ )
172
+
173
+
174
+ self.lm_head = nn.Linear(
175
+ config.embed_dim,
176
+ config.vocab_size,
177
+ bias=False
178
+ )
179
+
180
+ self.post_init()
181
+
182
+ def forward(
183
+ self,
184
+ input_ids,
185
+ labels=None,
186
+ **kwargs
187
+ ):
188
+
189
+ batch_size, seq_len = input_ids.shape
190
+
191
+ if seq_len > self.config.max_seq_len:
192
+ raise ValueError(
193
+ f"Sequence length ({seq_len}) "
194
+ f"cannot be greater than "
195
+ f"max_seq_len ({self.config.max_seq_len})"
196
+ )
197
+
198
+ positions = torch.arange(
199
+ seq_len,
200
+ device=input_ids.device
201
+ )
202
+
203
+ token_emb = self.token_embedding(
204
+ input_ids
205
+ )
206
+
207
+ pos_emb = self.position_embedding(
208
+ positions
209
+ )
210
+
211
+ x = token_emb + pos_emb
212
+
213
+ for block in self.transformer_blocks:
214
+ x = block(x)
215
+
216
+ x = self.ln_f(x)
217
+
218
+ logits = self.lm_head(x)
219
+
220
+ loss = None
221
+
222
+ if labels is not None:
223
+
224
+ shift_logits = logits[
225
+ :, :-1, :
226
+ ].contiguous()
227
+
228
+ shift_labels = labels[
229
+ :, 1:
230
+ ].contiguous()
231
+
232
+ loss = nn.functional.cross_entropy(
233
+ shift_logits.view(
234
+ -1,
235
+ shift_logits.size(-1)
236
+ ),
237
+ shift_labels.view(-1)
238
+ )
239
+
240
+ return CausalLMOutput(
241
+ loss=loss,
242
+ logits=logits
243
+ )
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e4bfeaa3c57234fef408a5c51c9463df89eb67ca3e3c8c7e7517d2d6d4bbd6de
3
+ size 179049920
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": "<|endoftext|>",
5
+ "eos_token": "<|endoftext|>",
6
+ "errors": "replace",
7
+ "is_local": false,
8
+ "local_files_only": false,
9
+ "model_max_length": 1024,
10
+ "pad_token": null,
11
+ "tokenizer_class": "GPT2Tokenizer",
12
+ "unk_token": "<|endoftext|>"
13
+ }