Text Generation
Transformers
Safetensors
English
qrax_ai
qraxai
gpt2-tokenizer
causal-lm
tiny-stories
from-scratch
custom_code
Instructions to use coderian/QraXAi-Basic-45M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use coderian/QraXAi-Basic-45M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="coderian/QraXAi-Basic-45M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("coderian/QraXAi-Basic-45M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use coderian/QraXAi-Basic-45M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "coderian/QraXAi-Basic-45M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coderian/QraXAi-Basic-45M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/coderian/QraXAi-Basic-45M
- SGLang
How to use coderian/QraXAi-Basic-45M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "coderian/QraXAi-Basic-45M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coderian/QraXAi-Basic-45M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "coderian/QraXAi-Basic-45M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "coderian/QraXAi-Basic-45M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use coderian/QraXAi-Basic-45M with Docker Model Runner:
docker model run hf.co/coderian/QraXAi-Basic-45M
QraXAi modeli eklendi
Browse files- LICENSE +21 -0
- README.md +163 -0
- config.json +20 -0
- configuration_qraxai.py +22 -0
- generation_config.json +9 -0
- model.py +243 -0
- model.safetensors +3 -0
- tokenizer.json +0 -0
- tokenizer_config.json +13 -0
LICENSE
ADDED
|
@@ -0,0 +1,21 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
MIT License
|
| 2 |
+
|
| 3 |
+
Copyright (c) 2026 coderian
|
| 4 |
+
|
| 5 |
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
| 6 |
+
of this software and associated documentation files (the "Software"), to deal
|
| 7 |
+
in the Software without restriction, including without limitation the rights
|
| 8 |
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
| 9 |
+
copies of the Software, and to permit persons to whom the Software is
|
| 10 |
+
furnished to do so, subject to the following conditions:
|
| 11 |
+
|
| 12 |
+
The above copyright notice and this permission notice shall be included in all
|
| 13 |
+
copies or substantial portions of the Software.
|
| 14 |
+
|
| 15 |
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
| 16 |
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
| 17 |
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
| 18 |
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
| 19 |
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
| 20 |
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
| 21 |
+
SOFTWARE.
|
README.md
ADDED
|
@@ -0,0 +1,163 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: transformers
|
| 3 |
+
pipeline_tag: text-generation
|
| 4 |
+
language:
|
| 5 |
+
- en
|
| 6 |
+
datasets:
|
| 7 |
+
- roneneldan/TinyStories
|
| 8 |
+
tags:
|
| 9 |
+
- qraxai
|
| 10 |
+
- gpt2-tokenizer
|
| 11 |
+
- causal-lm
|
| 12 |
+
- tiny-stories
|
| 13 |
+
- from-scratch
|
| 14 |
+
license: mit
|
| 15 |
+
---
|
| 16 |
+
|
| 17 |
+
# QraXAi-Basic-45M
|
| 18 |
+
|
| 19 |
+
A small, decoder-only Transformer (~44.75M parameters) trained **from scratch** on
|
| 20 |
+
[TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) with the GPT-2 BPE
|
| 21 |
+
tokenizer (50,257 tokens). QraXAi is a hand-written PyTorch GPT implementation
|
| 22 |
+
(`model.py` / `configuration_qraxai.py`, shipped with this repo), not a fine-tune of GPT-2 —
|
| 23 |
+
only the tokenizer is shared with GPT-2.
|
| 24 |
+
|
| 25 |
+
This is an experimental research model, intended for learning/demo purposes.
|
| 26 |
+
|
| 27 |
+
## Model details
|
| 28 |
+
|
| 29 |
+
| | |
|
| 30 |
+
|---|---|
|
| 31 |
+
| Architecture | Decoder-only Transformer (GPT-style), pre-norm |
|
| 32 |
+
| Parameters | **44,751,872** (~44.75M), all trainable |
|
| 33 |
+
| Layers | 24 |
|
| 34 |
+
| Hidden size | 256 |
|
| 35 |
+
| Attention heads | 8 (head dim 32) |
|
| 36 |
+
| Feed-forward | 4× hidden, GELU |
|
| 37 |
+
| Context length | 256 tokens (hard limit) |
|
| 38 |
+
| Vocabulary | 50,257 (GPT-2 BPE) |
|
| 39 |
+
| Position encoding | Learned absolute embeddings |
|
| 40 |
+
| Normalization | LayerNorm |
|
| 41 |
+
| Weight tying | No (`lm_head` is separate) |
|
| 42 |
+
| KV cache | No — generation recomputes the full context at every step |
|
| 43 |
+
| Weights | fp32, 179 MB (`model.safetensors`) |
|
| 44 |
+
| Special tokens | `bos = eos = <\|endoftext\|>` (id 50256) |
|
| 45 |
+
| Custom code | Yes — requires `trust_remote_code=True` |
|
| 46 |
+
|
| 47 |
+
Parameter breakdown: token embeddings 12.87M + position embeddings 0.07M +
|
| 48 |
+
24 × 0.79M transformer blocks (18.95M) + final norm + untied `lm_head` 12.87M.
|
| 49 |
+
|
| 50 |
+
## Quick start
|
| 51 |
+
|
| 52 |
+
```python
|
| 53 |
+
import torch
|
| 54 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 55 |
+
|
| 56 |
+
repo_id = "coderian/QraXAi-Basic-45M"
|
| 57 |
+
|
| 58 |
+
tokenizer = AutoTokenizer.from_pretrained(repo_id, trust_remote_code=True)
|
| 59 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 60 |
+
repo_id,
|
| 61 |
+
trust_remote_code=True,
|
| 62 |
+
dtype=torch.bfloat16 if torch.cuda.is_available() else torch.float32,
|
| 63 |
+
).to("cuda" if torch.cuda.is_available() else "cpu").eval()
|
| 64 |
+
|
| 65 |
+
prompt = "Once upon a time, there was a little girl named Lily"
|
| 66 |
+
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
| 67 |
+
|
| 68 |
+
with torch.inference_mode():
|
| 69 |
+
output = model.generate(
|
| 70 |
+
**inputs,
|
| 71 |
+
max_new_tokens=244, # prompt + new tokens must stay <= 256
|
| 72 |
+
do_sample=True,
|
| 73 |
+
temperature=0.8,
|
| 74 |
+
top_k=50,
|
| 75 |
+
top_p=0.95,
|
| 76 |
+
repetition_penalty=1.1,
|
| 77 |
+
pad_token_id=tokenizer.eos_token_id,
|
| 78 |
+
eos_token_id=tokenizer.eos_token_id,
|
| 79 |
+
)
|
| 80 |
+
|
| 81 |
+
print(tokenizer.decode(output[0], skip_special_tokens=True))
|
| 82 |
+
```
|
| 83 |
+
|
| 84 |
+
Notes for generation:
|
| 85 |
+
- **Context is a hard limit of 256 tokens.** The custom `forward` raises `ValueError`
|
| 86 |
+
if the sequence gets longer, so keep `len(prompt) + max_new_tokens <= 256`.
|
| 87 |
+
- The model has **no KV cache**: each new token re-runs the full context, so generation
|
| 88 |
+
cost grows quickly with sequence length.
|
| 89 |
+
- The custom `forward` does not use an attention mask for padding. Generate one prompt
|
| 90 |
+
at a time instead of batching.
|
| 91 |
+
- Generation usually stops at `<|endoftext|>`, but because the training data is a
|
| 92 |
+
continuous stream of stories, the model sometimes starts a new story instead.
|
| 93 |
+
|
| 94 |
+
## Example output
|
| 95 |
+
|
| 96 |
+
Prompt: `Once upon a time, there was a little girl named Lily`
|
| 97 |
+
(`temperature=0.8, top_k=50, top_p=0.95, repetition_penalty=1.1, max_new_tokens=244`):
|
| 98 |
+
|
| 99 |
+
> Once upon a time, there was a little girl named Lily who loved to play in the big, green field. One day, she found a shiny stone on top of her backyard. She picked it up and showed it to her mom.
|
| 100 |
+
>
|
| 101 |
+
> "Look mommy, I found a pretty mineral!" said Lily excitedly. "It's very pretty!"
|
| 102 |
+
>
|
| 103 |
+
> Her mom smiled and said, "That's right, sweetie. It'll make sure you touch it. But remember, be careful with it because you might find something else inside."
|
| 104 |
+
>
|
| 105 |
+
> Lily nodded her head and kept playing with the jewel until she noticed that the box had fallen into a hole. She felt sad for her mom, but then remembered what her mom said about when something is hurt.
|
| 106 |
+
>
|
| 107 |
+
> The next day, Lily went back to the park and saw that the unknown stone was broken. She asked her mom if they could try and fix it. Her mom told her that it's okay to ask for help and that sometimes you can't use it without asking permission. So, Lily listened to her mom and never touched the stone again.
|
| 108 |
+
> Once upon a time, there was a boy named Timmy. He loved to play with his toy car. One day, he went to
|
| 109 |
+
|
| 110 |
+
(The last line is a new story the model began, cut off by the token limit.)
|
| 111 |
+
|
| 112 |
+
## Training
|
| 113 |
+
|
| 114 |
+
| | |
|
| 115 |
+
|---|---|
|
| 116 |
+
| Dataset | `roneneldan/TinyStories` (train split, streaming), first 130,000 stories |
|
| 117 |
+
| Data size | 115.7M characters → 28.76M tokens → ~112,350 training blocks |
|
| 118 |
+
| Objective | Next-token prediction (causal LM), cross-entropy |
|
| 119 |
+
| Epochs | 1 (~7,000 optimizer steps) |
|
| 120 |
+
| Batch size | 16 |
|
| 121 |
+
| Block size | 256 |
|
| 122 |
+
| Optimizer | AdamW, lr 3e-4 |
|
| 123 |
+
| Gradient clipping | 1.0 |
|
| 124 |
+
| Mixed precision | bf16 (on CUDA) |
|
| 125 |
+
| Seed | 42 |
|
| 126 |
+
|
| 127 |
+
## Files in this repository
|
| 128 |
+
|
| 129 |
+
| File | Description |
|
| 130 |
+
|---|---|
|
| 131 |
+
| `config.json` | Model config (`GPTConfig` + `auto_map` for remote code) |
|
| 132 |
+
| `model.safetensors` | fp32 weights (179 MB) |
|
| 133 |
+
| `model.py` | `QraXAiForCausalLM` — custom modeling code |
|
| 134 |
+
| `configuration_qraxai.py` | `GPTConfig` — custom configuration |
|
| 135 |
+
| `tokenizer.json`, `tokenizer_config.json` | GPT-2 BPE tokenizer |
|
| 136 |
+
| `generation_config.json` | Default generation settings |
|
| 137 |
+
| `LICENSE` | MIT License |
|
| 138 |
+
|
| 139 |
+
## Limitations
|
| 140 |
+
|
| 141 |
+
- English only; trained exclusively on synthetic children's stories (TinyStories), so it
|
| 142 |
+
knows little about the real world.
|
| 143 |
+
- Trained for a single epoch: grammar is mostly coherent, but content can be repetitive,
|
| 144 |
+
inconsistent or nonsensical.
|
| 145 |
+
- Hard 256-token context; no sliding window, so long inputs must be truncated.
|
| 146 |
+
- Not instruction-tuned — it does not follow instructions and is not a chat model.
|
| 147 |
+
- No safety filtering or alignment of any kind. Do not use in production or for
|
| 148 |
+
user-facing applications.
|
| 149 |
+
- TinyStories is synthetic data; nothing prevents the model from producing odd or
|
| 150 |
+
inappropriate continuations.
|
| 151 |
+
|
| 152 |
+
## Acknowledgements
|
| 153 |
+
|
| 154 |
+
- Dataset and idea: **TinyStories: How Small Can Language Models Be and Still Speak
|
| 155 |
+
Coherent English?** — Ronen Eldan and Yuanzhi Li, 2023 ([arXiv:2305.07759](https://arxiv.org/abs/2305.07759)).
|
| 156 |
+
- Tokenizer: GPT-2 BPE (`gpt2`).
|
| 157 |
+
- Built with PyTorch and Hugging Face Transformers.
|
| 158 |
+
|
| 159 |
+
## License
|
| 160 |
+
|
| 161 |
+
Released under the [MIT License](LICENSE). Note that the training dataset
|
| 162 |
+
([TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)) has its own terms;
|
| 163 |
+
check them if you plan to redistribute the data or use the model commercially.
|
config.json
ADDED
|
@@ -0,0 +1,20 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"QraXAiForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"auto_map": {
|
| 6 |
+
"AutoConfig": "configuration_qraxai.GPTConfig",
|
| 7 |
+
"AutoModelForCausalLM": "model.QraXAiForCausalLM"
|
| 8 |
+
},
|
| 9 |
+
"bos_token_id": 50256,
|
| 10 |
+
"dtype": "float32",
|
| 11 |
+
"embed_dim": 256,
|
| 12 |
+
"eos_token_id": 50256,
|
| 13 |
+
"max_seq_len": 256,
|
| 14 |
+
"model_type": "qrax_ai",
|
| 15 |
+
"n_layers": 24,
|
| 16 |
+
"tie_word_embeddings": false,
|
| 17 |
+
"transformers_version": "5.17.0",
|
| 18 |
+
"use_cache": false,
|
| 19 |
+
"vocab_size": 50257
|
| 20 |
+
}
|
configuration_qraxai.py
ADDED
|
@@ -0,0 +1,22 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
from transformers import PretrainedConfig
|
| 2 |
+
|
| 3 |
+
class GPTConfig(PretrainedConfig):
|
| 4 |
+
|
| 5 |
+
model_type = "qrax_ai"
|
| 6 |
+
|
| 7 |
+
def __init__(
|
| 8 |
+
self,
|
| 9 |
+
vocab_size=10000,
|
| 10 |
+
n_layers=6,
|
| 11 |
+
max_seq_len=512,
|
| 12 |
+
embed_dim=256,
|
| 13 |
+
use_cache=False,
|
| 14 |
+
**kwargs
|
| 15 |
+
):
|
| 16 |
+
super().__init__(**kwargs)
|
| 17 |
+
|
| 18 |
+
self.use_cache = use_cache
|
| 19 |
+
self.vocab_size = vocab_size
|
| 20 |
+
self.n_layers = n_layers
|
| 21 |
+
self.max_seq_len = max_seq_len
|
| 22 |
+
self.embed_dim = embed_dim
|
generation_config.json
ADDED
|
@@ -0,0 +1,9 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_from_model_config": true,
|
| 3 |
+
"bos_token_id": 50256,
|
| 4 |
+
"eos_token_id": 50256,
|
| 5 |
+
"output_attentions": false,
|
| 6 |
+
"output_hidden_states": false,
|
| 7 |
+
"transformers_version": "5.17.0",
|
| 8 |
+
"use_cache": false
|
| 9 |
+
}
|
model.py
ADDED
|
@@ -0,0 +1,243 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
from transformers import GenerationMixin, PreTrainedModel
|
| 2 |
+
from transformers.modeling_outputs import CausalLMOutput
|
| 3 |
+
import torch.nn as nn
|
| 4 |
+
import torch
|
| 5 |
+
import math
|
| 6 |
+
|
| 7 |
+
try:
|
| 8 |
+
from .configuration_qraxai import GPTConfig
|
| 9 |
+
except ImportError:
|
| 10 |
+
from configuration_qraxai import GPTConfig
|
| 11 |
+
|
| 12 |
+
class CausalSelfAttention(nn.Module):
|
| 13 |
+
|
| 14 |
+
def __init__(self, embed_dim, num_heads):
|
| 15 |
+
super().__init__()
|
| 16 |
+
|
| 17 |
+
assert embed_dim % num_heads == 0, \
|
| 18 |
+
"embed_dim must be divisible by num_heads"
|
| 19 |
+
|
| 20 |
+
self.embed_dim = embed_dim
|
| 21 |
+
self.num_heads = num_heads
|
| 22 |
+
self.head_dim = embed_dim // num_heads
|
| 23 |
+
|
| 24 |
+
# Q, K, V projections
|
| 25 |
+
self.q_proj = nn.Linear(embed_dim, embed_dim)
|
| 26 |
+
self.k_proj = nn.Linear(embed_dim, embed_dim)
|
| 27 |
+
self.v_proj = nn.Linear(embed_dim, embed_dim)
|
| 28 |
+
|
| 29 |
+
self.o_proj = nn.Linear(embed_dim, embed_dim)
|
| 30 |
+
|
| 31 |
+
def forward(self, x):
|
| 32 |
+
batch_size, seq_len, embed_dim = x.shape
|
| 33 |
+
|
| 34 |
+
Q = self.q_proj(x)
|
| 35 |
+
V = self.v_proj(x)
|
| 36 |
+
K = self.k_proj(x)
|
| 37 |
+
|
| 38 |
+
Q = Q.view(
|
| 39 |
+
batch_size,
|
| 40 |
+
seq_len,
|
| 41 |
+
self.num_heads,
|
| 42 |
+
self.head_dim
|
| 43 |
+
)
|
| 44 |
+
|
| 45 |
+
K = K.view(
|
| 46 |
+
batch_size,
|
| 47 |
+
seq_len,
|
| 48 |
+
self.num_heads,
|
| 49 |
+
self.head_dim
|
| 50 |
+
)
|
| 51 |
+
|
| 52 |
+
V = V.view(
|
| 53 |
+
batch_size,
|
| 54 |
+
seq_len,
|
| 55 |
+
self.num_heads,
|
| 56 |
+
self.head_dim
|
| 57 |
+
)
|
| 58 |
+
|
| 59 |
+
Q = Q.transpose(1, 2)
|
| 60 |
+
K = K.transpose(1, 2)
|
| 61 |
+
V = V.transpose(1, 2)
|
| 62 |
+
|
| 63 |
+
|
| 64 |
+
scores = Q @ K.transpose(-2, -1)
|
| 65 |
+
|
| 66 |
+
scores = scores / math.sqrt(self.head_dim)
|
| 67 |
+
|
| 68 |
+
mask = torch.triu(
|
| 69 |
+
torch.ones(
|
| 70 |
+
seq_len,
|
| 71 |
+
seq_len,
|
| 72 |
+
device=x.device
|
| 73 |
+
),
|
| 74 |
+
diagonal=1
|
| 75 |
+
).bool()
|
| 76 |
+
|
| 77 |
+
scores = scores.masked_fill(
|
| 78 |
+
mask,
|
| 79 |
+
torch.finfo(scores.dtype).min
|
| 80 |
+
)
|
| 81 |
+
|
| 82 |
+
attention_w = torch.softmax(
|
| 83 |
+
scores,
|
| 84 |
+
dim=-1
|
| 85 |
+
)
|
| 86 |
+
|
| 87 |
+
output = attention_w @ V
|
| 88 |
+
|
| 89 |
+
output = output.transpose(1, 2)
|
| 90 |
+
|
| 91 |
+
output = output.contiguous().view(
|
| 92 |
+
batch_size,
|
| 93 |
+
seq_len,
|
| 94 |
+
embed_dim
|
| 95 |
+
)
|
| 96 |
+
|
| 97 |
+
# Final projection
|
| 98 |
+
output = self.o_proj(output)
|
| 99 |
+
|
| 100 |
+
return output
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
class TransformerBlock(nn.Module):
|
| 104 |
+
|
| 105 |
+
def __init__(
|
| 106 |
+
self,
|
| 107 |
+
embed_dim
|
| 108 |
+
):
|
| 109 |
+
super().__init__()
|
| 110 |
+
|
| 111 |
+
self.ln1 = nn.LayerNorm(embed_dim)
|
| 112 |
+
|
| 113 |
+
self.attention = CausalSelfAttention(embed_dim, 8)
|
| 114 |
+
|
| 115 |
+
self.ln2 = nn.LayerNorm(embed_dim)
|
| 116 |
+
|
| 117 |
+
# feed forward network
|
| 118 |
+
self.ffn = nn.Sequential(
|
| 119 |
+
nn.Linear(
|
| 120 |
+
in_features=embed_dim,
|
| 121 |
+
out_features=4*embed_dim
|
| 122 |
+
),
|
| 123 |
+
|
| 124 |
+
nn.GELU(),
|
| 125 |
+
|
| 126 |
+
nn.Linear(
|
| 127 |
+
in_features=4*embed_dim,
|
| 128 |
+
out_features=embed_dim
|
| 129 |
+
)
|
| 130 |
+
)
|
| 131 |
+
|
| 132 |
+
def forward(self, x):
|
| 133 |
+
x = x + self.attention(
|
| 134 |
+
self.ln1(x)
|
| 135 |
+
)
|
| 136 |
+
|
| 137 |
+
x = x + self.ffn(
|
| 138 |
+
self.ln2(x)
|
| 139 |
+
)
|
| 140 |
+
|
| 141 |
+
return x
|
| 142 |
+
|
| 143 |
+
class QraXAiForCausalLM(PreTrainedModel, GenerationMixin):
|
| 144 |
+
|
| 145 |
+
config_class = GPTConfig
|
| 146 |
+
|
| 147 |
+
def __init__(
|
| 148 |
+
self,
|
| 149 |
+
config
|
| 150 |
+
):
|
| 151 |
+
|
| 152 |
+
super().__init__(config)
|
| 153 |
+
|
| 154 |
+
self.token_embedding = nn.Embedding(
|
| 155 |
+
config.vocab_size,
|
| 156 |
+
config.embed_dim
|
| 157 |
+
)
|
| 158 |
+
|
| 159 |
+
self.position_embedding = nn.Embedding(
|
| 160 |
+
config.max_seq_len,
|
| 161 |
+
config.embed_dim
|
| 162 |
+
)
|
| 163 |
+
|
| 164 |
+
self.transformer_blocks = nn.ModuleList([
|
| 165 |
+
TransformerBlock(config.embed_dim)
|
| 166 |
+
for _ in range(config.n_layers)
|
| 167 |
+
])
|
| 168 |
+
|
| 169 |
+
self.ln_f = nn.LayerNorm(
|
| 170 |
+
config.embed_dim
|
| 171 |
+
)
|
| 172 |
+
|
| 173 |
+
|
| 174 |
+
self.lm_head = nn.Linear(
|
| 175 |
+
config.embed_dim,
|
| 176 |
+
config.vocab_size,
|
| 177 |
+
bias=False
|
| 178 |
+
)
|
| 179 |
+
|
| 180 |
+
self.post_init()
|
| 181 |
+
|
| 182 |
+
def forward(
|
| 183 |
+
self,
|
| 184 |
+
input_ids,
|
| 185 |
+
labels=None,
|
| 186 |
+
**kwargs
|
| 187 |
+
):
|
| 188 |
+
|
| 189 |
+
batch_size, seq_len = input_ids.shape
|
| 190 |
+
|
| 191 |
+
if seq_len > self.config.max_seq_len:
|
| 192 |
+
raise ValueError(
|
| 193 |
+
f"Sequence length ({seq_len}) "
|
| 194 |
+
f"cannot be greater than "
|
| 195 |
+
f"max_seq_len ({self.config.max_seq_len})"
|
| 196 |
+
)
|
| 197 |
+
|
| 198 |
+
positions = torch.arange(
|
| 199 |
+
seq_len,
|
| 200 |
+
device=input_ids.device
|
| 201 |
+
)
|
| 202 |
+
|
| 203 |
+
token_emb = self.token_embedding(
|
| 204 |
+
input_ids
|
| 205 |
+
)
|
| 206 |
+
|
| 207 |
+
pos_emb = self.position_embedding(
|
| 208 |
+
positions
|
| 209 |
+
)
|
| 210 |
+
|
| 211 |
+
x = token_emb + pos_emb
|
| 212 |
+
|
| 213 |
+
for block in self.transformer_blocks:
|
| 214 |
+
x = block(x)
|
| 215 |
+
|
| 216 |
+
x = self.ln_f(x)
|
| 217 |
+
|
| 218 |
+
logits = self.lm_head(x)
|
| 219 |
+
|
| 220 |
+
loss = None
|
| 221 |
+
|
| 222 |
+
if labels is not None:
|
| 223 |
+
|
| 224 |
+
shift_logits = logits[
|
| 225 |
+
:, :-1, :
|
| 226 |
+
].contiguous()
|
| 227 |
+
|
| 228 |
+
shift_labels = labels[
|
| 229 |
+
:, 1:
|
| 230 |
+
].contiguous()
|
| 231 |
+
|
| 232 |
+
loss = nn.functional.cross_entropy(
|
| 233 |
+
shift_logits.view(
|
| 234 |
+
-1,
|
| 235 |
+
shift_logits.size(-1)
|
| 236 |
+
),
|
| 237 |
+
shift_labels.view(-1)
|
| 238 |
+
)
|
| 239 |
+
|
| 240 |
+
return CausalLMOutput(
|
| 241 |
+
loss=loss,
|
| 242 |
+
logits=logits
|
| 243 |
+
)
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e4bfeaa3c57234fef408a5c51c9463df89eb67ca3e3c8c7e7517d2d6d4bbd6de
|
| 3 |
+
size 179049920
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"backend": "tokenizers",
|
| 4 |
+
"bos_token": "<|endoftext|>",
|
| 5 |
+
"eos_token": "<|endoftext|>",
|
| 6 |
+
"errors": "replace",
|
| 7 |
+
"is_local": false,
|
| 8 |
+
"local_files_only": false,
|
| 9 |
+
"model_max_length": 1024,
|
| 10 |
+
"pad_token": null,
|
| 11 |
+
"tokenizer_class": "GPT2Tokenizer",
|
| 12 |
+
"unk_token": "<|endoftext|>"
|
| 13 |
+
}
|