oddadmix commited on
Commit
2e48b3f
·
verified ·
1 Parent(s): 5298f01

Add Emhotob-25M: weights, tokenizer, and model card

Browse files
README.md ADDED
@@ -0,0 +1,80 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - ar
4
+ license: apache-2.0
5
+ library_name: transformers
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - arabic
9
+ - llama
10
+ - pretraining
11
+ - from-scratch
12
+ - emhotob
13
+ - small-language-model
14
+ datasets:
15
+ - kaust-generative-ai/fineweb-edu-ar
16
+ ---
17
+
18
+ # Emhotob-25M
19
+
20
+ **Emhotob** is a family of small Arabic language models pretrained **from scratch** on
21
+ Arabic web text. This is the **25M** rung of the ladder (~25.27M parameters),
22
+ part of a scaling series ranging from 500K to 25M parameters that all share the same
23
+ tokenizer, context length, and training recipe.
24
+
25
+ > ⚠️ These are tiny, research-scale models trained on a limited token budget. They are
26
+ > intended for scaling-law experiments, education, and Arabic NLP research — **not** for
27
+ > production use.
28
+
29
+ ## Model details
30
+
31
+ | Property | Value |
32
+ |---|---|
33
+ | Architecture | Llama (decoder-only, RoPE, GQA) |
34
+ | Parameters | 25,270,656 (~25.27M) |
35
+ | Hidden size | 384 |
36
+ | Layers | 8 |
37
+ | Attention heads | 6 (KV heads: 3) |
38
+ | Intermediate size | 1024 |
39
+ | Context length | 2048 |
40
+ | Vocabulary | 32,000 (custom Byte-Level BPE) |
41
+ | Tied embeddings | Yes |
42
+ | RoPE theta | 10,000 |
43
+ | Precision | bf16 |
44
+
45
+ ## Training
46
+
47
+ | Property | Value |
48
+ |---|---|
49
+ | Data | [`kaust-generative-ai/fineweb-edu-ar`](https://huggingface.co/datasets/kaust-generative-ai/fineweb-edu-ar) (Arabic) |
50
+ | Tokens seen | ~2.5B (1 epoch) |
51
+ | Optimizer | AdamW (fused), β=(0.9, 0.95), wd=0.1 |
52
+ | LR schedule | 6e-4, cosine, 2% warmup |
53
+ | Effective batch | 128 sequences × 2048 tokens |
54
+ | Grad clipping | 1.0 |
55
+
56
+ ## Usage
57
+
58
+ ```python
59
+ from transformers import AutoModelForCausalLM, AutoTokenizer
60
+ import torch
61
+
62
+ model_id = "oddadmix/Emhotob-25M"
63
+ tok = AutoTokenizer.from_pretrained(model_id)
64
+ model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype=torch.bfloat16)
65
+
66
+ prompt = "الذكاء الاصطناعي هو"
67
+ inputs = tok(prompt, return_tensors="pt")
68
+ out = model.generate(**inputs, max_new_tokens=50, do_sample=True, top_p=0.9, temperature=0.8)
69
+ print(tok.decode(out[0], skip_special_tokens=True))
70
+ ```
71
+
72
+ ## Limitations
73
+
74
+ Given its size and limited pretraining budget, Emhotob-25M has a narrow capability
75
+ range and will produce factually unreliable and sometimes incoherent text. It has not been
76
+ instruction-tuned or aligned, and no safety filtering has been applied. Use accordingly.
77
+
78
+ ---
79
+
80
+ *© SupraLabs 2026 — PROJECT EMHOTOB.*
config.json ADDED
@@ -0,0 +1,32 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "LlamaForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "bos_token_id": 0,
8
+ "dtype": "float32",
9
+ "eos_token_id": 2,
10
+ "head_dim": 64,
11
+ "hidden_act": "silu",
12
+ "hidden_size": 384,
13
+ "initializer_range": 0.02,
14
+ "intermediate_size": 1024,
15
+ "max_position_embeddings": 2048,
16
+ "mlp_bias": false,
17
+ "model_type": "llama",
18
+ "num_attention_heads": 6,
19
+ "num_hidden_layers": 8,
20
+ "num_key_value_heads": 3,
21
+ "pad_token_id": 1,
22
+ "pretraining_tp": 1,
23
+ "rms_norm_eps": 1e-06,
24
+ "rope_parameters": {
25
+ "rope_theta": 10000,
26
+ "rope_type": "default"
27
+ },
28
+ "tie_word_embeddings": true,
29
+ "transformers_version": "5.12.1",
30
+ "use_cache": false,
31
+ "vocab_size": 32000
32
+ }
generation_config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 0,
4
+ "eos_token_id": 2,
5
+ "output_attentions": false,
6
+ "output_hidden_states": false,
7
+ "pad_token_id": 1,
8
+ "transformers_version": "5.12.1",
9
+ "use_cache": true
10
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:dabeeb7865d60cd1c4eb7a8ba063420c04e4800ccb443973ffb122933b4139a3
3
+ size 101090696
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,9 @@
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "backend": "tokenizers",
3
+ "bos_token": "<s>",
4
+ "eos_token": "</s>",
5
+ "model_max_length": 1000000000000000019884624838656,
6
+ "pad_token": "<pad>",
7
+ "tokenizer_class": "TokenizersBackend",
8
+ "unk_token": "<unk>"
9
+ }
training_args.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:75108ea0d59ecfe2282ab5550f1cf0420c891ea4822c805a0c85e1facac2a16f
3
+ size 5201