GGUFGuy commited on
Commit
3ecd0cd
·
0 Parent(s):

Duplicate from Novi-AI/Novi-Nano-Instruct

Browse files
.gitattributes ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ banner.jpg filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,229 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ library_name: transformers
5
+ pipeline_tag: text-generation
6
+ tags:
7
+ - novi
8
+ - novi-nano
9
+ - novi-nano-instruct
10
+ - causal-lm
11
+ - gpt2
12
+ - from-scratch
13
+ - instruction-tuning
14
+ - chatml
15
+ datasets:
16
+ - Novi-AI/Novi-510x
17
+ ---
18
+
19
+ # Novi-Nano-Instruct
20
+
21
+ ![Novi-Nano Banner](banner.jpg)
22
+
23
+ **Novi-Nano-Instruct** is a tiny instruction-tuned causal language model developed by **Novi-AI**.
24
+
25
+ It is based on **Novi-Nano-Base** and fine-tuned on a small instruction dataset to experiment with instruction following and conversational behavior at an extremely small scale.
26
+
27
+ ⚡ **1.26M parameters · 500 training examples · 256-token context**
28
+
29
+ ## Model Details
30
+
31
+ ### Architecture
32
+
33
+ | Property | Value |
34
+ | --------------- | -----------------------: |
35
+ | Model type | Causal Language Model |
36
+ | Base model | `Novi-AI/Novi-Nano-Base` |
37
+ | Parameters | **1,258,848** |
38
+ | Vocabulary size | **8,195** |
39
+ | Context length | **256** |
40
+ | Embedding size | **96** |
41
+ | Layers | **4** |
42
+ | Attention heads | **4** |
43
+ | FFN size | **384** |
44
+ | Tensor type | **F32** |
45
+
46
+ ## Instruction Tuning
47
+
48
+ Novi-Nano-Instruct was trained from **Novi-Nano-Base** using a small instruction dataset containing **510 examples**.
49
+
50
+ ### Dataset
51
+
52
+ | Split | Examples |
53
+ | ---------- | -------: |
54
+ | Training | **500** |
55
+ | Validation | **10** |
56
+
57
+ The model uses a ChatML-style format with:
58
+
59
+ ```text
60
+ <|im_start|>
61
+ <|im_end|>
62
+ ```
63
+
64
+ Training loss was applied specifically to the assistant responses, allowing the model to focus on learning how to respond to user instructions.
65
+
66
+ ### Training Configuration
67
+
68
+ | Property | Value |
69
+ | ----------------------- | -------: |
70
+ | Epochs | **5** |
71
+ | Batch size | **16** |
72
+ | Gradient accumulation | **2** |
73
+ | Effective batch size | **32** |
74
+ | Maximum sequence length | **256** |
75
+ | Learning rate | **2e-5** |
76
+ | Precision | **FP32** |
77
+ | Device | **CPU** |
78
+
79
+ ## Training Statistics
80
+
81
+ The final training run produced:
82
+
83
+ | Metric | Result |
84
+ | --------------------------- | --------------: |
85
+ | Final validation loss | **5.153667** |
86
+ | Final validation perplexity | **173.0650** |
87
+ | Training examples | **500** |
88
+ | Validation examples | **10** |
89
+ | Training time | **~32 seconds** |
90
+
91
+ Because the validation set contains only **10 examples**, these metrics should be considered experimental rather than a comprehensive benchmark.
92
+
93
+ ## Tokenizer
94
+
95
+ Novi-Nano-Instruct uses the custom tokenizer developed for Novi-Nano.
96
+
97
+ The original tokenizer vocabulary was **8,192 tokens**, with additional tokens already present in the tokenizer.
98
+
99
+ Two ChatML tokens were added for instruction tuning:
100
+
101
+ * `<|im_start|>` — **8193**
102
+ * `<|im_end|>` — **8194**
103
+
104
+ The final tokenizer size is **8,195 tokens**.
105
+
106
+ The tokenizer was originally trained using data from:
107
+
108
+ * FineWeb-Edu
109
+ * FineWeb-HQ
110
+ * SmolLM-Cosmopedia
111
+
112
+ ## Intended Use
113
+
114
+ Novi-Nano-Instruct is primarily intended for:
115
+
116
+ * 🔬 Research and experimentation
117
+ * 🧪 Small-model instruction-tuning experiments
118
+ * 🎓 Educational purposes
119
+ * 💬 Tiny conversational-model experiments
120
+ * 💻 Lightweight local inference
121
+ * 🛠️ Experimenting with extremely small instruction-tuned models
122
+
123
+ As an **experimental 1.26M-parameter model**, it is not intended to compete with modern billion-parameter language models.
124
+
125
+ ## Limitations
126
+
127
+ Novi-Nano-Instruct is an extremely small experimental language model trained on only **500 instruction examples**.
128
+
129
+ Because of its size and limited training data, it may:
130
+
131
+ * Generate incoherent text
132
+ * Repeat phrases
133
+ * Produce unrelated responses
134
+ * Fail to follow instructions
135
+ * Produce factual errors
136
+ * Have very limited world knowledge
137
+ * Perform poorly on reasoning tasks
138
+ * Struggle with longer conversations
139
+ * Lose context beyond its 256-token window
140
+ * Produce malformed or unexpected responses
141
+
142
+ Generation quality is currently **highly experimental**. The model can generate text, but it does not yet consistently produce reliable assistant-style responses.
143
+
144
+ This model should be considered a **research and experimentation model**, rather than a production-ready conversational AI.
145
+
146
+ ## Usage
147
+
148
+ ```python
149
+ from transformers import AutoTokenizer, AutoModelForCausalLM
150
+
151
+ model_id = "Novi-AI/Novi-Nano-Instruct"
152
+
153
+ tokenizer = AutoTokenizer.from_pretrained(model_id)
154
+ model = AutoModelForCausalLM.from_pretrained(model_id)
155
+
156
+ messages = [
157
+ {
158
+ "role": "system",
159
+ "content": "You are Novi-Nano, a helpful AI assistant."
160
+ },
161
+ {
162
+ "role": "user",
163
+ "content": "Give a synonym for 'quiet'."
164
+ }
165
+ ]
166
+
167
+ prompt = tokenizer.apply_chat_template(
168
+ messages,
169
+ tokenize=False,
170
+ add_generation_prompt=True,
171
+ )
172
+
173
+ inputs = tokenizer(prompt, return_tensors="pt")
174
+
175
+ outputs = model.generate(
176
+ **inputs,
177
+ max_new_tokens=50,
178
+ )
179
+
180
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
181
+ ```
182
+
183
+ ## Chat Template
184
+
185
+ Novi-Nano-Instruct uses a ChatML-style conversation format:
186
+
187
+ ```text
188
+ <|im_start|>system
189
+ You are Novi-Nano, a helpful AI assistant.<|im_end|>
190
+ <|im_start|>user
191
+ Give a synonym for 'quiet'.<|im_end|>
192
+ <|im_start|>assistant
193
+ A synonym is 'silent'.<|im_end|>
194
+ ```
195
+
196
+ For generation, the assistant message is opened automatically by the chat template.
197
+
198
+ ## Project History
199
+
200
+ Novi AI follows the earlier **AppleMind** experiments, with Novi becoming the primary project for developing small language models.
201
+
202
+ **AppleMind → Novi AI → Novi-Nano → Novi-Nano-Instruct** 🚀
203
+
204
+ ## Acknowledgements
205
+
206
+ Novi-Nano was built using the open-source machine-learning ecosystem and datasets made available by the community.
207
+
208
+ Special thanks to:
209
+
210
+ * Hugging Face 🤗
211
+ * FineWeb
212
+ * SmolLM
213
+ * Cosmopedia
214
+
215
+ ## License
216
+
217
+ This model is released under the **Apache 2.0** license.
218
+
219
+ ---
220
+
221
+ ## 🧠 Novi AI
222
+
223
+ **Small models. Big experiments.**
224
+
225
+ Novi-Nano-Instruct explores instruction tuning at an extremely small scale, with just **1.26 million parameters** and **500 training examples**.
226
+
227
+ It is intentionally tiny — exploring how far instruction following can go with a fraction of the parameters used by modern LLMs.
228
+
229
+ *Novi AI 2026 — Project Kairo*
banner.jpg ADDED

Git LFS Details

  • SHA256: f6f75251f28bb4b99204c96e90c9754c246cd96d377dcb97134baeb4850d4625
  • Pointer size: 131 Bytes
  • Size of remote file: 468 kB
chat_template.jinja ADDED
@@ -0,0 +1 @@
 
 
1
+ {% for message in messages %}{{'<|im_start|>' + message['role'] + '\n' + message['content'] + '<|im_end|>' + '\n'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant\n' }}{% endif %}
config.json ADDED
@@ -0,0 +1,35 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "activation_function": "gelu_new",
3
+ "add_cross_attention": false,
4
+ "architectures": [
5
+ "GPT2LMHeadModel"
6
+ ],
7
+ "attn_pdrop": 0.0,
8
+ "bos_token_id": 8192,
9
+ "dtype": "float32",
10
+ "embd_pdrop": 0.0,
11
+ "eos_token_id": 8192,
12
+ "initializer_range": 0.02,
13
+ "layer_norm_epsilon": 1e-05,
14
+ "model_type": "gpt2",
15
+ "n_ctx": 256,
16
+ "n_embd": 96,
17
+ "n_head": 4,
18
+ "n_inner": 384,
19
+ "n_layer": 4,
20
+ "n_positions": 256,
21
+ "pad_token_id": 0,
22
+ "reorder_and_upcast_attn": false,
23
+ "resid_pdrop": 0.0,
24
+ "scale_attn_by_inverse_layer_idx": false,
25
+ "scale_attn_weights": true,
26
+ "summary_activation": null,
27
+ "summary_first_dropout": 0.1,
28
+ "summary_proj_to_labels": true,
29
+ "summary_type": "cls_index",
30
+ "summary_use_proj": true,
31
+ "tie_word_embeddings": true,
32
+ "transformers_version": "5.9.0",
33
+ "use_cache": false,
34
+ "vocab_size": 8195
35
+ }
generation_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "_from_model_config": true,
3
+ "bos_token_id": 8192,
4
+ "eos_token_id": [
5
+ 8192,
6
+ 2
7
+ ],
8
+ "output_attentions": false,
9
+ "output_hidden_states": false,
10
+ "pad_token_id": 0,
11
+ "transformers_version": "5.9.0",
12
+ "use_cache": false
13
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b52c50cc0b9bd8042ee674266ec2af4a8e34fd2f5b213b26d0db73d6dc59e739
3
+ size 5040360
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff
 
tokenizer_config.json ADDED
@@ -0,0 +1,17 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "backend": "tokenizers",
4
+ "bos_token": "<|endoftext|>",
5
+ "eos_token": "<|endoftext|>",
6
+ "errors": "replace",
7
+ "extra_special_tokens": [
8
+ "<|im_start|>",
9
+ "<|im_end|>"
10
+ ],
11
+ "is_local": false,
12
+ "local_files_only": false,
13
+ "model_max_length": 1000000000000000019884624838656,
14
+ "pad_token": "<|pad|>",
15
+ "tokenizer_class": "GPT2Tokenizer",
16
+ "unk_token": "<|endoftext|>"
17
+ }
training_args.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e4c534110bb4b95606560faca657721f10249028cc669c356a65e033b95b71ef
3
+ size 5265
training_info.json ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "model": "Novi-Nano-Instruct",
3
+ "base_model": "Novi-AI/Novi-Nano-Base",
4
+ "parameters": 1258848,
5
+ "trainable_parameters": 1258848,
6
+ "base_vocab_size": 8192,
7
+ "final_vocab_size": 8195,
8
+ "added_special_tokens": [
9
+ "<|im_start|>",
10
+ "<|im_end|>"
11
+ ],
12
+ "context_length": 256,
13
+ "epochs": 5,
14
+ "train_examples": 500,
15
+ "validation_examples": 10,
16
+ "batch_size": 16,
17
+ "gradient_accumulation_steps": 2,
18
+ "effective_batch_size": 32,
19
+ "learning_rate": 2e-05,
20
+ "weight_decay": 0.01,
21
+ "warmup_ratio": 0.1,
22
+ "scheduler": "cosine",
23
+ "bf16": false,
24
+ "fp16": false,
25
+ "validation_loss": 5.153667449951172,
26
+ "validation_perplexity": 173.0650352145914,
27
+ "chat_template": "{% for message in messages %}{{'<|im_start|>' + message['role'] + '\\n' + message['content'] + '<|im_end|>' + '\\n'}}{% endfor %}{% if add_generation_prompt %}{{ '<|im_start|>assistant\\n' }}{% endif %}"
28
+ }