File size: 3,729 Bytes
2fa9fef
f327c9d
 
2fa9fef
f327c9d
2fa9fef
 
 
 
f327c9d
2fa9fef
 
 
 
 
f327c9d
2fa9fef
f327c9d
2fa9fef
f327c9d
2fa9fef
f327c9d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2fa9fef
 
 
f327c9d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2fa9fef
f327c9d
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
---
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- novi
- novi-nano
- causal-lm
- gpt2
- from-scratch
---

# Novi-Nano-Base

![Novi-Nano Banner](banner.jpg)

**Novi-Nano-Base** is a tiny causal language model trained from scratch by **Novi-AI**.

With just **1,258,560 parameters**, Novi-Nano explores language modeling at an extremely small scale while remaining compatible with the Hugging Face Transformers ecosystem.

> ⚡ **1.26M parameters · 300M training tokens · 256-token context**

## Model Details

### Architecture

| Property        |                 Value |
| --------------- | --------------------: |
| Model type      | Causal Language Model |
| Parameters      |         **1,258,560** |
| Vocabulary size |             **8,192** |
| Context length  |               **256** |
| Embedding size  |                **96** |
| Layers          |                 **4** |
| Attention heads |                 **4** |
| FFN size        |               **384** |
| Tensor type     |               **F32** |

## Training

Novi-Nano-Base was trained from scratch using approximately **300 million training tokens**.

### Training Statistics

| Metric                      |          Result |
| --------------------------- | --------------: |
| Training tokens             | **300,023,808** |
| Best validation loss        |    **5.418699** |
| Final validation loss       |    **5.418699** |
| Final validation perplexity |    **225.5853** |

## Tokenizer

Novi-Nano uses a custom tokenizer with a vocabulary size of **8,192 tokens**.

The tokenizer was trained using data from:

* FineWeb-Edu
* FineWeb-HQ
* SmolLM-Cosmopedia

## Intended Use

Novi-Nano-Base is primarily intended for:

* 🔬 Research and experimentation
* 🧪 Small-model language-model experiments
* 🎓 Educational purposes
* 🛠️ Fine-tuning experiments
* 💻 Lightweight local inference

As a **base model**, it is not specifically instruction-tuned for following user commands or acting as a conversational assistant.

## Limitations

Novi-Nano-Base is an extremely small experimental language model.

Because of its size and short context window, it will have significant limitations compared with modern billion-parameter language models.

It may:

* Generate incoherent text
* Repeat phrases
* Produce factual errors
* Struggle with complex instructions
* Have limited world knowledge
* Perform poorly on reasoning tasks
* Lose context beyond its 256-token window

This model should be considered a **research and experimentation model**, rather than a production-ready general-purpose LLM.

## Usage

```python
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "Novi-AI/Novi-Nano-Base"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id)

prompt = "Hello, my name is"

inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(
    **inputs,
    max_new_tokens=50,
)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## Project History

Novi AI follows the earlier **AppleMind** experiments, with Novi becoming the primary project for developing small language models.

**AppleMind → Novi AI → Novi-Nano** 🚀

## Acknowledgements

Novi-Nano was built using the open-source machine-learning ecosystem and datasets made available by the community.

Special thanks to:

* Hugging Face 🤗
* FineWeb
* SmolLM
* Cosmopedia

## License

This model is released under the **Apache 2.0** license.

---

## 🧠 Novi AI

**Small models. Big experiments.**

Novi-Nano is intentionally tiny — exploring how far a language model can go with just a fraction of the parameters used by modern LLMs.