File size: 6,989 Bytes
a0f13dc
 
 
2277e2c
 
 
 
 
 
 
 
 
 
 
 
 
a0f13dc
 
2277e2c
a0f13dc
2277e2c
 
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
 
 
 
 
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
 
 
 
 
 
 
 
 
 
a0f13dc
2277e2c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
 
 
 
a0f13dc
2277e2c
 
 
 
 
a0f13dc
2277e2c
 
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
 
 
 
 
 
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
 
 
 
 
 
 
 
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
 
 
 
 
 
 
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
 
 
 
 
a0f13dc
2277e2c
a0f13dc
2277e2c
 
 
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
2277e2c
a0f13dc
48f09b0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
---
base_model: Qwen/Qwen2.5-0.5B-Instruct
library_name: peft
pipeline_tag: text-generation
language:
  - en
tags:
  - lora
  - peft
  - nigeria
  - nigerian-english
  - nigerian-pidgin
  - customer-service
  - scam-safety
  - business-writing
license: apache-2.0
---

# GaiaLab Naija Assistant v0.5

GaiaLab Naija Assistant v0.5 is an experimental LoRA adapter for
`Qwen/Qwen2.5-0.5B-Instruct`.

It is designed to explore small, accessible language models for Nigerian-context communication, including customer service, Nigerian English, basic Nigerian Pidgin, professional boundaries, business writing, and scam-safety guidance.

## Important status

This is an early research release.

The adapter was trained on a small, manually reviewed dataset of **47 examples**. It should not be treated as a production-ready general-purpose assistant.

The current release demonstrates a reproducible workflow for:

- creating training examples from CSV
- generating JSONL training data
- validating dataset structure
- detecting duplicate IDs and prompts
- calculating dataset statistics
- training a CPU-compatible LoRA adapter
- versioning model releases

## Model details

| Field | Value |
|---|---|
| Model | GaiaLab Naija Assistant v0.5 |
| Base model | `Qwen/Qwen2.5-0.5B-Instruct` |
| Fine-tuning method | LoRA / PEFT |
| Model type | Causal language model adapter |
| Primary language | English |
| Additional language variety | Nigerian English and basic Nigerian Pidgin |
| Training examples | 47 |
| Dataset health score | 95/100 |
| Developer | Oluwafemi Idiakhoa |
| Project | GaiaLab AI |

## Training-data categories

The v0.5 training dataset contained:

| Category | Examples |
|---|---:|
| Safety and scams | 13 |
| Professional boundaries | 12 |
| Customer service | 10 |
| Nigerian English | 10 |
| Business writing | 1 |
| Nigerian Pidgin | 1 |
| **Total** | **47** |

Risk-level distribution:

| Risk level | Examples |
|---|---:|
| High | 19 |
| Medium | 7 |
| Low | 21 |

## Dataset validation

The dataset pipeline reported:

- valid JSONL structure
- required fields present
- correct system, user, and assistant message order
- zero duplicate IDs
- zero duplicate prompts
- zero missing prompts
- zero missing responses
- dataset health score of 95/100

## Intended uses

This adapter may be useful for:

- research on Nigerian-context conversational AI
- educational demonstrations of LoRA fine-tuning
- Nigerian customer-service prototypes
- professional-message drafting experiments
- scam-awareness and credential-safety demonstrations
- Nigerian English and basic Pidgin experimentation
- CPU-friendly small-model research

## Out-of-scope uses

This model should not be used as the sole authority for:

- medical decisions
- legal advice
- financial decisions
- banking authentication
- emergency response
- employment decisions
- identity verification
- high-impact automated decision-making

Never provide passwords, PINs, one-time passwords, bank verification codes, private keys, or other sensitive credentials to the model.

## Installation

```bash
pip install torch transformers peft
```

## Usage

```python
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base_model_id = "Qwen/Qwen2.5-0.5B-Instruct"
adapter_id = "mgbam/gaialab-naija-adapter-v0.5"

tokenizer = AutoTokenizer.from_pretrained(
    base_model_id,
    trust_remote_code=True,
)

base_model = AutoModelForCausalLM.from_pretrained(
    base_model_id,
    torch_dtype=torch.float32,
    trust_remote_code=True,
)

model = PeftModel.from_pretrained(
    base_model,
    adapter_id,
)

messages = [
    {
        "role": "system",
        "content": (
            "You are GaiaLab Naija Assistant. Be helpful, concise, "
            "culturally aware, truthful, and safe."
        ),
    },
    {
        "role": "user",
        "content": "Write a polite reminder for a customer who has not paid.",
    },
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

inputs = tokenizer(text, return_tensors="pt")

with torch.no_grad():
    output = model.generate(
        **inputs,
        max_new_tokens=120,
        do_sample=False,
    )

generated_tokens = output[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(
    generated_tokens,
    skip_special_tokens=True,
)

print(response)
```

## Example areas

The adapter was trained on examples involving:

- suspicious requests for OTPs and PINs
- safe handling of account credentials
- professional customer responses
- Nigerian-style business communication
- polite payment reminders
- simple Nigerian English phrasing
- introductory Nigerian Pidgin translations
- maintaining appropriate professional boundaries

## Training approach

The adapter was trained with LoRA using the PEFT library.

The local training configuration included:

- LoRA rank: 16
- LoRA alpha: 32
- LoRA dropout: 0.05
- training epochs: 3
- batch size: 1
- gradient accumulation steps: 8
- maximum sequence length: 512
- optimizer: AdamW
- CPU-compatible float32 loading
- base model: `Qwen/Qwen2.5-0.5B-Instruct`

## Evaluation status

A formal side-by-side comparison between v0.4 and v0.5 has not yet been published.

Therefore, this model card does **not** claim that v0.5 performs better than v0.4.

Evaluation results will be added after both adapters are tested on the same held-out benchmark and reviewed using consistent criteria.

## Limitations

The training dataset is very small and unevenly distributed.

In particular:

- business writing has only one example
- Nigerian Pidgin has only one example
- the adapter may overfit specific phrasings
- responses may be inconsistent
- cultural coverage is narrow
- the model may hallucinate information
- safety behaviour has not been independently audited
- performance outside the training categories is unknown
- English and Pidgin quality may vary significantly

All important outputs should be reviewed by a person.

## Version history

| Version | Status |
|---|---|
| v0.1 | Initial experimental adapter |
| v0.2 | Early iterative release |
| v0.3 | Expanded experimental release |
| v0.4 | First formally reviewed and benchmarked development version |
| v0.5 | Reproducible dataset pipeline and corrective-example training release |

## Project links

- GitHub: `https://github.com/oluwafemidiakhoa/gaialab-naija-assistant`
- Model: `https://huggingface.co/mgbam/gaialab-naija-adapter-v0.5`
- GaiaLab AI: `https://www.gailabai.com`

## Responsible-use statement

GaiaLab Naija Assistant is an experimental research project. Users are responsible for reviewing generated content before relying on it or sending it to others.

Do not use this model to impersonate individuals, facilitate fraud, request confidential credentials, or make consequential decisions without qualified human oversight.

## Author

Developed by **Oluwafemi Idiakhoa** under the **GaiaLab AI** initiative.