File size: 1,631 Bytes
1d6f4c0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
---
license: apache-2.0
pipeline_tag: text-classification
tags:
- sentiment-analysis
- text-classification
- from-scratch
- transformer
- sst2
datasets:
- stanfordnlp/sst2
metrics:
- accuracy
---

# Vibe Check — a sentiment classifier trained from scratch

**Vibe Check** is a compact (~2M parameter) Transformer text classifier that predicts whether a
sentence is **Positive** or **Negative**. It was built **entirely from scratch** — my own word
tokenizer and my own model architecture, trained from **random initialization** (no pretrained
weights, no fine-tuning) on the public **Stanford Sentiment Treebank (SST-2)** dataset.

## Highlights
- **From scratch:** custom tokenizer + 2-layer Transformer encoder + mean-pooling classifier head.
- **~2.0M parameters**, trained from random init on ~67k public sentences.
- **~82% validation accuracy** on SST-2 (small model, no pretrained embeddings).
- Runs on CPU in milliseconds.

## Not a fine-tune
Unlike many models on the Hub, this is **not** a fine-tune of a large pretrained model — every weight
was learned from zero on public data. Training code, tokenizer, and config are included.

## Usage
```python
# see vibe.py in this repo for the model definition + predict_vibe()
from vibe import predict_vibe
print(predict_vibe("I absolutely love this!"))   # {'label': 'Positive', 'confidence': 98.5}
```

## Data & license
Trained on **[SST-2](https://huggingface.co/datasets/stanfordnlp/sst2)** (public research dataset).
Released under **Apache-2.0**. For entertainment/educational use; a small model, so expect ~1 in 5
predictions to be wrong on hard/ambiguous text.