sentiment / comparison.md
michaljach's picture
Publish v2: larger synthetic dataset, checkpoints and evaluation
88da524 verified
|
Raw History Blame Contribute Delete
896 Bytes
# Model comparison — sentiment
Test split identical across versions (fingerprint a9ff6039d984, n=51).
| | sentiment@v1 | sentiment@v2 |
|---|---|---|
| tier | encoder | encoder |
| base | `sentence-transformers/all-MiniLM-L6-v2` | `sentence-transformers/all-MiniLM-L6-v2` |
| **macro F1** | **0.882** | **0.881** |
| accuracy | 88.2% | 88.2% |
| coverage @ threshold | 82.4% (95.2% accurate) | 51.0% (96.2% accurate) |
| escalation rate | 17.6% | 49.0% |
| ECE after calibration | 0.066 | 0.093 |
| download | 23.7 MB | 24.31 MB |
| latency p95, python | 3.99 ms (cpu) | 3.57 ms (cpu) |
| latency p95, browser (best) | — | — |
| train time | 19 s | 120 s |
| F1 · positive | 0.875 | 0.919 |
| F1 · neutral | 0.914 | 0.875 |
| F1 · negative | 0.857 | 0.848 |
Prediction agreement on the test split: sentiment@v1 vs sentiment@v2: 88.2%
Targets: macro F1 ≥ 0.9, download ≤ 30.0 MB.