File size: 1,890 Bytes
dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa dbb5d78 a330cfa | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 | # Smartwatch LM v0.2 — Quality Benchmark
Quality evaluation for **Smartwatch LM v0.2** on a fixed set of 39 golden prompts covering all supported intents.
## Golden prompts
[`benchmark_prompts.json`](benchmark_prompts.json) holds 39 prompts — one per intent category plus multi-intent and conversational cases.
Each entry:
```json
{
"id": "get_steps_...",
"prompt": "How many steps today?",
"expected_intent": "GET_STEPS",
"expected_intents": ["GET_STEPS"],
"expected_slots": ["STEPS_TODAY", "STEP_GOAL"],
"source": "golden_set"
}
```
## Metrics
| Metric | Meaning |
|--------|---------|
| **Intent accuracy** | Predicted intent matches expected (or is in `expected_intents` for combo cases) |
| **Intent parse rate** | Reply contains a valid `<INTENT:...>` tag |
| **Clean output rate** | Raw decode has no BPE junk (`Ġ`, `Ċ`) before cleanup |
| **Slot presence** | Fraction of expected slot placeholders found in the cleaned template |
Generation settings used for evaluation: `max_new_tokens=40`, `temperature=0.5`, `top_k=40`, fresh history per prompt, fixed seed for reproducibility.
## v0.2 results
| Metric | Score |
|--------|------:|
| Intent accuracy | 100% |
| Intent parse rate | 100% |
| Clean output rate | 100% |
| Slot presence | 96.2% |
| Best val loss | 0.3243 |
Full per-prompt results are in [`report.json`](report.json).
## Charts
| Chart | Description |
|-------|-------------|
| `charts/overall_metrics.png` | Bar chart of all four quality metrics |
| `charts/per_intent_accuracy.png` | Accuracy for each intent |
Regenerate charts from the stored report:
```bash
pip install -r benchmark/requirements.txt
python benchmark/benchmark_charts.py --report benchmark/report.json
python benchmark/benchmark_charts.py --report benchmark/report.json --output-dir benchmark/charts
```
|