File size: 1,890 Bytes
dbb5d78
a330cfa
dbb5d78
a330cfa
 
 
dbb5d78
a330cfa
 
 
 
 
 
 
 
 
 
dbb5d78
a330cfa
 
 
dbb5d78
a330cfa
dbb5d78
 
 
 
 
 
a330cfa
dbb5d78
a330cfa
dbb5d78
a330cfa
dbb5d78
 
 
 
 
 
 
a330cfa
dbb5d78
a330cfa
dbb5d78
a330cfa
 
 
dbb5d78
 
a330cfa
dbb5d78
a330cfa
 
dbb5d78
a330cfa
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
# Smartwatch LM v0.2 — Quality Benchmark

Quality evaluation for **Smartwatch LM v0.2** on a fixed set of 39 golden prompts covering all supported intents.

## Golden prompts

[`benchmark_prompts.json`](benchmark_prompts.json) holds 39 prompts — one per intent category plus multi-intent and conversational cases.

Each entry:

```json

{

  "id": "get_steps_...",

  "prompt": "How many steps today?",

  "expected_intent": "GET_STEPS",

  "expected_intents": ["GET_STEPS"],

  "expected_slots": ["STEPS_TODAY", "STEP_GOAL"],

  "source": "golden_set"

}

```

## Metrics

| Metric | Meaning |
|--------|---------|
| **Intent accuracy** | Predicted intent matches expected (or is in `expected_intents` for combo cases) |
| **Intent parse rate** | Reply contains a valid `<INTENT:...>` tag |
| **Clean output rate** | Raw decode has no BPE junk (`Ġ`, `Ċ`) before cleanup |
| **Slot presence** | Fraction of expected slot placeholders found in the cleaned template |

Generation settings used for evaluation: `max_new_tokens=40`, `temperature=0.5`, `top_k=40`, fresh history per prompt, fixed seed for reproducibility.

## v0.2 results

| Metric | Score |
|--------|------:|
| Intent accuracy | 100% |
| Intent parse rate | 100% |
| Clean output rate | 100% |
| Slot presence | 96.2% |
| Best val loss | 0.3243 |

Full per-prompt results are in [`report.json`](report.json).

## Charts

| Chart | Description |
|-------|-------------|
| `charts/overall_metrics.png` | Bar chart of all four quality metrics |
| `charts/per_intent_accuracy.png` | Accuracy for each intent |

Regenerate charts from the stored report:

```bash

pip install -r benchmark/requirements.txt

python benchmark/benchmark_charts.py --report benchmark/report.json

python benchmark/benchmark_charts.py --report benchmark/report.json --output-dir benchmark/charts

```