# Smartwatch LM v0.2 — Quality Benchmark Quality evaluation for **Smartwatch LM v0.2** on a fixed set of 39 golden prompts covering all supported intents. ## Golden prompts [`benchmark_prompts.json`](benchmark_prompts.json) holds 39 prompts — one per intent category plus multi-intent and conversational cases. Each entry: ```json { "id": "get_steps_...", "prompt": "How many steps today?", "expected_intent": "GET_STEPS", "expected_intents": ["GET_STEPS"], "expected_slots": ["STEPS_TODAY", "STEP_GOAL"], "source": "golden_set" } ``` ## Metrics | Metric | Meaning | |--------|---------| | **Intent accuracy** | Predicted intent matches expected (or is in `expected_intents` for combo cases) | | **Intent parse rate** | Reply contains a valid `` tag | | **Clean output rate** | Raw decode has no BPE junk (`Ġ`, `Ċ`) before cleanup | | **Slot presence** | Fraction of expected slot placeholders found in the cleaned template | Generation settings used for evaluation: `max_new_tokens=40`, `temperature=0.5`, `top_k=40`, fresh history per prompt, fixed seed for reproducibility. ## v0.2 results | Metric | Score | |--------|------:| | Intent accuracy | 100% | | Intent parse rate | 100% | | Clean output rate | 100% | | Slot presence | 96.2% | | Best val loss | 0.3243 | Full per-prompt results are in [`report.json`](report.json). ## Charts | Chart | Description | |-------|-------------| | `charts/overall_metrics.png` | Bar chart of all four quality metrics | | `charts/per_intent_accuracy.png` | Accuracy for each intent | Regenerate charts from the stored report: ```bash pip install -r benchmark/requirements.txt python benchmark/benchmark_charts.py --report benchmark/report.json python benchmark/benchmark_charts.py --report benchmark/report.json --output-dir benchmark/charts ```