Make repo self-contained: rewrite docs, single-model benchmark, remove external references
dbb5d78 verified |
Download benchmark/README.md from prathamkode/smartwatch-lm-0.2: direct link, hf CLI and curl.
- Browser
- Download file 1.89 kB
-
https://huggingface.co/prathamkode/smartwatch-lm-0.2/resolve/dbb5d78326ae2eace84b67a293deaf2dcc815dab/benchmark/README.md
- Command line
-
hf download hf://prathamkode/smartwatch-lm-0.2@dbb5d78326ae2eace84b67a293deaf2dcc815dab/benchmark/README.md
-
curl -L -o README.md https://huggingface.co/prathamkode/smartwatch-lm-0.2/resolve/dbb5d78326ae2eace84b67a293deaf2dcc815dab/benchmark/README.md
1.89 kB
Smartwatch LM v0.2 — Quality Benchmark
Quality evaluation for Smartwatch LM v0.2 on a fixed set of 39 golden prompts covering all supported intents.
Golden prompts
benchmark_prompts.json holds 39 prompts — one per intent category plus multi-intent and conversational cases.
Each entry:
{
"id": "get_steps_...",
"prompt": "How many steps today?",
"expected_intent": "GET_STEPS",
"expected_intents": ["GET_STEPS"],
"expected_slots": ["STEPS_TODAY", "STEP_GOAL"],
"source": "golden_set"
}
Metrics
| Metric | Meaning |
|---|---|
| Intent accuracy | Predicted intent matches expected (or is in expected_intents for combo cases) |
| Intent parse rate | Reply contains a valid <INTENT:...> tag |
| Clean output rate | Raw decode has no BPE junk (Ġ, Ċ) before cleanup |
| Slot presence | Fraction of expected slot placeholders found in the cleaned template |
Generation settings used for evaluation: max_new_tokens=40, temperature=0.5, top_k=40, fresh history per prompt, fixed seed for reproducibility.
v0.2 results
| Metric | Score |
|---|---|
| Intent accuracy | 100% |
| Intent parse rate | 100% |
| Clean output rate | 100% |
| Slot presence | 96.2% |
| Best val loss | 0.3243 |
Full per-prompt results are in report.json.
Charts
| Chart | Description |
|---|---|
charts/overall_metrics.png |
Bar chart of all four quality metrics |
charts/per_intent_accuracy.png |
Accuracy for each intent |
Regenerate charts from the stored report:
pip install -r benchmark/requirements.txt
python benchmark/benchmark_charts.py --report benchmark/report.json
python benchmark/benchmark_charts.py --report benchmark/report.json --output-dir benchmark/charts