prathamkode's picture
Make repo self-contained: rewrite docs, single-model benchmark, remove external references
dbb5d78 verified
|
Raw History Blame
1.89 kB

Smartwatch LM v0.2 — Quality Benchmark

Quality evaluation for Smartwatch LM v0.2 on a fixed set of 39 golden prompts covering all supported intents.

Golden prompts

benchmark_prompts.json holds 39 prompts — one per intent category plus multi-intent and conversational cases.

Each entry:

{
  "id": "get_steps_...",
  "prompt": "How many steps today?",
  "expected_intent": "GET_STEPS",
  "expected_intents": ["GET_STEPS"],
  "expected_slots": ["STEPS_TODAY", "STEP_GOAL"],
  "source": "golden_set"
}

Metrics

Metric Meaning
Intent accuracy Predicted intent matches expected (or is in expected_intents for combo cases)
Intent parse rate Reply contains a valid <INTENT:...> tag
Clean output rate Raw decode has no BPE junk (Ġ, Ċ) before cleanup
Slot presence Fraction of expected slot placeholders found in the cleaned template

Generation settings used for evaluation: max_new_tokens=40, temperature=0.5, top_k=40, fresh history per prompt, fixed seed for reproducibility.

v0.2 results

Metric Score
Intent accuracy 100%
Intent parse rate 100%
Clean output rate 100%
Slot presence 96.2%
Best val loss 0.3243

Full per-prompt results are in report.json.

Charts

Chart Description
charts/overall_metrics.png Bar chart of all four quality metrics
charts/per_intent_accuracy.png Accuracy for each intent

Regenerate charts from the stored report:

pip install -r benchmark/requirements.txt
python benchmark/benchmark_charts.py --report benchmark/report.json
python benchmark/benchmark_charts.py --report benchmark/report.json --output-dir benchmark/charts