File size: 27,844 Bytes
792ca24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
98383f4
 
 
792ca24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
98383f4
 
 
792ca24
 
 
 
 
 
 
 
 
 
 
 
976a803
 
 
 
 
 
 
 
 
 
 
 
792ca24
 
 
 
a98a035
 
 
 
 
 
792ca24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
---
license: apache-2.0
language:
  - en
  - ro
library_name: mlx
pipeline_tag: text-classification
tags:
  - scam-detection
  - fraud-detection
  - smishing
  - phishing
  - on-device
  - mlx
  - gguf
  - qwen3
  - safety
  - romanian
base_model: Qwen/Qwen3-0.6B
datasets:
  - flowxai/scamguardbench
model-index:
  - name: scam-guard-qwen06b
    results:
      - task:
          type: text-classification
          name: Scam verdict classification (ScamGuardBench v0.2, 120-item slice)
        dataset:
          name: ScamGuardBench v0.2
          type: flowxai/scamguardbench
        metrics:
          - type: f1
            name: verdict micro-F1
            value: 0.958
          - type: f1
            name: verdict macro-F1
            value: 0.926
          - type: f1
            name: tactic macro-F1
            value: 0.951
          - type: recall
            name: evidence pass rate
            value: 0.980
          - type: false_positive_rate
            name: legit-confusable FP-rate (scam_likely on legit)
            value: 0.000
      - task:
          type: text-classification
          name: Out-of-distribution fresh CERT-pattern messages (20-item hand-authored set)
        dataset:
          name: ScamGuardBench v0.2
          type: flowxai/scamguardbench
        metrics:
          - type: accuracy
            name: verdict accuracy (correct / 20)
            value: 0.90
          - type: f1
            name: verdict macro-F1 (OOD)
            value: 0.614
          - type: false_positive_rate
            name: OOD legit false-alarm rate
            value: 0.000
---

## Inference contract

Running this model correctly requires its **frozen inference contract** β€” the exact
system prompt, output JSON schema, user-turn format, and constrained-decode spec it
was trained against. See [`inference_contract/`](./inference_contract):

- [`INFERENCE.md`](./inference_contract/INFERENCE.md) β€” wiring guide: system prompt, user turn `[channel: <tag>]\n<message>` (tag `sms`/`email`/`chat`), **constrained JSON decoding** (required β€” pins the enums), and the verbatim-evidence check.
- [`prompt_scamguard_sys_v1.txt`](./inference_contract/prompt_scamguard_sys_v1.txt) β€” the system prompt, verbatim.
- [`schema_scamguard_v1.json`](./inference_contract/schema_scamguard_v1.json) β€” output JSON Schema for constrained decoding.

Prompt version `scamguard_sys_v1`. Do not edit the prompt/schema; the weights are trained against them.

# scam-guard 0.6B (working name) β€” the on-device pick

**An on-device scam & fraud message detector for SMS, email, and chat text (English + Romanian).**

**How to read the numbers on this card.** The fine-tuned figures measured on the
synthetic benchmark are provisional: the benchmark is synthetic, not production traffic.
The out-of-distribution fresh-message results are separately measured and are included
below. Every fine-tuned figure here comes from the final CUDA 3-epoch run, which
supersedes an earlier MLX pass.

This is the **0.6B** model β€” the **smallest and fastest** of the two scam-guard
sizes, and the **on-device target** (~0.4 GB at int4). For higher out-of-distribution
accuracy at a larger footprint, see the sibling **[1.7B quality
pick](https://huggingface.co/flowxai/scam-guard-qwen17b)**.

Given one message, scam-guard returns a **3-level verdict**, the **manipulation
tactics** it found (each with a verbatim evidence span quoted from the message), a
**calm plain-language explanation**, and a **recommended safe action** from a fixed
list. It is built for everyday people β€” explicitly including **elderly and
non-technical users**, who are the most targeted.

The product insight that shapes everything: **the `suspicious` middle level exists
to be honest about uncertainty rather than force a binary.** A consumer safety tool
that must answer "scam or not" will either cry wolf or wave real scams through; a
third honest verdict β€” "this might be fine, verify through your own channel first"
β€” lets the model say *I'm not sure* instead of guessing. Relatedly, scam-guard
**never emits a probability**: an uncalibrated confidence number on a consumer
safety tool is worse than none, so we give you honest per-class behaviour instead
(see [Calibration](#calibration--we-dont-give-you-a-probability)).

- **`flowxai/scam-guard-qwen06b`** (this card, 0.6B) β€” smallest and fastest; the on-device target.
- **`flowxai/scam-guard-qwen17b`** (1.7B, sibling) β€” the more robust choice out-of-distribution (see [Size decision](#size-decision)).

Both are LoRA-fine-tuned from Apache-2.0 Qwen3 base models. On-device formats:
**GGUF** (llama.cpp, int8 `Q8_0` + int4 `Q4_K_M`) and **MLX-quantized** (int4 + int8;
mlx-swift runs these on iOS too).

---

## How do I use it?

Three copy-pasteable ways to turn a message into a verdict. All run **fully
on-device** β€” no network at inference, ever.

Real example input (a fresh Romanian courier-fee smishing message):

```
Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa
vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se
anuleaza in 48h.
```

### (a) llama.cpp / GGUF

Download a GGUF (int8 `Q8_0` recommended) and run the message through it. The model
emits a single strict JSON object.

**`llama-cli` (CPU-only, `-ngl 0`):**

```bash
llama-cli -m scam-guard-qwen06b-Q8_0.gguf -ngl 0 --temp 0 -no-cnv \
  -p "$(cat <<'EOF'
<system prompt: see src/scamguard/schema.py::SYSTEM_PROMPT>
[channel: sms]
Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se anuleaza in 48h.
EOF
)"
```

**`llama-cpp-python`:**

```python
from llama_cpp import Llama
from scamguard.schema import SYSTEM_PROMPT, ScamGuardOutput  # the fixed task prompt + schema

llm = Llama(model_path="scam-guard-qwen06b-Q8_0.gguf", n_gpu_layers=0)
msg = ("Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati "
       "taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata "
       "Livrarea se anuleaza in 48h.")

out = llm.create_chat_completion(
    messages=[
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"[channel: sms]\n{msg}"},
    ],
    temperature=0.0,
)
raw = out["choices"][0]["message"]["content"]
verdict = ScamGuardOutput.model_validate_json(raw)   # strict, extra="forbid"
```

### (b) MLX (Apple Silicon)

Off the MLX-quantized weights (int4/int8):

```bash
mlx_lm.generate --model scam-guard-qwen06b-mlx-int4 --temp 0 \
  --prompt "$(printf '[channel: sms]\nColetul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se anuleaza in 48h.')"
```

```python
from mlx_lm import load, generate
from scamguard.schema import SYSTEM_PROMPT

model, tok = load("scam-guard-qwen06b-mlx-int4")
msg = ("Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati "
       "taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata "
       "Livrarea se anuleaza in 48h.")
prompt = tok.apply_chat_template(
    [{"role": "system", "content": SYSTEM_PROMPT},
     {"role": "user", "content": f"[channel: sms]\n{msg}"}],
    add_generation_prompt=True, enable_thinking=False,
)
raw = generate(model, tok, prompt=prompt, max_tokens=256, verbose=False)
```

Either backend returns the **same strict JSON** for this message:

```json
{
  "verdict": "scam_likely",
  "tactics": [
    {
      "tactic": "subscription_trap",
      "evidence": "achitati taxa vamala de 3,20 lei",
      "explanation": "It asks you to pay a small fee to release a parcel."
    },
    {
      "tactic": "urgency_pressure",
      "evidence": "se anuleaza in 48h",
      "explanation": "It invents a 48-hour deadline to rush you."
    }
  ],
  "explanation": "This looks like a scam because it uses a fake fee to prompt a payment and pressures you with an artificial deadline; do not act on it, and check with the real organisation through a channel you already trust.",
  "recommended_action": "verify_via_official_app_or_site"
}
```

### (c) The demo verdict card (`demo/check.py`)

The reference demo renders that JSON as a card a family member can read. It
hard-enforces the no-network privacy promise:

```bash
echo "Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se anuleaza in 48h." | python demo/check.py
```

Rendered output (the actual card, from `reports/ood_fresh_demo.md`):

```
====================================================================
  scam-guard  β€”  message safety check
====================================================================
  [!]  VERDICT:  Likely a scam

  This message shows clear signs of a scam. You do not need to do
  anything it asks. Take your time β€” real organisations are fine
  with you checking first.
--------------------------------------------------------------------
  What we noticed:
    - Sending a fake renewal or invoice to make you call or click
        seen in: "achitati taxa vamala de 3,20 lei"
    - Rushing you with a deadline or threat
        seen in: "se anuleaza in 48h"
--------------------------------------------------------------------
  What to do:
    Check directly using the company's official app or website
    that you open yourself β€” not the link here.
--------------------------------------------------------------------
  In plain words:
  This looks like a scam because it uses a fake renewal invoice to
  prompt a call and it pressures you with an artificial deadline;
  do not act on it, and check with the real organisation through a
  channel you already trust.
====================================================================
  scam-guard is a helper, not a guarantee. When in doubt, verify
  through a channel you already trust. It never opens links.
====================================================================
```

`demo/check.py` defaults to the 0.6B MLX model (this one); `--backend gguf` and
`--model <path>` switch weights/backend, and `--size 1.7b` selects the sibling.
`--channel {sms,email,chat}` sets the channel tag.

---

## How it works

```mermaid
flowchart LR
  SMS  --> SG["scam-guard 0.6B"]
  Email --> SG
  Chat --> SG
  SG --> V[verdict]
  SG --> T[tactics]
  SG --> E[evidence]
  SG --> X[explanation]
  SG --> A[action]
```

Compact ASCII flow β€” three channels in, one local model, five fields out:

```
SMS
    \
Email ---> scam-guard 0.6B
Chat      /
              |
              + verdict
              + tactics
              + evidence
              + explanation
              + action
```

Under the hood, `scam-guard` is: `tokenizer β†’ Qwen3-0.6B (LoRA fine-tuned) β†’
constrained JSON decode β†’ evidence verifier (verbatim-substring kill-switch:
drops fabricated spans)`.

**The evidence kill-switch is the safety-critical stage.** Every tactic must cite a
span that is a **verbatim substring** of the input message (whitespace-normalized
only β€” no case/diacritic folding). A tactic whose evidence is not found verbatim is
**dropped and counted** as fabricated, so the model can never hallucinate a quote to
justify a warning.

---

## Output schema

A single strict JSON object (`extra="forbid"`, frozen β€” a spurious field like
`confidence` is rejected):

- `verdict` β€” `scam_likely` | `suspicious` | `no_indicators` (**never a probability**).
  - `scam_likely` β€” a clear scam mechanism is present and driven by tactics.
  - `suspicious` β€” the honest middle: signals present but plausibly legitimate, or
    only weak indicators (urgency alone, a link alone, an authority claim with no
    ask). "Verify through your own channel first."
  - `no_indicators` β€” no scam mechanism; legitimate messages can be urgent and contain links.
- `tactics[]` β€” each `{tactic, evidence, explanation}`, where `tactic` is one of 13
  fixed ids and `evidence` is a **verbatim substring** (see the kill-switch above).
- `explanation` β€” one or two calm, actionable sentences.
- `recommended_action` β€” one id from a fixed list of 10 safe actions; the model can
  never compose free-text advice that points back at the scammer's own channel.

The 13 tactics: `urgency_pressure`, `authority_impersonation`, `payment_redirect`,
`credential_phishing`, `courier_customs_fee`, `prize_lottery`, `investment_too_good`,
`romance_advance_fee`, `family_emergency_impersonation`, `tech_support`,
`link_obfuscation`, `refund_overpayment`, `subscription_trap`.

The 10 safe actions: `call_bank_official_number`, `do_not_click_link`,
`verify_via_official_app_or_site`, `call_family_member_known_number`,
`do_not_share_codes_or_credentials`, `do_not_send_money`, `ignore_and_delete`,
`report_to_authorities`, `check_sender_address`, `no_action_needed`.

---

## On-device privacy promise

**scam-guard makes no network calls at inference, ever.** It reasons over the
message text only, fully locally. URL handling is **lexical only** β€” it inspects the
visible URL string (lookalike domains, userinfo tricks, shorteners, punycode hints)
and **never fetches anything**. This is the whole point: it works on messages people
would never upload to a cloud service. The reference demo (`demo/check.py`)
hard-enforces this with a no-network guard. At ~0.4 GB (int4), the 0.6B is the
smallest footprint of the two sizes β€” the one most comfortable on a phone.

---

## Evaluation (0.6B)

Two evaluations, in order of what they tell you:

1. **Out-of-distribution (OOD) fresh messages** β€” the honest real-world signal.
2. **ScamGuardBench v0.2 synthetic bench** β€” a large in-distribution slice the model has
   a home-field advantage on (read the caveat).
3. **Calibration** β€” class-level behaviour, since there is no probability to calibrate.

All fine-tuned numbers below are the **FINAL CUDA 3-epoch** run (`qwen06b_cuda`),
which supersedes the MLX first pass.

### Out-of-distribution results (fresh CERT-pattern messages)

20 **fresh** messages β€” 10 realistic scam patterns modeled on current CERT/DNSC-style
alerts (RO+EN) and 10 genuinely legit messages β€” **hand-authored from public
alert-pattern descriptions, NOT run through the synthetic generator** the model
trained on, sanitized, and asserted text-disjoint from training + every ScamGuardBench
version (`tests/test_ood_fresh.py`). This is the honest generalization test. Full
write-up: `reports/ood_fresh_demo.md`.

| Model | Correct / 20 | Dangerous MISSES (scam→no_indicators) | FALSE ALARMS (legit→scam_likely) | verdict macro-F1 |
| --- | --- | --- | --- | --- |
| `flowxai/scam-guard-qwen06b` (0.6B, on-device) | **18/20 (90%)** | 1 | **0** | 0.614 |
| claude-haiku-4-5 (OOD reference) | **19/20 (95%)** | 0 | 0 | 0.649 |
| `flowxai/scam-guard-qwen17b` (1.7B, sibling) | **19/20 (95%)** | 1 | **0** | 0.967 |

**Bench β†’ OOD gap (the home-field advantage, quantified):** the 0.6B macro-F1 drops
**0.926 β†’ 0.614 (βˆ’0.31)** on fresh messages. On a 20-item set macro-F1 is noisy (one
slip on a rare class moves it a lot), so the plain-verdict accuracy (18/20) is the more
stable read. Legit-FP stays **0.000**.

Honest reading:

- **The one dangerous MISS:** the RO WhatsApp family-emergency scam *"Mama, am
  pierdut telefonul … poti sa imi trimiti 850 lei"*, waved through as
  `no_indicators`. Family-emergency framing without an explicit money-transfer
  keyword can slip past the model. This is a known RO-dominant pattern and the top
  recall gap to fix; the frontier reference (haiku) catches it. It **persists on the
  1.7B at 3 epochs too**, confirming it is a training-data gap, not a size/compute gap.
- **Zero false alarms** β€” no legit message was flagged `scam_likely`. The
  "disable-in-a-week" failure mode did not appear on fresh legit traffic. (The 0.6B
  soft-hedged one legit-adjacent scam to `suspicious` β€” the RO energy-subsidy scam β€”
  reported, not counted as a false alarm.)
- **Format robustness held:** JSON validity 1.000 and evidence pass 1.000 on fresh
  text (no repair/fallback needed).

**RO vs EN, OOD (plain per-language accuracy over the 20-item OOD set β€” 9 RO / 11 EN;
a per-language F1 is not computed in the reports):**

| Model | RO correct | EN correct |
| --- | --- | --- |
| `flowxai/scam-guard-qwen06b` | 7 / 9 | 10 / 11 |

RO is the first-class target language: the RO miss is the single family-emergency
scam (the 0.6B also soft-hedges the RO energy-subsidy scam to `suspicious`). On the
in-distribution bench the RO-dominant tactics score at ceiling
(`family_emergency_impersonation` F1 **1.000**, `courier_customs_fee` **1.000**,
`refund_overpayment` **0.889**), which is exactly why the OOD RO family-emergency
miss is the honest gap to close.

### ScamGuardBench v0.2 (synthetic, in-distribution) β€” with the home-field caveat

> **Honesty caveat β€” home-field advantage (read before citing these numbers).** The
> fine-tuned numbers **beat the frontier reference on this benchmark, and that does NOT
> mean the model is a better real-world scam detector.** ScamGuardBench v0.2 is built from the
> **same synthetic generator** the model trained on (held-out split,
> contamination-verified β€” no leakage of specific messages, but the **same
> distribution**: same output format, phrasing, tactic-to-message style). The
> fine-tune learned exactly that style. **The frontier reference is the honest upper
> reference for a cold-prompted generalist** β€” a better OOD proxy than the fine-tuned
> column β€” and the OOD results above are the real-world signal these bench numbers
> flatter.

Seeded **120-message** slice (18 `suspicious`, 39 legit-confusable). A false positive
is a `scam_likely` verdict on a legitimate message; `suspicious` is reported but
**not** counted as an FP. FP-rate on the legit-confusable subset is the **release gate**.

| Metric | **qwen06b (CUDA)** | claude-haiku-4-5 (frontier ref) | keyword (lower ref) |
| --- | --- | --- | --- |
| JSON validity (deployed decoder) | 1.000 | 1.000 | 1.000 |
| **raw** JSON validity (no repair) | 1.000 | n/a | n/a |
| verdict micro-F1 | 0.958 | 0.900 | 0.483 |
| verdict macro-F1 | **0.926** | 0.829 | 0.482 |
| tactic macro-F1 | 0.951 | 0.770 | 0.319 |
| evidence pass rate | 0.980 | 0.969 | 1.000 |
| **legit-confusable FP-rate** | **0.000** | 0.051 | 0.026 |

The reference `claude-haiku-4-5` **fails the FP gate** (legit-confusable FP 0.051,
2/39 β€” it over-flags genuine family money requests), while the fine-tune holds
**0.000**. That is the home-field-advantage signal, not a claim of superior
real-world judgment. `family_money_request` (a genuine "send me money" from family)
is the hardest legit class β€” every frontier model over-flags it β€” while this
fine-tune holds **0.000**.

> **Full-budget note (MLX β†’ CUDA).** The shipped 0.6B is the CUDA 3-epoch run. It
> supersedes the earlier MLX ~1-epoch first pass (macro-F1 0.903 β†’ **0.926**,
> micro 0.942 β†’ **0.958**), while **holding legit-FP at 0.000**. The full budget
> bought +2.3 macro-F1 points in-distribution.

### Calibration β€” we don't give you a probability

scam-guard **deliberately emits no confidence percentage**. An uncalibrated number
on a consumer safety tool is worse than none: it invites false precision on a
judgement that is genuinely uncertain. Instead of a probability to calibrate, we give
you **honest per-class behaviour** so you know where the model is weak and where it is
safe.

**Confusion matrix** (`qwen06b` CUDA, ScamGuardBench v0.2, 120-item slice; rows = gold,
cols = predicted):

| gold \ pred | scam_likely | suspicious | no_indicators | recall |
| --- | --- | --- | --- | --- |
| scam_likely | 63 | 0 | 0 | 1.000 (n=63) |
| suspicious | 0 | 13 | 5 | 0.722 (n=18) |
| no_indicators | 0 | 0 | 39 | 1.000 (n=39) |

**Per-class precision / recall / F1** (`reports/eval_frontier.md`) β€” micro-F1 0.958,
macro-F1 0.926:

| Verdict | Precision | Recall | F1 | Support |
| --- | --- | --- | --- | --- |
| scam_likely | 1.000 | 1.000 | 1.000 | 63 |
| suspicious | 1.000 | 0.722 | 0.839 | 18 |
| no_indicators | 0.886 | 1.000 | 0.940 | 39 |

**Where the model is weak, and where it is safe.** The weak class is `suspicious`
recall (0.722, 13/18) β€” the honest middle is the hardest to catch β€” but `suspicious`
**precision is 1.000** (it never over-calls the middle). **Crucially, every
`suspicious` miss bleeds into `no_indicators` (the safe direction), never into
`scam_likely`, and no legit item is ever flipped to `scam_likely`** β€” which is why
the legit-confusable FP-rate is **0.000**. The model under-warns on ambiguous
messages rather than over-warning on real ones; for a consumer guard that is the
failure mode you want.

### Size decision

- **This model (`flowxai/scam-guard-qwen06b`, 0.6B) β€” smallest and fastest, weaker
  OOD.** Passes all three release gates on the bench (JSON >99%, evidence >95%,
  legit-FP <3%) and is a genuinely defensible on-device ship at ~0.4 GB (int4), but
  drops hardest out-of-distribution (macro-F1 βˆ’0.31, 18/20).
- **The sibling `flowxai/scam-guard-qwen17b` (1.7B) β€” the more robust choice.** On
  fresh messages the extra capacity generalizes materially better (19/20, no macro-F1
  drop), a gap the in-distribution bench (both ~0.9+) did not surface. Where
  robustness matters more than size/latency, [ship the
  1.7B](https://huggingface.co/flowxai/scam-guard-qwen17b).
- **Both sizes share the RO family-emergency recall gap** (the one dangerous OOD
  miss) and both hold legit-FP at 0.000. That shared recall gap is the honest thing
  to fix before any release claim.

---

## Formats & on-device latency (0.6B)

**We targeted 150 ms. We measured ~1.1 s (best path). This target is currently not
met.** scam-guard emits a full multi-field JSON card (~138 tokens), not a single
label, so decode dominates latency.

The two headline configs for the 0.6B on the representative 300-char SMS (median,
M3 Max):

| Config | median | meets 150 ms? |
| --- | --- | --- |
| MLX int4, Apple-Silicon GPU (Metal) | ~1.1 s | no |
| GGUF int8 (`Q8_0`), CPU-only (`-ngl 0`) | ~3.3 s | no |

int4 GGUF (`Q4_K_M`) is a poor trade on this small model β€” on CPU it is *slower*
than int8 and degrades quality (a spot-check sample became invalid JSON) β€” so
**`Q8_0` is the recommended GGUF quant**, and MLX int4 is the fastest quality-holding
path.

<details>
<summary>Full per-quant numbers (0.6B)</summary>

GGUF (llama.cpp), CPU-only (`-ngl 0`), M3 Max β€” full JSON card, median over N=12
(`reports/benchmark_gguf.json`):

| quant | file size | median | JSON spot-check |
| --- | --- | --- | --- |
| int8 (Q8_0) | 639 MB | 3204 ms | 4/4 valid |
| int4 (Q4_K_M) | 397 MB | 4000 ms | 3/4 (degraded) |

MLX-quantized, Apple-Silicon GPU (Metal), M3 Max (`reports/benchmark_mlx_quant.json`):

| quant | weights size | median | JSON spot-check |
| --- | --- | --- | --- |
| int4 | 335 MB | 1153 ms | 4/4 valid |
| int8 | 633 MB | 1281 ms | 4/4 valid |

bf16 MLX-Metal latency/memory and per-input-length detail are in
`reports/benchmark.md`. **Core ML** (.mlpackage) was attempted; the LLM→Core ML
conversion is finicky (stateful KV-cache handling) and the attempt is documented in
`PROGRESS.md` rather than shipped as a fabricated artifact. GGUF and MLX are the
recommended on-device paths today.

</details>

A human decision is needed at the release STOP: accept the ~1–3 s latency, ship the
~1.1 s GPU/MLX path, or shrink the output contract to approach 150 ms.

---

## Intended use & limitations

**Intended use.** A **consumer triage aid** that explains *why* a message looks risky
and points you to *your own* trusted channel to verify. Runs on-device. Languages: EN
and RO at v1 (Romanian is a first-class citizen, not an afterthought); PL/HU planned
fast-follow through the same pipeline.

**Out of scope & limitations.**

- **Not a guarantee.** A verdict is a signal, not proof. Scammers adapt continuously;
  the benchmark is versioned because patterns rotate.
- **Verdicts can be wrong in both directions** β€” a real scam may score
  `no_indicators` (the OOD RO family-emergency miss is a documented example), and a
  legitimate message may score `suspicious`. The `suspicious` middle level exists to
  be honest about uncertainty rather than force a binary.
- **Known recall gap:** the RO family-emergency pattern (framing without an explicit
  money-transfer keyword) can slip past this model (and the 1.7B); it is a confirmed
  training-data gap, fixable with a data addition before any release claim.
- **The model never fetches URLs.** It cannot tell you where a link *actually*
  resolves, only what the visible string suggests. A lexically-clean URL can still be
  malicious.
- Not a replacement for a bank's fraud line, a national anti-fraud service, or human
  judgment. The recommended action always routes to *your own* channel.
- **Text only** at v1: no image/OCR, no audio, no attachment parsing, no
  email-header/routing analysis.

---

## Training data

- **Public seed layer (relabeled):** UCI SMS Spam Collection (CC BY 4.0), enron_ham
  (SetFit/enron_spam ham slice; no explicit license β†’ reference/research-use),
  phishing_email (zefang-liu; LGPL-3.0). Relabeled into the verdict+tactic scheme with
  the evidence kill-switch; per-source human spot-check kept relabel disagreement
  under the 10% gate.
- **Synthetic layer:** generated from specs (tactic Γ— channel Γ— language Γ— register),
  including a `suspicious` middle-ground tier and adversarial keyword-evasion
  paraphrases (the `hard` subset). Every bench-destined message passes a sanitizer
  audit (reserved domains, non-dialable phones).
- **Balance:** β‰₯45% legitimate messages, RO β‰₯35%, ~40% of RO diacritic-free,
  `suspicious` ~14.5% of the SFT train set.
- Split discipline: synthetic by spec-family, public by source-text hash; paraphrases
  follow their parent's split. **Train ∩ ScamGuardBench = βˆ…** (contamination-verified).

Both sizes trained from the **same data**; the size decision is evidence-based (above).

## Fine-tuning

**CUDA full-budget 3-epoch LoRA** (`training/train_cuda.py`, transformers + PEFT +
trl `SFTTrainer`), from `Qwen/Qwen3-0.6B`. LoRA rank 16 / alpha 32 / dropout 0.05,
all-linear modules, adamw, cosine 2e-5β†’2e-6 with 60-step warmup, effective batch 4,
seq 1280, bf16, gradient-checkpointing, seed 20260703; **3 real epochs** on an
on-demand NVIDIA L4 (~2h16m, 0 OOM), thinking disabled. This **supersedes** the MLX
first pass (~1 epoch, bounded by Metal stalls); the full budget bought +2.3 macro-F1
points in-distribution.

> **Base-model note.** The exact repo `Qwen3-0.6B-Instruct` does **not** exist on
> Hugging Face β€” Qwen3 merged instruct+thinking into the single base repo
> `Qwen/Qwen3-0.6B` (instruction-capable, Apache-2.0). We fine-tune it with thinking
> disabled.

---

## Dual-use statement

We release a **detector**, a **benchmark**, and **pattern-level** explanations. We do
**not** release the scam-variant generation prompts as a standalone tool. Public
benchmark scam texts carry no real dialable phone numbers and no working URLs
(reserved/example domains and clearly-fake numbers only). Output explanations describe
the manipulation *pattern*, never instructions for constructing one.

## Links

- **Sibling model (quality pick, 1.7B):** [`flowxai/scam-guard-qwen17b`](https://huggingface.co/flowxai/scam-guard-qwen17b) β€” higher OOD accuracy, larger footprint (~1.1 GB int4).
- **Benchmark / dataset:** [`flowxai/scamguardbench`](https://huggingface.co/datasets/flowxai/scamguardbench).

## License

Apache-2.0 (weights, code, and benchmark). Base model `Qwen/Qwen3-0.6B` is Apache-2.0.
</content>
</invoke>