Text Classification
MLX
Safetensors
GGUF
English
Romanian
qwen3
scam-detection
fraud-detection
smishing
phishing
on-device
safety
romanian
Eval Results (legacy)
conversational
Instructions to use flowxai/scam-guard-qwen06b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use flowxai/scam-guard-qwen06b with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download flowxai/scam-guard-qwen06b --local-dir scam-guard-qwen06b
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use flowxai/scam-guard-qwen06b with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf flowxai/scam-guard-qwen06b:Q8_0 # Run inference directly in the terminal: llama cli -hf flowxai/scam-guard-qwen06b:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf flowxai/scam-guard-qwen06b:Q8_0 # Run inference directly in the terminal: llama cli -hf flowxai/scam-guard-qwen06b:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf flowxai/scam-guard-qwen06b:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf flowxai/scam-guard-qwen06b:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf flowxai/scam-guard-qwen06b:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf flowxai/scam-guard-qwen06b:Q8_0
Use Docker
docker model run hf.co/flowxai/scam-guard-qwen06b:Q8_0
- LM Studio
- Jan
- Ollama
How to use flowxai/scam-guard-qwen06b with Ollama:
ollama run hf.co/flowxai/scam-guard-qwen06b:Q8_0
- Unsloth Desktop
- Pi
How to use flowxai/scam-guard-qwen06b with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "flowxai/scam-guard-qwen06b"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "flowxai/scam-guard-qwen06b" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use flowxai/scam-guard-qwen06b with Docker Model Runner:
docker model run hf.co/flowxai/scam-guard-qwen06b:Q8_0
- Lemonade
How to use flowxai/scam-guard-qwen06b with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull flowxai/scam-guard-qwen06b:Q8_0
Run and chat with the model
lemonade run user.scam-guard-qwen06b-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use flowxai/scam-guard-qwen06b with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "flowxai/scam-guard-qwen06b"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default flowxai/scam-guard-qwen06b
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use flowxai/scam-guard-qwen06b with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "flowxai/scam-guard-qwen06b"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "flowxai/scam-guard-qwen06b" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 27,844 Bytes
792ca24 98383f4 792ca24 98383f4 792ca24 976a803 792ca24 a98a035 792ca24 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 | ---
license: apache-2.0
language:
- en
- ro
library_name: mlx
pipeline_tag: text-classification
tags:
- scam-detection
- fraud-detection
- smishing
- phishing
- on-device
- mlx
- gguf
- qwen3
- safety
- romanian
base_model: Qwen/Qwen3-0.6B
datasets:
- flowxai/scamguardbench
model-index:
- name: scam-guard-qwen06b
results:
- task:
type: text-classification
name: Scam verdict classification (ScamGuardBench v0.2, 120-item slice)
dataset:
name: ScamGuardBench v0.2
type: flowxai/scamguardbench
metrics:
- type: f1
name: verdict micro-F1
value: 0.958
- type: f1
name: verdict macro-F1
value: 0.926
- type: f1
name: tactic macro-F1
value: 0.951
- type: recall
name: evidence pass rate
value: 0.980
- type: false_positive_rate
name: legit-confusable FP-rate (scam_likely on legit)
value: 0.000
- task:
type: text-classification
name: Out-of-distribution fresh CERT-pattern messages (20-item hand-authored set)
dataset:
name: ScamGuardBench v0.2
type: flowxai/scamguardbench
metrics:
- type: accuracy
name: verdict accuracy (correct / 20)
value: 0.90
- type: f1
name: verdict macro-F1 (OOD)
value: 0.614
- type: false_positive_rate
name: OOD legit false-alarm rate
value: 0.000
---
## Inference contract
Running this model correctly requires its **frozen inference contract** β the exact
system prompt, output JSON schema, user-turn format, and constrained-decode spec it
was trained against. See [`inference_contract/`](./inference_contract):
- [`INFERENCE.md`](./inference_contract/INFERENCE.md) β wiring guide: system prompt, user turn `[channel: <tag>]\n<message>` (tag `sms`/`email`/`chat`), **constrained JSON decoding** (required β pins the enums), and the verbatim-evidence check.
- [`prompt_scamguard_sys_v1.txt`](./inference_contract/prompt_scamguard_sys_v1.txt) β the system prompt, verbatim.
- [`schema_scamguard_v1.json`](./inference_contract/schema_scamguard_v1.json) β output JSON Schema for constrained decoding.
Prompt version `scamguard_sys_v1`. Do not edit the prompt/schema; the weights are trained against them.
# scam-guard 0.6B (working name) β the on-device pick
**An on-device scam & fraud message detector for SMS, email, and chat text (English + Romanian).**
**How to read the numbers on this card.** The fine-tuned figures measured on the
synthetic benchmark are provisional: the benchmark is synthetic, not production traffic.
The out-of-distribution fresh-message results are separately measured and are included
below. Every fine-tuned figure here comes from the final CUDA 3-epoch run, which
supersedes an earlier MLX pass.
This is the **0.6B** model β the **smallest and fastest** of the two scam-guard
sizes, and the **on-device target** (~0.4 GB at int4). For higher out-of-distribution
accuracy at a larger footprint, see the sibling **[1.7B quality
pick](https://huggingface.co/flowxai/scam-guard-qwen17b)**.
Given one message, scam-guard returns a **3-level verdict**, the **manipulation
tactics** it found (each with a verbatim evidence span quoted from the message), a
**calm plain-language explanation**, and a **recommended safe action** from a fixed
list. It is built for everyday people β explicitly including **elderly and
non-technical users**, who are the most targeted.
The product insight that shapes everything: **the `suspicious` middle level exists
to be honest about uncertainty rather than force a binary.** A consumer safety tool
that must answer "scam or not" will either cry wolf or wave real scams through; a
third honest verdict β "this might be fine, verify through your own channel first"
β lets the model say *I'm not sure* instead of guessing. Relatedly, scam-guard
**never emits a probability**: an uncalibrated confidence number on a consumer
safety tool is worse than none, so we give you honest per-class behaviour instead
(see [Calibration](#calibration--we-dont-give-you-a-probability)).
- **`flowxai/scam-guard-qwen06b`** (this card, 0.6B) β smallest and fastest; the on-device target.
- **`flowxai/scam-guard-qwen17b`** (1.7B, sibling) β the more robust choice out-of-distribution (see [Size decision](#size-decision)).
Both are LoRA-fine-tuned from Apache-2.0 Qwen3 base models. On-device formats:
**GGUF** (llama.cpp, int8 `Q8_0` + int4 `Q4_K_M`) and **MLX-quantized** (int4 + int8;
mlx-swift runs these on iOS too).
---
## How do I use it?
Three copy-pasteable ways to turn a message into a verdict. All run **fully
on-device** β no network at inference, ever.
Real example input (a fresh Romanian courier-fee smishing message):
```
Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa
vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se
anuleaza in 48h.
```
### (a) llama.cpp / GGUF
Download a GGUF (int8 `Q8_0` recommended) and run the message through it. The model
emits a single strict JSON object.
**`llama-cli` (CPU-only, `-ngl 0`):**
```bash
llama-cli -m scam-guard-qwen06b-Q8_0.gguf -ngl 0 --temp 0 -no-cnv \
-p "$(cat <<'EOF'
<system prompt: see src/scamguard/schema.py::SYSTEM_PROMPT>
[channel: sms]
Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se anuleaza in 48h.
EOF
)"
```
**`llama-cpp-python`:**
```python
from llama_cpp import Llama
from scamguard.schema import SYSTEM_PROMPT, ScamGuardOutput # the fixed task prompt + schema
llm = Llama(model_path="scam-guard-qwen06b-Q8_0.gguf", n_gpu_layers=0)
msg = ("Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati "
"taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata "
"Livrarea se anuleaza in 48h.")
out = llm.create_chat_completion(
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"[channel: sms]\n{msg}"},
],
temperature=0.0,
)
raw = out["choices"][0]["message"]["content"]
verdict = ScamGuardOutput.model_validate_json(raw) # strict, extra="forbid"
```
### (b) MLX (Apple Silicon)
Off the MLX-quantized weights (int4/int8):
```bash
mlx_lm.generate --model scam-guard-qwen06b-mlx-int4 --temp 0 \
--prompt "$(printf '[channel: sms]\nColetul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se anuleaza in 48h.')"
```
```python
from mlx_lm import load, generate
from scamguard.schema import SYSTEM_PROMPT
model, tok = load("scam-guard-qwen06b-mlx-int4")
msg = ("Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati "
"taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata "
"Livrarea se anuleaza in 48h.")
prompt = tok.apply_chat_template(
[{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"[channel: sms]\n{msg}"}],
add_generation_prompt=True, enable_thinking=False,
)
raw = generate(model, tok, prompt=prompt, max_tokens=256, verbose=False)
```
Either backend returns the **same strict JSON** for this message:
```json
{
"verdict": "scam_likely",
"tactics": [
{
"tactic": "subscription_trap",
"evidence": "achitati taxa vamala de 3,20 lei",
"explanation": "It asks you to pay a small fee to release a parcel."
},
{
"tactic": "urgency_pressure",
"evidence": "se anuleaza in 48h",
"explanation": "It invents a 48-hour deadline to rush you."
}
],
"explanation": "This looks like a scam because it uses a fake fee to prompt a payment and pressures you with an artificial deadline; do not act on it, and check with the real organisation through a channel you already trust.",
"recommended_action": "verify_via_official_app_or_site"
}
```
### (c) The demo verdict card (`demo/check.py`)
The reference demo renders that JSON as a card a family member can read. It
hard-enforces the no-network privacy promise:
```bash
echo "Coletul dumneavoastra nu a putut fi livrat. Pentru reprogramare achitati taxa vamala de 3,20 lei aici: http://colet-reprogramare.example.net/plata Livrarea se anuleaza in 48h." | python demo/check.py
```
Rendered output (the actual card, from `reports/ood_fresh_demo.md`):
```
====================================================================
scam-guard β message safety check
====================================================================
[!] VERDICT: Likely a scam
This message shows clear signs of a scam. You do not need to do
anything it asks. Take your time β real organisations are fine
with you checking first.
--------------------------------------------------------------------
What we noticed:
- Sending a fake renewal or invoice to make you call or click
seen in: "achitati taxa vamala de 3,20 lei"
- Rushing you with a deadline or threat
seen in: "se anuleaza in 48h"
--------------------------------------------------------------------
What to do:
Check directly using the company's official app or website
that you open yourself β not the link here.
--------------------------------------------------------------------
In plain words:
This looks like a scam because it uses a fake renewal invoice to
prompt a call and it pressures you with an artificial deadline;
do not act on it, and check with the real organisation through a
channel you already trust.
====================================================================
scam-guard is a helper, not a guarantee. When in doubt, verify
through a channel you already trust. It never opens links.
====================================================================
```
`demo/check.py` defaults to the 0.6B MLX model (this one); `--backend gguf` and
`--model <path>` switch weights/backend, and `--size 1.7b` selects the sibling.
`--channel {sms,email,chat}` sets the channel tag.
---
## How it works
```mermaid
flowchart LR
SMS --> SG["scam-guard 0.6B"]
Email --> SG
Chat --> SG
SG --> V[verdict]
SG --> T[tactics]
SG --> E[evidence]
SG --> X[explanation]
SG --> A[action]
```
Compact ASCII flow β three channels in, one local model, five fields out:
```
SMS
\
Email ---> scam-guard 0.6B
Chat /
|
+ verdict
+ tactics
+ evidence
+ explanation
+ action
```
Under the hood, `scam-guard` is: `tokenizer β Qwen3-0.6B (LoRA fine-tuned) β
constrained JSON decode β evidence verifier (verbatim-substring kill-switch:
drops fabricated spans)`.
**The evidence kill-switch is the safety-critical stage.** Every tactic must cite a
span that is a **verbatim substring** of the input message (whitespace-normalized
only β no case/diacritic folding). A tactic whose evidence is not found verbatim is
**dropped and counted** as fabricated, so the model can never hallucinate a quote to
justify a warning.
---
## Output schema
A single strict JSON object (`extra="forbid"`, frozen β a spurious field like
`confidence` is rejected):
- `verdict` β `scam_likely` | `suspicious` | `no_indicators` (**never a probability**).
- `scam_likely` β a clear scam mechanism is present and driven by tactics.
- `suspicious` β the honest middle: signals present but plausibly legitimate, or
only weak indicators (urgency alone, a link alone, an authority claim with no
ask). "Verify through your own channel first."
- `no_indicators` β no scam mechanism; legitimate messages can be urgent and contain links.
- `tactics[]` β each `{tactic, evidence, explanation}`, where `tactic` is one of 13
fixed ids and `evidence` is a **verbatim substring** (see the kill-switch above).
- `explanation` β one or two calm, actionable sentences.
- `recommended_action` β one id from a fixed list of 10 safe actions; the model can
never compose free-text advice that points back at the scammer's own channel.
The 13 tactics: `urgency_pressure`, `authority_impersonation`, `payment_redirect`,
`credential_phishing`, `courier_customs_fee`, `prize_lottery`, `investment_too_good`,
`romance_advance_fee`, `family_emergency_impersonation`, `tech_support`,
`link_obfuscation`, `refund_overpayment`, `subscription_trap`.
The 10 safe actions: `call_bank_official_number`, `do_not_click_link`,
`verify_via_official_app_or_site`, `call_family_member_known_number`,
`do_not_share_codes_or_credentials`, `do_not_send_money`, `ignore_and_delete`,
`report_to_authorities`, `check_sender_address`, `no_action_needed`.
---
## On-device privacy promise
**scam-guard makes no network calls at inference, ever.** It reasons over the
message text only, fully locally. URL handling is **lexical only** β it inspects the
visible URL string (lookalike domains, userinfo tricks, shorteners, punycode hints)
and **never fetches anything**. This is the whole point: it works on messages people
would never upload to a cloud service. The reference demo (`demo/check.py`)
hard-enforces this with a no-network guard. At ~0.4 GB (int4), the 0.6B is the
smallest footprint of the two sizes β the one most comfortable on a phone.
---
## Evaluation (0.6B)
Two evaluations, in order of what they tell you:
1. **Out-of-distribution (OOD) fresh messages** β the honest real-world signal.
2. **ScamGuardBench v0.2 synthetic bench** β a large in-distribution slice the model has
a home-field advantage on (read the caveat).
3. **Calibration** β class-level behaviour, since there is no probability to calibrate.
All fine-tuned numbers below are the **FINAL CUDA 3-epoch** run (`qwen06b_cuda`),
which supersedes the MLX first pass.
### Out-of-distribution results (fresh CERT-pattern messages)
20 **fresh** messages β 10 realistic scam patterns modeled on current CERT/DNSC-style
alerts (RO+EN) and 10 genuinely legit messages β **hand-authored from public
alert-pattern descriptions, NOT run through the synthetic generator** the model
trained on, sanitized, and asserted text-disjoint from training + every ScamGuardBench
version (`tests/test_ood_fresh.py`). This is the honest generalization test. Full
write-up: `reports/ood_fresh_demo.md`.
| Model | Correct / 20 | Dangerous MISSES (scamβno_indicators) | FALSE ALARMS (legitβscam_likely) | verdict macro-F1 |
| --- | --- | --- | --- | --- |
| `flowxai/scam-guard-qwen06b` (0.6B, on-device) | **18/20 (90%)** | 1 | **0** | 0.614 |
| claude-haiku-4-5 (OOD reference) | **19/20 (95%)** | 0 | 0 | 0.649 |
| `flowxai/scam-guard-qwen17b` (1.7B, sibling) | **19/20 (95%)** | 1 | **0** | 0.967 |
**Bench β OOD gap (the home-field advantage, quantified):** the 0.6B macro-F1 drops
**0.926 β 0.614 (β0.31)** on fresh messages. On a 20-item set macro-F1 is noisy (one
slip on a rare class moves it a lot), so the plain-verdict accuracy (18/20) is the more
stable read. Legit-FP stays **0.000**.
Honest reading:
- **The one dangerous MISS:** the RO WhatsApp family-emergency scam *"Mama, am
pierdut telefonul β¦ poti sa imi trimiti 850 lei"*, waved through as
`no_indicators`. Family-emergency framing without an explicit money-transfer
keyword can slip past the model. This is a known RO-dominant pattern and the top
recall gap to fix; the frontier reference (haiku) catches it. It **persists on the
1.7B at 3 epochs too**, confirming it is a training-data gap, not a size/compute gap.
- **Zero false alarms** β no legit message was flagged `scam_likely`. The
"disable-in-a-week" failure mode did not appear on fresh legit traffic. (The 0.6B
soft-hedged one legit-adjacent scam to `suspicious` β the RO energy-subsidy scam β
reported, not counted as a false alarm.)
- **Format robustness held:** JSON validity 1.000 and evidence pass 1.000 on fresh
text (no repair/fallback needed).
**RO vs EN, OOD (plain per-language accuracy over the 20-item OOD set β 9 RO / 11 EN;
a per-language F1 is not computed in the reports):**
| Model | RO correct | EN correct |
| --- | --- | --- |
| `flowxai/scam-guard-qwen06b` | 7 / 9 | 10 / 11 |
RO is the first-class target language: the RO miss is the single family-emergency
scam (the 0.6B also soft-hedges the RO energy-subsidy scam to `suspicious`). On the
in-distribution bench the RO-dominant tactics score at ceiling
(`family_emergency_impersonation` F1 **1.000**, `courier_customs_fee` **1.000**,
`refund_overpayment` **0.889**), which is exactly why the OOD RO family-emergency
miss is the honest gap to close.
### ScamGuardBench v0.2 (synthetic, in-distribution) β with the home-field caveat
> **Honesty caveat β home-field advantage (read before citing these numbers).** The
> fine-tuned numbers **beat the frontier reference on this benchmark, and that does NOT
> mean the model is a better real-world scam detector.** ScamGuardBench v0.2 is built from the
> **same synthetic generator** the model trained on (held-out split,
> contamination-verified β no leakage of specific messages, but the **same
> distribution**: same output format, phrasing, tactic-to-message style). The
> fine-tune learned exactly that style. **The frontier reference is the honest upper
> reference for a cold-prompted generalist** β a better OOD proxy than the fine-tuned
> column β and the OOD results above are the real-world signal these bench numbers
> flatter.
Seeded **120-message** slice (18 `suspicious`, 39 legit-confusable). A false positive
is a `scam_likely` verdict on a legitimate message; `suspicious` is reported but
**not** counted as an FP. FP-rate on the legit-confusable subset is the **release gate**.
| Metric | **qwen06b (CUDA)** | claude-haiku-4-5 (frontier ref) | keyword (lower ref) |
| --- | --- | --- | --- |
| JSON validity (deployed decoder) | 1.000 | 1.000 | 1.000 |
| **raw** JSON validity (no repair) | 1.000 | n/a | n/a |
| verdict micro-F1 | 0.958 | 0.900 | 0.483 |
| verdict macro-F1 | **0.926** | 0.829 | 0.482 |
| tactic macro-F1 | 0.951 | 0.770 | 0.319 |
| evidence pass rate | 0.980 | 0.969 | 1.000 |
| **legit-confusable FP-rate** | **0.000** | 0.051 | 0.026 |
The reference `claude-haiku-4-5` **fails the FP gate** (legit-confusable FP 0.051,
2/39 β it over-flags genuine family money requests), while the fine-tune holds
**0.000**. That is the home-field-advantage signal, not a claim of superior
real-world judgment. `family_money_request` (a genuine "send me money" from family)
is the hardest legit class β every frontier model over-flags it β while this
fine-tune holds **0.000**.
> **Full-budget note (MLX β CUDA).** The shipped 0.6B is the CUDA 3-epoch run. It
> supersedes the earlier MLX ~1-epoch first pass (macro-F1 0.903 β **0.926**,
> micro 0.942 β **0.958**), while **holding legit-FP at 0.000**. The full budget
> bought +2.3 macro-F1 points in-distribution.
### Calibration β we don't give you a probability
scam-guard **deliberately emits no confidence percentage**. An uncalibrated number
on a consumer safety tool is worse than none: it invites false precision on a
judgement that is genuinely uncertain. Instead of a probability to calibrate, we give
you **honest per-class behaviour** so you know where the model is weak and where it is
safe.
**Confusion matrix** (`qwen06b` CUDA, ScamGuardBench v0.2, 120-item slice; rows = gold,
cols = predicted):
| gold \ pred | scam_likely | suspicious | no_indicators | recall |
| --- | --- | --- | --- | --- |
| scam_likely | 63 | 0 | 0 | 1.000 (n=63) |
| suspicious | 0 | 13 | 5 | 0.722 (n=18) |
| no_indicators | 0 | 0 | 39 | 1.000 (n=39) |
**Per-class precision / recall / F1** (`reports/eval_frontier.md`) β micro-F1 0.958,
macro-F1 0.926:
| Verdict | Precision | Recall | F1 | Support |
| --- | --- | --- | --- | --- |
| scam_likely | 1.000 | 1.000 | 1.000 | 63 |
| suspicious | 1.000 | 0.722 | 0.839 | 18 |
| no_indicators | 0.886 | 1.000 | 0.940 | 39 |
**Where the model is weak, and where it is safe.** The weak class is `suspicious`
recall (0.722, 13/18) β the honest middle is the hardest to catch β but `suspicious`
**precision is 1.000** (it never over-calls the middle). **Crucially, every
`suspicious` miss bleeds into `no_indicators` (the safe direction), never into
`scam_likely`, and no legit item is ever flipped to `scam_likely`** β which is why
the legit-confusable FP-rate is **0.000**. The model under-warns on ambiguous
messages rather than over-warning on real ones; for a consumer guard that is the
failure mode you want.
### Size decision
- **This model (`flowxai/scam-guard-qwen06b`, 0.6B) β smallest and fastest, weaker
OOD.** Passes all three release gates on the bench (JSON >99%, evidence >95%,
legit-FP <3%) and is a genuinely defensible on-device ship at ~0.4 GB (int4), but
drops hardest out-of-distribution (macro-F1 β0.31, 18/20).
- **The sibling `flowxai/scam-guard-qwen17b` (1.7B) β the more robust choice.** On
fresh messages the extra capacity generalizes materially better (19/20, no macro-F1
drop), a gap the in-distribution bench (both ~0.9+) did not surface. Where
robustness matters more than size/latency, [ship the
1.7B](https://huggingface.co/flowxai/scam-guard-qwen17b).
- **Both sizes share the RO family-emergency recall gap** (the one dangerous OOD
miss) and both hold legit-FP at 0.000. That shared recall gap is the honest thing
to fix before any release claim.
---
## Formats & on-device latency (0.6B)
**We targeted 150 ms. We measured ~1.1 s (best path). This target is currently not
met.** scam-guard emits a full multi-field JSON card (~138 tokens), not a single
label, so decode dominates latency.
The two headline configs for the 0.6B on the representative 300-char SMS (median,
M3 Max):
| Config | median | meets 150 ms? |
| --- | --- | --- |
| MLX int4, Apple-Silicon GPU (Metal) | ~1.1 s | no |
| GGUF int8 (`Q8_0`), CPU-only (`-ngl 0`) | ~3.3 s | no |
int4 GGUF (`Q4_K_M`) is a poor trade on this small model β on CPU it is *slower*
than int8 and degrades quality (a spot-check sample became invalid JSON) β so
**`Q8_0` is the recommended GGUF quant**, and MLX int4 is the fastest quality-holding
path.
<details>
<summary>Full per-quant numbers (0.6B)</summary>
GGUF (llama.cpp), CPU-only (`-ngl 0`), M3 Max β full JSON card, median over N=12
(`reports/benchmark_gguf.json`):
| quant | file size | median | JSON spot-check |
| --- | --- | --- | --- |
| int8 (Q8_0) | 639 MB | 3204 ms | 4/4 valid |
| int4 (Q4_K_M) | 397 MB | 4000 ms | 3/4 (degraded) |
MLX-quantized, Apple-Silicon GPU (Metal), M3 Max (`reports/benchmark_mlx_quant.json`):
| quant | weights size | median | JSON spot-check |
| --- | --- | --- | --- |
| int4 | 335 MB | 1153 ms | 4/4 valid |
| int8 | 633 MB | 1281 ms | 4/4 valid |
bf16 MLX-Metal latency/memory and per-input-length detail are in
`reports/benchmark.md`. **Core ML** (.mlpackage) was attempted; the LLMβCore ML
conversion is finicky (stateful KV-cache handling) and the attempt is documented in
`PROGRESS.md` rather than shipped as a fabricated artifact. GGUF and MLX are the
recommended on-device paths today.
</details>
A human decision is needed at the release STOP: accept the ~1β3 s latency, ship the
~1.1 s GPU/MLX path, or shrink the output contract to approach 150 ms.
---
## Intended use & limitations
**Intended use.** A **consumer triage aid** that explains *why* a message looks risky
and points you to *your own* trusted channel to verify. Runs on-device. Languages: EN
and RO at v1 (Romanian is a first-class citizen, not an afterthought); PL/HU planned
fast-follow through the same pipeline.
**Out of scope & limitations.**
- **Not a guarantee.** A verdict is a signal, not proof. Scammers adapt continuously;
the benchmark is versioned because patterns rotate.
- **Verdicts can be wrong in both directions** β a real scam may score
`no_indicators` (the OOD RO family-emergency miss is a documented example), and a
legitimate message may score `suspicious`. The `suspicious` middle level exists to
be honest about uncertainty rather than force a binary.
- **Known recall gap:** the RO family-emergency pattern (framing without an explicit
money-transfer keyword) can slip past this model (and the 1.7B); it is a confirmed
training-data gap, fixable with a data addition before any release claim.
- **The model never fetches URLs.** It cannot tell you where a link *actually*
resolves, only what the visible string suggests. A lexically-clean URL can still be
malicious.
- Not a replacement for a bank's fraud line, a national anti-fraud service, or human
judgment. The recommended action always routes to *your own* channel.
- **Text only** at v1: no image/OCR, no audio, no attachment parsing, no
email-header/routing analysis.
---
## Training data
- **Public seed layer (relabeled):** UCI SMS Spam Collection (CC BY 4.0), enron_ham
(SetFit/enron_spam ham slice; no explicit license β reference/research-use),
phishing_email (zefang-liu; LGPL-3.0). Relabeled into the verdict+tactic scheme with
the evidence kill-switch; per-source human spot-check kept relabel disagreement
under the 10% gate.
- **Synthetic layer:** generated from specs (tactic Γ channel Γ language Γ register),
including a `suspicious` middle-ground tier and adversarial keyword-evasion
paraphrases (the `hard` subset). Every bench-destined message passes a sanitizer
audit (reserved domains, non-dialable phones).
- **Balance:** β₯45% legitimate messages, RO β₯35%, ~40% of RO diacritic-free,
`suspicious` ~14.5% of the SFT train set.
- Split discipline: synthetic by spec-family, public by source-text hash; paraphrases
follow their parent's split. **Train β© ScamGuardBench = β
** (contamination-verified).
Both sizes trained from the **same data**; the size decision is evidence-based (above).
## Fine-tuning
**CUDA full-budget 3-epoch LoRA** (`training/train_cuda.py`, transformers + PEFT +
trl `SFTTrainer`), from `Qwen/Qwen3-0.6B`. LoRA rank 16 / alpha 32 / dropout 0.05,
all-linear modules, adamw, cosine 2e-5β2e-6 with 60-step warmup, effective batch 4,
seq 1280, bf16, gradient-checkpointing, seed 20260703; **3 real epochs** on an
on-demand NVIDIA L4 (~2h16m, 0 OOM), thinking disabled. This **supersedes** the MLX
first pass (~1 epoch, bounded by Metal stalls); the full budget bought +2.3 macro-F1
points in-distribution.
> **Base-model note.** The exact repo `Qwen3-0.6B-Instruct` does **not** exist on
> Hugging Face β Qwen3 merged instruct+thinking into the single base repo
> `Qwen/Qwen3-0.6B` (instruction-capable, Apache-2.0). We fine-tune it with thinking
> disabled.
---
## Dual-use statement
We release a **detector**, a **benchmark**, and **pattern-level** explanations. We do
**not** release the scam-variant generation prompts as a standalone tool. Public
benchmark scam texts carry no real dialable phone numbers and no working URLs
(reserved/example domains and clearly-fake numbers only). Output explanations describe
the manipulation *pattern*, never instructions for constructing one.
## Links
- **Sibling model (quality pick, 1.7B):** [`flowxai/scam-guard-qwen17b`](https://huggingface.co/flowxai/scam-guard-qwen17b) β higher OOD accuracy, larger footprint (~1.1 GB int4).
- **Benchmark / dataset:** [`flowxai/scamguardbench`](https://huggingface.co/datasets/flowxai/scamguardbench).
## License
Apache-2.0 (weights, code, and benchmark). Base model `Qwen/Qwen3-0.6B` is Apache-2.0.
</content>
</invoke>
|