nora-4b-v3.2-GGUF / README.md
dmvevents's picture
Upload README.md with huggingface_hub
1873b74
|
Raw History Blame Contribute Delete
1.27 kB
metadata
license: apache-2.0
tags:
  - gguf
  - nora
  - trinidad-tobago
  - qwen3.5
  - quantized
base_model: dmvevents/nora-4b-v3.2

Nora v3.2 GGUF Quantizations

Laptop-deployable GGUF versions of dmvevents/nora-4b-v3.2.

File Quant Size Use
nora-v3.2-Q4_K_M.gguf Q4_K_M 2.9 GB Tightest RAM budgets
nora-v3.2-Q6_K.gguf Q6_K 3.8 GB Recommended on 16 GB laptops — ~half the quantization loss of Q4_K_M (quant damage concentrates in Creole + math), 4.0 GB peak RAM / 10 tok/s measured at 6 CPU threads
nora-v3.2-Q8_0.gguf Q8_0 4.9 GB Near-lossless if RAM allows
nora-v3.2-f16.gguf F16 9.1 GB Reference / further quantization

Run with repeat_penalty=1.0 (the production default): higher values pressure paraphrase of verbatim facts (phone numbers, fees).

Eval (underlying bf16 model)

  • 1,420-paraphrase eval: 89.2% Claude Sonnet 4.5 judge (v3.1 was 87.1%, v3 was 86.4%, v2 was 84.9%)
  • 143-base eval: 89.9% Claude
  • Targeted gains vs v3.1: safety +5.03pp, gov +3.62pp, creole +2.60pp
  • See dmvevents/tt-eval-v3.2-results for full data.