gyopak commited on
Commit
d8fdd02
·
verified ·
1 Parent(s): 7e84f19

Upload folder using huggingface_hub

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +94 -0
  3. emese-patak-Q4_K_M.gguf +3 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ emese-patak-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,94 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - hu
4
+ license: apache-2.0
5
+ library_name: gguf
6
+ base_model: utter-project/EuroLLM-9B
7
+ pipeline_tag: text-generation
8
+ tags:
9
+ - hungarian
10
+ - emese
11
+ - eurollm
12
+ - instruct
13
+ - chatml
14
+ - gguf
15
+ - llama.cpp
16
+ - quantized
17
+ ---
18
+
19
+ # Emese-Patak (9.15B) — GGUF Q4_K_M
20
+
21
+ **GGUF Q4_K_M** — a compact `llama.cpp`-compatible build of Patak, quantized from the model's
22
+ native q8 MLX artifact (`patak-mlx/`). See the `patak/` repo's README for full architecture,
23
+ CPT/SFT/DPO training details, and benchmarks — this file covers only the GGUF-specific notes.
24
+
25
+ | | |
26
+ |---|---|
27
+ | **Quantization** | Q4_K_M (llama.cpp k-quant) |
28
+ | **Size on disk** | ~5.2 GB |
29
+ | **Max context length** | 32,768 tokens (EuroLLM-9B's native context) |
30
+ | **Runtime** | `llama.cpp` / `llama-server` / `llama-cli` / any GGUF-compatible loader (LM Studio, Ollama, etc.) |
31
+
32
+ ## ⚠️ Tokenizer fix required — read this before using any other GGUF build of this model
33
+
34
+ A stock `convert_hf_to_gguf.py` export of this model family is **badly broken**: it types the
35
+ ChatML control tokens (`<|im_start|>`, `<|im_end|>`) as `NORMAL` instead of `CONTROL`, so
36
+ `<|im_start|>` gets shredded into 7 garbage sub-word tokens instead of being fed to the model as
37
+ the single trained token — a prompt shape the model never saw during training. It also writes a
38
+ flat placeholder BPE merge score for every token, corrupting subword-split priority. Together these
39
+ caused a severe, previously-misdiagnosed quality regression (early testing wrongly concluded it was
40
+ inherent to llama.cpp itself).
41
+
42
+ **This GGUF file has already been fixed** — `scripts/fix_gguf_tokenizer.py` (in the main repo) was
43
+ run on it after conversion/quantization to correct the special-token typing, BPE scores, and a
44
+ stray leading-space flag. Verified: 80/80 sampled bench prompts tokenize byte-identical to the
45
+ HF/MLX reference tokenization. **If you ever regenerate this GGUF from source yourself, you must
46
+ re-run that fix script** (or the equivalent metadata patch) — a plain `convert_hf_to_gguf.py` +
47
+ `llama-quantize` pipeline without it reproduces the old broken behavior.
48
+
49
+ ## Usage
50
+
51
+ ```bash
52
+ llama-server -m emese-patak-Q4_K_M.gguf -c 4096
53
+ ```
54
+
55
+ ```python
56
+ import requests
57
+ r = requests.post("http://127.0.0.1:8080/v1/chat/completions", json={
58
+ "messages": [{"role": "user", "content": "Mi Magyarország fővárosa?"}],
59
+ "temperature": 0.2, "repeat_penalty": 1.15, "stop": ["<|im_end|>"],
60
+ })
61
+ print(r.json()["choices"][0]["message"]["content"])
62
+ ```
63
+
64
+ **Decode:** temperature `0.2`, repeat_penalty `1.15`, stop on `<|im_end|>`, ChatML template
65
+ (`<|im_start|>role\n...<|im_end|>\n`).
66
+
67
+ ## Training
68
+
69
+ Same underlying weights as `patak-mlx/` (q8, the model's native training precision), just
70
+ re-quantized to GGUF Q4_K_M — no separate training. See `patak/README.md` for the full CPT
71
+ (5.1M tokens/5,000 iters), SFT (`instruct_v18b`, 1 epoch, rank16/scale32/lr1.5e-5), and DPO (36
72
+ alfa pairs, 120 iters, rank16/scale32/lr5e-6) recipe.
73
+
74
+ ## Benchmarks
75
+
76
+ **This exact Q4_K_M GGUF build (with the tokenizer fix applied) scored 384/500 (77%) on
77
+ emese-bench v1**, vs. 391/500 (78%) for the same fix's Q8_0 build and 413/500 (83%) for the
78
+ original MLX q8 artifact. Zero `<|im_start|>`/`<|im_end|>` leaks — the category-level strengths
79
+ and weaknesses closely track the MLX result (near-perfect reading/translation/safety/code; weak
80
+ on multi-step math and multi-constraint formatting). A residual, much smaller artifact distinct
81
+ from the tokenizer bug was found in a couple of spots: the model's own completion occasionally
82
+ leaks a fabricated literal `user` continuation into its response (a stop-condition timing issue,
83
+ not a tokenizer problem). See `emese-bench/results/patak-gguf-q4fix.md` for the full
84
+ category-by-category transcript and `emese-bench/README.md` for the benchmark's design and the
85
+ Q8_0 comparison point.
86
+
87
+ ## Limitations
88
+
89
+ - Can hallucinate specific facts (dates, attributions, biographical details) — verify critical
90
+ details. Two specific bench questions (about fictional/obscure Hungarian scientists) reliably
91
+ produce confidently-fabricated biographies across every tested variant of this model family.
92
+ - Hungarian-first; other-language quality inherited from EuroLLM-9B.
93
+ - Weak at multi-step math, spatial estimation, and strict multi-constraint formatting
94
+ (alphabetical ordering, exact word counts, banned letters) — consistent with the MLX original.
emese-patak-Q4_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:610a147b500b36ac04d744dbe8479ca6ccf3f6777ca92ef3e5ee49c7ca622791
3
+ size 5582837344