mmis1000 commited on
Commit
f3ad3c6
·
verified ·
1 Parent(s): 0815e22

Preserve pinned upstream MTP head in native-8k v0.2 GGUFs (unchanged adapters)

Browse files
README.md CHANGED
@@ -21,6 +21,8 @@ GGUF quantizations of a fine-tuned model for translating Japanese ASMR transcrip
21
 
22
  The model normalizes imperfect audio transcriptions, applies domain-specific glossaries, and translates character dialogue while retaining emotion and nuances.
23
 
 
 
24
  ## Echo Mode
25
 
26
  The model echoes the source Japanese text in an `"input"` field and records applied terms in a per-entry `"glossary"` object alongside the target translation. This provides an explicit source anchor that can reduce omitted or drifted segments, but it does not guarantee immunity to long-context repetition or noisy-ASR failures.
@@ -29,10 +31,10 @@ The model echoes the source Japanese text in an `"input"` field and records appl
29
 
30
  | Quantization | Filename | Size | Description |
31
  |---|---|---|---|
32
- | q4_k_m | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf` | 5.2 GB | Good balance of quality and size |
33
- | q6_k | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q6_k.gguf` | 6.9 GB | Higher quality, moderate size |
34
- | q8_0 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q8_0.gguf` | 8.9 GB | Near-lossless quality |
35
- | bf16 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-bf16.gguf` | 16.7 GB | Full BF16, no quantization loss |
36
 
37
  ## Prompt Example
38
 
@@ -98,6 +100,12 @@ llama-server -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 --port
98
  llama-cli -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 -p "<your prompt>" -n 2048
99
  ```
100
 
 
 
 
 
 
 
101
  ## Structured Decoding (Recommended)
102
 
103
  This model outputs JSON arrays. Using structured decoding (e.g. GBNF grammar or JSON schema constraints) avoids wasted computation on malformed output and guarantees valid JSON on every generation.
@@ -203,3 +211,8 @@ Treat long-context use of this variant as **experimental**. Prefer shorter windo
203
  ### Content Notice
204
 
205
  The training domain includes adult ASMR dialogue and may produce sexually explicit text. This model is intended for transcription translation and subtitle-processing workflows.
 
 
 
 
 
 
21
 
22
  The model normalizes imperfect audio transcriptions, applies domain-specific glossaries, and translates character dialogue while retaining emotion and nuances.
23
 
24
+ This variant preserves the upstream Qwen3.5 MTP / speculative-decoding head in GGUF format so it can be used with MTP-capable llama.cpp builds.
25
+
26
  ## Echo Mode
27
 
28
  The model echoes the source Japanese text in an `"input"` field and records applied terms in a per-entry `"glossary"` object alongside the target translation. This provides an explicit source anchor that can reduce omitted or drifted segments, but it does not guarantee immunity to long-context repetition or noisy-ASR failures.
 
31
 
32
  | Quantization | Filename | Size | Description |
33
  |---|---|---|---|
34
+ | q4_k_m | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf` | 5.4 GB | Good balance of quality and size |
35
+ | q6_k | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q6_k.gguf` | 7.0 GB | Higher quality, moderate size |
36
+ | q8_0 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q8_0.gguf` | 9.1 GB | Near-lossless quality |
37
+ | bf16 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-bf16.gguf` | 17.1 GB | Full BF16, no quantization loss |
38
 
39
  ## Prompt Example
40
 
 
100
  llama-cli -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 -p "<your prompt>" -n 2048
101
  ```
102
 
103
+ ### llama-cli with MTP
104
+
105
+ ```bash
106
+ llama-cli -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 --spec-type draft-mtp --spec-draft-n-max 6 -p "<your prompt>" -n 2048
107
+ ```
108
+
109
  ## Structured Decoding (Recommended)
110
 
111
  This model outputs JSON arrays. Using structured decoding (e.g. GBNF grammar or JSON schema constraints) avoids wasted computation on malformed output and guarantees valid JSON on every generation.
 
211
  ### Content Notice
212
 
213
  The training domain includes adult ASMR dialogue and may produce sexually explicit text. This model is intended for transcription translation and subtitle-processing workflows.
214
+
215
+
216
+ ## MTP-preserving export
217
+
218
+ This revision preserves and verifies all 15 upstream MTP tensors against the pinned base checkpoint after the LoRA merge. The native-8192 adapter is unchanged; no retraining was performed. The MTP head itself is not fine-tuned. Earlier revisions of this v0.2 repository omitted MTP. Filenames and translation weights remain compatible; use a fresh repository revision to avoid cached non-MTP files. All four quantizations have verified MTP metadata and tensor inventory. Q4_K_M was GPU-smoke-tested with llama.cpp b9247, context 8192, and `--spec-type draft-mtp --spec-draft-n-max 3`. This is a runtime smoke test, not a new quality evaluation or a speedup guarantee. See `mtp-preservation.json` and `mtp-release-verification.json` for provenance and checks.
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-bf16.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:d195f732d70e053e84ed48fbbafbe00bf006e258a324a15cdde4186241e5dda4
3
- size 17920696864
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b2fd49a58f5f829276ba4da8715875113064b45dc971bb403db3bbe4284e54d9
3
+ size 18407321056
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:a867434b5bde577e170582624a0ed0d178f56c16c07712e30b780fa9bad25105
3
- size 5629108768
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e185134cb504cbb93e42c9da03461331b063155bcec3530daa87bc0e610fee77
3
+ size 5780090336
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q6_k.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:c1a98bdac4fab31bf109228c75376ccff3cfb9201a10e745845d713dce4801ff
3
- size 7359259168
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4bcab83434cf001ac7da9a98e2ba4d4809e735aa2f02a71df4efb3408244ea3f
3
+ size 7558901216
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q8_0.gguf CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:62a1e33b8c78360259d6d9f3f6be4049861551c4e98b328dcc9586d90a263050
3
- size 9527501344
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:49a46cb38baee3f3b87814198a6a6da41b120c208de560ce02b6e9ae6c43ec5b
3
+ size 9786060256
mtp-preservation.json ADDED
@@ -0,0 +1,23 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "base_revision": "005429cee5cb648998cf2b70eebdd83175989c9a",
3
+ "base_model": "unsloth/Qwen3.5-9B",
4
+ "restored": [],
5
+ "tensor_sha256": {
6
+ "mtp.layers.0.input_layernorm.weight": "c8913bfe7cef186fb59b7f1ca80d391ba5eaec45c9c8db54f54e63f6410ae633",
7
+ "mtp.layers.0.post_attention_layernorm.weight": "63500aa0ed93918f6f3d5f1bfd1656921543931392a06267f81c0b3ad680fdcf",
8
+ "mtp.layers.0.self_attn.k_norm.weight": "8f5a8071649fa9a7b6b6afcea6d5404cd849640be2c4f90e6759c4c0a60120ac",
9
+ "mtp.layers.0.self_attn.k_proj.weight": "1b0129a3c6f4add6c8417d67e84050f8595f96f7693511beaace9ef56d1ad30a",
10
+ "mtp.layers.0.self_attn.o_proj.weight": "31140d974d94190127d0f3d908a18bafe5fba0d8c3d7f522e012013795c89924",
11
+ "mtp.layers.0.self_attn.q_norm.weight": "10e6c9fa42ceb72373c207422c305663a89dfd08f5332cbbe6d1057301fe57d1",
12
+ "mtp.layers.0.self_attn.v_proj.weight": "c21332947c44203ca197d014de23559bf6e88d7c0debca8d4ecad735d717e2c9",
13
+ "mtp.norm.weight": "7e4daf06ad25b834b3b95c9592fa690b75a4a90fb9f4128ccee8537d7f15988e",
14
+ "mtp.pre_fc_norm_embedding.weight": "8b9bca62497def4783b20dfd58ddf32573a255b44a79f0acf8ca9c611b7bea15",
15
+ "mtp.pre_fc_norm_hidden.weight": "a7410736f9962dd6ebd8ac2b4c355898d057296a03c47795223b7298cac8988c",
16
+ "mtp.fc.weight": "a4639d8f4b81cbdc65c61f1cba82816ae1534ec494a95a0b677dba98f04f4017",
17
+ "mtp.layers.0.self_attn.q_proj.weight": "17d3aac05cb017e9ef98f93fe567f1c1ffae08fbe5001b80e67829bda189a584",
18
+ "mtp.layers.0.mlp.down_proj.weight": "52c72564f7da59c25233b2194a79239cc1e69dd6694130aae87d5fffc472707c",
19
+ "mtp.layers.0.mlp.gate_proj.weight": "dd3e6d05e9c519ebb16c5eee64b4d4d217d7efb8ae7c88e8dfa96f6e6f7d3eac",
20
+ "mtp.layers.0.mlp.up_proj.weight": "79e02e92d4cb775120f48cd523577345b277dacb749e56e2d52532583f92800e"
21
+ },
22
+ "mtp_layers": 1
23
+ }
mtp-release-verification.json ADDED
@@ -0,0 +1,148 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "adapter_sha256": "8000e9e54726e1943fd0c63c7f1249a351f462c2902274539e9245f483443740",
3
+ "files": {
4
+ "q4_k_m": {
5
+ "sha256": "e185134cb504cbb93e42c9da03461331b063155bcec3530daa87bc0e610fee77",
6
+ "size": 5780090336,
7
+ "metadata": {
8
+ "qwen35.block_count": 33,
9
+ "qwen35.nextn_predict_layers": 1
10
+ },
11
+ "mtp_tensors": [
12
+ "blk.32.attn_k.weight",
13
+ "blk.32.attn_k_norm.weight",
14
+ "blk.32.attn_norm.weight",
15
+ "blk.32.attn_output.weight",
16
+ "blk.32.attn_q.weight",
17
+ "blk.32.attn_q_norm.weight",
18
+ "blk.32.attn_v.weight",
19
+ "blk.32.ffn_down.weight",
20
+ "blk.32.ffn_gate.weight",
21
+ "blk.32.ffn_up.weight",
22
+ "blk.32.nextn.eh_proj.weight",
23
+ "blk.32.nextn.enorm.weight",
24
+ "blk.32.nextn.hnorm.weight",
25
+ "blk.32.nextn.shared_head_norm.weight",
26
+ "blk.32.post_attention_norm.weight"
27
+ ],
28
+ "tensor_count": 442
29
+ },
30
+ "q6_k": {
31
+ "sha256": "4bcab83434cf001ac7da9a98e2ba4d4809e735aa2f02a71df4efb3408244ea3f",
32
+ "size": 7558901216,
33
+ "metadata": {
34
+ "qwen35.block_count": 33,
35
+ "qwen35.nextn_predict_layers": 1
36
+ },
37
+ "mtp_tensors": [
38
+ "blk.32.attn_k.weight",
39
+ "blk.32.attn_k_norm.weight",
40
+ "blk.32.attn_norm.weight",
41
+ "blk.32.attn_output.weight",
42
+ "blk.32.attn_q.weight",
43
+ "blk.32.attn_q_norm.weight",
44
+ "blk.32.attn_v.weight",
45
+ "blk.32.ffn_down.weight",
46
+ "blk.32.ffn_gate.weight",
47
+ "blk.32.ffn_up.weight",
48
+ "blk.32.nextn.eh_proj.weight",
49
+ "blk.32.nextn.enorm.weight",
50
+ "blk.32.nextn.hnorm.weight",
51
+ "blk.32.nextn.shared_head_norm.weight",
52
+ "blk.32.post_attention_norm.weight"
53
+ ],
54
+ "tensor_count": 442
55
+ },
56
+ "q8_0": {
57
+ "sha256": "49a46cb38baee3f3b87814198a6a6da41b120c208de560ce02b6e9ae6c43ec5b",
58
+ "size": 9786060256,
59
+ "metadata": {
60
+ "qwen35.block_count": 33,
61
+ "qwen35.nextn_predict_layers": 1
62
+ },
63
+ "mtp_tensors": [
64
+ "blk.32.attn_k.weight",
65
+ "blk.32.attn_k_norm.weight",
66
+ "blk.32.attn_norm.weight",
67
+ "blk.32.attn_output.weight",
68
+ "blk.32.attn_q.weight",
69
+ "blk.32.attn_q_norm.weight",
70
+ "blk.32.attn_v.weight",
71
+ "blk.32.ffn_down.weight",
72
+ "blk.32.ffn_gate.weight",
73
+ "blk.32.ffn_up.weight",
74
+ "blk.32.nextn.eh_proj.weight",
75
+ "blk.32.nextn.enorm.weight",
76
+ "blk.32.nextn.hnorm.weight",
77
+ "blk.32.nextn.shared_head_norm.weight",
78
+ "blk.32.post_attention_norm.weight"
79
+ ],
80
+ "tensor_count": 442
81
+ },
82
+ "bf16": {
83
+ "sha256": "b2fd49a58f5f829276ba4da8715875113064b45dc971bb403db3bbe4284e54d9",
84
+ "size": 18407321056,
85
+ "metadata": {
86
+ "qwen35.block_count": 33,
87
+ "qwen35.nextn_predict_layers": 1
88
+ },
89
+ "mtp_tensors": [
90
+ "blk.32.ffn_down.weight",
91
+ "blk.32.ffn_gate.weight",
92
+ "blk.32.ffn_up.weight",
93
+ "blk.32.nextn.eh_proj.weight",
94
+ "blk.32.attn_q.weight",
95
+ "blk.32.attn_norm.weight",
96
+ "blk.32.post_attention_norm.weight",
97
+ "blk.32.attn_k_norm.weight",
98
+ "blk.32.attn_k.weight",
99
+ "blk.32.attn_output.weight",
100
+ "blk.32.attn_q_norm.weight",
101
+ "blk.32.attn_v.weight",
102
+ "blk.32.nextn.shared_head_norm.weight",
103
+ "blk.32.nextn.enorm.weight",
104
+ "blk.32.nextn.hnorm.weight"
105
+ ],
106
+ "tensor_count": 442
107
+ }
108
+ },
109
+ "runtime": {
110
+ "binary": "/root/llama-b9247-rocm/llama-server",
111
+ "command": [
112
+ "/root/llama-b9247-rocm/llama-server",
113
+ "-m",
114
+ "/root/asmr-one-dump/train/publish-mtp-v0.2/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2/q4_k_m/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf",
115
+ "-ngl",
116
+ "99",
117
+ "-c",
118
+ "8192",
119
+ "-np",
120
+ "1",
121
+ "--port",
122
+ "18942",
123
+ "--host",
124
+ "127.0.0.1",
125
+ "--spec-type",
126
+ "draft-mtp",
127
+ "--spec-draft-n-max",
128
+ "3",
129
+ "--no-warmup",
130
+ "--verbose"
131
+ ],
132
+ "log_sha256": "4824e0025b2f74c1cf1af460e9f6e9eb2ada0bed0d7f82b2cf4184c98365638d",
133
+ "response_sha256": "cff93c67da9ee3e444356c07866c107106e601adbd4de1d0b61e37cd27a77e07",
134
+ "timings": {
135
+ "cache_n": 0,
136
+ "prompt_n": 24,
137
+ "prompt_ms": 378.189,
138
+ "prompt_per_token_ms": 15.757875,
139
+ "prompt_per_second": 63.460333325400796,
140
+ "predicted_n": 8,
141
+ "predicted_ms": 810.715,
142
+ "predicted_per_token_ms": 101.339375,
143
+ "predicted_per_second": 9.86783271556589,
144
+ "draft_n": 12,
145
+ "draft_n_accepted": 5
146
+ }
147
+ }
148
+ }