Instructions to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M # Run inference directly in the terminal: llama cli -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M # Run inference directly in the terminal: llama cli -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Use Docker
docker model run hf.co/mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
- Ollama
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with Ollama:
ollama run hf.co/mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
- Unsloth Desktop
- Pi
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with Docker Model Runner:
docker model run hf.co/mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
- Lemonade
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Run and chat with the model
lemonade run user.asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mmis1000/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Preserve pinned upstream MTP head in native-8k v0.2 GGUFs (unchanged adapters)
Browse files
README.md
CHANGED
|
@@ -21,6 +21,8 @@ GGUF quantizations of a fine-tuned model for translating Japanese ASMR transcrip
|
|
| 21 |
|
| 22 |
The model normalizes imperfect audio transcriptions, applies domain-specific glossaries, and translates character dialogue while retaining emotion and nuances.
|
| 23 |
|
|
|
|
|
|
|
| 24 |
## Echo Mode
|
| 25 |
|
| 26 |
The model echoes the source Japanese text in an `"input"` field and records applied terms in a per-entry `"glossary"` object alongside the target translation. This provides an explicit source anchor that can reduce omitted or drifted segments, but it does not guarantee immunity to long-context repetition or noisy-ASR failures.
|
|
@@ -29,10 +31,10 @@ The model echoes the source Japanese text in an `"input"` field and records appl
|
|
| 29 |
|
| 30 |
| Quantization | Filename | Size | Description |
|
| 31 |
|---|---|---|---|
|
| 32 |
-
| q4_k_m | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf` | 5.
|
| 33 |
-
| q6_k | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q6_k.gguf` |
|
| 34 |
-
| q8_0 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q8_0.gguf` |
|
| 35 |
-
| bf16 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-bf16.gguf` |
|
| 36 |
|
| 37 |
## Prompt Example
|
| 38 |
|
|
@@ -98,6 +100,12 @@ llama-server -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 --port
|
|
| 98 |
llama-cli -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 -p "<your prompt>" -n 2048
|
| 99 |
```
|
| 100 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 101 |
## Structured Decoding (Recommended)
|
| 102 |
|
| 103 |
This model outputs JSON arrays. Using structured decoding (e.g. GBNF grammar or JSON schema constraints) avoids wasted computation on malformed output and guarantees valid JSON on every generation.
|
|
@@ -203,3 +211,8 @@ Treat long-context use of this variant as **experimental**. Prefer shorter windo
|
|
| 203 |
### Content Notice
|
| 204 |
|
| 205 |
The training domain includes adult ASMR dialogue and may produce sexually explicit text. This model is intended for transcription translation and subtitle-processing workflows.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 21 |
|
| 22 |
The model normalizes imperfect audio transcriptions, applies domain-specific glossaries, and translates character dialogue while retaining emotion and nuances.
|
| 23 |
|
| 24 |
+
This variant preserves the upstream Qwen3.5 MTP / speculative-decoding head in GGUF format so it can be used with MTP-capable llama.cpp builds.
|
| 25 |
+
|
| 26 |
## Echo Mode
|
| 27 |
|
| 28 |
The model echoes the source Japanese text in an `"input"` field and records applied terms in a per-entry `"glossary"` object alongside the target translation. This provides an explicit source anchor that can reduce omitted or drifted segments, but it does not guarantee immunity to long-context repetition or noisy-ASR failures.
|
|
|
|
| 31 |
|
| 32 |
| Quantization | Filename | Size | Description |
|
| 33 |
|---|---|---|---|
|
| 34 |
+
| q4_k_m | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf` | 5.4 GB | Good balance of quality and size |
|
| 35 |
+
| q6_k | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q6_k.gguf` | 7.0 GB | Higher quality, moderate size |
|
| 36 |
+
| q8_0 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q8_0.gguf` | 9.1 GB | Near-lossless quality |
|
| 37 |
+
| bf16 | `asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-bf16.gguf` | 17.1 GB | Full BF16, no quantization loss |
|
| 38 |
|
| 39 |
## Prompt Example
|
| 40 |
|
|
|
|
| 100 |
llama-cli -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 -p "<your prompt>" -n 2048
|
| 101 |
```
|
| 102 |
|
| 103 |
+
### llama-cli with MTP
|
| 104 |
+
|
| 105 |
+
```bash
|
| 106 |
+
llama-cli -m asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf -c 8192 --spec-type draft-mtp --spec-draft-n-max 6 -p "<your prompt>" -n 2048
|
| 107 |
+
```
|
| 108 |
+
|
| 109 |
## Structured Decoding (Recommended)
|
| 110 |
|
| 111 |
This model outputs JSON arrays. Using structured decoding (e.g. GBNF grammar or JSON schema constraints) avoids wasted computation on malformed output and guarantees valid JSON on every generation.
|
|
|
|
| 211 |
### Content Notice
|
| 212 |
|
| 213 |
The training domain includes adult ASMR dialogue and may produce sexually explicit text. This model is intended for transcription translation and subtitle-processing workflows.
|
| 214 |
+
|
| 215 |
+
|
| 216 |
+
## MTP-preserving export
|
| 217 |
+
|
| 218 |
+
This revision preserves and verifies all 15 upstream MTP tensors against the pinned base checkpoint after the LoRA merge. The native-8192 adapter is unchanged; no retraining was performed. The MTP head itself is not fine-tuned. Earlier revisions of this v0.2 repository omitted MTP. Filenames and translation weights remain compatible; use a fresh repository revision to avoid cached non-MTP files. All four quantizations have verified MTP metadata and tensor inventory. Q4_K_M was GPU-smoke-tested with llama.cpp b9247, context 8192, and `--spec-type draft-mtp --spec-draft-n-max 3`. This is a runtime smoke test, not a new quality evaluation or a speedup guarantee. See `mtp-preservation.json` and `mtp-release-verification.json` for provenance and checks.
|
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-bf16.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:b2fd49a58f5f829276ba4da8715875113064b45dc971bb403db3bbe4284e54d9
|
| 3 |
+
size 18407321056
|
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e185134cb504cbb93e42c9da03461331b063155bcec3530daa87bc0e610fee77
|
| 3 |
+
size 5780090336
|
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q6_k.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4bcab83434cf001ac7da9a98e2ba4d4809e735aa2f02a71df4efb3408244ea3f
|
| 3 |
+
size 7558901216
|
asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q8_0.gguf
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:49a46cb38baee3f3b87814198a6a6da41b120c208de560ce02b6e9ae6c43ec5b
|
| 3 |
+
size 9786060256
|
mtp-preservation.json
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"base_revision": "005429cee5cb648998cf2b70eebdd83175989c9a",
|
| 3 |
+
"base_model": "unsloth/Qwen3.5-9B",
|
| 4 |
+
"restored": [],
|
| 5 |
+
"tensor_sha256": {
|
| 6 |
+
"mtp.layers.0.input_layernorm.weight": "c8913bfe7cef186fb59b7f1ca80d391ba5eaec45c9c8db54f54e63f6410ae633",
|
| 7 |
+
"mtp.layers.0.post_attention_layernorm.weight": "63500aa0ed93918f6f3d5f1bfd1656921543931392a06267f81c0b3ad680fdcf",
|
| 8 |
+
"mtp.layers.0.self_attn.k_norm.weight": "8f5a8071649fa9a7b6b6afcea6d5404cd849640be2c4f90e6759c4c0a60120ac",
|
| 9 |
+
"mtp.layers.0.self_attn.k_proj.weight": "1b0129a3c6f4add6c8417d67e84050f8595f96f7693511beaace9ef56d1ad30a",
|
| 10 |
+
"mtp.layers.0.self_attn.o_proj.weight": "31140d974d94190127d0f3d908a18bafe5fba0d8c3d7f522e012013795c89924",
|
| 11 |
+
"mtp.layers.0.self_attn.q_norm.weight": "10e6c9fa42ceb72373c207422c305663a89dfd08f5332cbbe6d1057301fe57d1",
|
| 12 |
+
"mtp.layers.0.self_attn.v_proj.weight": "c21332947c44203ca197d014de23559bf6e88d7c0debca8d4ecad735d717e2c9",
|
| 13 |
+
"mtp.norm.weight": "7e4daf06ad25b834b3b95c9592fa690b75a4a90fb9f4128ccee8537d7f15988e",
|
| 14 |
+
"mtp.pre_fc_norm_embedding.weight": "8b9bca62497def4783b20dfd58ddf32573a255b44a79f0acf8ca9c611b7bea15",
|
| 15 |
+
"mtp.pre_fc_norm_hidden.weight": "a7410736f9962dd6ebd8ac2b4c355898d057296a03c47795223b7298cac8988c",
|
| 16 |
+
"mtp.fc.weight": "a4639d8f4b81cbdc65c61f1cba82816ae1534ec494a95a0b677dba98f04f4017",
|
| 17 |
+
"mtp.layers.0.self_attn.q_proj.weight": "17d3aac05cb017e9ef98f93fe567f1c1ffae08fbe5001b80e67829bda189a584",
|
| 18 |
+
"mtp.layers.0.mlp.down_proj.weight": "52c72564f7da59c25233b2194a79239cc1e69dd6694130aae87d5fffc472707c",
|
| 19 |
+
"mtp.layers.0.mlp.gate_proj.weight": "dd3e6d05e9c519ebb16c5eee64b4d4d217d7efb8ae7c88e8dfa96f6e6f7d3eac",
|
| 20 |
+
"mtp.layers.0.mlp.up_proj.weight": "79e02e92d4cb775120f48cd523577345b277dacb749e56e2d52532583f92800e"
|
| 21 |
+
},
|
| 22 |
+
"mtp_layers": 1
|
| 23 |
+
}
|
mtp-release-verification.json
ADDED
|
@@ -0,0 +1,148 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"adapter_sha256": "8000e9e54726e1943fd0c63c7f1249a351f462c2902274539e9245f483443740",
|
| 3 |
+
"files": {
|
| 4 |
+
"q4_k_m": {
|
| 5 |
+
"sha256": "e185134cb504cbb93e42c9da03461331b063155bcec3530daa87bc0e610fee77",
|
| 6 |
+
"size": 5780090336,
|
| 7 |
+
"metadata": {
|
| 8 |
+
"qwen35.block_count": 33,
|
| 9 |
+
"qwen35.nextn_predict_layers": 1
|
| 10 |
+
},
|
| 11 |
+
"mtp_tensors": [
|
| 12 |
+
"blk.32.attn_k.weight",
|
| 13 |
+
"blk.32.attn_k_norm.weight",
|
| 14 |
+
"blk.32.attn_norm.weight",
|
| 15 |
+
"blk.32.attn_output.weight",
|
| 16 |
+
"blk.32.attn_q.weight",
|
| 17 |
+
"blk.32.attn_q_norm.weight",
|
| 18 |
+
"blk.32.attn_v.weight",
|
| 19 |
+
"blk.32.ffn_down.weight",
|
| 20 |
+
"blk.32.ffn_gate.weight",
|
| 21 |
+
"blk.32.ffn_up.weight",
|
| 22 |
+
"blk.32.nextn.eh_proj.weight",
|
| 23 |
+
"blk.32.nextn.enorm.weight",
|
| 24 |
+
"blk.32.nextn.hnorm.weight",
|
| 25 |
+
"blk.32.nextn.shared_head_norm.weight",
|
| 26 |
+
"blk.32.post_attention_norm.weight"
|
| 27 |
+
],
|
| 28 |
+
"tensor_count": 442
|
| 29 |
+
},
|
| 30 |
+
"q6_k": {
|
| 31 |
+
"sha256": "4bcab83434cf001ac7da9a98e2ba4d4809e735aa2f02a71df4efb3408244ea3f",
|
| 32 |
+
"size": 7558901216,
|
| 33 |
+
"metadata": {
|
| 34 |
+
"qwen35.block_count": 33,
|
| 35 |
+
"qwen35.nextn_predict_layers": 1
|
| 36 |
+
},
|
| 37 |
+
"mtp_tensors": [
|
| 38 |
+
"blk.32.attn_k.weight",
|
| 39 |
+
"blk.32.attn_k_norm.weight",
|
| 40 |
+
"blk.32.attn_norm.weight",
|
| 41 |
+
"blk.32.attn_output.weight",
|
| 42 |
+
"blk.32.attn_q.weight",
|
| 43 |
+
"blk.32.attn_q_norm.weight",
|
| 44 |
+
"blk.32.attn_v.weight",
|
| 45 |
+
"blk.32.ffn_down.weight",
|
| 46 |
+
"blk.32.ffn_gate.weight",
|
| 47 |
+
"blk.32.ffn_up.weight",
|
| 48 |
+
"blk.32.nextn.eh_proj.weight",
|
| 49 |
+
"blk.32.nextn.enorm.weight",
|
| 50 |
+
"blk.32.nextn.hnorm.weight",
|
| 51 |
+
"blk.32.nextn.shared_head_norm.weight",
|
| 52 |
+
"blk.32.post_attention_norm.weight"
|
| 53 |
+
],
|
| 54 |
+
"tensor_count": 442
|
| 55 |
+
},
|
| 56 |
+
"q8_0": {
|
| 57 |
+
"sha256": "49a46cb38baee3f3b87814198a6a6da41b120c208de560ce02b6e9ae6c43ec5b",
|
| 58 |
+
"size": 9786060256,
|
| 59 |
+
"metadata": {
|
| 60 |
+
"qwen35.block_count": 33,
|
| 61 |
+
"qwen35.nextn_predict_layers": 1
|
| 62 |
+
},
|
| 63 |
+
"mtp_tensors": [
|
| 64 |
+
"blk.32.attn_k.weight",
|
| 65 |
+
"blk.32.attn_k_norm.weight",
|
| 66 |
+
"blk.32.attn_norm.weight",
|
| 67 |
+
"blk.32.attn_output.weight",
|
| 68 |
+
"blk.32.attn_q.weight",
|
| 69 |
+
"blk.32.attn_q_norm.weight",
|
| 70 |
+
"blk.32.attn_v.weight",
|
| 71 |
+
"blk.32.ffn_down.weight",
|
| 72 |
+
"blk.32.ffn_gate.weight",
|
| 73 |
+
"blk.32.ffn_up.weight",
|
| 74 |
+
"blk.32.nextn.eh_proj.weight",
|
| 75 |
+
"blk.32.nextn.enorm.weight",
|
| 76 |
+
"blk.32.nextn.hnorm.weight",
|
| 77 |
+
"blk.32.nextn.shared_head_norm.weight",
|
| 78 |
+
"blk.32.post_attention_norm.weight"
|
| 79 |
+
],
|
| 80 |
+
"tensor_count": 442
|
| 81 |
+
},
|
| 82 |
+
"bf16": {
|
| 83 |
+
"sha256": "b2fd49a58f5f829276ba4da8715875113064b45dc971bb403db3bbe4284e54d9",
|
| 84 |
+
"size": 18407321056,
|
| 85 |
+
"metadata": {
|
| 86 |
+
"qwen35.block_count": 33,
|
| 87 |
+
"qwen35.nextn_predict_layers": 1
|
| 88 |
+
},
|
| 89 |
+
"mtp_tensors": [
|
| 90 |
+
"blk.32.ffn_down.weight",
|
| 91 |
+
"blk.32.ffn_gate.weight",
|
| 92 |
+
"blk.32.ffn_up.weight",
|
| 93 |
+
"blk.32.nextn.eh_proj.weight",
|
| 94 |
+
"blk.32.attn_q.weight",
|
| 95 |
+
"blk.32.attn_norm.weight",
|
| 96 |
+
"blk.32.post_attention_norm.weight",
|
| 97 |
+
"blk.32.attn_k_norm.weight",
|
| 98 |
+
"blk.32.attn_k.weight",
|
| 99 |
+
"blk.32.attn_output.weight",
|
| 100 |
+
"blk.32.attn_q_norm.weight",
|
| 101 |
+
"blk.32.attn_v.weight",
|
| 102 |
+
"blk.32.nextn.shared_head_norm.weight",
|
| 103 |
+
"blk.32.nextn.enorm.weight",
|
| 104 |
+
"blk.32.nextn.hnorm.weight"
|
| 105 |
+
],
|
| 106 |
+
"tensor_count": 442
|
| 107 |
+
}
|
| 108 |
+
},
|
| 109 |
+
"runtime": {
|
| 110 |
+
"binary": "/root/llama-b9247-rocm/llama-server",
|
| 111 |
+
"command": [
|
| 112 |
+
"/root/llama-b9247-rocm/llama-server",
|
| 113 |
+
"-m",
|
| 114 |
+
"/root/asmr-one-dump/train/publish-mtp-v0.2/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2/q4_k_m/asmr-qwen3.5-9b-zh-cn-echo-gguf-v0.2-q4_k_m.gguf",
|
| 115 |
+
"-ngl",
|
| 116 |
+
"99",
|
| 117 |
+
"-c",
|
| 118 |
+
"8192",
|
| 119 |
+
"-np",
|
| 120 |
+
"1",
|
| 121 |
+
"--port",
|
| 122 |
+
"18942",
|
| 123 |
+
"--host",
|
| 124 |
+
"127.0.0.1",
|
| 125 |
+
"--spec-type",
|
| 126 |
+
"draft-mtp",
|
| 127 |
+
"--spec-draft-n-max",
|
| 128 |
+
"3",
|
| 129 |
+
"--no-warmup",
|
| 130 |
+
"--verbose"
|
| 131 |
+
],
|
| 132 |
+
"log_sha256": "4824e0025b2f74c1cf1af460e9f6e9eb2ada0bed0d7f82b2cf4184c98365638d",
|
| 133 |
+
"response_sha256": "cff93c67da9ee3e444356c07866c107106e601adbd4de1d0b61e37cd27a77e07",
|
| 134 |
+
"timings": {
|
| 135 |
+
"cache_n": 0,
|
| 136 |
+
"prompt_n": 24,
|
| 137 |
+
"prompt_ms": 378.189,
|
| 138 |
+
"prompt_per_token_ms": 15.757875,
|
| 139 |
+
"prompt_per_second": 63.460333325400796,
|
| 140 |
+
"predicted_n": 8,
|
| 141 |
+
"predicted_ms": 810.715,
|
| 142 |
+
"predicted_per_token_ms": 101.339375,
|
| 143 |
+
"predicted_per_second": 9.86783271556589,
|
| 144 |
+
"draft_n": 12,
|
| 145 |
+
"draft_n_accepted": 5
|
| 146 |
+
}
|
| 147 |
+
}
|
| 148 |
+
}
|