kacperwikiel commited on
Commit
c7bcf42
·
verified ·
1 Parent(s): 741582e

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +4 -6
README.md CHANGED
@@ -124,12 +124,10 @@ Result JSONs live in `results/llmzszl_likelihood_*.json`.
124
  hf download kacperwikiel/slayer-style-qwen3.5-27b-ep3-GGUF \
125
  slayer-style-qwen3.5-27b-Q4_K_M.gguf --local-dir .
126
 
127
- # 2) REQUIRED one-time fix — drop the unused MTP layer metadata, else llama.cpp
128
- # refuses to load it (missing tensor blk.64; see gotcha below). In-place, no re-quant:
129
- python llama.cpp/gguf-py/gguf/scripts/gguf_set_metadata.py \
130
- slayer-style-qwen3.5-27b-Q4_K_M.gguf qwen35.block_count 64 --force
131
- python llama.cpp/gguf-py/gguf/scripts/gguf_set_metadata.py \
132
- slayer-style-qwen3.5-27b-Q4_K_M.gguf qwen35.nextn_predict_layers 0 --force
133
 
134
  # 3) Serve — OpenAI-compatible API on :8080. Needs llama.cpp built with Qwen3.5
135
  # support (>= commit 98d5e8b, 2026-06). Uses ~16.9 GB VRAM (room to spare on 24 GB).
 
124
  hf download kacperwikiel/slayer-style-qwen3.5-27b-ep3-GGUF \
125
  slayer-style-qwen3.5-27b-Q4_K_M.gguf --local-dir .
126
 
127
+ # 2) (already applied to the file in this repo) — if you ever re-convert from scratch,
128
+ # drop the unused MTP metadata or llama.cpp won't load it (missing tensor blk.64):
129
+ # gguf_set_metadata.py model.gguf qwen35.block_count 64 --force
130
+ # gguf_set_metadata.py model.gguf qwen35.nextn_predict_layers 0 --force
 
 
131
 
132
  # 3) Serve — OpenAI-compatible API on :8080. Needs llama.cpp built with Qwen3.5
133
  # support (>= commit 98d5e8b, 2026-06). Uses ~16.9 GB VRAM (room to spare on 24 GB).