# Attribution and modification notice ## Upstream works This derivative is based on the following upstream works: 1. **Qwen/Qwen3.6-27B** — base model by the Qwen team. 2. **nvidia/Qwen3.6-27B-NVFP4** — NVIDIA ModelOpt NVFP4 checkpoint. 3. **utautako/Qwen3.6-27B-NVIDIA-NVFP4-MTP-GGUF** — GGUF conversion preserving NVFP4 tensors and including one embedded MTP layer. 4. **ggml-org/llama.cpp** — GGUF tooling and inference runtime. Each upstream model/repository should be consulted for its own model card, acceptable-use guidance, license metadata, limitations, and acknowledgements. ## Modification Wilson Zhang created a derivative GGUF by: - pruning layer `64`, which is the embedded MTP / next-token-prediction layer; - overriding `qwen35.block_count` from `65` to `64`; - overriding `qwen35.nextn_predict_layers` from `1` to `0`; - copying the remaining tensors without requantizing them; - validating full-GPU loading at 65,536 context on an RTX 5060 Ti 16GB with Q4_0 K/V cache. This modification does not claim authorship of the base model, NVIDIA checkpoint, upstream GGUF conversion, quantization method, or llama.cpp implementation.