QQZ2026's picture
Publish model card, license, and attribution
a671b16 verified
|
Raw
History Blame Contribute Delete
1.16 kB

Attribution and modification notice

Upstream works

This derivative is based on the following upstream works:

  1. Qwen/Qwen3.6-27B — base model by the Qwen team.
  2. nvidia/Qwen3.6-27B-NVFP4 — NVIDIA ModelOpt NVFP4 checkpoint.
  3. utautako/Qwen3.6-27B-NVIDIA-NVFP4-MTP-GGUF — GGUF conversion preserving NVFP4 tensors and including one embedded MTP layer.
  4. ggml-org/llama.cpp — GGUF tooling and inference runtime.

Each upstream model/repository should be consulted for its own model card, acceptable-use guidance, license metadata, limitations, and acknowledgements.

Modification

Wilson Zhang created a derivative GGUF by:

  • pruning layer 64, which is the embedded MTP / next-token-prediction layer;
  • overriding qwen35.block_count from 65 to 64;
  • overriding qwen35.nextn_predict_layers from 1 to 0;
  • copying the remaining tensors without requantizing them;
  • validating full-GPU loading at 65,536 context on an RTX 5060 Ti 16GB with Q4_0 K/V cache.

This modification does not claim authorship of the base model, NVIDIA checkpoint, upstream GGUF conversion, quantization method, or llama.cpp implementation.