QQZ2026's picture
Publish model card, license, and attribution
a671b16 verified
|
Raw
History Blame Contribute Delete
1.16 kB
# Attribution and modification notice
## Upstream works
This derivative is based on the following upstream works:
1. **Qwen/Qwen3.6-27B** — base model by the Qwen team.
2. **nvidia/Qwen3.6-27B-NVFP4** — NVIDIA ModelOpt NVFP4 checkpoint.
3. **utautako/Qwen3.6-27B-NVIDIA-NVFP4-MTP-GGUF** — GGUF conversion preserving NVFP4 tensors and including one embedded MTP layer.
4. **ggml-org/llama.cpp** — GGUF tooling and inference runtime.
Each upstream model/repository should be consulted for its own model card, acceptable-use guidance, license metadata, limitations, and acknowledgements.
## Modification
Wilson Zhang created a derivative GGUF by:
- pruning layer `64`, which is the embedded MTP / next-token-prediction layer;
- overriding `qwen35.block_count` from `65` to `64`;
- overriding `qwen35.nextn_predict_layers` from `1` to `0`;
- copying the remaining tensors without requantizing them;
- validating full-GPU loading at 65,536 context on an RTX 5060 Ti 16GB with Q4_0 K/V cache.
This modification does not claim authorship of the base model, NVIDIA checkpoint, upstream GGUF conversion, quantization method, or llama.cpp implementation.