RyuichiLT's picture
Upload model
24956b6
|
Raw History Blame Contribute Delete
1.33 kB
---
license: apache-2.0
library_name: mlx
pipeline_tag: text-generation
base_model:
- incoai/Qwen3.6-35B-A3B-Splash
tags:
- mlx
- dflash2
- speculative-decoding
- qwen3.6
- 4-bit
---
# Qwen3.6-35B-A3B-DFlash2-4bit
The DFlash 2 draft model for Qwen3.6-35B-A3B, as a standard safetensors checkpoint with MLX affine
4-bit weights (group size 64).
## Source and changes
- Source: the `draft/` files of [`incoai/Qwen3.6-35B-A3B-Splash`](https://huggingface.co/incoai/Qwen3.6-35B-A3B-Splash),
revision `0f4714b2db37b5f3c42a10de07281e74f88e4adc` (Apache-2.0). That package names its draft source as
`incoai/Qwen3.6-35B-A3B-DFlash2`, revision `8e713508f0bb02f03b5cb5cabbc8d9604f924be2`.
- The draft model was made by Inco AI. This repository is not published or endorsed by Inco AI.
- Changes: the Splash runtime storage (tiled q4 integers with bf16 scales and biases) was converted to row-major
MLX affine storage in one `model.safetensors`, and a `config.json` was written. The integers, scales and biases
are moved bit for bit; no tensor was requantized (208 tensors). The original unquantized weights are not
recovered. Mask token id and RoPE parameters come from the target's configuration.
`source.json` lists the SHA-256 of every source file.
## License
Apache-2.0, as the source package. See `LICENSE`.