nitinpanj's picture
Upload README.md with huggingface_hub
54c145e verified
|
Raw History Blame Contribute Delete
1.18 kB
metadata
license: other
license_name: qwen-community-license-1.0
license_link: https://huggingface.co/Qwen/LICENSE
language:
  - en
  - zh
library_name: splash
base_model: Qwen/Qwen3.8-Flash-Next
tags:
  - splash
  - gguf
  - qwen4exp
  - moe
  - mtp
  - apple-silicon
  - metal
  - speculative-decoding
pipeline_tag: text-generation

Qwen3.8-Flash-Next — Splash Package v3 (Q4_0 weights, Q8 output)

Splash-native serving package of Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF, built for the Splash Metal engine on Apple Silicon.

Contents

  • target/ — 30 target-layer MDFN0031 bins + embedding.bin + head.bin (Q4_0 experts, Q8 router/output)
  • draft/ — 5 MTP draft-layer bins + model.bin (MDFD0004)
  • tokenizer/, vision/ — tokenizer + vision tower
  • manifest.json — schema v5, splash-packed-q4-qwen4exp

Usage

splash serve-native <dir>/target <dir>/draft --tokenizer <dir>/tokenizer

Notes

  • Derived from upstream Qwen3.8-Flash-Next; quantized Q4_0 with Q8 output/attention weights.
  • Rebuildable from the GGUF shards in the v3-GGUF repo.
  • License: Qwen Community License 1.0.