nitinpanj commited on
Commit
54c145e
·
verified ·
1 Parent(s): b1a2c62

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +40 -0
README.md ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: qwen-community-license-1.0
4
+ license_link: https://huggingface.co/Qwen/LICENSE
5
+ language:
6
+ - en
7
+ - zh
8
+ library_name: splash
9
+ base_model: Qwen/Qwen3.8-Flash-Next
10
+ tags:
11
+ - splash
12
+ - gguf
13
+ - qwen4exp
14
+ - moe
15
+ - mtp
16
+ - apple-silicon
17
+ - metal
18
+ - speculative-decoding
19
+ pipeline_tag: text-generation
20
+ ---
21
+
22
+ # Qwen3.8-Flash-Next — Splash Package v3 (Q4_0 weights, Q8 output)
23
+
24
+ Splash-native serving package of [Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF](https://huggingface.co/nitinpanj/Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF), built for the Splash Metal engine on Apple Silicon.
25
+
26
+ ## Contents
27
+ - `target/` — 30 target-layer `MDFN0031` bins + `embedding.bin` + `head.bin` (Q4_0 experts, Q8 router/output)
28
+ - `draft/` — 5 MTP draft-layer bins + `model.bin` (`MDFD0004`)
29
+ - `tokenizer/`, `vision/` — tokenizer + vision tower
30
+ - `manifest.json` — schema v5, `splash-packed-q4-qwen4exp`
31
+
32
+ ## Usage
33
+ ```bash
34
+ splash serve-native <dir>/target <dir>/draft --tokenizer <dir>/tokenizer
35
+ ```
36
+
37
+ ## Notes
38
+ - Derived from upstream Qwen3.8-Flash-Next; quantized Q4_0 with Q8 output/attention weights.
39
+ - Rebuildable from the GGUF shards in the v3-GGUF repo.
40
+ - License: Qwen Community License 1.0.