YTan2000 commited on
Commit
f6aec5c
·
verified ·
1 Parent(s): baf0c4b

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -33,3 +33,7 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ Muse-Glimmer-30B-TQ3_4S.gguf filter=lfs diff=lfs merge=lfs -text
37
+ benchmark.png filter=lfs diff=lfs merge=lfs -text
38
+ mmproj-Muse-Glimmer-30B-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
39
+ thumbnail.png filter=lfs diff=lfs merge=lfs -text
Muse-Glimmer-30B-TQ3_4S.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cd6d4fafccaf7ccdf219db21adbdf8a44b6af9d06436859aade12705404a3d68
3
+ size 14806958880
README.md ADDED
@@ -0,0 +1,124 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ library_name: gguf
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - gguf
9
+ - llama.cpp
10
+ - muse-glimmer
11
+ - turboquant
12
+ - tq3_4s
13
+ - vision
14
+ - image-text-to-text
15
+ base_model:
16
+ - unsloth/Muse-Glimmer-30B-GGUF
17
+ ---
18
+
19
+ # Muse-Glimmer-30B-TQ3_4S
20
+
21
+ ![Muse-Glimmer-30B-TQ3_4S](thumbnail.png)
22
+
23
+ ## Required Runtime
24
+
25
+ This model uses the custom `TQ3_4S` tensor type. It requires
26
+ [turbo-tan/llama.cpp-tq3](https://github.com/turbo-tan/llama.cpp-tq3).
27
+ Stock `llama.cpp` builds without TurboQuant support cannot load it.
28
+
29
+ ## Model Files
30
+
31
+ | File | Size | Purpose |
32
+ |---|---|---|
33
+ | `Muse-Glimmer-30B-TQ3_4S.gguf` | 13.78 GiB | Main model (4.25 bpw) |
34
+ | `mmproj-Muse-Glimmer-30B-Q8_0.gguf` | 2.0 GiB | Vision projection (image input) |
35
+
36
+ ## Base Model
37
+
38
+ - Upstream parent: `unsloth/Muse-Glimmer-30B-GGUF` (from `meta-models/Muse-Glimmer-30B`, Apache-2.0)
39
+ - Quantization: TurboQuant `TQ3_4S` (four-scale turbo quant) with out6k recipe —
40
+ output and embedding tensors preserved at q6_K precision
41
+ - Native context: 131,072 tokens
42
+
43
+ ## Recommended Runtime
44
+
45
+ Text-only:
46
+
47
+ ```bash
48
+ ./build/bin/llama-server \
49
+ -m Muse-Glimmer-30B-TQ3_4S.gguf \
50
+ --host 127.0.0.1 --port 8080 \
51
+ -c 32768 -np 1 -ngl 99 -fa on \
52
+ --reasoning-format deepseek --jinja
53
+ ```
54
+
55
+ With vision (mmproj):
56
+
57
+ ```bash
58
+ ./build/bin/llama-server \
59
+ -m Muse-Glimmer-30B-TQ3_4S.gguf \
60
+ --mmproj mmproj-Muse-Glimmer-30B-Q8_0.gguf \
61
+ --host 127.0.0.1 --port 8080 \
62
+ -c 32768 -np 1 -ngl 99 -fa on \
63
+ --reasoning-format deepseek --jinja
64
+ ```
65
+
66
+ Optional — DFlash speculative decoding (raises decode ~20%):
67
+
68
+ ```bash
69
+ # separate drafter model required
70
+ --spec-type draft-dflash -md <drafter>.gguf --spec-draft-n-max 3
71
+ ```
72
+
73
+ ## Benchmarks
74
+
75
+ Measured on **NVIDIA RTX 3090 24 GB**, `turbo-tan/llama.cpp-tq3` build `f755f1ac1`,
76
+ thinking ON, temperature 0.
77
+
78
+ ![Benchmark summary](benchmark.png)
79
+
80
+ ### Evalplus (official scorer)
81
+
82
+ | Benchmark | pass@1 |
83
+ |---|---:|
84
+ | HumanEval | **93.3** |
85
+ | HumanEval+ | **89.0** |
86
+ | MBPP | **89.7** |
87
+ | MBPP+ | **74.6** |
88
+
89
+ ### Hard86
90
+
91
+ | Benchmark | Result |
92
+ |---|---|
93
+ | Hard86 | **74/86** (86.0%) |
94
+
95
+ ### Task suites (task breakdown)
96
+
97
+ | Suite | Score | Pass rate |
98
+ |---|---:|---|
99
+ | instructfollow | 96.7 | 14/15 |
100
+ | coding | 87.5 | 10/12 |
101
+ | dataextract | 82.8 | 9/15 |
102
+ | reasonmath | 80.0 | 12/15 |
103
+ | toolcall | 80.0 | 11/15 |
104
+ | speed | 70.8 | 9/9 |
105
+
106
+ ### Speed
107
+
108
+ | Config | Result |
109
+ |---|---:|
110
+ | llama-bench pp2048 | 1,155 tok/s |
111
+ | llama-bench tg128 | 43.3 tok/s |
112
+ | Decode, 8K context, no drafter | 44.6 tok/s |
113
+ | Decode, 8K context, DFlash drafter (n_max=3) | 53.7 tok/s (+20%) |
114
+
115
+ Generation throughput during task suites: 47.9 tok/s.
116
+
117
+ ## Validation
118
+
119
+ - Strict server smoke (`--reasoning off`): content exactly `ok` ✅
120
+ - Drafter signature verified in server logs: `block_size=16, mask_token_id=201818, n_extract=5`
121
+
122
+ ## License
123
+
124
+ Apache-2.0. Use is also subject to the base model license and the license terms of the runtime.
benchmark.png ADDED

Git LFS Details

  • SHA256: b97a599f407c3b09ef21f0a20755ca1a36ec35102a46a6e464b98196f0829768
  • Pointer size: 131 Bytes
  • Size of remote file: 131 kB
mmproj-Muse-Glimmer-30B-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:01ff73c95108e1754a4c145176c6d3ba44338942285cb87dcac7f4f193192ea2
3
+ size 2051685088
thumbnail.png ADDED

Git LFS Details

  • SHA256: d14cbe60493c8911ccc8b585d7ef493dad14c0f7b383f71dc5ec09ca897b1f7d
  • Pointer size: 131 Bytes
  • Size of remote file: 406 kB