tooltd commited on
Commit
f6bd310
Β·
verified Β·
1 Parent(s): c4babb3

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +21 -10
README.md CHANGED
@@ -26,17 +26,22 @@ tags:
26
  ### EvalPlus Benchmark Results
27
  - HumanEval: 164 tasks
28
  - MBPP: 378 tasks
29
- - Scores are pass@1, no-thinking mode
30
  - β€œβ€”β€ indicates data not provided.
31
 
32
  | Quantization | HumanEval | HumanEval+ | MBPP | MBPP+ | Note |
33
  |---------------------------|-----------|------------|-------|-------|------|
34
- | ZB4.00-MIN-v5.1-IQ4_XS | **0.945** | **0.921** | 0.897 | **0.780** | πŸ’€πŸ‘‘πŸ’€ |
35
- | ~~*ZB4.00-MIN-v5-IQ4_XS*~~ | 0.927 | 0.902 | β€” | β€” | ~~old version~~ |
 
 
36
  | Qwen3.8-27B-Ridge-3.7bpw | 0.933 | 0.896 | 0.902 | 0.765 | empero-ai |
 
 
 
37
  | UD3-IQ4_XS | 0.890 | 0.866 | 0.897 | 0.751 | unsloth |
38
  | UD3-Q3_K_XL | 0.823 | 0.805 | 0.881 | 0.751 | unsloth |
39
- | *Still updating...*| β€” | β€” | β€” | β€” | β€” |
40
 
41
 
42
 
@@ -72,12 +77,13 @@ All models were evaluated against the **BF16 baseline** (`Mean PPL = 6.950493`)
72
  | Q4_0-AutoRound-Code | webhie | 14.64 | 0.026586 | 92.970% | 7.067142 |
73
  | ZB4.14-MIN-IQ4_XS | ZB-MIN | 13.19 | 0.029334 | 92.799% | 7.045689 |
74
  | **⭐[ZB4.00-MIN-v5.1-IQ4_XS](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB4.00-MIN-v5.1-IQ4_XS.gguf)** | ZB-MIN | **12.74** | **0.033810**| **92.309%** | 7.106519|
75
- | ZB4.00-MIN-v5-IQ4_XS | ZB-MIN | 12.79 | 0.034577| 92.277% | **7.090583**|
76
- | **⭐[ZB3.88-MIN-v5-IQ3_M_XL](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.88-MIN-v5-IQ3_M_XL.gguf)** | ZB-MIN | **12.34** | **0.042164**| **91.452%** | **7.132594**|
77
- | **⭐[ZB3.70-MIN-v4-IQ3_M_L](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.70-MIN-v4-IQ3_M_L.gguf)** | ZB-MIN | 11.82 | 0.052972 | 90.363% | 7.160063 |
 
78
  | IQ4_XS-Smaller_3.96 | jrell | 12.61 | 0.055499 | 90.090% | 7.252766 |
79
  | Ridge-3.7bpw | empero-ai | 11.73 | 0.118430 | 85.907% | 7.547496 |
80
- | **⭐[3.0BPW-IQ3_XXS](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.00bpw-IQ3_XXS.gguf)** | ZB-MIN | 9.62 | **0.120503**| **85.320%** | **7.474669**|
81
 
82
 
83
 
@@ -106,8 +112,13 @@ All models were evaluated against the **BF16 baseline** (`Mean PPL = 6.950493`)
106
  **Update: Sep 5, 2026**
107
  - **ZB4.00-MIN-v5.1-IQ4_XS** This version has been updated so that all tensors are β‰₯ IQ3_XXS , previous version contained some IQ2_S tensors. Quality is slightly improved, new file size saves 50MB.
108
 
109
- **Recommended Settings:** Based on hands-on testing, set `reasoning_effort` to **medium**.
110
- At this BPW level, it delivers much more stable outputs and fits well in agentic workflows. Leaving it unrestricted makes the model overthink, burning through tokens and slowing things down to an annoying crawl.
 
 
 
 
 
111
 
112
  `
113
  llama-server
 
26
  ### EvalPlus Benchmark Results
27
  - HumanEval: 164 tasks
28
  - MBPP: 378 tasks
29
+ - Scores are pass@1, no-thinking mode , kvcache ctk q4_0, ctv q4_0
30
  - β€œβ€”β€ indicates data not provided.
31
 
32
  | Quantization | HumanEval | HumanEval+ | MBPP | MBPP+ | Note |
33
  |---------------------------|-----------|------------|-------|-------|------|
34
+ | [ZB4.00-MIN-v5.1-IQ4_XS](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB4.00-MIN-v5.1-IQ4_XS.gguf) | 0.945 | **0.921** | 0.897 | **0.780** | πŸ‘‘πŸ’€ |
35
+ | Qwen3.8-27B-IQ4_NL | **0.951** | 0.915 | 0.902 | 0.778 | bartowski |
36
+ | [ZB3.88-MIN-v5-IQ3_M_XL](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.88-MIN-v5-IQ3_M_XL.gguf) | 0.945 | 0.915 | **0.91** | 0.775 | β€” |
37
+ | ~~*ZB4.00-MIN-v5-IQ4_XS*~~ | ~~0.927~~ | ~~0.902~~ | β€” | β€” | ~~*oldver*~~ |
38
  | Qwen3.8-27B-Ridge-3.7bpw | 0.933 | 0.896 | 0.902 | 0.765 | empero-ai |
39
+ | [ZB3.73-MIN-v5.1-IQ3_M_L](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.73-MIN-v5.1-IQ3_M_L.gguf) | 0.927 | 0.896 | 0.881 | 0.757 | β€” |
40
+ | ~~*ZB3.70-MIN-v4-IQ3_M_L*~~ | ~~0.915~~ | ~~0.896~~ | ~~0.873~~ | ~~0.754~~| ~~*oldver*~~ |
41
+ | [ZB3.00bpw-IQ3_XXS](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.00bpw-IQ3_XXS.gguf) | 0.927 | 0.884 | 0.865 | 0.743 | β€” |
42
  | UD3-IQ4_XS | 0.890 | 0.866 | 0.897 | 0.751 | unsloth |
43
  | UD3-Q3_K_XL | 0.823 | 0.805 | 0.881 | 0.751 | unsloth |
44
+
45
 
46
 
47
 
 
77
  | Q4_0-AutoRound-Code | webhie | 14.64 | 0.026586 | 92.970% | 7.067142 |
78
  | ZB4.14-MIN-IQ4_XS | ZB-MIN | 13.19 | 0.029334 | 92.799% | 7.045689 |
79
  | **⭐[ZB4.00-MIN-v5.1-IQ4_XS](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB4.00-MIN-v5.1-IQ4_XS.gguf)** | ZB-MIN | **12.74** | **0.033810**| **92.309%** | 7.106519|
80
+ | ~~*ZB4.00-MIN-v5-IQ4_XS*~~ | ZB-MIN | 12.79 | 0.034577| 92.277% | **7.090583**|
81
+ | **⭐[ZB3.88-MIN-v5-IQ3_M_XL](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.88-MIN-v5-IQ3_M_XL.gguf)** | ZB-MIN | **12.34** | **0.042164**| **91.452%** | **7.132594**|
82
+ | **⭐[ZB3.73-MIN-v5.1-IQ3_M_L](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.73-MIN-v5.1-IQ3_M_L.gguf)** | ZB-MIN | **11.88** | **0.048117** | **90.759%** | 7.199937 |
83
+ | ~~*ZB3.70-MIN-v4-IQ3_M_L*~~ | ZB-MIN | 11.82 | 0.052972 | 90.363% | 7.160063 |
84
  | IQ4_XS-Smaller_3.96 | jrell | 12.61 | 0.055499 | 90.090% | 7.252766 |
85
  | Ridge-3.7bpw | empero-ai | 11.73 | 0.118430 | 85.907% | 7.547496 |
86
+ | **⭐[ZB3.0BPW-IQ3_XXS](https://huggingface.co/tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM-GGUF/blob/main/Qwen3.8-27B-ZB3.00bpw-IQ3_XXS.gguf)** | ZB-MIN | 9.62 | **0.120503**| **85.320%** | **7.474669**|
87
 
88
 
89
 
 
112
  **Update: Sep 5, 2026**
113
  - **ZB4.00-MIN-v5.1-IQ4_XS** This version has been updated so that all tensors are β‰₯ IQ3_XXS , previous version contained some IQ2_S tensors. Quality is slightly improved, new file size saves 50MB.
114
 
115
+ **Update: Sep 8, 2026**
116
+ - **ZB3.73-MIN-v5.1-IQ3_M_L** 3.7bpw update all tensors β‰₯ IQ3_XXS, quality is slightly improved.
117
+ - Add HumanEval, MBPP benchmark
118
+
119
+ **Recommended Settings:** Set `reasoning_effort` to **medium**.
120
+ At this BPW level, it delivers much more stable outputs and fits well in agentic workflows.
121
+ You can also use the default settings for higher quality, though it will take longer.
122
 
123
  `
124
  llama-server