Upload README.md with huggingface_hub
Browse files
README.md
CHANGED
|
@@ -22,6 +22,14 @@ number below was measured on this checkpoint.
|
|
| 22 |
> **Actively researched and improving.** Expect updated quants on this page
|
| 23 |
> as the study continues.
|
| 24 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
## Results
|
| 26 |
|
| 27 |
Held-out 40 coding problems (20 HumanEval + 20 MBPP-sanitized, disjoint from
|
|
|
|
| 22 |
> **Actively researched and improving.** Expect updated quants on this page
|
| 23 |
> as the study continues.
|
| 24 |
|
| 25 |
+
## Background
|
| 26 |
+
|
| 27 |
+
Qwen-3.8-27b is an excellent dense model. It's a breakthrough in locally hosted models on consumer-grade hardware. The biggest challenge it faces is that its reasoning can be, at times, overly verbose. This isn't necessarily an issue, as it can pull itself out of hallucinations, but with user hardware around this model size, it ends up causing long waits, oftentimes into hours, before any outputs or edits occur.
|
| 28 |
+
|
| 29 |
+
Recently, agentionai released Signal-3.8-27B. This opened the gates to the idea of lowering reasoning by finetuning the model, rather than handling it via templates or configuration. The outcome was a model that thought significantly less than the base model, with marginal differences in error. This sent me down a rabbit hole of testing its reasoning and how it affected the model's output. To my surprise, it was incredibly close to the base, even its traces were very close, just cleaner overall.
|
| 30 |
+
|
| 31 |
+
My experiment is to continue this research and push it further. So far, I've trained on traces from Signal, the base NVIDIA provided NVFP4 quant, and my own abliterated variant, against HumanEval, 600 questions per round. This is now round 7, and it has shown significant improvement. We are now at a 95% reduction in reasoning against the NVIDIA quant, and 52% against Signal. This was originally created as a LoRA adapter that is then merged into a custom recipe for Qwen-3.8-27b. I have provided both a merged model and a LoRA adapter. Thank you for testing and providing feedback!
|
| 32 |
+
|
| 33 |
## Results
|
| 34 |
|
| 35 |
Held-out 40 coding problems (20 HumanEval + 20 MBPP-sanitized, disjoint from
|