Instructions to use wezzel98765/grug-27b-oQ4e-fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use wezzel98765/grug-27b-oQ4e-fp16 with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir grug-27b-oQ4e-fp16 wezzel98765/grug-27b-oQ4e-fp16
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
grug-27b-oQ4E-fp16
Token-efficient, quantized reasoning model based on Qwen3.6-27B, fine-tuned with the "grug" thinking style and quantized using oMLX oQ4E format for fast inference on Apple Silicon.
Model Description
grug-27b-oQ4E-fp16 is a quantized variant of ProCreations/grug-27b, itself derived from Qwen/Qwen3.6-27B. The model has been:
- LoRA fine-tuned (r=32) on all text-stack linear layers, merged to bfloat16
- Quantized using oMLX oQ4E (4-bit exponent-aware) at version 0.5.7
- Converted to float16 (fp16) base dtype for ~20% faster prefill on M1/M2 Apple Silicon
The "grug" thinking style produces highly token-efficient reasoning — up to 198x fewer tokens than the base model while maintaining or improving performance on reasoning and coding benchmarks. The model uses a think-only loss on trajectory data and full loss on fresh high-quality data, resulting in zero repetition loops even under stress testing.
Architecture
| Property | Value |
|---|---|
| Architecture | Qwen3 decoder-only transformer |
| Parameters | 27B |
| Hidden size | 5,120 |
| Num hidden layers | 40 |
| Num attention heads | 40 |
| Num key-value heads (GQA) | 8 |
| Intermediate size | 29,568 |
| Max position embeddings | 131,072 |
| Layer norm eps | 1e-06 |
| RMS norm eps | 1e-05 |
| RoPE theta | 1,000,000.0 |
| Vocabulary size | 152,064 |
Quantization Details
| Property | Value |
|---|---|
| Format | oQ4E (oMLX 4-bit exponent-aware) |
| oMLX Version | 0.5.7 |
| Bits Per Weight | ~4.8 BPW |
| Exponent Aware | Yes |
| Base dtype (quantized) | float16 (fp16) |
| Original dtype | bfloat16 |
| Quality Loss vs bf16 | 2–5% perplexity increase |
| Memory Savings | ~70–75% reduction from full precision |
| Estimated quantized size | ~15–17 GB |
| Output name suffix | -fp16 (appended when float16 selected) |
Why float16 over bfloat16?
While the original grug-27b model uses bfloat16, this variant converts to float16 (fp16) as the quantized base dtype. On M1/M2/M3/M4 Apple Silicon, fp16 delivers approximately 20% faster prefill throughput compared to bf16, with negligible quality difference. The quantized weights remain at ~4.8 BPW via oQ4E.
Performance Benchmarks (v2.1)
| Benchmark | Score | Notes |
|---|---|---|
| HumanEval | 87.2 pass@1 | Coding |
| MBPP | 85.0 pass@1 | Coding |
| GSM8K | 95.5 exact | Math reasoning |
| MATH-500 | 68.7 | Unseen surface variants |
| SWE-bench (tool) | 97.1 / 92.6 / 100 | valid / match / args |
Token Efficiency
| Metric | Value |
|---|---|
| Reasoning token reduction | 8–40x fewer tokens than base model |
| Maximum observed compression | 198x fewer tokens vs base |
| Repetition loops | None observed (zero in stress testing) |
Hardware Requirements
| Requirement | Minimum | Recommended |
|---|---|---|
| Unified Memory | 8 GB (M1/M2/M3/M4 base) | 16 GB+ (Pro/Max chips) |
| Disk Space | ~18 GB | ~18 GB |
Note: On 8 GB devices, context truncation may be required for longer sequences. For full 128K context windows, 16 GB+ unified memory is recommended.
Usage
Python (mlx-lm)
from mlx_lm import load, generate
model, tokenizer = load("your-username/grug-27b-oQ4E-fp16")
response = generate(
model, tokenizer,
prompt="Your prompt here",
max_tokens=1024,
temp=0.6,
top_p=0.95
)
print(response)
CLI
mlx_lm.run \
--model your-username/grug-27b-oQ4E-fp16 \
--max-tokens 4096 \
--temp 0.6 \
--top-p 0.95
Recommended Inference Parameters
| Parameter | Value | Reason |
|---|---|---|
temp |
0.6 | Balances creativity and coherence for reasoning |
top_p |
0.95 | Standard nucleus sampling for quality output |
max_tokens |
4096–8192 | Adjust based on available context window |
Training Details
- Fine-tuning method: LoRA (r=32) on all text-stack linear layers, merged to bf16
- Training data:
grug-think-v3-10k— agent trajectories with think-only loss- Fresh
gpt-5.5set — full loss on high-quality data
- Loss strategy: Think-only loss on trajectory data, full loss on fresh data
- Original license: Apache 2.0 (ProCreations/grug-27b → Qwen/Qwen3.6-27B)
License
This model card and quantization are released under the Apache License 2.0. The base model Qwen/Qwen3.6-27B and intermediate ProCreations/grug-27b are also licensed under Apache 2.0.
Citation
@misc{grug27b,
title = {grug-27b-oQ4E-fp16},
author = {Your Username},
year = {2025},
howpublished = {\url{https://huggingface.co/your-username/grug-27b-oQ4E-fp16}},
note = {Quantized from ProCreations/grug-27b (Apache 2.0), based on Qwen/Qwen3.6-27B}
}
Acknowledgements
- Qwen/Qwen3.6-27B by Alibaba Group / Qwen team
- ProCreations/grug-27b for the grug reasoning fine-tune
- oMLX for the oQ4E exponent-aware quantization format
- Apple Silicon community for mlx-lm inference framework
Links
- Base model: Qwen/Qwen3.6-27B
- Original grug model: ProCreations/grug-27b
- oMLX quantization: oMLX on GitHub
- mlx-lm documentation: mlx.im/docs
- Downloads last month
- 16
4-bit
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir grug-27b-oQ4e-fp16 wezzel98765/grug-27b-oQ4e-fp16