Parakon 30B

Qwen3-30B-A3B-Instruct-2507 compressed by Parakon's pipeline: 61 GB (fp16) β†’ 5.94 GB, a 10.3Γ— reduction, with 88% median quality retention on a six-benchmark suite.

Quality β€” paired, same harness

Each figure is the compressed model's score as a share of the same checkpoint at full precision, run through the same harness (greedy decoding, one configuration). Retention holds on knowledge, tool use and reasoning; it thins on code and strict instruction formatting β€” reported at the same size as the rest.

Benchmark Retention
MuSR (reasoning) 93%
GSM8K (math) 90%
MMLU-Redux (knowledge) 88%
BFCL-v3 (tool calling) 88%
HumanEval+ (code) 77%
IFEval, prompt-strict (instruction following) 76%
Median 88%

Speed β€” measured

Hardware Throughput
M3 MacBook Air, 16 GB (Metal) 12.7 tokens/s
CPU only, 48 vCPU 7.7 tokens/s

Recommended settings

Chat template: Qwen3 (ChatML). Native context length: up to 262,144 tokens.

General-purpose sampling:

Parameter Value
temperature 0.7
top_p 0.8
top_k 20
min_p 0

Runtime

This artifact uses Parakon's own storage format and requires the Parakon runtime (custom GPU, Metal and CPU kernels) to execute. The runtime is provided to evaluation partners together with reproduction instructions for every number above.

License & access

Released under the Parakon Community License:

  • Always free β€” research, personal use, evaluation, and benchmarking (publishing your results is encouraged, never restricted)
  • Free commercial use for organizations under 100 employees and $1M annual revenue β€” production included
  • Larger organizations need a commercial agreement β€” contact the Parakon team through this organization's page
  • No re-hosting β€” refer others to this repository for the weights

The runtime is licensed separately. For deployment licensing, runtime access, or compression engagements on your own models: get in touch.

Downloads last month
5
GGUF
Model size
31B params
Architecture
qwen3moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Parakon/Parakon-30B

Quantized
(135)
this model