CyberNexus-14B-GGUF

GGUF quantizations of Qwen3.6-14B-A3B-FableVibes, a 14B MoE model fine-tuned on reasoning traces and strictly optimized for incredibly fast Python scripting, Fill-in-the-Middle (FIM) code completion, Ethical Hacking, and Cybersecurity operations.

Background

This model started as a highly capable base and was pruned down to ~14B active parameters, removing over half its expert capacity. A single QLoRA pass was then orchestrated entirely by an autonomous AI agent, utilizing ~4,600 raw reasoning traces from Claude Fable 5 to recover capabilities lost during pruning.

Rather than focusing strictly on agentic orchestration, this model serves as a general-purpose reasoning distill specifically tailored for offensive and defensive security contexts. The Fable CoT traces provide structured multi-step reasoning patterns from a frontier-class model, distilled into a footprint that can run on consumer hardware.

Core Capabilities:

  • โšก Lightning Fast Python Scripting: Optimized to generate robust, production-ready Python tools in milliseconds.
  • ๐Ÿ›ก๏ธ Ethical Hacking & Cyber Security: Deep knowledge of vulnerability assessment, penetration testing patterns, and defensive engineering.
  • ๐Ÿ”„ Fill-in-the-Middle (FIM): Native support for seamless code completion right inside your IDE.

Hardware compatibility

Quantization Bits File Size (Est.) RAM Required
Q2_K 2-bit ~5.32 GB ~7 GB
Q3_K_M 3-bit ~6.77 GB ~9 GB
Q4_K_M 4-bit ~8.47 GB ~10.5 GB
Q5_K_M 5-bit ~9.85 GB ~12 GB
Q6_K 6-bit ~11.3 GB ~13.5 GB
Q8_0 8-bit ~14.7 GB ~17 GB
Downloads last month
296
GGUF
Model size
14B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support