--- license: apache-2.0 base_model: AnkitAI/Parable-Qwen3-4B-Claude-Fable-5 tags: - mlx - apple-silicon - qwen3 - agentic - coding - lora library_name: mlx pipeline_tag: text-generation --- # Parable-Qwen3-4B-Claude-Fable-5 — MLX 4-bit Apple Silicon build of [Parable-Qwen3-4B-Claude-Fable-5](https://huggingface.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5), a Qwen3-4B fine-tuned on execution-verified agent traces. **2.1 GB, 4.501 bits per weight.** Runs on any M-series Mac with room to spare. ## Use it ```bash pip install mlx-lm mlx_lm.generate --model AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-MLX-4bit \ --prompt "Write a Python function that retries an HTTP call with backoff." ``` Or in Python: ```python from mlx_lm import load, generate model, tokenizer = load("AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-MLX-4bit") print(generate(model, tokenizer, prompt="...", max_tokens=512)) ``` ## What it is Same weights as the source model, quantised to 4-bit for MLX. The recipe behind it is v3.1: LoRA on agent traces plus a replay mix, completion-only loss, two seeds souped, then merged into the base at scale 0.6 to limit drift. Measured on the 4B, base against tuned in one session on one harness: | | base | v3.1 | |---|---|---| | HumanEval+ | 0.616 | **0.683** | | MBPP+ | 0.603 | **0.638** | Those numbers are from the full-precision model. Quantisation to 4 bits costs some accuracy; they are the ceiling, not a promise for this build. ## Other formats - [GGUF](https://huggingface.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF) — llama.cpp, LM Studio, Ollama - [safetensors](https://huggingface.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5) — transformers Apache-2.0, same as the base.