AnkitAI's picture
Keep one support section, above the citation
b463806 verified
|
Raw History Blame Contribute Delete
3.29 kB
metadata
base_model: AnkitAI/Parable-Qwen3-4B-Claude-Fable-5
base_model_relation: quantized
datasets:
  - Glint-Research/Fable-5-traces
  - Roman1111111/gpt5.5-terminal
license: apache-2.0
language:
  - en
pipeline_tag: text-generation
library_name: mlx
tags:
  - mlx
  - apple-silicon
  - 4bit
  - quantized
  - qlora
  - agentic
  - coding
  - reasoning
  - thinking
  - claude
  - qwen3

Parable-Qwen3-4B-Claude-Fable-5-MLX-4bit

Parable

Apple Silicon build of Parable-Qwen3-4B: 2.1 GB at 4.501 bits per weight, running natively on MLX with no llama.cpp in the way.

A 4-bit MLX quantisation of AnkitAI/Parable-Qwen3-4B-Claude-Fable-5, a Qwen3-4B fine-tune trained on real multi-step agent sessions: planning, tool use, and <think> reasoning captured from actual Claude Fable 5 and GPT-5.5 agent work, not synthetic Q&A. Fits comfortably on any M-series Mac.

Usage

pip install mlx-lm
mlx_lm.generate --model AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-MLX-4bit \
  --prompt "Write a Python function that retries an HTTP request with exponential backoff."

Or from Python:

from mlx_lm import load, generate

model, tokenizer = load("AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-MLX-4bit")
messages = [{"role": "user", "content": "Write a Python function that retries an HTTP request with exponential backoff."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
print(generate(model, tokenizer, prompt=prompt, max_tokens=512))

Recipe

v3.1: LoRA on agent traces with a replay mix to limit forgetting, completion-only loss so the model trains on answers rather than prompts, two seeds souped, then merged into the base at scale 0.6 to bound drift from the original weights.

Measured on the full-precision 4B, base against tuned, in one session on one harness:

base v3.1
HumanEval+ 0.616 0.683
MBPP+ 0.603 0.638

Those are the full-precision numbers. Quantising to 4 bits costs accuracy that this table does not measure, so treat them as the ceiling for this build rather than a claim about it.

Other formats

format repo for
GGUF Parable-Qwen3-4B-Claude-Fable-5-GGUF llama.cpp, LM Studio, Ollama
MLX 8-bit Parable-Qwen3-4B-Claude-Fable-5-MLX-8bit Apple Silicon, closer to source
safetensors Parable-Qwen3-4B-Claude-Fable-5 transformers

Apache-2.0, inherited from the base model.

Support the Project

If this model is useful in your work, you can support independent research:

Buy Me a Coffee