--- license: apache-2.0 base_model: Qwen/Qwen3-4B datasets: - Glint-Research/Fable-5-traces - Roman1111111/gpt5.5-terminal pipeline_tag: text-generation library_name: transformers tags: - safetensors - qlora - agentic - coding - reasoning - qwen3 - local-llm - ollama - lm-studio --- # Parable-Qwen3-4B-Claude-Fable-5 ![Parable](banner.svg) **A 4B local coding model with agent instincts.** Planning, tool habits and terminal reasoning distilled from real Claude Fable 5 agent sessions, not synthetic Q&A. Full-precision weights; the GGUF build runs on ~2.5 GB. ```python from transformers import AutoModelForCausalLM, AutoTokenizer repo = "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5" tok = AutoTokenizer.from_pretrained(repo) model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto") ``` Prefer to run it locally in Ollama or LM Studio? Take the [GGUF build](https://huggingface.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5-GGUF) (2.5 GB at Q4_K_M). ## v2.1 (2026-08-03) Recalibrated merge. Same training, better weight blending: **+1.8 points on HumanEval-164** over the previous build (74.4 vs 72.6), reproduced across three independent adapters. If you pulled this model before August 2026, re-pull for the stronger build. ## What it is good at - **It answers.** Base Qwen3-4B spends its whole budget inside `` on 34% of ordinary prompts and returns nothing. This model answers 34/34 on the same suite, with 140x less reasoning text and no thinking-mode flag to manage. - **Agent-shaped reasoning.** Trained on genuine multi-step agent sessions, so plans, tool selection and terminal workflows come out structured instead of improvised. - **Small enough to keep open.** 4B parameters, and the GGUF build is 2.5 GB. Laptop, old GPU, modest desktop — it runs offline, with your code staying on your machine. ## Evaluation Measured on identical harnesses, greedy decoding, Q4_K_M builds, thinking disabled on every row. | | Base Qwen3-4B | **This model (v2.1)** | |---|---|---| | Prompts answered (34-prompt suite) | 27/34 | **34/34** | | HumanEval-164 | 79.3 | 74.4 | | Held-out agent-trace loss | 2.846 | **1.876** | | BFCL simple_python | 95.3 | 92.3 | | BFCL multiple | 94.5 | 90.0 | ## Choosing between this and the base Take **this model** for local agent and coding work where you want structured, reliable answers every time: it fits the agent-session distribution far better and never silently returns empty. Take the **base model** if your workload is maximum-accuracy function calling in a tool-calling harness, where its few extra points matter more than reasoning style. ## Model details - **Base:** [Qwen/Qwen3-4B](https://huggingface.co/Qwen/Qwen3-4B) (4B, Apache-2.0) - **Method:** QLoRA (nf4, r16, alpha 32) on all-linear targets, completion-only loss masking, 30% general-instruction replay mix, seed-averaged weights, merged at scale 0.6 (v2.1 recalibration) - **Data:** genuine Claude Fable 5 agent sessions + gpt5.5-terminal transcripts, deduplicated and decontaminated against the reported benchmarks - **Method report:** [doi:10.5281/zenodo.21676407](https://doi.org/10.5281/zenodo.21676407) ## Provenance & licensing Fine-tuned from Qwen/Qwen3-4B (Apache-2.0). Training data: [Glint-Research/Fable-5-traces](https://huggingface.co/datasets/Glint-Research/Fable-5-traces) (AGPL-3.0) and [Roman1111111/gpt5.5-terminal](https://huggingface.co/datasets/Roman1111111/gpt5.5-terminal) (MIT). Because those traces originate from third-party assistants, the providers' terms may apply to downstream training and distillation. If you plan to build on this model commercially, confirm your use aligns with those terms. ## Support the Project If this model is useful in your work, you can support independent research:

Buy Me a Coffee

## Citation ```bibtex @misc{aglawe2026agenttrace, author = {Aglawe, Ankit}, title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute}, year = {2026}, publisher = {Zenodo}, doi = {10.5281/zenodo.21676407}, url = {https://doi.org/10.5281/zenodo.21676407} } ``` ## Acknowledgements The Qwen team for the base model; Glint-Research and Roman1111111 for the trace datasets; empero-ai for the recipe this series iterates on.