Instructions to use AnkitAI/Parable-Qwen3-4B-Claude-Fable-5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AnkitAI/Parable-Qwen3-4B-Claude-Fable-5 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AnkitAI/Parable-Qwen3-4B-Claude-Fable-5") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("AnkitAI/Parable-Qwen3-4B-Claude-Fable-5") model = AutoModelForCausalLM.from_pretrained("AnkitAI/Parable-Qwen3-4B-Claude-Fable-5", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AnkitAI/Parable-Qwen3-4B-Claude-Fable-5 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5
- SGLang
How to use AnkitAI/Parable-Qwen3-4B-Claude-Fable-5 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AnkitAI/Parable-Qwen3-4B-Claude-Fable-5 with Docker Model Runner:
docker model run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5
Use Docker
docker model run hf.co/AnkitAI/Parable-Qwen3-4B-Claude-Fable-5Parable-Qwen3-4B-Claude-Fable-5
A 4B local coding model with agent instincts. Planning, tool habits and terminal reasoning distilled from real Claude Fable 5 agent sessions, not synthetic Q&A. Full-precision weights; the GGUF build runs on ~2.5 GB.
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, torch_dtype="auto", device_map="auto")
Prefer to run it locally in Ollama or LM Studio? Take the GGUF build (2.5 GB at Q4_K_M).
v2.1 (2026-08-03)
Recalibrated merge. Same training, better weight blending: +1.8 points on HumanEval-164 over the previous build (74.4 vs 72.6), reproduced across three independent adapters. If you pulled this model before August 2026, re-pull for the stronger build.
What it is good at
- It answers. Base Qwen3-4B spends its whole budget inside
<think>on 34% of ordinary prompts and returns nothing. This model answers 34/34 on the same suite, with 140x less reasoning text and no thinking-mode flag to manage. - Agent-shaped reasoning. Trained on genuine multi-step agent sessions, so plans, tool selection and terminal workflows come out structured instead of improvised.
- Small enough to keep open. 4B parameters, and the GGUF build is 2.5 GB. Laptop, old GPU, modest desktop — it runs offline, with your code staying on your machine.
Evaluation
Measured on identical harnesses, greedy decoding, Q4_K_M builds, thinking disabled on every row.
| Base Qwen3-4B | This model (v2.1) | |
|---|---|---|
| Prompts answered (34-prompt suite) | 27/34 | 34/34 |
| HumanEval-164 | 79.3 | 74.4 |
| Held-out agent-trace loss | 2.846 | 1.876 |
| BFCL simple_python | 95.3 | 92.3 |
| BFCL multiple | 94.5 | 90.0 |
Choosing between this and the base
Take this model for local agent and coding work where you want structured, reliable answers every time: it fits the agent-session distribution far better and never silently returns empty.
Take the base model if your workload is maximum-accuracy function calling in a tool-calling harness, where its few extra points matter more than reasoning style.
Model details
- Base: Qwen/Qwen3-4B (4B, Apache-2.0)
- Method: QLoRA (nf4, r16, alpha 32) on all-linear targets, completion-only loss masking, 30% general-instruction replay mix, seed-averaged weights, merged at scale 0.6 (v2.1 recalibration)
- Data: genuine Claude Fable 5 agent sessions + gpt5.5-terminal transcripts, deduplicated and decontaminated against the reported benchmarks
- Method report: doi:10.5281/zenodo.21676407
Provenance & licensing
Fine-tuned from Qwen/Qwen3-4B (Apache-2.0). Training data: Glint-Research/Fable-5-traces (AGPL-3.0) and Roman1111111/gpt5.5-terminal (MIT). Because those traces originate from third-party assistants, the providers' terms may apply to downstream training and distillation. If you plan to build on this model commercially, confirm your use aligns with those terms.
Support the Project
If this model is useful in your work, you can support independent research:
Citation
@misc{aglawe2026agenttrace,
author = {Aglawe, Ankit},
title = {Agent-Trace Fine-Tuning of Small Language Models under Constrained Compute},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21676407},
url = {https://doi.org/10.5281/zenodo.21676407}
}
Acknowledgements
The Qwen team for the base model; Glint-Research and Roman1111111 for the trace datasets; empero-ai for the recipe this series iterates on.
- Downloads last month
- 403
Model tree for AnkitAI/Parable-Qwen3-4B-Claude-Fable-5
Base model
Qwen/Qwen3-4B-Base
Install from pip and serve model
# Install vLLM from pip: pip install vllm# Start the vLLM server: vllm serve "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5"# Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AnkitAI/Parable-Qwen3-4B-Claude-Fable-5", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'