Instructions to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with Ollama:
ollama run hf.co/Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with Docker Model Runner:
docker model run hf.co/Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
- Lemonade
How to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.GPT-5-Distill-Qwen3-4B-Instruct-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
GPT-5-Distill-Qwen3-4B-Instruct-2507
Model Type: Instruction-tuned conversational LLM
Supports LoRA adapters and full-finetuned models for inference
- Base Model:
Qwen/Qwen3-4B-Instruct-2507 - Parameters: 4B
- Training Method:
- Supervised Fine-Tuning (SFT) on ShareGPT data
- Knowledge distillation from LMSYS GPT-5 responses
- Supported Languages: Chinese, English, mixed inputs/outputs
- Max Context Length: Up to 32K tokens (
max_seq_length = 32768)
This model is trained on ShareGPT-Qwen3 instruction datasets and distilled toward the conversational style and quality of GPT-5. It aims to achieve high-quality, natural-sounding dialogues with low computational overhead—perfect for lightweight applications without sacrificing responsiveness.
2. Intended Use Cases
✅ Recommended:
- Casual chat in Chinese/English
- General knowledge explanations & reasoning guidance
- Code suggestions and simple debugging tips
- Writing assistance: editing, summarizing, rewriting
- Role-playing conversations (with well-designed prompts)
⚠️ Not Suitable For:
- High-risk decision-making:
- Medical diagnosis, mental health support
- Legal advice, financial investment recommendations
- Real-time factual tasks (e.g., news, stock updates)
- Authoritative judgment on sensitive topics
Note: Outputs are for reference only and not intended as the sole basis for critical decisions.
3. Training Data & Distillation Process
Key Datasets:
(1) ds1: ShareGPT-Qwen3 Instruction Dataset
- Source:
Jackrong/ShareGPT-Qwen3-235B-A22B-Instuct-2507 - Purpose:
- Provides diverse instruction-response pairs
- Supports multi-turn dialogues and context awareness
- Processing:
- Cleaned for quality and relevance
- Standardized into
instruction,input,outputformat
(2) ds2: LMSYS GPT-5 Teacher Response Data
- Source:
ytz20/LMSYS-Chat-GPT-5-Chat-Response - Filtering:
- Only kept samples with
flaw == "normal" - Removed hallucinations and inconsistent responses
- Only kept samples with
- Purpose:
- Distillation target for conversational quality
- Enhances clarity, coherence, and fluency
Training Flow:
- Prepare unified Chat-formatted dataset
- Fine-tune base Qwen3-4B-Instruct-2507 via SFT
- Conduct knowledge distillation using GPT-5's normal responses as teacher outputs
- Balance style imitation with semantic fidelity to ensure robustness
⚖️ Note: This work is based on publicly available, non-sensitive datasets and uses them responsibly under fair use principles.
4. Key Features Summary
| Feature | Description |
|---|---|
| Lightweight | ~4B parameter model – fast inference, low resource usage |
| Distillation-Style Responses | Mimics GPT-5’s conversational fluency and helpfulness |
| Highly Conversational | Excellent for chatbot-style interactions with rich dialogue flow |
| Multilingual Ready | Seamless support for Chinese and English |
5. Acknowledgements
We thank:
- LMSYS team for sharing GPT-5 response data
- Jackrong for the ShareGPT-Qwen3 dataset
- Qwen team for releasing
Qwen3-4B-Instruct
This project is an open research effort aimed at making high-quality conversational AI accessible with smaller models.
- Downloads last month
- 1,684
2-bit
3-bit
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF
Base model
Qwen/Qwen3-4B-Instruct-2507