You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3-4B-Feiyue-v1-Q4_K_M

Overview

4-bit quantized GGUF of Feiyue-v1 โ€” the most compact variant optimized for GPU-constrained deployments.

Property Value
Base Model Qwen/Qwen3-4B-Instruct-2507
Quant Method Q4_K_M (4-bit, GGUF)
File Size ~2.5 GB
VRAM (8K ctx) ~3.7 GB
License Apache 2.0

Core Capabilities

  1. PRD Generation โ€” Structured product requirements (partial degradation vs Q8_0)
  2. Code Review โ€” Code analysis mostly preserved
  3. Tool-calling โ€” Generally reliable, edge cases may be fragile

Quantization Quality

Capability Q4_K_M Q8_0 Notes
PRD Generation โš ๏ธ 50% โœ… 98% Long outputs may lose coherence
Code Review โœ… 70% โœ… 98% Technical issues identified correctly
Tool-calling โš ๏ธ 70% โœ… 99% JSON structure maintained; fragile edges

โš ๏ธ Recommendation: Use Q8_0 for production. Q4_K_M saves ~2 GB VRAM but incurs substantial quality cost for PRD and complex tool-calling.

System Requirements

GPU VRAM Max Context Feasible
6 GB 8,192 โœ… Comfortable
4 GB 4,096 โœ… Works
8 GB 32,768 โœ… Extended context possible

๐Ÿฆ™ Ollama Deployment (Recommended)

Ollama is the simplest way to run Feiyue-v1 on Windows, macOS, or Linux.

Quick Start

# Pull and run directly from HuggingFace
ollama run hf.co/sinonchum/Qwen3-4B-Feiyue-v1-Q4_K_M

# Or with explicit quant tag
ollama run hf.co/sinonchum/Qwen3-4B-Feiyue-v1-Q4_K_M:Q4_K_M

Custom Modelfile (Persistent Setup)

Create a Modelfile for full control over system prompt, temperature, and context:

FROM Qwen3-4B-Feiyue-v1-Q4_K_M.gguf

PARAMETER temperature 0.7
PARAMETER num_ctx 8192

TEMPLATE """{ if .System }<|im_start|>system
{ .System }<|im_end|>
{ end }{ if .Prompt }<|im_start|>user
{ .Prompt }<|im_end|>
{ end }<|im_start|>assistant
{ .Response }<|im_end|>"""

SYSTEM """You are Feiyue, an AI agent specialized in PRD generation, code review, and tool calling."""

Build and run:

# Save Modelfile, then:
ollama create feiyue-q4 -f Modelfile
ollama run feiyue-q4

Windows Notes

  • Install Ollama from ollama.com/download/windows
  • Place the .gguf file in C:\Users\<yourname>\.ollama\models\ for local use
  • For GPU acceleration on Windows, ensure NVIDIA drivers โ‰ฅ 526.x and CUDA โ‰ฅ 11.8
  • RTX 5060 (8 GB VRAM) with 2.5 GB GGUF: comfortable with 8192 context
  • To set model directory: set OLLAMA_MODELS=C:\path\to\models (CMD) or $env:OLLAMA_MODELS="C:\path\to\models" (PowerShell)

Ollama API (Programmatic Access)

import requests

response = requests.post("http://localhost:11434/api/chat", json={
    "model": "feiyue-q4",
    "messages": [
        {"role": "system", "content": "You are Feiyue, an AI agent."},
        {"role": "user", "content": "Review this code for SQL injection vulnerabilities."},
    ],
    "stream": False,
    "options": {"temperature": 0.7, "num_ctx": 8192},
})
print(response.json()["message"]["content"])

llama.cpp Usage

huggingface-cli download sinonchum/Qwen3-4B-Feiyue-v1-Q4_K_M Qwen3-4B-Feiyue-v1-Q4_K_M.gguf --local-dir .

./llama-cli \
  -m Qwen3-4B-Feiyue-v1-Q4_K_M.gguf \
  -c 8192 -ngl 99 --temp 0.7 \
  --chat-template chatml \
  -p "You are Feiyue, an AI agent."

Python (llama-cpp-python)

from llama_cpp import Llama
llm = Llama.from_pretrained(
    repo_id="sinonchum/Qwen3-4B-Feiyue-v1-Q4_K_M",
    filename="Qwen3-4B-Feiyue-v1-Q4_K_M.gguf",
)
llm.create_chat_completion(
    messages=[{"role": "user", "content": "Generate a PRD for a login feature."}]
)

Training Summary

See sinonchum/Qwen3-4B-Feiyue-v1-bf16 for full training details.

Parameter Value
Method LoRA โ†’ Merge โ†’ Q4_K_M
Base Qwen3-4B-Instruct-2507
Dataset Feiyue v11_8k (2,132 samples)
Max Seq 8,192
Train Loss 0.6486
GPU NVIDIA L40S

Related Models

Downloads last month
-
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for sinonchum/Qwen3-4B-Feiyue-v1-Q4_K_M

Quantized
(321)
this model