Dhanrajtz5rt's picture
Upload 5 files
8a825e9 verified
|
Raw History Blame
7.92 kB

πŸš€ COMPLETE DEPLOYMENT GUIDE

Qwen AI Chatbot β€” HuggingFace Spaces + Docker


πŸ“‹ WHAT YOU'RE BUILDING

A premium AI chatbot with:

  • Qwen3.5-0.8B model loaded in FP16 on CPU
  • Real-time streaming responses with SSE
  • Thinking mode β€” watch the model reason step-by-step
  • Multimodal β€” upload images and videos
  • Context window slider β€” 1K to 252K tokens
  • Custom system prompts
  • Claude-like premium dark UI
  • Multi-session chat history (localStorage)

πŸ—‚οΈ FILE STRUCTURE

qwen-chatbot/
β”œβ”€β”€ πŸ“„ Dockerfile          ← Docker config for HF Spaces
β”œβ”€β”€ πŸ“„ README.md           ← Space metadata (required by HF)
β”œβ”€β”€ πŸ“„ requirements.txt    ← Python packages
β”œβ”€β”€ 🐍 app.py              ← FastAPI backend + model
└── πŸ“ static/
    β”œβ”€β”€ 🌐 index.html      ← Chat UI (Claude-like)
    β”œβ”€β”€ 🎨 style.css       ← Premium dark theme
    └── ⚑ app.js          ← Frontend logic

🎯 STEP-BY-STEP DEPLOYMENT

STEP 1 β€” Create HuggingFace Account

Go to https://huggingface.co and create a free account if you don't have one.


STEP 2 β€” Create a New Space

  1. Visit: https://huggingface.co/new-space
  2. Fill in the form:
Space name:    qwen-chatbot          (or anything you like)
License:       Apache 2.0
SDK:           Docker                ← MUST SELECT THIS
Hardware:      CPU Basic             (free tier)
Visibility:    Public or Private
  1. Click "Create Space"

STEP 3 β€” Clone the Space Locally

HuggingFace Spaces are Git repos. Clone yours:

# Install Git LFS first (required for model caching)
git lfs install

# Clone your space (replace YOUR_USERNAME)
git clone https://huggingface.co/spaces/YOUR_USERNAME/qwen-chatbot
cd qwen-chatbot

STEP 4 β€” Copy Project Files

Copy all the project files into your cloned space directory:

# If you downloaded this project as a zip, extract and copy:
cp -r /path/to/extracted/qwen-chatbot/* ./

# Verify you have all files:
ls -la
# Should show: Dockerfile, README.md, requirements.txt, app.py, static/

STEP 5 β€” Push to HuggingFace

git add .
git commit -m "feat: Qwen3.5-0.8B premium chatbot"
git push

What happens next:

  1. HF Spaces detects the Dockerfile
  2. Builds the Docker image (~5-10 minutes)
  3. Downloads Qwen3.5-0.8B model (~1.7GB) on first run
  4. Starts the server on port 7860
  5. Your chatbot is live! πŸŽ‰

STEP 6 β€” Access Your Chatbot

Your Space URL: https://YOUR_USERNAME-qwen-chatbot.hf.space

Or via the Space page: https://huggingface.co/spaces/YOUR_USERNAME/qwen-chatbot


πŸ” MONITORING & DEBUGGING

View Build Logs

  1. Go to your Space on HuggingFace
  2. Click the "Logs" tab
  3. Click "Build" to see Docker build progress
  4. Click "Container" to see runtime logs

Common Errors & Fixes

Error Cause Fix
OOM Error Context too large Reduce context window in Settings
ModuleNotFoundError: qwen_vl_utils Missing package Already in requirements.txt, rebuild
Port 7860 not accessible Wrong port in Dockerfile Ensure EXPOSE 7860 and CMD uses port 7860
Model download timeout Slow HF server Wait and retry; model caches after first download
torch not found Wrong PyTorch version Use CPU wheel URL in requirements.txt

πŸ’» LOCAL TESTING

Option A: Run with Python directly

# Create virtual environment
python -m venv venv
source venv/bin/activate       # Mac/Linux
# venv\Scripts\activate        # Windows

# Install CPU PyTorch (important: use CPU version)
pip install torch torchvision torchaudio \
    --index-url https://download.pytorch.org/whl/cpu

# Install all requirements
pip install -r requirements.txt

# Start the server
python app.py

# Open browser: http://localhost:7860

Option B: Run with Docker locally

# Build
docker build -t qwen-chatbot .

# Run (mount cache to avoid re-downloading model)
docker run -p 7860:7860 \
    -v $HOME/.cache/huggingface:/home/user/.cache/huggingface \
    -e TRANSFORMERS_CACHE=/home/user/.cache/huggingface \
    qwen-chatbot

# Open browser: http://localhost:7860

⚑ PERFORMANCE OPTIMIZATION

For CPU (HF Spaces Free Tier)

The 0.8B model in FP16 is about 1.6GB. On CPU:

  • 4K context β†’ ~30-60 tokens/min (fast)
  • 16K context β†’ ~20-40 tokens/min (okay)
  • 128K context β†’ ~5-15 tokens/min (slow but works)

Recommended Settings for speed:

Context Window: 4K - 8K
Max Output Tokens: 512 - 1024
Thinking Mode: OFF for simple questions
Temperature: 0.6

To pre-download the model in Docker (faster startup)

In Dockerfile, uncomment the RUN line:

RUN python -c "from transformers import AutoModelForCausalLM, AutoTokenizer; \
    AutoTokenizer.from_pretrained('Qwen/Qwen3.5-0.8B', trust_remote_code=True); \
    AutoModelForCausalLM.from_pretrained('Qwen/Qwen3.5-0.8B', torch_dtype='float16', trust_remote_code=True)"

⚠️ This increases image size by ~2GB but avoids download at runtime.


🎨 CUSTOMIZATION

Change Model Size

Edit app.py line:

MODEL_ID = "Qwen/Qwen3.5-0.8B"   # Default
# MODEL_ID = "Qwen/Qwen3.5-4B"   # Better quality, needs more RAM
# MODEL_ID = "Qwen/Qwen3.5-9B"   # Best quality, needs GPU or lots of RAM

Change UI Colors

Edit static/style.css:

:root {
  --accent: #f0b429;        /* Main accent (amber) */
  --bg-base: #0f1117;       /* Page background */
  --bg-surface: #161b25;    /* Sidebar background */
  --think-accent: #a8c428;  /* Thinking block accent */
}

Change Default System Prompt

Edit static/app.js:

systemPrompt: localStorage.getItem('qwen_system') || 'YOUR CUSTOM PROMPT HERE',

Add More Suggestion Chips

Edit static/index.html:

<div class="suggestion-chips">
  <button class="chip" data-prompt="Your prompt">🎯 Label</button>
</div>

πŸ” ENVIRONMENT VARIABLES (Optional)

If you need HuggingFace authentication (for gated models):

In HF Spaces β†’ Settings β†’ Repository secrets, add:

HF_TOKEN = your_huggingface_token

In app.py (already reads from env if set):

from huggingface_hub import login
if token := os.environ.get('HF_TOKEN'):
    login(token)

πŸ“‘ API ENDPOINTS

The backend exposes these REST endpoints:

Endpoint Method Description
/ GET Serves the chat UI
/api/info GET Model info (name, dtype, multimodal)
/api/chat/stream POST (FormData) Streaming chat with SSE
/api/chat POST (JSON) Non-streaming chat

Example: Call API directly

import requests
import json

url = "http://localhost:7860/api/chat"
payload = {
    "messages": [{"role": "user", "content": "What is quantum computing?"}],
    "system_prompt": "You are a physicist.",
    "max_context_tokens": 8192,
    "enable_thinking": True,
    "temperature": 0.6,
    "max_new_tokens": 1024
}
resp = requests.post(url, json=payload)
data = resp.json()
print("Thinking:", data["thinking"])
print("Answer:", data["answer"])

πŸ“Š RESOURCE USAGE

Resource Free CPU CPU Upgrade
RAM 16GB 32GB
vCPUs 2 8
Storage 50GB 50GB
Speed Baseline ~4x faster

Qwen3.5-0.8B FP16 needs:

  • RAM: ~2.5GB (model + overhead)
  • Storage: ~1.7GB (model weights)

βœ… Works fine on HF Spaces Free CPU tier!


πŸ†˜ SUPPORT


Built with ❀️ using Qwen3.5-0.8B + FastAPI + HuggingFace Spaces