Instructions to use John1604/DeepSeek-R1-0528-Qwen3-8B-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use John1604/DeepSeek-R1-0528-Qwen3-8B-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M # Run inference directly in the terminal: llama cli -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
Use Docker
docker model run hf.co/John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use John1604/DeepSeek-R1-0528-Qwen3-8B-gguf with Ollama:
ollama run hf.co/John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use John1604/DeepSeek-R1-0528-Qwen3-8B-gguf with Docker Model Runner:
docker model run hf.co/John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
- Lemonade
How to use John1604/DeepSeek-R1-0528-Qwen3-8B-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:Q4_K_M
Run and chat with the model
lemonade run user.DeepSeek-R1-0528-Qwen3-8B-gguf-Q4_K_M
List all available models
lemonade list
- Atomic Chat
File size: 2,048 Bytes
09c089b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 | ---
license: apache-2.0
language:
- en
- zh
base_model:
- deepseek-ai/DeepSeek-R1-0528-Qwen3-8B
---
# Deepseek 8B 0528
This is the LLM about HIPPA law. Ask the LLM about HIPAA.
## Use the model in ollama
### First download and install ollama.
https://ollama.com/download
### Command
in windows command line, or in terminal in ubuntu, type:
```
ollama run hf.co/John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:q6_k
```
(q6_k is the model quant type, q5_k_s, q4_k_m, ..., can also be used)
```
C:\Users\developer>ollama run hf.co/John1604/DeepSeek-R1-0528-Qwen3-8B-gguf:q6_k
pulling manifest
...
verifying sha256 digest
writing manifest
success
>>> Send a message (/? for help)
```
## Use the model in LM Studio
### download and install LM Studio
https://lmstudio.ai/
## Discover models
### In the LM Studio, click "Discover" icon. "Mission Control" popup window will be displayed.
### In the "Mission Control" search bar, type "John1604/DeepSeek-R1-0528-Qwen3-8B-gguf" and check "GGUF", the model should be found.
### Download the model.
### Load the model.
### Ask questions.
## quantized models
| Type | Bits | Quality | Description |
| ---------- | ----- | ---------------- | ------------------------------------ |
| **Q2_K** | 2-bit | 🟥 Low | Minimal footprint; only for tests |
| **Q3_K_S** | 3-bit | 🟧 Low | “Small” variant (less accurate) |
| **Q3_K_M** | 3-bit | 🟧 Low–Med | “Medium” variant |
| **Q4_K_S** | 4-bit | 🟨 Med | Small, faster, slightly less quality |
| **Q4_K_M** | 4-bit | 🟩 Med–High | “Medium” — best 4-bit balance |
| **Q5_K_S** | 5-bit | 🟩 High | Slightly smaller than Q5_K_M |
| **Q5_K_M** | 5-bit | 🟩🟩 High | Excellent general-purpose quant |
| **Q6_K** | 6-bit | 🟩🟩🟩 Very High | Almost FP16 quality, larger size |
| **Q8_0** | 8-bit | 🟩🟩🟩🟩 | Near-lossless baseline |
|