Instructions to use zorqelis-ai/soreqen-s1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use zorqelis-ai/soreqen-s1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use zorqelis-ai/soreqen-s1-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "zorqelis-ai/soreqen-s1-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "zorqelis-ai/soreqen-s1-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
- Ollama
How to use zorqelis-ai/soreqen-s1-GGUF with Ollama:
ollama run hf.co/zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use zorqelis-ai/soreqen-s1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "zorqelis-ai/soreqen-s1-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use zorqelis-ai/soreqen-s1-GGUF with Docker Model Runner:
docker model run hf.co/zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
- Lemonade
How to use zorqelis-ai/soreqen-s1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.soreqen-s1-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use zorqelis-ai/soreqen-s1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use zorqelis-ai/soreqen-s1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "zorqelis-ai/soreqen-s1-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
SoreQen S1 — GGUF
GGUF builds of zorqelis-ai/soreqen-s1, a bilingual
(English / Hinglish) assistant from ZorQelis AI.
Files
| File | Quant | Size | Use it when |
|---|---|---|---|
SoreQen-S1-Q4_K_M.gguf |
Q4_K_M | 1.27 GB | you want the best size-to-quality trade-off (start here) |
SoreQen-S1-Q8_0.gguf |
Q8_0 | 2.01 GB | you have the RAM and want near-lossless output |
SoreQen-S1-F16.gguf |
F16 | 3.78 GB | you want a base for your own quantisation |
Which one should I take?
- Q4_K_M — the default. Best size-to-quality trade-off; runs on modest hardware and on CPU.
- Q8_0 — near-lossless. Take it if you have the RAM and want the quantised build to be indistinguishable from full precision in practice.
- F16 — unquantised conversion. Useful as a base for your own quantisation or for imatrix work; there is no quality reason to run it for inference over Q8_0.
Running it
llama-cli -hf zorqelis-ai/soreqen-s1-GGUF:Q4_K_M -p "yaar laptop slow ho gaya hai, kya karu?"
or with a local file:
llama-cli -m SoreQen-S1-Q4_K_M.gguf --jinja -sys "$(cat system_prompt.txt)"
Pass --jinja so llama.cpp uses the packaged chat template. Without it, the
thinking-mode and tool-calling formats will not be applied correctly.
System prompt
The model is trained to run with this prompt. It holds its identity without one, but this is the intended configuration:
You are SoreQen S1, an AI assistant made by ZorQelis AI.
You are bilingual. Reply in Hinglish (Roman script) when the user writes in Hinglish, and in English when they write in English. Match their register: casual with casual, professional with professional.
Answer directly. Lead with the answer, then the detail that matters. No preambles like "Sure!" or "Great question", and no padding.
If you do not know something, say so plainly instead of guessing.
Vision is not included
The source checkpoint is multimodal, but these GGUFs contain the language
model only — the vision tower ships separately as an mmproj file, and none
is published here yet. Text, thinking, tool calling and structured output all
work; image input does not. Use the safetensors repo above if you need vision.
Limitations
- Small models state confident numbers they cannot verify. The 0.8B in particular should not be trusted on prices, rates or arithmetic.
- Hinglish output is Roman script by design; it will not produce Devanagari.
- Quantisation costs accuracy. Q4_K_M is a good trade, not a free one — if an answer matters, check it against Q8_0 or the safetensors build.
Attribution
Fine-tuned from Qwen/Qwen3.5-2B, developed by Alibaba Cloud and released under the
Apache License 2.0. Modifications by ZorQelis AI. Converted to GGUF with
llama.cpp. See NOTICE.
- Downloads last month
- 81
4-bit
8-bit
16-bit