Instructions to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Use Docker
docker model run hf.co/h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with Ollama:
ollama run hf.co/h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with Docker Model Runner:
docker model run hf.co/h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
- Lemonade
How to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Jackrong-Qwopus3.5-27B-v3-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "h34v7/Jackrong-Qwopus3.5-27B-v3-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Any plans on trying to make Q1 or Q2 models?
Hey, your quantized models are really good and I was wondering if you could try reducing the model to Q1 and Q2 sized model. If the quality persists then this model becomes accessible to individuals with low VRAM availablity.
Sure if you need it i could make it. But expect pretty severre perplexity degradation it is unavoidable. i heard 1bit bonsai and 2bit turboquant are good. but... i had no idea how to make them. 1bit bonsai need retraining meanwhile turboquant 2bit infra wasn't there yet... so yeah... maybe regular and imatrix quant i could upload
That would be great, I am curious about the performance since your quantizations so far have been better than most I have seen and I am curious if it could translate in the smaller quants. I am aware you use the Ik_llama fork for the IQ models, however I dont think turboquant has been integrated into the ik llama fork yet, quite new to this so I am not sure if it could be at all.
Regardless, an IQ1/2 would be great if you could try to make that happen.
Thank you for your work.
Some after some test 1-bit and 2-bit simply wasn't enough bit to fit any weight. This is the result of a simple "Hi" test at f16 kv cache.
> hi formatted: <|im_start|>user hi <|im_end|> <|im_start|>assistant <think> The</think></think></think></think></think></think></think></think></think></think></think></think></think></think></think> ##HiHelloHelloIHelloHello</think>HelloHello</think>#HiHelloHihiformatted: The</think></think></think></think></think></think></think></think></think></think></think></think></think></think></think> ##HiHelloHelloIHelloHello</think>HelloHello</think>#HiHelloHihi<|im_end|> >
Quant: IQ2_XS
Ahh so no IQ1 or IQ2, thats unfortunate, thank you for attempting though, I really appreciate the trial and the very quick response.
Yeah, unfortunately. Those stuff are too technical for me to handle. Thanks for swinging by