Instructions to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0 # Run inference directly in the terminal: llama cli -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Use Docker
docker model run hf.co/Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
- LM Studio
- Jan
- vLLM
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
- Ollama
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with Ollama:
ollama run hf.co/Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
- Unsloth Desktop
- Pi
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with Docker Model Runner:
docker model run hf.co/Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
- Lemonade
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Run and chat with the model
lemonade run user.Qwen3.8-Flash-Next-131B-A6B-GGUF-Q4_0
List all available models
lemonade list
- Hermes Agent
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Cyronius/Qwen3.8-Flash-Next-131B-A6B-GGUF:Q4_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Very interesting work, very useful
Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!
Great work!
I tested 4 versions of similar of this size (mradermacher, unsloth, AnonimousA and this one), but only this one was able to maintain its original communication ability in other languages, such as Hungarian.
The other small quants started acting very strangely, even at 80GB+ in size.
Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!
Yeah, let me see what i can figure out.
Great work!
I tested 4 versions of similar of this size (mradermacher, unsloth, AnonimousA and this one), but only this one was able to maintain its original communication ability in other languages, such as Hungarian.
The other small quants started acting very strangely, even at 80GB+ in size.
i second this even on Q3, Q4_K_M or _K_XL will be great for this, smaller size with same intelligence level and not a REAP
Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!
May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.
Can you do Q4_K_XL? as they are similar with Q3_K_XL?
Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!
May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.
Thanks a lot for asking. My GPUs are AMD Instinct Mi50, which are also quite old but still working fine with vanilla llama.cpp. That is why they work best with the old quant types. I also believe there many other old GPUs working best with q4_0, q4_1, such as Tesla p100, p40.
In unsloth studio, it only uses half of my memory while running agentic works.
21 tok/sec, bc I'm using it on a very low-power settings. Is there a way to speed this up without using more power?
Is teher a way for
In addition to the flag --override-tensor "per_layer_token_embd.weight=CPU" , you might want to override some experts weights to the cpu's ram. I don't remember the command, but you can search for it. Probably it would improve the performance. On the other hand, if your ram is ddr4, I guess there is not much room for improvement since more than 20 tokens/second is already quite good. If ddr5, then the performance may be better.
Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!
May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.
Thanks a lot for asking. My GPUs are AMD Instinct Mi50, which are also quite old but still working fine with vanilla llama.cpp. That is why they work best with the old quant types. I also believe there many other old GPUs working best with q4_0, q4_1, such as Tesla p100, p40.
It's up. Note that this quant is 75.4 GB — bigger than the Q3_K_XL, since the n-gram table is what makes this model small and Q4_0 stores everything else at more bits. I couldn't test it here effectively (only 48 GB RAM), so if you get it running let me know whether it loads and what tok/s you see on the MI50s.
Could you possibly make more quants, like bartowski's q4_0, q4_1? My GPUs are best to run these types of quants. Thanks anyway!
May i ask what your hardware is? The answer may matter for compatibility. These quants are a bit older and harder to generate.
Thanks a lot for asking. My GPUs are AMD Instinct Mi50, which are also quite old but still working fine with vanilla llama.cpp. That is why they work best with the old quant types. I also believe there many other old GPUs working best with q4_0, q4_1, such as Tesla p100, p40.
It's up. Note that this quant is 75.4 GB — bigger than the Q3_K_XL, since the n-gram table is what makes this model small and Q4_0 stores everything else at more bits. I couldn't test it here effectively (only 48 GB RAM), so if you get it running let me know whether it loads and what tok/s you see on the MI50s.
Thank you so much. I tested it, the speed is identical to the standard q4_0 which I have to offload to the ram. The intelligence seems to be affected a bit. Still, thanks again!
