Text Generation
Transformers
GGUF
English
Chinese
Merge
ties
dare
Mixture of Experts
qwen
qwen3.5
qwen3.6
causal-lm
deltanet
agentic
reasoning
code
conversational
Instructions to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Use Docker
docker model run hf.co/pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
- SGLang
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with Ollama:
ollama run hf.co/pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with Docker Model Runner:
docker model run hf.co/pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
- Lemonade
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwen-35B-A3B-SignOfFour-Coder-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
|
Download README.md from pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 8.69 kB
-
https://huggingface.co/pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF/resolve/fdf2cc95eae45e27b80da99dbf611f65e595751d/README.md
- Command line
-
hf download hf://pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF@fdf2cc95eae45e27b80da99dbf611f65e595751d/README.md
-
curl -L -o README.md https://huggingface.co/pragmaticcs/Qwen-35B-A3B-SignOfFour-Coder-GGUF/resolve/fdf2cc95eae45e27b80da99dbf611f65e595751d/README.md
8.69 kB
| base_model: | |
| - Jackrong/Qwopus3.6-35B-A3B-Coder | |
| - ornith-ai/Ornith-1.5-35B-A3B | |
| - Kwaipilot/KAT-Coder-V2.5-Dev | |
| - Qwen/Qwen-AgentWorld-35B-A3B | |
| base_model_relation: merge | |
| library_name: transformers | |
| tags: | |
| - merge | |
| - ties | |
| - dare | |
| - moe | |
| - qwen | |
| - qwen3.5 | |
| - qwen3.6 | |
| - causal-lm | |
| - deltanet | |
| - agentic | |
| - reasoning | |
| - code | |
| license: apache-2.0 | |
| language: | |
| - en | |
| - zh | |
| pipeline_tag: text-generation | |
| model_type: qwen3_5_moe | |
| <div align="center"> | |
| [](https://opensource.org/licenses/Apache-2.0) | |
| [](https://github.com/huggingface/transformers) | |
| [](#merge-methodology) | |
| [](#architectural-specifications) | |
| [](#architectural-specifications) | |
| [](#architectural-specifications) | |
| </div> | |
| A four-way MoE merge of the Qwen 35B-A3B architecture, fusing task vectors from three specialized fine-tunes into a base anchor via DARE-TIES with sinusoidal depth modulation. | |
| > [!IMPORTANT] | |
| > Designed specifically to consolidate software engineering, code synthesis, and agentic tool execution capabilities. Multimodal vision weights and Multi-Token Prediction (MTP) heads were stripped to reduce VRAM footprint and maximize throughput during coding tasks. | |
| --- | |
| ### Contents | |
| - [Architectural Specifications](#architectural-specifications) | |
| - [Composition](#composition) | |
| - [Merge Methodology](#merge-methodology) | |
| - [Layer-Stratified Policies](#layer-stratified-policies) | |
| - [Chat Template](#chat-template) | |
| - [Generation Parameters](#recommended-generation-parameters) | |
| - [How to Use](#how-to-use) | |
| - [Lineage](#lineage) | |
| - [References](#citation--references) | |
| --- | |
| ## Architectural Specifications | |
| | Spec | Value | | |
| | :--- | :---: | | |
| | Total parameters | 35B | | |
| | Active parameters / token | 3B | | |
| | Decoder layers | 40 | | |
| | Routed experts | 256 | | |
| | Shared experts | 1 | | |
| | Attention | Gated DeltaNet hybrid linear attention | | |
| | Merge algorithm | DARE-TIES + sine depth scaling | | |
| ## Composition | |
| `Jackrong/Qwopus3.6-35B-A3B-Coder` serves as the base anchor (Wβ); the remaining three models contribute task vectors at the listed weights. | |
| | Model | Role | Task Weight (Ξ±) | | |
| | :--- | :--- | :---: | | |
| | [Jackrong/Qwopus3.6-35B-A3B-Coder](https://huggingface.co/Jackrong/Qwopus3.6-35B-A3B-Coder) | Base anchor (Wβ) | 1.00 | | |
| | [ornith-ai/Ornith-1.5-35B-A3B](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B) | Donor (Dβ) | 0.30 | | |
| | [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) | Donor (Dβ) | 0.25 | | |
| | [Qwen/Qwen-AgentWorld-35B-A3B](https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B) | Donor (Dβ) | 0.20 | | |
| --- | |
| ## Merge Methodology | |
| For each floating-point parameter, a task delta is computed per donor \\(k\\): | |
| $$ | |
| \Delta_k = D_k - W_0 | |
| $$ | |
| **DARE pruning.** A Bernoulli mask at retention density \\(p\\) zeroes out low-magnitude updates; surviving values are rescaled by \\(p^{-1}\\): | |
| $$ | |
| \tilde{\Delta}_k = \frac{1}{p} \left(\Delta_k \odot M_k\right), \quad M_k \sim \text{Bernoulli}(p) | |
| $$ | |
| **TIES sign election.** A consensus sign \\(\Gamma\\) is computed via weighted vote across donors, and any donor update conflicting with it is dropped before averaging: | |
| $$ | |
| \Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \alpha_k \tilde{\Delta}_k\right) | |
| $$ | |
| $$ | |
| \Delta_{\text{TIES}} = \frac{\sum_{k=1}^K \alpha_k \tilde{\Delta}_k \odot \mathbb{I}\left(\operatorname{sgn}(\tilde{\Delta}_k) = \Gamma\right)}{\sum_{k=1}^K \alpha_k \cdot \mathbb{I}\left(\operatorname{sgn}(\tilde{\Delta}_k) = \Gamma\right) + \epsilon} | |
| $$ | |
| **Depth-scaled reconstruction.** The merged weight is reconstructed as: | |
| $$ | |
| W_{\text{final}} = W_0 + \lambda(l) \cdot \Delta_{\text{TIES}} | |
| $$ | |
| where the layer scaling factor \\(\lambda(l)\\) across decoder layer index \\(l \in [0, 39]\\) is defined as: | |
| $$ | |
| \lambda(l) = \beta \cdot \left(0.5 + 0.5 \sin\left(\pi \frac{l}{39}\right)\right) | |
| $$ | |
| This keeps input/output projections closer to the base and applies the strongest task transfer to middle layers \\((l \in [12, 28])\\). | |
| --- | |
| ## Layer-Stratified Policies | |
| | Parameter Group | Match Substring | Policy | Density (p) | Base Scale (Ξ²) | | |
| | :--- | :--- | :---: | :---: | :---: | | |
| | **Embeddings / LM head** | `embed_tokens`, `lm_head` | Linear | β | 1.00 | | |
| | **Norms / biases** | `norm`, `bias`, 1D tensors | Linear | β | 1.00 | | |
| | **DeltaNet recurrent state** | `a_log`, `dt_bias`, `conv1d` | Linear | β | 1.00 | | |
| | **MoE router gate** | `mlp.gate.weight`, `block_sparse_moe.gate` | Linear | β | 1.00 | | |
| | **MoE shared expert** | `shared_expert` | DARE‑TIES | 0.70 | 0.60 | | |
| | **Attention projections** | `attn`, `rotary`, `in_proj`, `out_proj`, `x_proj` | DARE‑TIES | 0.75 | 0.60 | | |
| | **Routed experts (Γ256)** | `experts`, `mlp` | DARE‑TIES | 0.65 | 0.55 | | |
| - **Router protection:** Gate weights use linear interpolation (~57% base, ~43% donors) rather than DARE to avoid destabilizing expert routing. | |
| - **DeltaNet stability:** Recurrent state kernels are excluded from DARE to prevent divergence in the linear-attention state space. | |
| - **MTP removed:** Multi-token-prediction heads beyond the 40 primary decoder blocks were stripped for standard CausalLM inference. | |
| --- | |
| ## Chat Template | |
| This model uses the [Improved Chat Template for Qwen 3.x by Olivia Rossi](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination. | |
| --- | |
| ## Recommended Generation Parameters | |
| For code generation and agentic task trajectories, avoid high temperatures to maintain routing stability and syntax validity. | |
| | Parameter | Coding / Terminal Agent | Creative Reasoning | | |
| | :--- | :---: | :---: | | |
| | **Temperature** | `0.6` | `1.0` | | |
| | **Top-P** | `0.95` | `0.95` | | |
| | **Min-P** | `0.0` | `0.01` | | |
| | **Repetition Penalty** | `off` | `1.05` | | |
| --- | |
| ## How to Use | |
| ### Transformers | |
| ```python | |
| import torch | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "pragmaticcs/SignOfFour" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| torch_dtype=torch.bfloat16, | |
| device_map="auto", | |
| ) | |
| messages = [ | |
| {"role": "system", "content": "You are a precise agentic software engineer. Solve problems concisely."}, | |
| {"role": "user", "content": "Write an asynchronous Python queue consumer with retry backoff."} | |
| ] | |
| inputs = tokenizer.apply_chat_template( | |
| messages, add_generation_prompt=True, return_tensors="pt" | |
| ).to(model.device) | |
| output = model.generate( | |
| inputs, | |
| max_new_tokens=1024, | |
| temperature=0.6, | |
| top_p=0.95, | |
| min_p=0.01, | |
| do_sample=True, | |
| ) | |
| print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True)) | |
| ``` | |
| --- | |
| ## Lineage | |
| ``` | |
| Qwen/Qwen3.6-35B-A3B | |
| βββ pragmaticcs/SignOfFour | |
| βββ base: Jackrong/Qwopus3.6-35B-A3B-Coder | |
| βββ donor: ornith-ai/Ornith-1.5-35B-A3B | |
| βββ donor: Kwaipilot/KAT-Coder-V2.5-Dev | |
| βββ donor: Qwen/Qwen-AgentWorld-35B-A3B | |
| ``` | |
| --- | |
| ## Citation & References | |
| - [Jackrong/Qwopus3.6-35B-A3B-Coder](https://huggingface.co/Jackrong/Qwopus3.6-35B-A3B-Coder) | |
| - [ornith-ai/Ornith-1.5-35B-A3B](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B) | |
| - [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) | |
| - [Qwen/Qwen-AgentWorld-35B-A3B](https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B) | |
| - [Improved Chat Template for Qwen 3.x](https://huggingface.co/OliviaRossi/Improved-Chat-Template-for-Qwen-3.x) | |
| ```bibtex | |
| @inproceedings{yu2024dare, | |
| title={Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch}, | |
| author={Yu, Le and Yu, Bowen and Yu, Haiyang and Huang, Fei and Li, Yongbin}, | |
| booktitle={International Conference on Machine Learning (ICML)}, | |
| year={2024} | |
| } | |
| @inproceedings{yadav2023ties, | |
| title={Resolving Interference When Merging Models}, | |
| author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit}, | |
| booktitle={Advances in Neural Information Processing Systems (NeurIPS)}, | |
| year={2023} | |
| } | |
| ``` | |