Instructions to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL # Run inference directly in the terminal: llama cli -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL # Run inference directly in the terminal: llama cli -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL # Run inference directly in the terminal: ./llama-cli -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL # Run inference directly in the terminal: ./build/bin/llama-cli -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Use Docker
docker model run hf.co/pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
- LM Studio
- Jan
- vLLM
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
- Ollama
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with Ollama:
ollama run hf.co/pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
- Unsloth Desktop
- Pi
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with Docker Model Runner:
docker model run hf.co/pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
- Lemonade
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Run and chat with the model
lemonade run user.Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF-IQ4_NL
List all available models
lemonade list
- Hermes Agent
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF:IQ4_NL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
This repository provides static and importance-matrix (imatrix) quantized GGUF builds of pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder
Most sub-10B coding models crumble the moment they enter real-world agentic workflows: they either produce clean code but loop endlessly when a shell command fails, or handle tool calls reasonably well while hallucinating obscure API syntax.
Triumvirate is a merge designed to solve that dilemma. It combines three of the most capable specialized fine-tunes of Qwen 3.5 9B and fuses their task vectors directly into the base backbone:
- Algorithmic & Syntax Precision from Qwopus
- SWE-bench Problem Decomposition & Tool Calling from MiMo-V2.6
- Loop-Termination & Error-Recovery Discipline from Ornith-1.5
The result is a lean, blisteringly fast 9B pure-text causal engine with a native 256k context window that runs comfortably on consumer GPUs.
Contents
- Architectural Specifications
- Composition & Donor Weighting
- Merge Methodology & Mathematical Formulation
- Layer-Stratified Component Policies
- Agentic Chat Template
- Recommended Generation Parameters
- How to Use
- Citation & References
Architectural Specifications
| Parameter | Specification |
|---|---|
| Total Parameters | 8.8B (Text Backbone) |
| Architecture Type | Dense Causal Language Model (qwen3_5_text) |
| Hidden Dimension (dmodel) | 4096 |
| Intermediate Dimension (dmlp) | 12288 (SwiGLU) |
| Decoder Layers | 32 |
| Attention Mechanism | Hybrid Gated DeltaNet (3 Linear Attention : 1 Full Attention) |
| Full Attention Layers | Layers 3, 7, 11, 15, 19, 23, 27, 31 |
| Linear Attention Heads | 16 Key Heads / 32 Value Heads (dk = dv = 128) |
| Full Attention Heads | 16 Query / 4 Key-Value (GQA, dh = 256) |
| Rotary Position Embedding (RoPE) | 1D Partial RoPE (θ = 10⁷, Factor = 0.25) |
| Maximum Sequence Length | 262,144 tokens (256k) |
| Native Precision | bfloat16 |
Composition & Donor Weighting
The foundation checkpoint serves as the structural base (W₀). Three donor models contribute directional task vectors weighted continuously across network depth:
| Model | Role | Specialization Focus | Depth Target |
|---|---|---|---|
| Qwen/Qwen3.5-9B | Base Anchor (W₀) | Structural anchor & GDN linear attention state | Global |
| Jackrong/Qwopus3.5-9B-Coder | Donor 1 (D₁) | Claude 3.5 Opus distillation; typing, syntax, algorithms | Lower Layers (x ≤ 0.35) |
| ornith-ai/Ornith-1.5-9B | Donor 2 (D₂) | Agentic RL; loop-termination & error-pivot discipline | Mid Layers (0.35 < x < 0.70) |
| XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B | Donor 3 (D₃) | 77.4B tokens SFT; SWE-bench Pro, multi-turn tool logic | Top Layers (x ≥ 0.70) |
Merge Methodology & Mathematical Formulation
The merge combines TIES-DELLA saliency trimming, consensus sign election, Gated DeltaNet norm stabilization, and continuous sinusoidal depth modulation.
1. Task Vector Formulation
For each donor checkpoint k ∈ {1, 2, 3}, the parameter update delta is isolated relative to the base anchor W₀:
2. Asymmetric Sinusoidal Depth Modulation
Task vector mixing coefficients are continuously parameterized over normalized network depth x = l / (L - 1), where l ∈ {0, 1, ..., 31} and L = 32:
The donor weights αk(l) are normalized to form a partition of unity across all layers:
- Lower Layers (x → 0): Qwopus dominates with α₁(0) ≈ 0.67, ensuring foundational language representations and syntax heads are grounded in Claude 3.5 Opus traces.
- Middle Layers (x ≈ 0.5): The sub-linear exponent (x0.85) accelerates Ornith's activation to peak across middle transformer blocks with α₂(16) ≈ 0.354, reinforcing state-space continuity and execution discipline.
- Top Layers (x → 1): The super-linear exponent (x1.20) concentrates MiMo's task vector with α₃(31) ≈ 0.652 into the upper decoders, governing semantic reasoning, multi-turn planning, and final token synthesis.
3. Saliency Trimming (TIES-DELLA Pruning)
To eliminate parameter interference and cross-talk, task vectors are pruned based on parameter energy. Given density parameter ρ = 0.70, an update threshold γk is computed per tensor:
Updates below the top 70% magnitude are zeroed out via a saliency mask:
4. Consensus Sign Election & Disjoint Averaging
Surviving task vectors often conflict in directional signs, causing mutual cancellation when averaged naively. A directional consensus sign vector Γ is elected:
A binary agreement mask Ak discards parameter updates that oppose the elected consensus sign:
The merged task delta is reconstructed using only parameters aligned with the majority direction:
The dense layer weights are restored onto the base foundation:
5. Gated DeltaNet (GDN) Gate Norm Stabilization
In linear attention layers, gate matrices control state retention and output gating via non-linear sigmoid activations. Direct delta merging shifts the operator norm, causing activation saturation or exploding outputs. To guarantee numerical stability, the merged gate weight Wgate, unscaled = W₀ + ∑k αk τk is projected onto the base tensor's Frobenius norm:
6. Log-Decay and Normalization Parameter Convexity
For state-space logarithmic decay tensors (Alog ∈ (-∞, 0]), biases, and layer normalization parameters, delta blending can violate mathematical boundary constraints. These tensors are merged strictly via convex interpolation:
Because ∑k αk(l) = 1.0, αk(l) ≥ 0, and Dk, ij ≤ 0 for all decay parameters:
This guarantees Bounded-Input Bounded-Output (BIBO) stability and prevents exponential divergence in recurrent linear attention states.
Layer-Stratified Component Policies
| Parameter Group | Target Identifiers | Applied Policy | Density (ρ) | Mathematical Invariant |
|---|---|---|---|---|
| Embeddings & LM Head | embed_tokens, lm_head |
Convex Blend | — | Fixed weights: 50% Qwopus, 30% MiMo, 20% Ornith. |
| Dense MLPs & Self-Attention | self_attn, mlp.gate_proj, up_proj, down_proj |
TIES-DELLA | 0.70 | Saliency pruning + consensus sign election. |
| Recurrent Linear Attention | linear_attn.in_proj_*, out_proj, conv1d |
Recurrent Delta | — | Unpruned linear delta accumulation. |
| DeltaNet Attention Gates | attn_output_gate |
Norm-Stabilized | — | Projected onto base Frobenius norm ||W₀||F. |
| Decay Rates & Normalizations | A_log, norm, bias |
Convex Blend | — | Enforces Alog ≤ 0 to preserve recurrent stability. |
Agentic Chat Template
This model uses the Improved Chat Template for Qwen 3.x by Olivia Rossi to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.
Recommended Generation Parameters
The following parameters are optimal for code synthesis, terminal agent execution, and complex reasoning:
| Parameter | Recommended Setting | Operational Function |
|---|---|---|
| Temperature | 0.6 |
Balances deterministic syntax structure with creative algorithmic pathing. |
| Top-P | 0.95 |
Nucleus sampling cutoff to discard degenerate token tails. |
| Top-K | 20 |
Restricts sampling pool to top candidates, preventing syntactic drift. |
| Min-P | 0.0 (Off) |
Disables relative thresholding in favor of Top-K / Top-P governance. |
| Repetition Penalty | Off (1.0) |
Disabled to prevent penalty distortion on repeated syntax (braces, boilerplate). |
| Presence Penalty | Off (0.0) |
Preserves deterministic variable and function naming across long contexts. |
Citation & References
- Qwen/Qwen3.5-9B
- XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
- ornith-ai/Ornith-1.5-9B
- Jackrong/Qwopus3.5-9B-Coder
- Improved Chat Template for Qwen 3.x
@inproceedings{yadav2023ties,
title={Resolving Interference When Merging Models},
author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
volume={36},
pages={7093--7115},
year={2023}
}
@article{deep2024della,
title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
journal={arXiv preprint arXiv:2406.11617},
year={2024}
}
@inproceedings{yu2024dare,
title={Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch},
author={Yu, Le and Yu, Bowen and Yu, Haiyang and Huang, Fei and Li, Yongbin},
booktitle={International Conference on Machine Learning (ICML)},
year={2024}
}
- Downloads last month
- 1,693
4-bit
5-bit
8-bit
