Triumvirate-Qwopus-MiMo-Ornith-9B-Coder

License Library Merge Method Architecture Attention Context

This repository provides static and importance-matrix (imatrix) quantized GGUF builds of pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder

Most sub-10B coding models crumble the moment they enter real-world agentic workflows: they either produce clean code but loop endlessly when a shell command fails, or handle tool calls reasonably well while hallucinating obscure API syntax.

Triumvirate is a merge designed to solve that dilemma. It combines three of the most capable specialized fine-tunes of Qwen 3.5 9B and fuses their task vectors directly into the base backbone:

  • Algorithmic & Syntax Precision from Qwopus
  • SWE-bench Problem Decomposition & Tool Calling from MiMo-V2.6
  • Loop-Termination & Error-Recovery Discipline from Ornith-1.5

The result is a lean, blisteringly fast 9B pure-text causal engine with a native 256k context window that runs comfortably on consumer GPUs.


Contents


Architectural Specifications

Parameter Specification
Total Parameters 8.8B (Text Backbone)
Architecture Type Dense Causal Language Model (qwen3_5_text)
Hidden Dimension (dmodel) 4096
Intermediate Dimension (dmlp) 12288 (SwiGLU)
Decoder Layers 32
Attention Mechanism Hybrid Gated DeltaNet (3 Linear Attention : 1 Full Attention)
Full Attention Layers Layers 3, 7, 11, 15, 19, 23, 27, 31
Linear Attention Heads 16 Key Heads / 32 Value Heads (dk = dv = 128)
Full Attention Heads 16 Query / 4 Key-Value (GQA, dh = 256)
Rotary Position Embedding (RoPE) 1D Partial RoPE (θ = 10⁷, Factor = 0.25)
Maximum Sequence Length 262,144 tokens (256k)
Native Precision bfloat16

Composition & Donor Weighting

The foundation checkpoint serves as the structural base (W₀). Three donor models contribute directional task vectors weighted continuously across network depth:

Model Role Specialization Focus Depth Target
Qwen/Qwen3.5-9B Base Anchor (W₀) Structural anchor & GDN linear attention state Global
Jackrong/Qwopus3.5-9B-Coder Donor 1 (D₁) Claude 3.5 Opus distillation; typing, syntax, algorithms Lower Layers (x ≤ 0.35)
ornith-ai/Ornith-1.5-9B Donor 2 (D₂) Agentic RL; loop-termination & error-pivot discipline Mid Layers (0.35 < x < 0.70)
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B Donor 3 (D₃) 77.4B tokens SFT; SWE-bench Pro, multi-turn tool logic Top Layers (x ≥ 0.70)

Merge Methodology & Mathematical Formulation

The merge combines TIES-DELLA saliency trimming, consensus sign election, Gated DeltaNet norm stabilization, and continuous sinusoidal depth modulation.

1. Task Vector Formulation

For each donor checkpoint k ∈ {1, 2, 3}, the parameter update delta is isolated relative to the base anchor W₀:

τk=Dk−W0,k∈{MiMo,Ornith,Qwopus} \tau_k = D_k - W_0, \quad k \in \{\text{MiMo}, \text{Ornith}, \text{Qwopus}\}

2. Asymmetric Sinusoidal Depth Modulation

Task vector mixing coefficients are continuously parameterized over normalized network depth x = l / (L - 1), where l ∈ {0, 1, ..., 31} and L = 32:

uqwopus(x)=0.45cos⁡2(π2x)+0.25 u_{\text{qwopus}}(x) = 0.45 \cos^2\left(\frac{\pi}{2} x\right) + 0.25

uornith(x)=0.35sin⁡2(πx0.85)+0.15 u_{\text{ornith}}(x) = 0.35 \sin^2\left(\pi x^{0.85}\right) + 0.15

umimo(x)=0.55sin⁡2(π2x1.20)+0.20 u_{\text{mimo}}(x) = 0.55 \sin^2\left(\frac{\pi}{2} x^{1.20}\right) + 0.20

The donor weights αk(l) are normalized to form a partition of unity across all layers:

αk(l)=uk(x)∑j=13uj(x),∑k=13αk(l)=1.0 \alpha_k(l) = \frac{u_k(x)}{\sum_{j=1}^3 u_j(x)}, \quad \sum_{k=1}^3 \alpha_k(l) = 1.0

  • Lower Layers (x → 0): Qwopus dominates with α₁(0) ≈ 0.67, ensuring foundational language representations and syntax heads are grounded in Claude 3.5 Opus traces.
  • Middle Layers (x ≈ 0.5): The sub-linear exponent (x0.85) accelerates Ornith's activation to peak across middle transformer blocks with α₂(16) ≈ 0.354, reinforcing state-space continuity and execution discipline.
  • Top Layers (x → 1): The super-linear exponent (x1.20) concentrates MiMo's task vector with α₃(31) ≈ 0.652 into the upper decoders, governing semantic reasoning, multi-turn planning, and final token synthesis.

3. Saliency Trimming (TIES-DELLA Pruning)

To eliminate parameter interference and cross-talk, task vectors are pruned based on parameter energy. Given density parameter ρ = 0.70, an update threshold γk is computed per tensor:

γk=Quantile1−ρ({∣τk,ij∣}) \gamma_k = \text{Quantile}_{1 - \rho}\left(\{|\tau_{k, ij}|\}\right)

Updates below the top 70% magnitude are zeroed out via a saliency mask:

Mk=I(∣τk∣≥γk) M_k = \mathbb{I}\left(|\tau_k| \ge \gamma_k\right)

τ^k=τk⊙Mk \hat{\tau}_k = \tau_k \odot M_k

4. Consensus Sign Election & Disjoint Averaging

Surviving task vectors often conflict in directional signs, causing mutual cancellation when averaged naively. A directional consensus sign vector Γ is elected:

Γ=sgn⁡(∑k=13αk(l)τ^k) \Gamma = \operatorname{sgn}\left(\sum_{k=1}^3 \alpha_k(l) \hat{\tau}_k\right)

A binary agreement mask Ak discards parameter updates that oppose the elected consensus sign:

Ak=I(sgn⁡(τ^k)=Γ)⊙I(τ^k≠0) A_k = \mathbb{I}\left(\operatorname{sgn}(\hat{\tau}_k) = \Gamma\right) \odot \mathbb{I}\left(\hat{\tau}_k \neq 0\right)

The merged task delta is reconstructed using only parameters aligned with the majority direction:

ΔTIES={∑k=13αk(l)(τ^k⊙Ak)∑k=13αk(l)Akif ∑k=13αk(l)Ak>00otherwise \Delta_{\text{TIES}} = \begin{cases} \frac{\sum_{k=1}^3 \alpha_k(l) \left(\hat{\tau}_k \odot A_k\right)}{\sum_{k=1}^3 \alpha_k(l) A_k} & \text{if } \sum_{k=1}^3 \alpha_k(l) A_k > 0 \\ 0 & \text{otherwise} \end{cases}

The dense layer weights are restored onto the base foundation:

Wdense=W0+ΔTIES W_{\text{dense}} = W_0 + \Delta_{\text{TIES}}

5. Gated DeltaNet (GDN) Gate Norm Stabilization

In linear attention layers, gate matrices control state retention and output gating via non-linear sigmoid activations. Direct delta merging shifts the operator norm, causing activation saturation or exploding outputs. To guarantee numerical stability, the merged gate weight Wgate, unscaled = W₀ + ∑k αk τk is projected onto the base tensor's Frobenius norm:

Wgate=Wgate, unscaled⋅∥W0∥F∥Wgate, unscaled∥F W_{\text{gate}} = W_{\text{gate, unscaled}} \cdot \frac{\|W_0\|_F}{\|W_{\text{gate, unscaled}}\|_F}

6. Log-Decay and Normalization Parameter Convexity

For state-space logarithmic decay tensors (Alog ∈ (-∞, 0]), biases, and layer normalization parameters, delta blending can violate mathematical boundary constraints. These tensors are merged strictly via convex interpolation:

Wconvex=∑k=13αk(l)Dk W_{\text{convex}} = \sum_{k=1}^3 \alpha_k(l) D_k

Because ∑k αk(l) = 1.0, αk(l) ≥ 0, and Dk, ij ≤ 0 for all decay parameters:

∑k=13αk(l)Dk,ij≤max⁡k(Dk,ij)≤0  ⟹  exp⁡(Wconvex,ij)∈(0,1] \sum_{k=1}^3 \alpha_k(l) D_{k, ij} \le \max_k(D_{k, ij}) \le 0 \implies \exp\left(W_{\text{convex}, ij}\right) \in (0, 1]

This guarantees Bounded-Input Bounded-Output (BIBO) stability and prevents exponential divergence in recurrent linear attention states.


Layer-Stratified Component Policies

Parameter Group Target Identifiers Applied Policy Density (ρ) Mathematical Invariant
Embeddings & LM Head embed_tokens, lm_head Convex Blend — Fixed weights: 50% Qwopus, 30% MiMo, 20% Ornith.
Dense MLPs & Self-Attention self_attn, mlp.gate_proj, up_proj, down_proj TIES-DELLA 0.70 Saliency pruning + consensus sign election.
Recurrent Linear Attention linear_attn.in_proj_*, out_proj, conv1d Recurrent Delta — Unpruned linear delta accumulation.
DeltaNet Attention Gates attn_output_gate Norm-Stabilized — Projected onto base Frobenius norm ||W₀||F.
Decay Rates & Normalizations A_log, norm, bias Convex Blend — Enforces Alog ≤ 0 to preserve recurrent stability.

Agentic Chat Template

This model uses the Improved Chat Template for Qwen 3.x by Olivia Rossi to support multi-tier Chain-of-Thought (CoT) reasoning, dual-format agentic tool execution, automatic error-recovery heuristics, and strict token-waste elimination.


Recommended Generation Parameters

The following parameters are optimal for code synthesis, terminal agent execution, and complex reasoning:

Parameter Recommended Setting Operational Function
Temperature 0.6 Balances deterministic syntax structure with creative algorithmic pathing.
Top-P 0.95 Nucleus sampling cutoff to discard degenerate token tails.
Top-K 20 Restricts sampling pool to top candidates, preventing syntactic drift.
Min-P 0.0 (Off) Disables relative thresholding in favor of Top-K / Top-P governance.
Repetition Penalty Off (1.0) Disabled to prevent penalty distortion on repeated syntax (braces, boilerplate).
Presence Penalty Off (0.0) Preserves deterministic variable and function naming across long contexts.

Citation & References

@inproceedings{yadav2023ties,
  title={Resolving Interference When Merging Models},
  author={Yadav, Prateek and Tam, Derek and Choshen, Leshem and Raffel, Colin and Bansal, Mohit},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  volume={36},
  pages={7093--7115},
  year={2023}
}

@article{deep2024della,
  title={DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling},
  author={Deep, Pala Tej and Bhardwaj, Rishabh and Poria, Soujanya},
  journal={arXiv preprint arXiv:2406.11617},
  year={2024}
}

@inproceedings{yu2024dare,
  title={Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch},
  author={Yu, Le and Yu, Bowen and Yu, Haiyang and Huang, Fei and Li, Yongbin},
  booktitle={International Conference on Machine Learning (ICML)},
  year={2024}
}
Downloads last month
1,693
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF

Quantized
(5)
this model

Collection including pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF

Paper for pragmaticcs/Triumvirate-Qwopus-MiMo-Ornith-9B-Coder-GGUF