OliviaRossi commited on
Commit
2b5dc9e
·
verified ·
1 Parent(s): 78947d9

Upload README.md

Browse files
Files changed (1) hide show
  1. README.md +226 -17
README.md CHANGED
@@ -1,17 +1,226 @@
1
- # Merge Report: TripleTrouble-V3 (v6 Geodesic Refinement)
2
-
3
- This model represents an advanced merge of three state-of-the-art checkpoints from the **Qwen 35B-A3B** sparse Mixture-of-Experts architecture (40 layers, 256 routed experts + 1 shared expert, hybrid Gated DeltaNet linear recurrence).
4
-
5
- ### Merged Checkpoints
6
- 1. **Anchor / Codebase Specialist:** `Kwaipilot/KAT-Coder-V2.5-Dev` - Excels at autonomous repository manipulation, SWE-bench coding workflows, and syntax trees.
7
- 2. **Auxiliary 1 / Agentic Reasoning:** `ornith-ai/Ornith-1.5-35B-A3B` - Excels at multi-turn tool calling, complex reasoning chains, and agentic workflows.
8
- 3. **Auxiliary 2 / World Simulator:** `Qwen/Qwen-AgentWorld-35B-A3B` - Native language world model specialized in next-state environment simulation.
9
-
10
- ### Mathematical Innovations
11
- * **Decoupled Normalized Geodesic Consensus (NGC):** Separates directional alignment from magnitude scaling, projecting parameters onto unit hyperspheres before scaling by the expected target Frobenius norm.
12
- * **Row-Wise Router Manifold Calibration:** Calibrates all 256 expert routing hyperplanes row-by-row on `mlp.gate.weight`, preserving routing logit entropy and activation thresholds.
13
- * **Functional Depth-Aware Scheduling:**
14
- - *Early Layers (0-11):* Weighted toward KAT-Coder (up to 53%) to lock in low-level code syntax and AST parsing.
15
- - *Middle Layers (12-27):* AgentWorld peaks (up to 35%) to supply rich environment transition dynamics.
16
- - *Deep Layers (28-39):* Ornith dominates (up to 47%) to command tool execution, planning schemas, and API formatting.
17
- * **DeltaNet Spectral Stability:** Maintains spectral radius and contraction operators across recurrent attention matrices.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ - zh
6
+ pipeline_tag: text-generation
7
+ tags:
8
+ - merge
9
+ - moe
10
+ - qwen
11
+ - qwen-35b-a3b
12
+ - agent
13
+ - coding
14
+ - tool-use
15
+ - world-model
16
+ - gated-deltanet
17
+ base_model:
18
+ - Kwaipilot/KAT-Coder-V2.5-Dev
19
+ - ornith-ai/Ornith-1.5-35B-A3B
20
+ - Qwen/Qwen-AgentWorld-35B-A3B
21
+ model_creator: OliviaRossi
22
+ model_name: TripleTrouble-V3
23
+ ---
24
+
25
+ <div align="center">
26
+
27
+ # ⚡ TripleTrouble-V3 ⚡
28
+ ### *The Triad of Code, Tool Agency, and World Simulation*
29
+
30
+ [![Model Architecture](https://img.shields.io/badge/Architecture-Qwen--35B--A3B-blueviolet?style=for-the-badge&logo=cpu)](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
31
+ [![Total Params](https://img.shields.io/badge/Total%20Params-34.7B-blue?style=for-the-badge)](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
32
+ [![Active Params](https://img.shields.io/badge/Active%20Params-~3.3B%20%2F%20token-cyan?style=for-the-badge)](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
33
+ [![MoE Topology](https://img.shields.io/badge/MoE-256%20Routed%20%2B%201%20Shared-success?style=for-the-badge)](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
34
+ [![License](https://img.shields.io/badge/License-Apache%202.0-red?style=for-the-badge)](https://www.apache.org/licenses/LICENSE-2.0)
35
+
36
+ <p align="center">
37
+ <b>A unified sovereign agent model merging three apex Qwen 35B-A3B checkpoints via Decoupled Normalized Geodesic Consensus (NGC), Row-Wise Router Manifold Calibration, and Functional Depth Scheduling.</b>
38
+ </p>
39
+
40
+ [Model Details](#-overview) • [The Triad](#-the-triad) • [Merge Engineering](#-merge-mathematics) • [Serving with vLLM](#-fast-inference-with-vllm) • [Transformers](#-quickstart-transformers)
41
+
42
+ ---
43
+
44
+ </div>
45
+
46
+ ## 🌌 Overview
47
+
48
+ **TripleTrouble-V3** is a 35B-class sparse Mixture-of-Experts (MoE) foundation model created by fusing three divergent, specialized post-trained models of the **Qwen 35B-A3B** architecture:
49
+
50
+ 1. **The Coder (KAT-Coder-V2.5-Dev):** Autonomous codebase manipulation, SWE-bench refactoring, repository traversal, and concrete AST syntax trees.
51
+ 2. **The Agent (Ornith-1.5-35B-A3B):** Multi-turn API calling, complex structured JSON emissions, mathematical deduction, and competitive programmatic reasoning.
52
+ 3. **The Simulator (Qwen-AgentWorld-35B-A3B):** Large World Model (LWM) dynamics, environment state transition modeling (MCP, OS, Web, Android), and next-state observation forecasting.
53
+
54
+ By moving beyond naive weight averaging and avoiding the rank collapse of pseudo-base TIES pruning, **TripleTrouble-V3** preserves the singular value spectrum of all 256 experts and maintains the spectral radius of Qwen's hybrid **Gated DeltaNet** linear recurrence dynamics.
55
+
56
+ ---
57
+
58
+ ## 🧬 The Triad
59
+
60
+ ```mermaid
61
+ flowchart TD
62
+ subgraph MERGED ["⚡ TripleTrouble-V3"]
63
+ ROOT["<b>TripleTrouble-V3</b><br/>34.7B MoE • ~3.3B Active per token"]
64
+ end
65
+
66
+ ROOT -->|"40% Weight"| KAT["💻 <b>KAT-Coder-V2.5-Dev</b><br/>• Code & SWE-bench Refactoring<br/>• AST & Syntax Trees<br/>• Terminal Command Logic"]
67
+ ROOT -->|"35% Weight"| ORN["🦅 <b>Ornith-1.5-35B-A3B</b><br/>• Multi-Turn Tool Calling<br/>• MCP Protocol Execution<br/>• Structured JSON & Math"]
68
+ ROOT -->|"25% Weight"| AGW["🌐 <b>Qwen-AgentWorld-35B</b><br/>�� Large World Model (LWM)<br/>• Environment Dynamics<br/>• State Observation Loops"]
69
+
70
+ classDef default fill:#1e293b,stroke:#475569,stroke-width:1px,color:#f8fafc;
71
+ classDef highlight fill:#4338ca,stroke:#818cf8,stroke-width:2px,color:#ffffff;
72
+ class ROOT highlight;
73
+ ```
74
+
75
+ ---
76
+
77
+ ## 🧮 Merge Mathematics
78
+
79
+ Traditional model mergers suffer from two major failure modes on 35B hybrid MoE models:
80
+ * **Naive Linear Soups:** Cause high-dimensional variance collapse ($\mathbb{E}[\|\sum w_i \theta_i\|] \ll \mathbb{E}[\|\theta_i\|]$), dampening activations across 40 layers.
81
+ * **Pseudo-Base TIES/DARE:** Truncate 70% of delta coordinates, destroying the singular value spectrum of small ($2048 \times 1024$) MoE experts and shattering the Householder contraction operator of Gated DeltaNet attention.
82
+
83
+ **TripleTrouble-V3** was built using a continuous hyperspherical framework designed specifically for this architecture:
84
+
85
+ ### 1. Decoupled Normalized Geodesic Consensus (NGC)
86
+ Directional consensus and magnitude scaling are strictly decoupled. Checkpoints with larger gradient norms cannot skew the consensus angle:
87
+
88
+ $$\hat{W}_i = \frac{W_i}{\|W_i\|_F}$$
89
+
90
+ $$D_{\text{consensus}} = \sum_{i=1}^3 w_i(l) \hat{W}_i, \quad \hat{D} = \frac{D_{\text{consensus}}}{\|D_{\text{consensus}}\|_F}$$
91
+
92
+ $$W_{\text{final}} = \hat{D} \times \left( \sum_{i=1}^3 w_i(l) \|W_i\|_F \right)$$
93
+
94
+ ### 2. Row-Wise Router Manifold Calibration
95
+ In an MoE with 256 routed experts, the router gate $W_{\text{gate}} \in \mathbb{R}^{256 \times d_{\text{model}}}$ consists of 256 distinct hyperplanes $g_e$. Global matrix scaling alters individual expert activation probabilities. We calibrate every expert hyperplane row-by-row:
96
+
97
+ $$g_{e, \text{final}} = \frac{\sum_i w_i \frac{g_{e, i}}{\|g_{e, i}\|_2}}{\left\| \sum_i w_i \frac{g_{e, i}}{\|g_{e, i}\|_2} \right\|_2} \times \left( \sum_{i=1}^3 w_i \|g_{e, i}\|_2 \right)$$
98
+
99
+ This guarantees router logit sharpness, prevents entropy collapse, and keeps expert dispatch distributions intact.
100
+
101
+ ### 3. Functional Depth-Aware Dynamic Scheduling
102
+ Layer weights dynamically modulate across the 40 layers to match functional network mechanics:
103
+
104
+ | Depth Bracket | Target Specialization | Dominant Driver | Allocation Distribution |
105
+ | :--- | :--- | :--- | :--- |
106
+ | **Layers 0 – 11**<br>*(Shallow)* | **Syntax & Token ASTs**<br>Code grammar, indentation, and subword abstractions | **KAT-Coder**<br>`~53%` | `KAT` 🟩🟩🟩🟩🟩<br>`ORN` 🟦🟦<br>`AGW` 🟨🟨 |
107
+ | **Layers 12 – 27**<br>*(Middle)* | **Latent World Simulation**<br>Environment state tracking and entity transition dynamics | **AgentWorld**<br>`~35% (Peak)` | `KAT` 🟩🟩🟩<br>`ORN` 🟦🟦🟦<br>`AGW` 🟨🟨🟨🟨 |
108
+ | **Layers 28 – 39**<br>*(Deep)* | **Action Policy & Tool Execution**<br>Function schemas, JSON formatting, and multi-turn planning | **Ornith**<br>`~47%` | `KAT` 🟩🟩🟩<br>`ORN` 🟦🟦🟦🟦🟦<br>`AGW` 🟨🟨 |
109
+
110
+ ---
111
+
112
+ ## ⚙️ Architectural Specifications
113
+
114
+ | Parameter | Value |
115
+ | :--- | :--- |
116
+ | **Total Parameters** | **34.7 Billion** |
117
+ | **Active Parameters per Token** | **~3.3 Billion** |
118
+ | **Layers** | **40** |
119
+ | **Routed Experts** | **256** (Top-8 active per token) |
120
+ | **Shared Experts** | **1** (Always active) |
121
+ | **Attention Mechanism** | **Hybrid Gated DeltaNet (Linear Attention) + GQA** |
122
+ | **Attention Heads** | **16 Query Heads / 2 Key-Value Heads** |
123
+ | **Hidden Dimension ($d_{\text{model}}$)** | **2048** |
124
+ | **Intermediate Dimension ($d_{\text{ffn}}$)** | **1024 (per routed expert)** |
125
+ | **Vocabulary Size** | **248,320** |
126
+ | **Context Window** | **131,072 tokens** |
127
+
128
+ ---
129
+
130
+ ## 🚀 Fast Inference with vLLM
131
+
132
+ Thanks to its sparse MoE architecture, **TripleTrouble-V3** runs at the inference speed of a ~3.3B dense model while delivering 35B-scale reasoning.
133
+
134
+ ### Installation
135
+ ```bash
136
+ pip install vllm>=0.6.0
137
+ ```
138
+
139
+ ### Launch an OpenAI-Compatible API Server
140
+ ```bash
141
+ vllm serve OliviaRossi/TripleTrouble-V3 \
142
+ --tensor-parallel-size 2 \
143
+ --max-model-len 32768 \
144
+ --gpu-memory-utilization 0.90 \
145
+ --trust-remote-code
146
+ ```
147
+ *(For a single 80GB GPU, run using FP8 or AWQ quantization: `--quantization fp8`)*.
148
+
149
+ ---
150
+
151
+ ## 💻 Quickstart: Transformers
152
+
153
+ ```python
154
+ import torch
155
+ from transformers import AutoModelForCausalLM, AutoTokenizer
156
+
157
+ model_id = "OliviaRossi/TripleTrouble-V3"
158
+
159
+ tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
160
+ model = AutoModelForCausalLM.from_pretrained(
161
+ model_id,
162
+ torch_dtype=torch.bfloat16,
163
+ device_map="auto",
164
+ trust_remote_code=True
165
+ )
166
+
167
+ messages = [
168
+ {
169
+ "role": "system",
170
+ "content": (
171
+ "You are TripleTrouble, an expert agent combining deep codebase mastery, "
172
+ "rigorous multi-turn tool planning, and environment simulation capabilities."
173
+ )
174
+ },
175
+ {
176
+ "role": "user",
177
+ "content": "Analyze the concurrency bottlenecks in an asynchronous Python producer-consumer queue and write a resilient implementation using asyncio.Queue."
178
+ }
179
+ ]
180
+
181
+ prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
182
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
183
+
184
+ outputs = model.generate(
185
+ **inputs,
186
+ max_new_tokens=2048,
187
+ temperature=0.6,
188
+ top_p=0.9,
189
+ repetition_penalty=1.05
190
+ )
191
+
192
+ response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
193
+ print(response)
194
+ ```
195
+
196
+ ---
197
+
198
+ ## 🛠️ Recommended Sampling Parameters
199
+
200
+ | Workload | Temperature | Top-P | Repetition Penalty | Notes |
201
+ | :--- | :---: | :---: | :---: | :--- |
202
+ | **Code Synthesis & Bug Fixing** | `0.2` | `0.85` | `1.02` | Maximizes syntax precision and strict AST adherence. |
203
+ | **Tool Calling & MCP Agents** | `0.3` | `0.90` | `1.00` | Ensures strict valid JSON schemas and tool-call formatting. |
204
+ | **General Problem Solving** | `0.6` | `0.92` | `1.05` | Balances creative reasoning with grounded problem deduction. |
205
+ | **World Simulation & Planning** | `0.7` | `0.95` | `1.08` | Explores diverse trajectory states and action paths. |
206
+
207
+ ---
208
+
209
+ ## ⚖️ License & Attribution
210
+
211
+ This model is released under the **Apache 2.0 License**.
212
+
213
+ ### Source Models:
214
+ * **KAT-Coder-V2.5-Dev** by [Kwaipilot](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev)
215
+ * **Ornith-1.5-35B-A3B** by [ornith-ai](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B)
216
+ * **Qwen-AgentWorld-35B-A3B** by [Qwen / Alibaba Cloud](https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B)
217
+
218
+ ```bibtex
219
+ @misc{tripletrouble2026,
220
+ author = {Olivia Rossi},
221
+ title = {TripleTrouble-V3: A Geodesic Manifold Merge of Code, Tool Agency, and World Simulation},
222
+ year = {2026},
223
+ publisher = {Hugging Face},
224
+ howpublished = {\url{https://huggingface.co/OliviaRossi/TripleTrouble-V3}}
225
+ }
226
+ ```