Upload README.md
Browse files
README.md
CHANGED
|
@@ -1,17 +1,226 @@
|
|
| 1 |
-
|
| 2 |
-
|
| 3 |
-
|
| 4 |
-
|
| 5 |
-
|
| 6 |
-
|
| 7 |
-
|
| 8 |
-
|
| 9 |
-
|
| 10 |
-
|
| 11 |
-
|
| 12 |
-
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
|
| 17 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
language:
|
| 4 |
+
- en
|
| 5 |
+
- zh
|
| 6 |
+
pipeline_tag: text-generation
|
| 7 |
+
tags:
|
| 8 |
+
- merge
|
| 9 |
+
- moe
|
| 10 |
+
- qwen
|
| 11 |
+
- qwen-35b-a3b
|
| 12 |
+
- agent
|
| 13 |
+
- coding
|
| 14 |
+
- tool-use
|
| 15 |
+
- world-model
|
| 16 |
+
- gated-deltanet
|
| 17 |
+
base_model:
|
| 18 |
+
- Kwaipilot/KAT-Coder-V2.5-Dev
|
| 19 |
+
- ornith-ai/Ornith-1.5-35B-A3B
|
| 20 |
+
- Qwen/Qwen-AgentWorld-35B-A3B
|
| 21 |
+
model_creator: OliviaRossi
|
| 22 |
+
model_name: TripleTrouble-V3
|
| 23 |
+
---
|
| 24 |
+
|
| 25 |
+
<div align="center">
|
| 26 |
+
|
| 27 |
+
# ⚡ TripleTrouble-V3 ⚡
|
| 28 |
+
### *The Triad of Code, Tool Agency, and World Simulation*
|
| 29 |
+
|
| 30 |
+
[](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
|
| 31 |
+
[](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
|
| 32 |
+
[](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
|
| 33 |
+
[](https://huggingface.co/OliviaRossi/TripleTrouble-V3)
|
| 34 |
+
[](https://www.apache.org/licenses/LICENSE-2.0)
|
| 35 |
+
|
| 36 |
+
<p align="center">
|
| 37 |
+
<b>A unified sovereign agent model merging three apex Qwen 35B-A3B checkpoints via Decoupled Normalized Geodesic Consensus (NGC), Row-Wise Router Manifold Calibration, and Functional Depth Scheduling.</b>
|
| 38 |
+
</p>
|
| 39 |
+
|
| 40 |
+
[Model Details](#-overview) • [The Triad](#-the-triad) • [Merge Engineering](#-merge-mathematics) • [Serving with vLLM](#-fast-inference-with-vllm) • [Transformers](#-quickstart-transformers)
|
| 41 |
+
|
| 42 |
+
---
|
| 43 |
+
|
| 44 |
+
</div>
|
| 45 |
+
|
| 46 |
+
## 🌌 Overview
|
| 47 |
+
|
| 48 |
+
**TripleTrouble-V3** is a 35B-class sparse Mixture-of-Experts (MoE) foundation model created by fusing three divergent, specialized post-trained models of the **Qwen 35B-A3B** architecture:
|
| 49 |
+
|
| 50 |
+
1. **The Coder (KAT-Coder-V2.5-Dev):** Autonomous codebase manipulation, SWE-bench refactoring, repository traversal, and concrete AST syntax trees.
|
| 51 |
+
2. **The Agent (Ornith-1.5-35B-A3B):** Multi-turn API calling, complex structured JSON emissions, mathematical deduction, and competitive programmatic reasoning.
|
| 52 |
+
3. **The Simulator (Qwen-AgentWorld-35B-A3B):** Large World Model (LWM) dynamics, environment state transition modeling (MCP, OS, Web, Android), and next-state observation forecasting.
|
| 53 |
+
|
| 54 |
+
By moving beyond naive weight averaging and avoiding the rank collapse of pseudo-base TIES pruning, **TripleTrouble-V3** preserves the singular value spectrum of all 256 experts and maintains the spectral radius of Qwen's hybrid **Gated DeltaNet** linear recurrence dynamics.
|
| 55 |
+
|
| 56 |
+
---
|
| 57 |
+
|
| 58 |
+
## 🧬 The Triad
|
| 59 |
+
|
| 60 |
+
```mermaid
|
| 61 |
+
flowchart TD
|
| 62 |
+
subgraph MERGED ["⚡ TripleTrouble-V3"]
|
| 63 |
+
ROOT["<b>TripleTrouble-V3</b><br/>34.7B MoE • ~3.3B Active per token"]
|
| 64 |
+
end
|
| 65 |
+
|
| 66 |
+
ROOT -->|"40% Weight"| KAT["💻 <b>KAT-Coder-V2.5-Dev</b><br/>• Code & SWE-bench Refactoring<br/>• AST & Syntax Trees<br/>• Terminal Command Logic"]
|
| 67 |
+
ROOT -->|"35% Weight"| ORN["🦅 <b>Ornith-1.5-35B-A3B</b><br/>• Multi-Turn Tool Calling<br/>• MCP Protocol Execution<br/>• Structured JSON & Math"]
|
| 68 |
+
ROOT -->|"25% Weight"| AGW["🌐 <b>Qwen-AgentWorld-35B</b><br/>�� Large World Model (LWM)<br/>• Environment Dynamics<br/>• State Observation Loops"]
|
| 69 |
+
|
| 70 |
+
classDef default fill:#1e293b,stroke:#475569,stroke-width:1px,color:#f8fafc;
|
| 71 |
+
classDef highlight fill:#4338ca,stroke:#818cf8,stroke-width:2px,color:#ffffff;
|
| 72 |
+
class ROOT highlight;
|
| 73 |
+
```
|
| 74 |
+
|
| 75 |
+
---
|
| 76 |
+
|
| 77 |
+
## 🧮 Merge Mathematics
|
| 78 |
+
|
| 79 |
+
Traditional model mergers suffer from two major failure modes on 35B hybrid MoE models:
|
| 80 |
+
* **Naive Linear Soups:** Cause high-dimensional variance collapse ($\mathbb{E}[\|\sum w_i \theta_i\|] \ll \mathbb{E}[\|\theta_i\|]$), dampening activations across 40 layers.
|
| 81 |
+
* **Pseudo-Base TIES/DARE:** Truncate 70% of delta coordinates, destroying the singular value spectrum of small ($2048 \times 1024$) MoE experts and shattering the Householder contraction operator of Gated DeltaNet attention.
|
| 82 |
+
|
| 83 |
+
**TripleTrouble-V3** was built using a continuous hyperspherical framework designed specifically for this architecture:
|
| 84 |
+
|
| 85 |
+
### 1. Decoupled Normalized Geodesic Consensus (NGC)
|
| 86 |
+
Directional consensus and magnitude scaling are strictly decoupled. Checkpoints with larger gradient norms cannot skew the consensus angle:
|
| 87 |
+
|
| 88 |
+
$$\hat{W}_i = \frac{W_i}{\|W_i\|_F}$$
|
| 89 |
+
|
| 90 |
+
$$D_{\text{consensus}} = \sum_{i=1}^3 w_i(l) \hat{W}_i, \quad \hat{D} = \frac{D_{\text{consensus}}}{\|D_{\text{consensus}}\|_F}$$
|
| 91 |
+
|
| 92 |
+
$$W_{\text{final}} = \hat{D} \times \left( \sum_{i=1}^3 w_i(l) \|W_i\|_F \right)$$
|
| 93 |
+
|
| 94 |
+
### 2. Row-Wise Router Manifold Calibration
|
| 95 |
+
In an MoE with 256 routed experts, the router gate $W_{\text{gate}} \in \mathbb{R}^{256 \times d_{\text{model}}}$ consists of 256 distinct hyperplanes $g_e$. Global matrix scaling alters individual expert activation probabilities. We calibrate every expert hyperplane row-by-row:
|
| 96 |
+
|
| 97 |
+
$$g_{e, \text{final}} = \frac{\sum_i w_i \frac{g_{e, i}}{\|g_{e, i}\|_2}}{\left\| \sum_i w_i \frac{g_{e, i}}{\|g_{e, i}\|_2} \right\|_2} \times \left( \sum_{i=1}^3 w_i \|g_{e, i}\|_2 \right)$$
|
| 98 |
+
|
| 99 |
+
This guarantees router logit sharpness, prevents entropy collapse, and keeps expert dispatch distributions intact.
|
| 100 |
+
|
| 101 |
+
### 3. Functional Depth-Aware Dynamic Scheduling
|
| 102 |
+
Layer weights dynamically modulate across the 40 layers to match functional network mechanics:
|
| 103 |
+
|
| 104 |
+
| Depth Bracket | Target Specialization | Dominant Driver | Allocation Distribution |
|
| 105 |
+
| :--- | :--- | :--- | :--- |
|
| 106 |
+
| **Layers 0 – 11**<br>*(Shallow)* | **Syntax & Token ASTs**<br>Code grammar, indentation, and subword abstractions | **KAT-Coder**<br>`~53%` | `KAT` 🟩🟩🟩🟩🟩<br>`ORN` 🟦🟦<br>`AGW` 🟨🟨 |
|
| 107 |
+
| **Layers 12 – 27**<br>*(Middle)* | **Latent World Simulation**<br>Environment state tracking and entity transition dynamics | **AgentWorld**<br>`~35% (Peak)` | `KAT` 🟩🟩🟩<br>`ORN` 🟦🟦🟦<br>`AGW` 🟨🟨🟨🟨 |
|
| 108 |
+
| **Layers 28 – 39**<br>*(Deep)* | **Action Policy & Tool Execution**<br>Function schemas, JSON formatting, and multi-turn planning | **Ornith**<br>`~47%` | `KAT` 🟩🟩🟩<br>`ORN` 🟦🟦🟦🟦🟦<br>`AGW` 🟨🟨 |
|
| 109 |
+
|
| 110 |
+
---
|
| 111 |
+
|
| 112 |
+
## ⚙️ Architectural Specifications
|
| 113 |
+
|
| 114 |
+
| Parameter | Value |
|
| 115 |
+
| :--- | :--- |
|
| 116 |
+
| **Total Parameters** | **34.7 Billion** |
|
| 117 |
+
| **Active Parameters per Token** | **~3.3 Billion** |
|
| 118 |
+
| **Layers** | **40** |
|
| 119 |
+
| **Routed Experts** | **256** (Top-8 active per token) |
|
| 120 |
+
| **Shared Experts** | **1** (Always active) |
|
| 121 |
+
| **Attention Mechanism** | **Hybrid Gated DeltaNet (Linear Attention) + GQA** |
|
| 122 |
+
| **Attention Heads** | **16 Query Heads / 2 Key-Value Heads** |
|
| 123 |
+
| **Hidden Dimension ($d_{\text{model}}$)** | **2048** |
|
| 124 |
+
| **Intermediate Dimension ($d_{\text{ffn}}$)** | **1024 (per routed expert)** |
|
| 125 |
+
| **Vocabulary Size** | **248,320** |
|
| 126 |
+
| **Context Window** | **131,072 tokens** |
|
| 127 |
+
|
| 128 |
+
---
|
| 129 |
+
|
| 130 |
+
## 🚀 Fast Inference with vLLM
|
| 131 |
+
|
| 132 |
+
Thanks to its sparse MoE architecture, **TripleTrouble-V3** runs at the inference speed of a ~3.3B dense model while delivering 35B-scale reasoning.
|
| 133 |
+
|
| 134 |
+
### Installation
|
| 135 |
+
```bash
|
| 136 |
+
pip install vllm>=0.6.0
|
| 137 |
+
```
|
| 138 |
+
|
| 139 |
+
### Launch an OpenAI-Compatible API Server
|
| 140 |
+
```bash
|
| 141 |
+
vllm serve OliviaRossi/TripleTrouble-V3 \
|
| 142 |
+
--tensor-parallel-size 2 \
|
| 143 |
+
--max-model-len 32768 \
|
| 144 |
+
--gpu-memory-utilization 0.90 \
|
| 145 |
+
--trust-remote-code
|
| 146 |
+
```
|
| 147 |
+
*(For a single 80GB GPU, run using FP8 or AWQ quantization: `--quantization fp8`)*.
|
| 148 |
+
|
| 149 |
+
---
|
| 150 |
+
|
| 151 |
+
## 💻 Quickstart: Transformers
|
| 152 |
+
|
| 153 |
+
```python
|
| 154 |
+
import torch
|
| 155 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 156 |
+
|
| 157 |
+
model_id = "OliviaRossi/TripleTrouble-V3"
|
| 158 |
+
|
| 159 |
+
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
|
| 160 |
+
model = AutoModelForCausalLM.from_pretrained(
|
| 161 |
+
model_id,
|
| 162 |
+
torch_dtype=torch.bfloat16,
|
| 163 |
+
device_map="auto",
|
| 164 |
+
trust_remote_code=True
|
| 165 |
+
)
|
| 166 |
+
|
| 167 |
+
messages = [
|
| 168 |
+
{
|
| 169 |
+
"role": "system",
|
| 170 |
+
"content": (
|
| 171 |
+
"You are TripleTrouble, an expert agent combining deep codebase mastery, "
|
| 172 |
+
"rigorous multi-turn tool planning, and environment simulation capabilities."
|
| 173 |
+
)
|
| 174 |
+
},
|
| 175 |
+
{
|
| 176 |
+
"role": "user",
|
| 177 |
+
"content": "Analyze the concurrency bottlenecks in an asynchronous Python producer-consumer queue and write a resilient implementation using asyncio.Queue."
|
| 178 |
+
}
|
| 179 |
+
]
|
| 180 |
+
|
| 181 |
+
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
| 182 |
+
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
| 183 |
+
|
| 184 |
+
outputs = model.generate(
|
| 185 |
+
**inputs,
|
| 186 |
+
max_new_tokens=2048,
|
| 187 |
+
temperature=0.6,
|
| 188 |
+
top_p=0.9,
|
| 189 |
+
repetition_penalty=1.05
|
| 190 |
+
)
|
| 191 |
+
|
| 192 |
+
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
|
| 193 |
+
print(response)
|
| 194 |
+
```
|
| 195 |
+
|
| 196 |
+
---
|
| 197 |
+
|
| 198 |
+
## 🛠️ Recommended Sampling Parameters
|
| 199 |
+
|
| 200 |
+
| Workload | Temperature | Top-P | Repetition Penalty | Notes |
|
| 201 |
+
| :--- | :---: | :---: | :---: | :--- |
|
| 202 |
+
| **Code Synthesis & Bug Fixing** | `0.2` | `0.85` | `1.02` | Maximizes syntax precision and strict AST adherence. |
|
| 203 |
+
| **Tool Calling & MCP Agents** | `0.3` | `0.90` | `1.00` | Ensures strict valid JSON schemas and tool-call formatting. |
|
| 204 |
+
| **General Problem Solving** | `0.6` | `0.92` | `1.05` | Balances creative reasoning with grounded problem deduction. |
|
| 205 |
+
| **World Simulation & Planning** | `0.7` | `0.95` | `1.08` | Explores diverse trajectory states and action paths. |
|
| 206 |
+
|
| 207 |
+
---
|
| 208 |
+
|
| 209 |
+
## ⚖️ License & Attribution
|
| 210 |
+
|
| 211 |
+
This model is released under the **Apache 2.0 License**.
|
| 212 |
+
|
| 213 |
+
### Source Models:
|
| 214 |
+
* **KAT-Coder-V2.5-Dev** by [Kwaipilot](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev)
|
| 215 |
+
* **Ornith-1.5-35B-A3B** by [ornith-ai](https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B)
|
| 216 |
+
* **Qwen-AgentWorld-35B-A3B** by [Qwen / Alibaba Cloud](https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B)
|
| 217 |
+
|
| 218 |
+
```bibtex
|
| 219 |
+
@misc{tripletrouble2026,
|
| 220 |
+
author = {Olivia Rossi},
|
| 221 |
+
title = {TripleTrouble-V3: A Geodesic Manifold Merge of Code, Tool Agency, and World Simulation},
|
| 222 |
+
year = {2026},
|
| 223 |
+
publisher = {Hugging Face},
|
| 224 |
+
howpublished = {\url{https://huggingface.co/OliviaRossi/TripleTrouble-V3}}
|
| 225 |
+
}
|
| 226 |
+
```
|