OliviaRossi commited on
Commit
9146ed6
·
verified ·
1 Parent(s): fe77fb8

Upload 9 files

Browse files
.gitattributes CHANGED
@@ -34,3 +34,11 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
 
 
 
 
 
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ assets/dark/banner.png filter=lfs diff=lfs merge=lfs -text
38
+ assets/dark/chart.png filter=lfs diff=lfs merge=lfs -text
39
+ assets/dark/pipeline.png filter=lfs diff=lfs merge=lfs -text
40
+ assets/dark/vectors.png filter=lfs diff=lfs merge=lfs -text
41
+ assets/light/banner.png filter=lfs diff=lfs merge=lfs -text
42
+ assets/light/chart.png filter=lfs diff=lfs merge=lfs -text
43
+ assets/light/pipeline.png filter=lfs diff=lfs merge=lfs -text
44
+ assets/light/vectors.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -1,319 +1,281 @@
1
- ---
2
- license: apache-2.0
3
- base_model:
4
- - XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
5
- - ornith-ai/Ornith-1.5-9B
6
- tags:
7
- - merge
8
- - agsi
9
- - qwen
10
- - qwen3_5
11
- - reasoning
12
- - coding
13
- - agentic
14
- - terminal-use
15
- - swe-bench
16
- - tool-use
17
- language:
18
- - en
19
- - zh
20
- pipeline_tag: text-generation
21
- library_name: transformers
22
- ---
23
-
24
- <div align="center">
25
-
26
- # 🌌 MiMo-Ornith-9B-AGSI
27
- ### High-Order Manifold Fusion for Autonomous Coding, Reasoning & Terminal Execution
28
-
29
- [![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
30
- [![Architecture](https://img.shields.io/badge/Architecture-Qwen%209B%20Hybrid-orange.svg)](https://huggingface.co/models?other=qwen)
31
- [![Method](https://img.shields.io/badge/Merge%20Method-AGSI%20(DoRA--SLERP)-blueviolet.svg)](#-the-mathematics-of-agsi)
32
- [![Context Length](https://img.shields.io/badge/Context-131%2C072%20Tokens-green.svg)](#-deployment--inference)
33
-
34
- </div>
35
-
36
- ---
37
-
38
- ## 📌 Executive Summary
39
-
40
- **MiMo-Ornith-9B-AGSI** is a non-linear parameter-space synthesis of two leading fine-tuned models derived from the Qwen 9B architecture:
41
-
42
- * **[XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)**: A distillation checkpoint optimized for dense mathematical deduction, SWE-bench verified programmatic problem-solving, and long-horizon chain-of-thought (CoT) reasoning.
43
- * **[ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B)**: A reinforcement-learning-driven agentic model specializing in terminal/CLI mastery, autonomous bash execution, self-debugging loops, and GrandCode algorithmic generation.
44
-
45
- Rather than relying on naive linear averaging (LERP) or global spherical interpolation (SLERP)—which flatten parameter tensors into uncalibrated vectors—this model was fused using **Adaptive Geodesic Spectral Interpolation (AGSI)**. AGSI executes row-wise hyperspherical geodesics on decoupled directional manifolds, enforces second-order spectral energy conservation, shields against anti-phase gradient interference, and dynamically modulates parameters across depth via a $C^2$-continuous quintic smoothstep curve.
46
-
47
- ---
48
-
49
- ## 🔬 The Mathematics of AGSI
50
-
51
- Standard model interpolation techniques often suffer from:
52
- 1. **Frobenius Attenuation**: Convex parameter averaging causes systematic shrinkage of matrix norms ($\|(1-t)W_A + tW_B\|_F < \|W\|_F$), resulting in signal degradation across 32 transformer layers.
53
- 2. **Isotropic Collapsing**: Flattening multi-head projections into a single 1D vector treats distinct semantic subspaces as an isotropic sphere, corrupting specialized attention head alignments.
54
- 3. **Anti-Phase Annihilation**: Conflicting updates between reinforcement learning (Ornith) and knowledge distillation (MiMo) cause destructive cancellation when vectors point in opposing directions ($\cos\theta < 0$).
55
-
56
- AGSI resolves these issues through a five-stage manifold interpolation framework:
57
-
58
- ```
59
- ┌─────────────────────────────────────────┐
60
- │W_A, W_B (Weight Matrices in R^(DxK)) │
61
- └────────────────────┬────────────────────┘
62
- │
63
- [ 1. DoRA-Style Decoupling ]
64
- │
65
- ┌────────────────────┴────────────────────────┐
66
- ▼ ▼
67
- m_A, m_B = ||W||_row U_A, U_B = W / ||W||_row
68
- (Radial Magnitudes) (Unit Hypersphere S^(D-1))
69
- │ │
70
- │ [ 2. Row-wise Geodesic SLERP ]
71
- │ │
72
- │ θ_i = arccos(clamp(u_A · u_B))
73
- │ │
74
- │ [ 3. Anti-Phase Dominance Gate ]
75
- │ │
76
- │ If cos(θ) < -0.05:
77
- │ Route to dominant feature
78
- │ │
79
- │ ▼
80
- │ U_fused ∈ S^(D-1)
81
- │ │
82
- └──────────────────────┬──────────────────────┘
83
- │
84
- [ 4. Quadratic Spectral Energy Invariant ]
85
- m_target = sqrt((1-t)m_A^2 + t*m_B^2)
86
- │
87
- ▼
88
- W_cand = m_target ⊙ U_fused
89
- │
90
- [ 5. Global Frobenius Norm Conservation ]
91
- W_final = W_cand * (F_target / F_cand)
92
- ```
93
-
94
- ---
95
-
96
- ### 1. Direction-Magnitude (DoRA) Decoupling
97
-
98
- Linear layers compute transformations $y = x W^T$, where individual rows $W_i \in \mathbb{R}^{D_{\text{in}}}$ represent the hyperplanes of specific neurons. AGSI isolates radial feature scale from directional orientation:
99
-
100
- $$m^{(i)} = \|W^{(i)}\|_2 = \sqrt{\sum_{j=1}^{D_{\text{in}}} (W_{ij})^2}$$
101
-
102
- $$u^{(i)} = \frac{W^{(i)}}{m^{(i)} + \epsilon}, \quad \text{where } u^{(i)} \in S^{D_{\text{in}}-1}$$
103
-
104
- By constraining directional updates to the unit hypersphere $S^{D_{\text{in}}-1}$, the model preserves the angular separation of neuron receptive fields.
105
-
106
- ---
107
-
108
- ### 2. Row-Wise Hyperspherical Geodesics ($S^{D-1}$)
109
-
110
- For each individual neuron row $i$, the angular geodesic distance between checkpoint trajectories is calculated directly in its tangent space:
111
-
112
- $$\theta_i = \arccos\left(\text{clamp}\left(\langle u_A^{(i)}, u_B^{(i)} \rangle, -1 + \delta, 1 - \delta\right)\right)$$
113
-
114
- Rather than using Euclidean displacement, the directional basis moves along the great-circle arc:
115
-
116
- $$u_{\text{fused}}^{(i)} = c_A^{(i)} u_A^{(i)} + c_B^{(i)} u_B^{(i)}$$
117
-
118
- where the geodesic velocity coefficients are defined as:
119
-
120
- $$c_A^{(i)} = \frac{\sin((1 - t)\theta_i)}{\sin\theta_i + \epsilon}, \quad c_B^{(i)} = \frac{\sin(t \theta_i)}{\sin\theta_i + \epsilon}$$
121
-
122
- If $\sin\theta_i \to 0$ (quasi-collinear vectors), the interpolation smoothly transitions to normalized linear blending:
123
-
124
- $$\lim_{\theta \to 0} u_{\text{fused}}^{(i)} = (1 - t)u_A^{(i)} + t u_B^{(i)}$$
125
-
126
- ---
127
-
128
- ### 3. Anti-Phase Interference Shielding
129
-
130
- When RL optimization (Ornith) and distillation gradients (MiMo) pull in opposing directions ($\theta_i > 90^\circ, \cos\theta_i < 0$), standard spherical combination results in substantial cancellation:
131
-
132
- $$\|u_A + u_B\| = \sqrt{2 + 2\cos\theta} < \sqrt{2} \approx 1.414 \quad (\text{vs. } 2.0 \text{ when aligned})$$
133
-
134
- AGSI monitors the directional dot product against a critical interference bound ($\tau = -0.05$, corresponding to $\theta \approx 92.86^\circ$). If destructive cancellation occurs:
135
-
136
- $$\text{Conflict Condition: } \langle u_A^{(i)}, u_B^{(i)} \rangle < \tau$$
137
-
138
- $$\gamma_A^{(i)} = \frac{m_A^{(i)}}{m_A^{(i)} + m_B^{(i)} + \epsilon}$$
139
-
140
- $$u_{\text{fused}}^{(i)} = \begin{cases}
141
- 0.85\, u_A^{(i)} + 0.15\, u_B^{(i)}, & \text{if } \gamma_A^{(i)} \ge 0.5 \\
142
- 0.15\, u_A^{(i)} + 0.85\, u_B^{(i)}, & \text{if } \gamma_A^{(i)} < 0.5
143
- \end{cases}$$
144
-
145
- This dominance gate projects conflicting features toward the checkpoint exhibiting higher parameter variance, preventing the formation of dormant neurons.
146
-
147
- ---
148
-
149
- ### 4. Quadratic Spectral Energy Invariant
150
-
151
- In deep autoregressive networks normalized by RMSNorm, layer output variance relies heavily on parameter norm preservation:
152
-
153
- $$\mathbb{E}[\|W x\|_2^2] \approx \frac{1}{D_{\text{in}}} \|W\|_F^2 \, \mathbb{E}[\|x\|_2^2]$$
154
-
155
- To maintain stable forward-pass activations without logit saturation or signal decay, AGSI scales the interpolated direction vector by the root-mean-square energy:
156
-
157
- $$m_{\text{target}}^{(i)} = \sqrt{(1 - t)(m_A^{(i)})^2 + t(m_B^{(i)})^2}$$
158
-
159
- $$W_{\text{candidate}}^{(i)} = m_{\text{target}}^{(i)} \cdot \frac{u_{\text{fused}}^{(i)}}{\|u_{\text{fused}}^{(i)}\|_2}$$
160
-
161
- Finally, a global Frobenius norm correction is applied across the full matrix:
162
-
163
- $$W_{\text{final}} = W_{\text{candidate}} \times \left( \frac{\sqrt{(1 - t)\|W_A\|_F^2 + t\|W_B\|_F^2}}{\|W_{\text{candidate}}\|_F + \epsilon} \right)$$
164
-
165
- ---
166
-
167
- ### 5. Quintic Smoothstep Depth & Functional Block Routing
168
-
169
- The mixing coefficient $t$ is non-static. It is governed by a $C^2$-continuous quintic polynomial function across the 32 transformer layers, augmented by functional block biases:
170
-
171
- $$\xi = \frac{l}{L - 1} \in [0, 1], \quad l \in \{0, 1, \dots, 31\}$$
172
-
173
- $$S(\xi) = 6\xi^5 - 15\xi^4 + 10\xi^3$$
174
-
175
- $$t(l, \text{block}) = \text{clamp}\left(t_{\text{base}} + \Delta t \cdot \left(S(\xi) - 0.5\right) + \beta_{\text{block}}, \, 0.0, \, 1.0\right)$$
176
-
177
- ```
178
- Mixing Ratio (t) -> Weight of Ornith-1.5
179
- 1.00 ┤
180
- 0.75 ┤
181
- 0.55 ┤ ╭────────────── (Ornith Terminal Routing)
182
- 0.48 ┤ ╭──────────────────╯ (Base Balance)
183
- 0.41 ┤ ────────────────╯ (MiMo Reasoning Anchor)
184
- 0.00 ┼───┬──────────────┬──────────────────┬──────────────┬
185
- L0 L7 L16 L24 L31 (Depth)
186
- ```
187
-
188
- #### Layer Profile Breakdown:
189
- * **Layers 0–7 ($t \approx 0.41 - 0.44$) — Syntactic Anchoring**: Biased toward MiMo-V2.6 to preserve input parsing, stable token embeddings, and baseline representations.
190
- * **Layers 8–23 ($t \approx 0.45 - 0.51$) — Cognitive & Algorithmic Core**: Near-equilibrium blending. Mathematical deductions and algorithm planning synthesize with Ornith's tool-use reasoning.
191
- * **Layers 24–31 ($t \approx 0.52 - 0.55$) — Action Policy & Output Heads**: Biased toward Ornith-1.5 to prioritize terminal command generation, tool-calling syntax, and execution policies.
192
- * **Attention vs. MLP Decoupling**:
193
- * **Attention Projections (`q, k, v, o, linear_attn`)**: $\beta_{\text{attn}} = +0.04$ (Ornith priority for context selection and environment state tracking).
194
- * **MLP / Feed-Forward (`gate, up, down_proj`)**: $\beta_{\text{mlp}} = -0.04$ (MiMo priority for factual retention and associative coding memories).
195
-
196
- ---
197
-
198
- ## ⚙️ AGSI Configuration Profile
199
-
200
- ```python
201
- # The Golden SOTA Preset applied during parameter synthesis
202
- BASE_ORNITH_RATIO = 0.48 # Foundational balance (52% MiMo / 48% Ornith)
203
- DEPTH_MODULATION = 0.14 # Dynamic span across transformer depth
204
- ATTN_ROUTING_BIAS = 0.04 # Attention block offset (favors Ornith)
205
- MLP_ROUTING_BIAS = -0.04 # MLP block offset (favors MiMo)
206
- ANTI_PHASE_THRESHOLD = -0.05 # Critical angle cutoff (θ = 92.86°)
207
- SPECTRAL_ENERGY_EXPONENT = 2.0 # Second-order L2 moment conservation
208
- ```
209
-
210
- ---
211
-
212
- ## 🚀 Deployment & Inference
213
-
214
- This model is compatible with systems supporting the Qwen 9B architecture (e.g., standard Hugging Face `transformers`, `vLLM`, `SGLang`).
215
-
216
- ### Serving with vLLM (Recommended)
217
-
218
- To run the model with vLLM, including support for streaming reasoning tokens (`<think>`) and agentic tool calls:
219
-
220
- ```bash
221
- vllm serve DEST_REPO_ID \
222
- --port 8000 \
223
- --model DEST_REPO_ID \
224
- --trust-remote-code \
225
- --tensor-parallel-size 1 \
226
- --max-model-len 131072 \
227
- --gpu-memory-utilization 0.92 \
228
- --reasoning-parser qwen3 \
229
- --tool-call-parser hermes
230
- ```
231
-
232
- ### Python Inference via Transformers
233
-
234
- ```python
235
- import torch
236
- from transformers import AutoModelForCausalLM, AutoTokenizer
237
-
238
- model_id = "DEST_REPO_ID"
239
-
240
- tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
241
- model = AutoModelForCausalLM.from_pretrained(
242
- model_id,
243
- torch_dtype=torch.bfloat16,
244
- device_map="auto",
245
- trust_remote_code=True
246
- )
247
-
248
- messages = [
249
- {
250
- "role": "system",
251
- "content": "You are a master systems engineer and competitive programmer. Solve problems using precise step-by-step reasoning enclosed in <think> tags, then execute terminal commands or output production code."
252
- },
253
- {
254
- "role": "user",
255
- "content": "Write an optimized eBPF program in C that traces network socket latency outliers (>100ms) and provides a Python BCC analysis script."
256
- }
257
- ]
258
-
259
- prompt = tokenizer.apply_chat_template(
260
- messages,
261
- tokenize=False,
262
- add_generation_prompt=True
263
- )
264
-
265
- inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
266
- outputs = model.generate(
267
- **inputs,
268
- max_new_tokens=4096,
269
- temperature=0.6,
270
- top_p=0.95,
271
- repetition_penalty=1.05
272
- )
273
-
274
- response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
275
- print(response)
276
- ```
277
-
278
- ---
279
-
280
- ## 🎯 Prompt Formatting & Chat Template
281
-
282
- This model uses the unified Qwen chat template, configured to separate internal reasoning steps from final outputs:
283
-
284
- ```xml
285
- <|im_start|>system
286
- You are a helpful assistant.<|im_end|>
287
- <|im_start|>user
288
- Write a bash script to monitor memory usage.<|im_end|>
289
- <|im_start|>assistant
290
- <think>
291
- 1. Identify target metrics: available vs. used memory.
292
- 2. Use /proc/meminfo or 'free -m' for portability.
293
- 3. Handle logging and alert thresholds.
294
- </think>
295
- Here is the monitoring script:
296
- ```
297
-
298
- ```bash
299
- #!/usr/bin/env bash
300
- set -euo pipefail
301
-
302
- THRESHOLD=85
303
- CURRENT=$(free | awk '/Mem:/ {printf("%.0f"), $3/$2 * 100}')
304
-
305
- if [ "$CURRENT" -gt "$THRESHOLD" ]; then
306
- echo "WARNING: Memory usage at ${CURRENT}%"
307
- fi
308
- ```
309
- `<|im_end|>`
310
-
311
- ---
312
-
313
- ## ⚖️ License & Attribution
314
-
315
- * **Base Model Checkpoints**:
316
- * Checkpoint A: [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)
317
- * Checkpoint B: [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B)
318
- * **Underlying Architecture**: Qwen Series (Alibaba Cloud)
319
- * **License**: Apache 2.0
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model:
4
+ - XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
5
+ - ornith-ai/Ornith-1.5-9B
6
+ tags:
7
+ - merge
8
+ - agsi
9
+ - qwen
10
+ - qwen3_5
11
+ - reasoning
12
+ - coding
13
+ - agentic
14
+ - terminal-use
15
+ - swe-bench
16
+ - tool-use
17
+ language:
18
+ - en
19
+ - zh
20
+ pipeline_tag: text-generation
21
+ library_name: transformers
22
+ ---
23
+
24
+ ![MiMo-Ornith-9B-AGSI](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/dark/banner.png#hf-dark-mode-only)
25
+ ![MiMo-Ornith-9B-AGSI](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/light/banner.png#hf-light-mode-only)
26
+
27
+ <div align="center">
28
+
29
+ [![License](https://img.shields.io/badge/License-Apache%202.0-blue.svg)](https://opensource.org/licenses/Apache-2.0)
30
+ [![Architecture](https://img.shields.io/badge/Architecture-Qwen%209B%20Hybrid-orange.svg)](https://huggingface.co/models?other=qwen)
31
+ [![Method](https://img.shields.io/badge/Merge%20Method-AGSI%20(DoRA--SLERP)-blueviolet.svg)](#-the-mathematics-of-agsi)
32
+ [![Context Length](https://img.shields.io/badge/Context-131%2C072%20Tokens-green.svg)](#-deployment--inference)
33
+
34
+ </div>
35
+
36
+ ---
37
+
38
+ ## 📌 Executive Summary
39
+
40
+ **MiMo-Ornith-9B-AGSI** is a non-linear parameter-space synthesis of two leading fine-tuned models derived from the Qwen 9B architecture:
41
+
42
+ * **[XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)**: A distillation checkpoint optimized for dense mathematical deduction, SWE-bench verified programmatic problem-solving, and long-horizon chain-of-thought (CoT) reasoning.
43
+ * **[ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B)**: A reinforcement-learning-driven agentic model specializing in terminal/CLI mastery, autonomous bash execution, self-debugging loops, and GrandCode algorithmic generation.
44
+
45
+ Rather than relying on naive linear averaging (LERP) or global spherical interpolation (SLERP)—which flatten parameter tensors into uncalibrated vectors—this model was fused using **Adaptive Geodesic Spectral Interpolation (AGSI)**. AGSI executes row-wise hyperspherical geodesics on decoupled directional manifolds, enforces second-order spectral energy conservation, shields against anti-phase gradient interference, and dynamically modulates parameters across depth via a $C^2$-continuous quintic smoothstep curve.
46
+
47
+ ---
48
+
49
+ ## 🔬 The Mathematics of AGSI
50
+
51
+ Standard model interpolation techniques often suffer from:
52
+ 1. **Frobenius Attenuation**: Convex parameter averaging causes systematic shrinkage of matrix norms ($\|(1-t)W_A + tW_B\|_F < \|W\|_F$), resulting in signal degradation across 32 transformer layers.
53
+ 2. **Isotropic Collapsing**: Flattening multi-head projections into a single 1D vector treats distinct semantic subspaces as an isotropic sphere, corrupting specialized attention head alignments.
54
+ 3. **Anti-Phase Annihilation**: Conflicting updates between reinforcement learning (Ornith) and knowledge distillation (MiMo) cause destructive cancellation when vectors point in opposing directions ($\cos\theta < 0$).
55
+
56
+ AGSI resolves these issues through a five-stage manifold interpolation framework:
57
+
58
+ ![AGSI pipeline](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/dark/pipeline.png#hf-dark-mode-only)
59
+ ![AGSI pipeline](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/light/pipeline.png#hf-light-mode-only)
60
+
61
+ ---
62
+
63
+ ### 1. Direction-Magnitude (DoRA) Decoupling
64
+
65
+ Linear layers compute transformations $y = x W^T$, where individual rows $W_i \in \mathbb{R}^{D_{\text{in}}}$ represent the hyperplanes of specific neurons. AGSI isolates radial feature scale from directional orientation:
66
+
67
+ $$m^{(i)} = \|W^{(i)}\|_2 = \sqrt{\sum_{j=1}^{D_{\text{in}}} (W_{ij})^2}$$
68
+
69
+ $$u^{(i)} = \frac{W^{(i)}}{m^{(i)} + \epsilon}, \quad \text{where } u^{(i)} \in S^{D_{\text{in}}-1}$$
70
+
71
+ By constraining directional updates to the unit hypersphere $S^{D_{\text{in}}-1}$, the model preserves the angular separation of neuron receptive fields.
72
+
73
+ ---
74
+
75
+ ### 2. Row-Wise Hyperspherical Geodesics ($S^{D-1}$)
76
+
77
+ For each individual neuron row $i$, the angular geodesic distance between checkpoint trajectories is calculated directly in its tangent space:
78
+
79
+ $$\theta_i = \arccos\left(\text{clamp}\left(\langle u_A^{(i)}, u_B^{(i)} \rangle, -1 + \delta, 1 - \delta\right)\right)$$
80
+
81
+ Rather than using Euclidean displacement, the directional basis moves along the great-circle arc:
82
+
83
+ $$u_{\text{fused}}^{(i)} = c_A^{(i)} u_A^{(i)} + c_B^{(i)} u_B^{(i)}$$
84
+
85
+ where the geodesic velocity coefficients are defined as:
86
+
87
+ $$c_A^{(i)} = \frac{\sin((1 - t)\theta_i)}{\sin\theta_i + \epsilon}, \quad c_B^{(i)} = \frac{\sin(t \theta_i)}{\sin\theta_i + \epsilon}$$
88
+
89
+ If $\sin\theta_i \to 0$ (quasi-collinear vectors), the interpolation smoothly transitions to normalized linear blending:
90
+
91
+ $$\lim_{\theta \to 0} u_{\text{fused}}^{(i)} = (1 - t)u_A^{(i)} + t u_B^{(i)}$$
92
+
93
+ ---
94
+
95
+ ### 3. Anti-Phase Interference Shielding
96
+
97
+ When RL optimization (Ornith) and distillation gradients (MiMo) pull in opposing directions ($\theta_i > 90^\circ, \cos\theta_i < 0$), standard spherical combination results in substantial cancellation:
98
+
99
+ $$\|u_A + u_B\| = \sqrt{2 + 2\cos\theta} < \sqrt{2} \approx 1.414 \quad (\text{vs. } 2.0 \text{ when aligned})$$
100
+
101
+ AGSI monitors the directional dot product against a critical interference bound ($\tau = -0.05$, corresponding to $\theta \approx 92.86^\circ$). If destructive cancellation occurs:
102
+
103
+ $$\text{Conflict Condition: } \langle u_A^{(i)}, u_B^{(i)} \rangle < \tau$$
104
+
105
+ $$\gamma_A^{(i)} = \frac{m_A^{(i)}}{m_A^{(i)} + m_B^{(i)} + \epsilon}$$
106
+
107
+ $$u_{\text{fused}}^{(i)} = \begin{cases}
108
+ 0.85\, u_A^{(i)} + 0.15\, u_B^{(i)}, & \text{if } \gamma_A^{(i)} \ge 0.5 \\
109
+ 0.15\, u_A^{(i)} + 0.85\, u_B^{(i)}, & \text{if } \gamma_A^{(i)} < 0.5
110
+ \end{cases}$$
111
+
112
+ This dominance gate projects conflicting features toward the checkpoint exhibiting higher parameter variance, preventing the formation of dormant neurons.
113
+
114
+ ![Geodesic SLERP vs anti-phase dominance gate](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/dark/vectors.png#hf-dark-mode-only)
115
+ ![Geodesic SLERP vs anti-phase dominance gate](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/light/vectors.png#hf-light-mode-only)
116
+
117
+ ---
118
+
119
+ ### 4. Quadratic Spectral Energy Invariant
120
+
121
+ In deep autoregressive networks normalized by RMSNorm, layer output variance relies heavily on parameter norm preservation:
122
+
123
+ $$\mathbb{E}[\|W x\|_2^2] \approx \frac{1}{D_{\text{in}}} \|W\|_F^2 \, \mathbb{E}[\|x\|_2^2]$$
124
+
125
+ To maintain stable forward-pass activations without logit saturation or signal decay, AGSI scales the interpolated direction vector by the root-mean-square energy:
126
+
127
+ $$m_{\text{target}}^{(i)} = \sqrt{(1 - t)(m_A^{(i)})^2 + t(m_B^{(i)})^2}$$
128
+
129
+ $$W_{\text{candidate}}^{(i)} = m_{\text{target}}^{(i)} \cdot \frac{u_{\text{fused}}^{(i)}}{\|u_{\text{fused}}^{(i)}\|_2}$$
130
+
131
+ Finally, a global Frobenius norm correction is applied across the full matrix:
132
+
133
+ $$W_{\text{final}} = W_{\text{candidate}} \times \left( \frac{\sqrt{(1 - t)\|W_A\|_F^2 + t\|W_B\|_F^2}}{\|W_{\text{candidate}}\|_F + \epsilon} \right)$$
134
+
135
+ ---
136
+
137
+ ### 5. Quintic Smoothstep Depth & Functional Block Routing
138
+
139
+ The mixing coefficient $t$ is non-static. It is governed by a $C^2$-continuous quintic polynomial function across the 32 transformer layers, augmented by functional block biases:
140
+
141
+ $$\xi = \frac{l}{L - 1} \in [0, 1], \quad l \in \{0, 1, \dots, 31\}$$
142
+
143
+ $$S(\xi) = 6\xi^5 - 15\xi^4 + 10\xi^3$$
144
+
145
+ $$t(l, \text{block}) = \text{clamp}\left(t_{\text{base}} + \Delta t \cdot \left(S(\xi) - 0.5\right) + \beta_{\text{block}}, \, 0.0, \, 1.0\right)$$
146
+
147
+ ![Mixing ratio across transformer depth](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/dark/chart.png#hf-dark-mode-only)
148
+ ![Mixing ratio across transformer depth](https://huggingface.co/OliviaRossi/MiMo-Ornith-9B-AGSI/resolve/main/assets/light/chart.png#hf-light-mode-only)
149
+
150
+ #### Layer Profile Breakdown:
151
+ * **Layers 0–7 ($t \approx 0.41 - 0.44$) — Syntactic Anchoring**: Biased toward MiMo-V2.6 to preserve input parsing, stable token embeddings, and baseline representations.
152
+ * **Layers 8–23 ($t \approx 0.45 - 0.51$) — Cognitive & Algorithmic Core**: Near-equilibrium blending. Mathematical deductions and algorithm planning synthesize with Ornith's tool-use reasoning.
153
+ * **Layers 24–31 ($t \approx 0.52 - 0.55$) — Action Policy & Output Heads**: Biased toward Ornith-1.5 to prioritize terminal command generation, tool-calling syntax, and execution policies.
154
+ * **Attention vs. MLP Decoupling**:
155
+ * **Attention Projections (`q, k, v, o, linear_attn`)**: $\beta_{\text{attn}} = +0.04$ (Ornith priority for context selection and environment state tracking).
156
+ * **MLP / Feed-Forward (`gate, up, down_proj`)**: $\beta_{\text{mlp}} = -0.04$ (MiMo priority for factual retention and associative coding memories).
157
+
158
+ ---
159
+
160
+ ## ⚙️ AGSI Configuration Profile
161
+
162
+ ```python
163
+ # The Golden SOTA Preset applied during parameter synthesis
164
+ BASE_ORNITH_RATIO = 0.48 # Foundational balance (52% MiMo / 48% Ornith)
165
+ DEPTH_MODULATION = 0.14 # Dynamic span across transformer depth
166
+ ATTN_ROUTING_BIAS = 0.04 # Attention block offset (favors Ornith)
167
+ MLP_ROUTING_BIAS = -0.04 # MLP block offset (favors MiMo)
168
+ ANTI_PHASE_THRESHOLD = -0.05 # Critical angle cutoff (θ = 92.86°)
169
+ SPECTRAL_ENERGY_EXPONENT = 2.0 # Second-order L2 moment conservation
170
+ ```
171
+
172
+ ---
173
+
174
+ ## 🚀 Deployment & Inference
175
+
176
+ This model is compatible with systems supporting the Qwen 9B architecture (e.g., standard Hugging Face `transformers`, `vLLM`, `SGLang`).
177
+
178
+ ### Serving with vLLM (Recommended)
179
+
180
+ To run the model with vLLM, including support for streaming reasoning tokens (`<think>`) and agentic tool calls:
181
+
182
+ ```bash
183
+ vllm serve DEST_REPO_ID \
184
+ --port 8000 \
185
+ --model DEST_REPO_ID \
186
+ --trust-remote-code \
187
+ --tensor-parallel-size 1 \
188
+ --max-model-len 131072 \
189
+ --gpu-memory-utilization 0.92 \
190
+ --reasoning-parser qwen3 \
191
+ --tool-call-parser hermes
192
+ ```
193
+
194
+ ### Python Inference via Transformers
195
+
196
+ ```python
197
+ import torch
198
+ from transformers import AutoModelForCausalLM, AutoTokenizer
199
+
200
+ model_id = "DEST_REPO_ID"
201
+
202
+ tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
203
+ model = AutoModelForCausalLM.from_pretrained(
204
+ model_id,
205
+ torch_dtype=torch.bfloat16,
206
+ device_map="auto",
207
+ trust_remote_code=True
208
+ )
209
+
210
+ messages = [
211
+ {
212
+ "role": "system",
213
+ "content": "You are a master systems engineer and competitive programmer. Solve problems using precise step-by-step reasoning enclosed in <think> tags, then execute terminal commands or output production code."
214
+ },
215
+ {
216
+ "role": "user",
217
+ "content": "Write an optimized eBPF program in C that traces network socket latency outliers (>100ms) and provides a Python BCC analysis script."
218
+ }
219
+ ]
220
+
221
+ prompt = tokenizer.apply_chat_template(
222
+ messages,
223
+ tokenize=False,
224
+ add_generation_prompt=True
225
+ )
226
+
227
+ inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
228
+ outputs = model.generate(
229
+ **inputs,
230
+ max_new_tokens=4096,
231
+ temperature=0.6,
232
+ top_p=0.95,
233
+ repetition_penalty=1.05
234
+ )
235
+
236
+ response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
237
+ print(response)
238
+ ```
239
+
240
+ ---
241
+
242
+ ## 🎯 Prompt Formatting & Chat Template
243
+
244
+ This model uses the unified Qwen chat template, configured to separate internal reasoning steps from final outputs:
245
+
246
+ ```xml
247
+ <|im_start|>system
248
+ You are a helpful assistant.<|im_end|>
249
+ <|im_start|>user
250
+ Write a bash script to monitor memory usage.<|im_end|>
251
+ <|im_start|>assistant
252
+ <think>
253
+ 1. Identify target metrics: available vs. used memory.
254
+ 2. Use /proc/meminfo or 'free -m' for portability.
255
+ 3. Handle logging and alert thresholds.
256
+ </think>
257
+ Here is the monitoring script:
258
+ ```
259
+
260
+ ```bash
261
+ #!/usr/bin/env bash
262
+ set -euo pipefail
263
+
264
+ THRESHOLD=85
265
+ CURRENT=$(free | awk '/Mem:/ {printf("%.0f"), $3/$2 * 100}')
266
+
267
+ if [ "$CURRENT" -gt "$THRESHOLD" ]; then
268
+ echo "WARNING: Memory usage at ${CURRENT}%"
269
+ fi
270
+ ```
271
+ `<|im_end|>`
272
+
273
+ ---
274
+
275
+ ## ⚖️ License & Attribution
276
+
277
+ * **Base Model Checkpoints**:
278
+ * Checkpoint A: [XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B)
279
+ * Checkpoint B: [ornith-ai/Ornith-1.5-9B](https://huggingface.co/ornith-ai/Ornith-1.5-9B)
280
+ * **Underlying Architecture**: Qwen Series (Alibaba Cloud)
281
+ * **License**: Apache 2.0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
assets/dark/banner.png ADDED

Git LFS Details

  • SHA256: ddf16b14cdc76025da2bd85c21e014ac81ff098e4d024c9be6a5f88e70ea3ced
  • Pointer size: 131 Bytes
  • Size of remote file: 318 kB
assets/dark/chart.png ADDED

Git LFS Details

  • SHA256: 104b34f69980bc6c2f241818807f221783bd55d2d53f808ef5d6c5f55b54a98f
  • Pointer size: 131 Bytes
  • Size of remote file: 172 kB
assets/dark/pipeline.png ADDED

Git LFS Details

  • SHA256: ada86134b66b38e72ebc115b05e808e408c1ee333a658f2a30b6b9301643bc65
  • Pointer size: 131 Bytes
  • Size of remote file: 388 kB
assets/dark/vectors.png ADDED

Git LFS Details

  • SHA256: 0c8ab09ad13db2cf1ec9fb898716aa6a178aeb3d59c08ac118cac7e4c7196abd
  • Pointer size: 131 Bytes
  • Size of remote file: 205 kB
assets/light/banner.png ADDED

Git LFS Details

  • SHA256: 3b6fad89d717309a823e7deb719826312e3ddcddfb43dbfeea55bd9f135c7d13
  • Pointer size: 131 Bytes
  • Size of remote file: 238 kB
assets/light/chart.png ADDED

Git LFS Details

  • SHA256: 08f3863df70273c416c9ce2e00b3f2565ae90a35d2c71da1c89f6e66a75d5a86
  • Pointer size: 131 Bytes
  • Size of remote file: 171 kB
assets/light/pipeline.png ADDED

Git LFS Details

  • SHA256: 6158a0f8701d4e684cd789b52353954110411f5d3b0105d5777b26af755f4937
  • Pointer size: 131 Bytes
  • Size of remote file: 382 kB
assets/light/vectors.png ADDED

Git LFS Details

  • SHA256: 179948fee347766fdb7c077c492ea43e071d212e967ebd3c7b57ead6f860f471
  • Pointer size: 131 Bytes
  • Size of remote file: 193 kB