pragmaticcs commited on
Commit
fdf2cc9
·
1 Parent(s): 7815aa8

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +27 -15
README.md CHANGED
@@ -87,43 +87,55 @@ A four-way MoE merge of the Qwen 35B-A3B architecture, fusing task vectors from
87
 
88
  ## Merge Methodology
89
 
90
- For each floating-point parameter, a task delta is computed per donor $k$:
91
 
92
- $$\Delta_k = D_k - W_0$$
 
 
93
 
94
- **DARE pruning.** A Bernoulli mask at retention density $p$ zeroes out low-magnitude updates; surviving values are rescaled by $p^{-1}$:
95
 
96
- $$\tilde{\Delta}_k = \frac{1}{p} \left(\Delta_k \odot M_k\right), \quad M_k \sim \text{Bernoulli}(p)$$
 
 
97
 
98
- **TIES sign election.** A consensus sign $\Gamma$ is computed via weighted vote across donors, and any donor update conflicting with it is dropped before averaging:
99
 
100
- $$\Gamma = \text{sgn}\left(\sum_{k=1}^K \alpha_k \tilde{\Delta}_k\right)$$
 
 
101
 
102
- $$\Delta_{\text{TIES}} = \frac{\sum_{k=1}^K \alpha_k \tilde{\Delta}_k \odot \mathbb{I}\left(\text{sgn}(\tilde{\Delta}_k) = \Gamma\right)}{\sum_{k=1}^K \alpha_k \cdot \mathbb{I}\left(\text{sgn}(\tilde{\Delta}_k) = \Gamma\right) + \epsilon}$$
 
 
103
 
104
  **Depth-scaled reconstruction.** The merged weight is reconstructed as:
105
 
106
- $$W_{\text{final}} = W_0 + \lambda(l) \cdot \Delta_{\text{TIES}}$$
 
 
107
 
108
- where the layer scaling factor $\lambda(l)$ across decoder layer index $l \in [0, 39]$ is defined as:
109
 
110
- $$\lambda(l) = \beta \cdot \left(0.5 + 0.5 \sin\left(\pi \frac{l}{39}\right)\right)$$
 
 
111
 
112
- This keeps input/output projections closer to the base and applies the strongest task transfer to middle layers ($l \in [12, 28]$).
113
 
114
  ---
115
 
116
  ## Layer-Stratified Policies
117
 
118
- | Parameter Group | Match Substring | Policy | Density ($p$) | Base Scale ($\beta$) |
119
  | :--- | :--- | :---: | :---: | :---: |
120
  | **Embeddings / LM head** | `embed_tokens`, `lm_head` | Linear | — | 1.00 |
121
  | **Norms / biases** | `norm`, `bias`, 1D tensors | Linear | — | 1.00 |
122
  | **DeltaNet recurrent state** | `a_log`, `dt_bias`, `conv1d` | Linear | — | 1.00 |
123
  | **MoE router gate** | `mlp.gate.weight`, `block_sparse_moe.gate` | Linear | — | 1.00 |
124
- | **MoE shared expert** | `shared_expert` | DARE-TIES | 0.70 | 0.60 |
125
- | **Attention projections** | `attn`, `rotary`, `in_proj`, `out_proj`, `x_proj` | DARE-TIES | 0.75 | 0.60 |
126
- | **Routed experts (×256)** | `experts`, `mlp` | DARE-TIES | 0.65 | 0.55 |
127
 
128
  - **Router protection:** Gate weights use linear interpolation (~57% base, ~43% donors) rather than DARE to avoid destabilizing expert routing.
129
  - **DeltaNet stability:** Recurrent state kernels are excluded from DARE to prevent divergence in the linear-attention state space.
 
87
 
88
  ## Merge Methodology
89
 
90
+ For each floating-point parameter, a task delta is computed per donor \\(k\\):
91
 
92
+ $$
93
+ \Delta_k = D_k - W_0
94
+ $$
95
 
96
+ **DARE pruning.** A Bernoulli mask at retention density \\(p\\) zeroes out low-magnitude updates; surviving values are rescaled by \\(p^{-1}\\):
97
 
98
+ $$
99
+ \tilde{\Delta}_k = \frac{1}{p} \left(\Delta_k \odot M_k\right), \quad M_k \sim \text{Bernoulli}(p)
100
+ $$
101
 
102
+ **TIES sign election.** A consensus sign \\(\Gamma\\) is computed via weighted vote across donors, and any donor update conflicting with it is dropped before averaging:
103
 
104
+ $$
105
+ \Gamma = \operatorname{sgn}\left(\sum_{k=1}^K \alpha_k \tilde{\Delta}_k\right)
106
+ $$
107
 
108
+ $$
109
+ \Delta_{\text{TIES}} = \frac{\sum_{k=1}^K \alpha_k \tilde{\Delta}_k \odot \mathbb{I}\left(\operatorname{sgn}(\tilde{\Delta}_k) = \Gamma\right)}{\sum_{k=1}^K \alpha_k \cdot \mathbb{I}\left(\operatorname{sgn}(\tilde{\Delta}_k) = \Gamma\right) + \epsilon}
110
+ $$
111
 
112
  **Depth-scaled reconstruction.** The merged weight is reconstructed as:
113
 
114
+ $$
115
+ W_{\text{final}} = W_0 + \lambda(l) \cdot \Delta_{\text{TIES}}
116
+ $$
117
 
118
+ where the layer scaling factor \\(\lambda(l)\\) across decoder layer index \\(l \in [0, 39]\\) is defined as:
119
 
120
+ $$
121
+ \lambda(l) = \beta \cdot \left(0.5 + 0.5 \sin\left(\pi \frac{l}{39}\right)\right)
122
+ $$
123
 
124
+ This keeps input/output projections closer to the base and applies the strongest task transfer to middle layers \\((l \in [12, 28])\\).
125
 
126
  ---
127
 
128
  ## Layer-Stratified Policies
129
 
130
+ | Parameter Group | Match Substring | Policy | Density (p) | Base Scale (β) |
131
  | :--- | :--- | :---: | :---: | :---: |
132
  | **Embeddings / LM head** | `embed_tokens`, `lm_head` | Linear | — | 1.00 |
133
  | **Norms / biases** | `norm`, `bias`, 1D tensors | Linear | — | 1.00 |
134
  | **DeltaNet recurrent state** | `a_log`, `dt_bias`, `conv1d` | Linear | — | 1.00 |
135
  | **MoE router gate** | `mlp.gate.weight`, `block_sparse_moe.gate` | Linear | — | 1.00 |
136
+ | **MoE shared expert** | `shared_expert` | DARE‑TIES | 0.70 | 0.60 |
137
+ | **Attention projections** | `attn`, `rotary`, `in_proj`, `out_proj`, `x_proj` | DARE‑TIES | 0.75 | 0.60 |
138
+ | **Routed experts (×256)** | `experts`, `mlp` | DARE‑TIES | 0.65 | 0.55 |
139
 
140
  - **Router protection:** Gate weights use linear interpolation (~57% base, ~43% donors) rather than DARE to avoid destabilizing expert routing.
141
  - **DeltaNet stability:** Recurrent state kernels are excluded from DARE to prevent divergence in the linear-attention state space.