UraionLabs commited on
Commit
cf82f55
·
verified ·
1 Parent(s): 635b0b4

Supersede unsupported legacy artifact

Browse files

Preserve the original card while withdrawing unsupported evaluation and runtime claims. Current work is FinStruct.

Files changed (2) hide show
  1. LEGACY_CARD.md +845 -0
  2. README.md +22 -827
LEGACY_CARD.md ADDED
@@ -0,0 +1,845 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: Qwen/Qwen2.5-7B-Instruct
3
+ base_model_relation: finetune
4
+ library_name: transformers
5
+ license: apache-2.0
6
+ language:
7
+ - en
8
+ pipeline_tag: text-generation
9
+ tags:
10
+ - agent
11
+ - function-calling
12
+ - tool-use
13
+ - h-res
14
+ - manifold-steering
15
+ - peft
16
+ - uraion-labs
17
+ - uraion
18
+ - iclr-2026
19
+ - associative-memory
20
+ - hopfield
21
+ - neural-collapse
22
+ - qwen2.5
23
+ - sft
24
+ - trl
25
+ - hermes-function-calling
26
+ - apigen
27
+ - xlam
28
+ - toolace
29
+ datasets:
30
+ - NousResearch/hermes-function-calling-v1
31
+ - Salesforce/xlam-function-calling-60k
32
+ - mlabonne/FineTome-100k
33
+ - Salesforce/APIGen-MT-5k
34
+ - glaiveai/glaive-function-calling-v2
35
+ - Team-ACE/ToolACE
36
+ inference:
37
+ parameters:
38
+ temperature: 0.7
39
+ top_p: 0.95
40
+ max_new_tokens: 4096
41
+ ---
42
+
43
+ <p align="center">
44
+ <picture>
45
+ <source media="(prefers-color-scheme: dark)" srcset="https://uraionlabs.com/public/icons/icon-192.png">
46
+ <img src="https://uraionlabs.com/public/icons/icon-192.png" alt="Uraion Labs" width="64" height="64">
47
+ </picture>
48
+ </p>
49
+
50
+ <p align="center">
51
+ <strong style="font-family: 'Instrument Serif', Georgia, serif; font-size: 2rem; color: #F7F4ED; letter-spacing: -0.02em;">
52
+ Uraion Labs
53
+ </strong>
54
+ <br>
55
+ <span style="font-family: 'Inter', sans-serif; font-size: 0.875rem; color: #8A8478;">Foundational systems research.</span>
56
+ </p>
57
+
58
+ <p align="center">
59
+ <strong style="font-family: 'Inter', sans-serif; font-size: 1.15rem; color: #E45A1A;">
60
+ Uraion-Agent-Steer
61
+ </strong>
62
+ <br>
63
+ <span style="font-family: 'Inter', sans-serif; font-size: 0.875rem; color: #8A8478;">
64
+ Agentic LLM fine-tuned via Hierarchical Residual Steering (H-Res) — steers activations, not weights.
65
+ </span>
66
+ </p>
67
+
68
+ ---
69
+
70
+ **Uraion-Agent-Steer** is a 7-billion parameter model adapted from [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) using **H-Res (Hierarchical Residual Steering)** — a novel PEFT method from ["Parallel Manifold Steering"](https://arxiv.org/abs/2606.24396) (ICLR Workshop 2026). Rather than modifying model weights (LoRA) or injecting synthetic tokens (VPT/Prefix Tuning), H-Res learns a **state-dependent vector field** that steers hidden activations into task-specific attractors — preserving the foundation model's associative memory while adapting it for agentic tool use.
71
+
72
+ This is a research artifact in Uraion Labs' systems-first approach: studying novel adaptation mechanisms, the harness layer, evaluation, and deployment of agent-capable models. It is the first publicly available model trained with the full H-Res method.
73
+
74
+ **Intelligence is a systems problem.** This model is one piece of that system — and the adaptation method itself is part of the research.
75
+
76
+ ---
77
+
78
+ ## The H-Res Method
79
+
80
+ ### The problem with existing PEFT
81
+
82
+ | Method | Mechanism | Fatal flaw |
83
+ |--------|-----------|------------|
84
+ | **LoRA** | Modifies weights globally | Catastrophic interference — distorts retrieval dynamics of pre-trained memories |
85
+ | **VPT / Prefix Tuning** | Appends synthetic tokens to input | Buffer congestion — dilutes attention probability mass, weakens associative recall |
86
+ | **H-Res** | Steers activations via vector field | *None of the above* — operates orthogonal to weights and input buffer |
87
+
88
+ ### How H-Res works
89
+
90
+ H-Res frames Transformer adaptation as a **control problem on the activation manifold**. Each layer `l` receives a state-dependent residual:
91
+
92
+ ```
93
+ z_{l+1} = Attn(z_l) + FFN(z_l) + λ · H_θ(z_l)
94
+
95
+ where H_θ(x) = W_up · GeLU(W_down · x)
96
+ ```
97
+
98
+ - **W_down ∈ ℝ^{d×r}** — projects to a low-rank "control manifold" (bottleneck)
99
+ - **W_up ∈ ℝ^{r×d}** — projects the steering signal back to activation space
100
+ - **W_up initialized to zero** — no initialization shock; training starts from the pre-trained energy minimum
101
+ - **λ** — learnable per-layer scaling factor
102
+ - **Applied parallel to self-attention** — via forward hooks, orthogonal to the frozen backbone
103
+
104
+ ### Theoretical guarantees (from the paper)
105
+
106
+ | Property | Proof |
107
+ |----------|-------|
108
+ | **Attention entropy preserved** | No synthetic tokens → constant sequence length → H(A_cls) minimal |
109
+ | **Neural Collapse facilitated** | Residual adapter acts as Maxwell's Demon, filtering task-irrelevant noise |
110
+ | **Zero initialization** | W_up = 0 → H_θ(z) = 0 at t=0 → training starts from global energy minimum |
111
+ | **SSM-compatible** | Operates entirely in residual stream — compatible with Mamba, S4, DeltaNet |
112
+ | **Multi-task orthogonality** | Null-Space Projection of gradients across tasks (Eq. 6 in paper) |
113
+
114
+ ---
115
+
116
+ ## Contents
117
+
118
+ - [Model Details](#model-details)
119
+ - [H-Res Architecture (Deep Dive)](#h-res-architecture-deep-dive)
120
+ - [Intended Uses & Limitations](#intended-uses--limitations)
121
+ - [Training Data](#training-data)
122
+ - [Training Procedure](#training-procedure)
123
+ - [Hyperparameters](#hyperparameters)
124
+ - [Training Loss](#training-loss)
125
+ - [Quickstart](#quickstart)
126
+ - [H-Res Adapter Analysis](#h-res-adapter-analysis)
127
+ - [Hardware & Infrastructure](#hardware--infrastructure)
128
+ - [GGUF Availability](#gguf-availability)
129
+ - [Ethical Considerations](#ethical-considerations)
130
+ - [Citations](#citations)
131
+
132
+ ---
133
+
134
+ ## Model Details
135
+
136
+ | Property | Value |
137
+ |----------|-------|
138
+ | **Base model** | [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) |
139
+ | **Architecture** | Qwen2.5ForCausalLM — 28-layer pure Transformer (RoPE, SwiGLU, RMSNorm) |
140
+ | **Adaptation method** | **H-Res (Hierarchical Residual Steering)** — state-dependent vector field |
141
+ | **Context length** | 32,768 tokens (native, inherited) |
142
+ | **Parameters** | ~7.6B total, 12.8M H-Res trainable (0.17%) |
143
+ | **H-Res rank** | r = 64 per layer |
144
+ | **H-Res layers** | 28/28 injected (all layers compatible) |
145
+ | **Precision** | BF16 (full precision — no quantization of base model) |
146
+ | **License** | Apache 2.0 (inherited from Qwen2.5) |
147
+ | **On-disk size** | ~15.3 GB (BF16 safetensors) |
148
+ | **Paper** | [arXiv:2606.24396](https://arxiv.org/abs/2606.24396) — ICLR Workshop 2026 |
149
+
150
+ ### Architecture choice
151
+
152
+ Qwen2.5-7B-Instruct was chosen for this H-Res implementation because:
153
+
154
+ 1. **Pure Transformer** — 28 identical decoder layers with standard `input_layernorm` + `self_attn` + `post_attention_layernorm` + `mlp` — cleanest architecture for H-Res hook injection
155
+ 2. **Apache 2.0 license** — no gated access, no approval required, fully open
156
+ 3. **Strong instruct base** — already instruction-tuned, providing a solid foundation for agentic adaptation
157
+ 4. **7B weight class** — punches above its weight on agent benchmarks while fitting comfortably on A100-40GB
158
+
159
+ ---
160
+
161
+ ## H-Res Architecture (Deep Dive)
162
+
163
+ ### Injection mechanism
164
+
165
+ H-Res adapters are injected into each transformer layer via **PyTorch forward hooks** — no monkey-patching of forward methods, no model code modification:
166
+
167
+ ```
168
+ Layer forward (simplified):
169
+ ┌─────────────────────────────────────────────┐
170
+ │ residual = hidden_states │
171
+ │ normed = input_layernorm(hidden_states) │
172
+ │ │
173
+ │ attn_out = self_attn(normed) ← frozen │
174
+ │ hres_out = hres(normed) ← trained │ ← Hook: captures normed, adds to attn output
175
+ │ │
176
+ │ hidden_states = residual + attn_out + hres_out │
177
+ │ hidden_states = hidden_states + mlp(norm(hidden_states)) │
178
+ └─────────────────────────────────────────────┘
179
+ ```
180
+
181
+ ### Per-layer H-Res parameters
182
+
183
+ Each of the 28 layers contains:
184
+
185
+ ```
186
+ HResAdapter:
187
+ W_down: Linear(3584 → 64, bias=False) 228,544 params
188
+ W_up: Linear(64 → 3584, bias=False) 228,544 params
189
+ scale: scalar (learnable) 1 param
190
+ ─────────────────────────────────────────────────────
191
+ Total per layer: 457,089 params
192
+ Total (28 layers): 12,798,492 params
193
+ % of base model (7.6B): 0.17%
194
+ ```
195
+
196
+ ### Initialization (per paper Section 2.3)
197
+
198
+ ```python
199
+ W_down ~ N(0, 1/d_model) # Normal with σ = 1/√3584
200
+ W_up = 0 # Zero — preserves pre-trained energy minimum
201
+ scale = 0.1 # Small constant — gentle ramp-up
202
+ ```
203
+
204
+ At initialization, H_θ(x) = 0 for all x → the model behaves identically to the frozen base. Training gradually "turns on" the steering field.
205
+
206
+ ### What H-Res is NOT
207
+
208
+ - **NOT LoRA** — doesn't modify frozen weights; computes input-dependent residuals
209
+ - **NOT an adapter** — doesn't sit sequentially after attention/MLP; runs *parallel* to self-attention
210
+ - **NOT a prompt method** — doesn't add tokens to the input sequence
211
+ - **NOT a mixture-of-experts** — all layers are always active; the "expertise" is in the learned vector field
212
+
213
+ ---
214
+
215
+ ## Intended Uses & Limitations
216
+
217
+ ### Intended use
218
+
219
+ - **Tool-calling agents** — function calling, API orchestration, multi-turn tool use
220
+ - **Agent frameworks** — drop-in replacement for agent runtimes (OpenAI-compatible via vLLM)
221
+ - **Systems research** — studying the H-Res adaptation mechanism, its properties, and its limits
222
+ - **Associative retrieval tasks** — the H-Res method specifically excels at retrieval (26% better than LoRA on SQuAD per the paper)
223
+
224
+ ### Out-of-scope
225
+
226
+ - **Production deployment without validation** — research artifact; evaluate on your specific use case
227
+ - **High-stakes decision making** — not intended for medical, legal, or financial advice without human oversight
228
+ - **Unsupported languages** — trained exclusively on English data
229
+ - **Multimodal tasks** — text-only fine-tune
230
+
231
+ ### Limitations
232
+
233
+ - **Trained for 1 epoch** on ~35K examples. More data/epochs would improve tool-calling reliability.
234
+ - **H-Res is a research method** — this is the first public deployment; edge cases may exist.
235
+ - **GGUF conversion** — H-Res adapters are state-dependent (nonlinear), so they can't be directly merged into base weights for standard GGUF conversion. A LoRA-distilled GGUF version is available separately.
236
+ - **May produce malformed tool calls** in edge cases — validate output before execution.
237
+ - **7B weight class** — while punching above its weight, has inherent capacity limits compared to larger models.
238
+
239
+ ---
240
+
241
+ ## Training Data
242
+
243
+ Six datasets were curated for agentic capability — prioritizing function-calling and tool-use signal over raw instruction volume:
244
+
245
+ | Dataset | Type | Samples | Focus |
246
+ |---------|------|---------|-------|
247
+ | [NousResearch/hermes-function-calling-v1](https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1) | Function calling | 1,893 | Single-turn and multi-turn tool use conversations (MIT) |
248
+ | [Salesforce/xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k) | Function calling | 10,000 | Diverse API function calling (sampled from 60K, MIT) |
249
+ | [mlabonne/FineTome-100k](https://huggingface.co/datasets/mlabonne/FineTome-100k) | Instruction following | 20,000 | General instruct/chat data (sampled from 100K, MIT) |
250
+ | [Salesforce/APIGen-MT-5k](https://huggingface.co/datasets/Salesforce/APIGen-MT-5k) | API generation | 5,000 | Multi-turn API call generation across diverse APIs (MIT) |
251
+ | [glaiveai/glaive-function-calling-v2](https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2) | Function calling | 8,000 | Multi-turn tool-use conversations (MIT) |
252
+ | [Team-ACE/ToolACE](https://huggingface.co/datasets/Team-ACE/ToolACE) | Tool use | 8,000 | Agentic tool-use conversations (Apache 2.0) |
253
+ | **Total** | | **52,893 raw → 34,893 filtered** | |
254
+
255
+ All data formatted via `tokenizer.apply_chat_template()` with the Qwen2.5 ChatML template. Examples without a `user` role were filtered. Sequence length capped at 2,048 tokens.
256
+
257
+ ---
258
+
259
+ ## Training Procedure
260
+
261
+ ### Framework
262
+
263
+ - **Training**: HuggingFace TRL `SFTTrainer` with `SFTConfig`
264
+ - **Adaptation**: H-Res — custom `HResAdapter` injected via forward hooks (no PEFT library dependency for the core method)
265
+ - **Quantization**: None — full BF16 precision for base model (H-Res adds only 0.17% trainable params)
266
+ - **Attention**: PyTorch SDPA (`attn_implementation="sdpa"`)
267
+ - **Loss**: Standard causal language modeling (no packing)
268
+
269
+ ### Pipeline
270
+
271
+ 1. **Model loading**: BF16 full precision via `AutoModelForCausalLM.from_pretrained()`
272
+ 2. **H-Res injection**: Forward hooks on `input_layernorm` (capture) + `self_attn` (inject)
273
+ 3. **Base model freeze**: `model.requires_grad_(False)` — only H-Res params trainable
274
+ 4. **Dataset processing**: ShareGPT → ChatML → filtered → concatenated → shuffled
275
+ 5. **Training**: `SFTTrainer` with `dataset_text_field="text"`, `packing=False`, `gradient_checkpointing=True`
276
+ 6. **Export**: `model.save_pretrained(safe_serialization=True)` — H-Res adapters embedded in model state dict
277
+ 7. **Upload**: `HfApi.upload_folder()` → `UraionLabs/Uraion-Agent-Steer`
278
+
279
+ ### Novel aspects
280
+
281
+ This training represents the **first public implementation** of the full H-Res method:
282
+
283
+ - **Hook-based injection** — no model code modification; works with any HuggingFace Transformer
284
+ - **Full BF16 precision** — no quantization noise; H-Res is parameter-efficient enough to not need it
285
+ - **Learnable scale parameter λ** — per-layer, initialized at 0.1, allowing layers to independently adjust steering intensity
286
+ - **Architecture-agnostic** — the same injection code works on Llama, Mistral, Qwen2/3, Gemma, and Phi
287
+
288
+ ---
289
+
290
+ ## Hyperparameters
291
+
292
+ ### H-Res
293
+
294
+ | Parameter | Value |
295
+ |-----------|-------|
296
+ | `r` (bottleneck rank) | 64 |
297
+ | `d_model` (hidden size) | 3584 |
298
+ | `W_down init` | N(0, 1/d_model) |
299
+ | `W_up init` | 0 (zero) |
300
+ | `scale init` | 0.1 |
301
+ | `activation` | GeLU |
302
+ | `bias` | None |
303
+
304
+ ### Training
305
+
306
+ | Parameter | Value |
307
+ |-----------|-------|
308
+ | **Sequence length** | 2048 |
309
+ | **Effective batch size** | 32 |
310
+ | **Per-device batch** | 2 |
311
+ | **Gradient accumulation** | 16 |
312
+ | **Learning rate** | 1×10⁻⁴ |
313
+ | **LR scheduler** | Cosine with warmup |
314
+ | **Warmup ratio** | 0.03 |
315
+ | **Optimizer** | AdamW 8-bit |
316
+ | **Epochs** | 1 |
317
+ | **Max steps** | 1,091 |
318
+ | **Weight decay** | 0.0 |
319
+ | **Gradient checkpointing** | True (non-reentrant) |
320
+ | **Precision** | BF16 |
321
+ | **Logging steps** | 10 |
322
+ | **Save steps** | 50 |
323
+ | **Save total limit** | 3 |
324
+
325
+ ---
326
+
327
+ ## Training Loss
328
+
329
+ | Step | Loss | Δ from start | Notes |
330
+ |------|------|-------------|-------|
331
+ | 10 | 1.310 | — | Initial — H-Res scale still ramping |
332
+ | 20 | 1.264 | ↓ 3.5% | W_up beginning to activate |
333
+ | 50 | 1.013 | ↓ 22.7% | First checkpoint saved; steering field forming |
334
+ | 100 | 0.879 | ↓ 32.9% | Rapid convergence phase |
335
+ | 200 | 0.741 | ↓ 43.4% | Entering fine-tuning regime |
336
+ | 300 | 0.745 | ↓ 43.1% | Stable convergence |
337
+ | 400 | 0.699 | ↓ 46.6% | Steady improvement |
338
+ | 500 | 0.689 | ↓ 47.4% | Approaching plateau |
339
+ | 600 | 0.645 | ↓ 50.8% | Best single-step loss |
340
+ | 700 | 0.688 | ↓ 47.5% | Minor oscillation — normal |
341
+ | 800 | 0.646 | ↓ 50.7% | Consistent low-loss regime |
342
+ | 900 | 0.663 | ↓ 49.4% | Stable |
343
+ | 1000 | 0.67 | ↓ 48.9% | Final stretch |
344
+ | **1091** | **0.657** | **↓ 49.8%** | **Final — 50% loss reduction** |
345
+
346
+ **Key observations:**
347
+ - **Rapid early convergence** — 22.7% loss reduction by step 50 (first 4.6% of training)
348
+ - **Smooth learning curve** — no spikes, no divergence, consistent downward trend
349
+ - **50% total loss reduction** — from 1.310 to 0.657
350
+ - **H-Res's zero-initialization advantage** — no "initialization shock" means the model starts from a good place and improves monotonically
351
+
352
+ ---
353
+
354
+ ## Local Inference Guide
355
+
356
+ This model uses **safetensors with H-Res adapters embedded** — no extra adapter files needed. Load it like any standard Transformers model. Below are instructions for every major local inference tool.
357
+
358
+ ### Contents
359
+ - [Transformers (Python)](#transformers-python) — full quality, recommended
360
+ - [vLLM (OpenAI-compatible server)](#vllm-openai-compatible-server) — production serving
361
+ - [Unsloth (further fine-tuning)](#unsloth-further-fine-tuning) — continue training
362
+ - [Ollama](#ollama) — import from safetensors
363
+ - [LM Studio](#lm-studio) — local desktop inference
364
+ - [llama.cpp](#llamacpp) — GGUF note
365
+ - [text-generation-webui (Oobabooga)](#text-generation-webui-oobabooga)
366
+
367
+ ---
368
+
369
+ ### Transformers (Python)
370
+
371
+ The simplest way — loads H-Res adapters automatically.
372
+
373
+ ```python
374
+ import torch
375
+ from transformers import AutoModelForCausalLM, AutoTokenizer
376
+
377
+ model_name = "UraionLabs/Uraion-Agent-Steer"
378
+
379
+ tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
380
+ model = AutoModelForCausalLM.from_pretrained(
381
+ model_name,
382
+ torch_dtype=torch.bfloat16,
383
+ device_map="auto",
384
+ trust_remote_code=True,
385
+ )
386
+
387
+ # H-Res adapters are embedded — no extra loading needed
388
+ messages = [
389
+ {"role": "system", "content": "You are Uraion-Agent-Steer, an agent with tool-use capabilities."},
390
+ {"role": "user", "content": "What's the weather in Tokyo? Should I bring an umbrella?"},
391
+ ]
392
+
393
+ text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
394
+ inputs = tokenizer(text, return_tensors="pt").to(model.device)
395
+
396
+ outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7, top_p=0.95, do_sample=True)
397
+ response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
398
+ print(response)
399
+ ```
400
+
401
+ **VRAM requirement:** ~16 GB (BF16). Works on RTX 3090/4090, A4000, A5000, A100, or any 24GB+ consumer GPU.
402
+
403
+ **With 12GB GPUs** (RTX 3080/4070): use 8-bit quantization:
404
+
405
+ ```python
406
+ model = AutoModelForCausalLM.from_pretrained(
407
+ model_name,
408
+ load_in_8bit=True,
409
+ device_map="auto",
410
+ trust_remote_code=True,
411
+ )
412
+ ```
413
+
414
+ **With `pipeline` (simpler):**
415
+
416
+ ```python
417
+ from transformers import pipeline
418
+
419
+ pipe = pipeline(
420
+ "text-generation",
421
+ model="UraionLabs/Uraion-Agent-Steer",
422
+ torch_dtype=torch.bfloat16,
423
+ device_map="auto",
424
+ trust_remote_code=True,
425
+ )
426
+
427
+ messages = [{"role": "user", "content": "Search for the latest AI research papers."}]
428
+ output = pipe(messages, max_new_tokens=512, temperature=0.7, top_p=0.95)
429
+ print(output[0]["generated_text"])
430
+ ```
431
+
432
+ ---
433
+
434
+ ### vLLM (OpenAI-compatible server)
435
+
436
+ Best for production agent deployments. vLLM loads safetensors directly with full H-Res adapter support.
437
+
438
+ ```bash
439
+ # Install vLLM
440
+ pip install vllm
441
+
442
+ # Serve with OpenAI-compatible API
443
+ vllm serve UraionLabs/Uraion-Agent-Steer \
444
+ --trust-remote-code \
445
+ --host 0.0.0.0 \
446
+ --port 8000 \
447
+ --max-model-len 8192 \
448
+ --gpu-memory-utilization 0.90
449
+ ```
450
+
451
+ **OpenAI client (Python):**
452
+
453
+ ```python
454
+ from openai import OpenAI
455
+
456
+ client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
457
+
458
+ response = client.chat.completions.create(
459
+ model="UraionLabs/Uraion-Agent-Steer",
460
+ messages=[{"role": "user", "content": "What's 2+2?"}],
461
+ temperature=0.7,
462
+ )
463
+ print(response.choices[0].message.content)
464
+ ```
465
+
466
+ **With tool calling:**
467
+
468
+ ```python
469
+ tools = [{
470
+ "type": "function",
471
+ "function": {
472
+ "name": "get_weather",
473
+ "description": "Get current weather for a location",
474
+ "parameters": {
475
+ "type": "object",
476
+ "properties": {
477
+ "location": {"type": "string", "description": "City name"}
478
+ },
479
+ "required": ["location"]
480
+ }
481
+ }
482
+ }]
483
+
484
+ response = client.chat.completions.create(
485
+ model="UraionLabs/Uraion-Agent-Steer",
486
+ messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
487
+ tools=tools,
488
+ temperature=0.0,
489
+ )
490
+ print(response.choices[0].message.tool_calls)
491
+ ```
492
+
493
+ **VRAM:** ~18 GB for vLLM serving. Recommended: A10, A100, or 24GB consumer GPU (RTX 3090/4090).
494
+
495
+ ---
496
+
497
+ ### Unsloth (further fine-tuning)
498
+
499
+ Continue training Uraion-Agent-Steer with Unsloth for 2× faster, 70% less memory fine-tuning.
500
+
501
+ ```python
502
+ from unsloth import FastLanguageModel
503
+ from unsloth.chat_templates import get_chat_template
504
+ import torch
505
+
506
+ # Load Uraion-Agent-Steer with Unsloth acceleration
507
+ model, tokenizer = FastLanguageModel.from_pretrained(
508
+ model_name="UraionLabs/Uraion-Agent-Steer",
509
+ max_seq_length=2048,
510
+ dtype=torch.bfloat16,
511
+ load_in_4bit=True, # 4-bit for further QLoRA training
512
+ trust_remote_code=True,
513
+ )
514
+
515
+ # Apply QLoRA for continued training
516
+ model = FastLanguageModel.get_peft_model(
517
+ model,
518
+ r=32,
519
+ lora_alpha=32,
520
+ target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
521
+ "gate_proj", "up_proj", "down_proj"],
522
+ lora_dropout=0,
523
+ bias="none",
524
+ )
525
+
526
+ # Continue training with your data...
527
+ # model = ... your training loop ...
528
+
529
+ # Save LoRA adapters
530
+ model.save_pretrained("uraion-agent-steer-continued")
531
+ ```
532
+
533
+ > **Note:** The H-Res adapters remain frozen alongside the base model during Unsloth QLoRA training. The new LoRA adapters learn on top of the H-Res steering field — a "steer + adapt" stack.
534
+
535
+ **VRAM:** ~8 GB with 4-bit QLoRA via Unsloth.
536
+
537
+ ---
538
+
539
+ ### Ollama
540
+
541
+ Ollama uses GGUF format. Since H-Res adapters can't be directly merged into base weights, you have two options:
542
+
543
+ **Option A: Import from safetensors (Ollama 0.3.0+)**
544
+
545
+ ```bash
546
+ # Create Modelfile
547
+ cat > Modelfile << 'EOF'
548
+ FROM UraionLabs/Uraion-Agent-Steer
549
+ TEMPLATE """{{ if .System }}<|im_start|>system
550
+ {{ .System }}<|im_end|>
551
+ {{ end }}{{ if .Prompt }}<|im_start|>user
552
+ {{ .Prompt }}<|im_end|>
553
+ {{ end }}<|im_start|>assistant
554
+ """
555
+ PARAMETER stop "<|im_start|>"
556
+ PARAMETER stop "<|im_end|>"
557
+ PARAMETER temperature 0.7
558
+ EOF
559
+
560
+ # Import into Ollama
561
+ ollama create uraion-agent-steer -f Modelfile
562
+
563
+ # Run
564
+ ollama run uraion-agent-steer
565
+ ```
566
+
567
+ **Option B: Use the GGUF release (when available)**
568
+
569
+ ```bash
570
+ # Coming soon — LoRA-distilled GGUF version
571
+ ollama run hf.co/UraionLabs/Uraion-Agent-Steer-GGUF:Q4_K_M
572
+ ```
573
+
574
+ > **Note for Option A:** Ollama's safetensors import loads the full BF16 model (~15 GB download, ~16 GB VRAM). If you're VRAM-constrained, wait for the GGUF release or use 8-bit Transformers.
575
+
576
+ ---
577
+
578
+ ### LM Studio
579
+
580
+ LM Studio works with GGUF models. For safetensors models, use one of these approaches:
581
+
582
+ **Approach 1: Wait for GGUF release**
583
+
584
+ The LoRA-distilled GGUF version (`UraionLabs/Uraion-Agent-Steer-GGUF`) will be importable directly in LM Studio's model browser. Pick a quant (Q4_K_M recommended) and download.
585
+
586
+ **Approach 2: Use MLX (Apple Silicon)**
587
+
588
+ If you're on a Mac with Apple Silicon, MLX can load safetensors directly:
589
+
590
+ ```bash
591
+ pip install mlx mlx-lm
592
+
593
+ # Convert to MLX format
594
+ mlx_lm.convert --hf-path UraionLabs/Uraion-Agent-Steer --mlx-path ./uraion-agent-steer-mlx
595
+
596
+ # Run inference
597
+ mlx_lm.generate --model ./uraion-agent-steer-mlx --prompt "What is tool calling?"
598
+ ```
599
+
600
+ **Approach 3: Use vLLM or llama.cpp server**
601
+
602
+ Run the model locally via vLLM (see above), then connect LM Studio to it via the "Local Server" option in LM Studio's settings.
603
+
604
+ ---
605
+
606
+ ### llama.cpp
607
+
608
+ llama.cpp requires GGUF format. Since H-Res adapters are state-dependent, direct GGUF conversion isn't possible. Two paths:
609
+
610
+ **Path 1: Use the GGUF-distilled release**
611
+
612
+ ```bash
613
+ # Coming soon
614
+ llama-server -hf UraionLabs/Uraion-Agent-Steer-GGUF:Q4_K_M --host 0.0.0.0 --port 8000
615
+ ```
616
+
617
+ **Path 2: Use Transformers server + llama.cpp client**
618
+
619
+ ```bash
620
+ # Server side (Transformers with H-Res, fast)
621
+ python -c "
622
+ from transformers import AutoModelForCausalLM, AutoTokenizer
623
+ import torch
624
+ # ... serve with FastAPI or vLLM
625
+ "
626
+
627
+ # Client side (any OpenAI-compatible client)
628
+ curl http://localhost:8000/v1/chat/completions \
629
+ -H "Content-Type: application/json" \
630
+ -d '{"model":"uraion-agent-steer","messages":[{"role":"user","content":"Hello"}]}'
631
+ ```
632
+
633
+ ---
634
+
635
+ ### text-generation-webui (Oobabooga)
636
+
637
+ Load directly in Oobabooga's Transformers loader:
638
+
639
+ 1. Go to the **Model** tab
640
+ 2. In "Download model or LoRA", enter: `UraionLabs/Uraion-Agent-Steer`
641
+ 3. Click **Download**
642
+ 4. After download, select the model and set:
643
+ - **Loader:** Transformers
644
+ - **trust_remote_code:** ✓
645
+ - **dtype:** bfloat16
646
+ 5. Click **Load**
647
+
648
+ For lower VRAM, enable `load_in_8bit` or `load_in_4bit` in the loader settings.
649
+
650
+ **Chat template** (if not auto-detected): `chatml` (Qwen2.5 ChatML format).
651
+
652
+ ---
653
+
654
+ ### VRAM Reference
655
+
656
+ | GPU | VRAM | Config | Notes |
657
+ |-----|------|--------|-------|
658
+ | RTX 4090 (24GB) | 24 GB | BF16 | Full quality, fits comfortably |
659
+ | RTX 3090 (24GB) | 24 GB | BF16 | Same as above |
660
+ | RTX 4080 (16GB) | 16 GB | BF16 | Tight — use 8-bit for safety |
661
+ | RTX 3080 (10GB) | 10 GB | 8-bit | Works with `load_in_8bit=True` |
662
+ | RTX 4070 (12GB) | 12 GB | 8-bit | Works with `load_in_8bit=True` |
663
+ | A100 (40GB) | 40 GB | BF16 | Full quality, plenty of room |
664
+ | A10 (24GB) | 24 GB | BF16 | Full quality |
665
+ | T4 (16GB) | 16 GB | 8-bit | Use `load_in_8bit=True` |
666
+ | Apple M2/M3 (16GB+) | Unified | MLX | Convert with `mlx_lm.convert` |
667
+
668
+ ---
669
+
670
+ ## H-Res Adapter Analysis
671
+
672
+ After training, we inspected the learned H-Res adapters across all 28 layers:
673
+
674
+ | Layer | Scale (λ) | ‖W_up‖ | ‖W_down‖ | Steering activity |
675
+ |-------|-----------|--------|----------|-------------------|
676
+ | 0 (early) | 0.1001 | 0.0000 | 7.94 | **Silent** — shallow layers don't steer |
677
+ | 8 (mid) | 0.1001 | 2.12 | 8.45 | Moderate steering |
678
+ | 16 (mid-deep) | 0.1001 | 2.87 | 9.12 | Active steering |
679
+ | 24 (deep) | 0.1001 | 3.12 | 9.56 | Strong steering |
680
+ | 27 (final) | 0.1001 | **3.72** | **9.69** | **Maximum steering** |
681
+
682
+ **Key finding:** Steering intensity increases monotonically with layer depth. Early layers (0–3) have W_up ≈ 0 — the adapter is effectively dormant. Deep layers (20–27) have the strongest steering activity. This aligns with the paper's theoretical prediction: H-Res acts primarily on high-level semantic representations in deeper layers, while preserving low-level features in early layers.
683
+
684
+ The scale parameter λ stayed at ~0.1 across all layers — the model preferred to learn through W_up/W_down rather than adjusting the global scaling factor.
685
+
686
+ ---
687
+
688
+ ## Hardware & Infrastructure
689
+
690
+ | Component | Detail |
691
+ |-----------|--------|
692
+ | **Provisioning** | Google Colab CLI (`colab-cli`) via OAuth2 |
693
+ | **GPU** | 1× NVIDIA A100-SXM4-40GB |
694
+ | **Runtime** | `colab run --gpu A100 --keep --timeout 28800` |
695
+ | **Training time** | ~3 hours (1,091 steps at ~10s/step) |
696
+ | **VRAM usage** | ~35 GB (7.6B BF16 base + 12.8M H-Res + activations + optimizer) |
697
+ | **Setup** | Self-installing dependencies via pip |
698
+ | **Session lifecycle** | `colab run` → auto-execute → `--keep` → training → auto-upload → session release |
699
+
700
+ Training dependencies auto-installed on Colab: `transformers>=4.57`, `trl>=0.21`, `datasets`, `accelerate`, `safetensors`, `huggingface_hub`.
701
+
702
+ ---
703
+
704
+ ## GGUF Availability
705
+
706
+ H-Res adapters are **state-dependent** (nonlinear function of the input), so they can't be directly merged into base weights for standard GGUF/llama.cpp conversion. For GGUF inference:
707
+
708
+ | Option | How | VRAM | Quality |
709
+ |--------|-----|------|---------|
710
+ | **[Ollama safetensors import](#ollama)** | `FROM UraionLabs/Uraion-Agent-Steer` in Modelfile | ~16 GB | Full H-Res quality |
711
+ | **[MLX conversion](#lm-studio)** | `mlx_lm.convert` on Apple Silicon | ~16 GB unified | Full H-Res quality |
712
+ | **LoRA-distilled GGUF** | `UraionLabs/Uraion-Agent-Steer-GGUF` (coming soon) | 4–8 GB | LoRA-approximated |
713
+ | **8-bit Transformers** | `load_in_8bit=True` with Transformers | ~8 GB | Near-full quality |
714
+
715
+ The LoRA-distilled GGUF release is in progress (Colab GPU quota recovery). For maximum quality TODAY, use Ollama's safetensors import or vLLM.
716
+
717
+ ---
718
+
719
+ ## Ethical Considerations
720
+
721
+ This model is a fine-tune of Qwen2.5-7B-Instruct and inherits its base capabilities and biases:
722
+
723
+ - Training data includes user-generated content from HuggingFace datasets, which may contain biases.
724
+ - Function-calling capabilities could automate actions without human oversight — always validate tool calls before execution.
725
+ - The model has not undergone safety alignment beyond the base model's existing safeguards.
726
+ - The H-Res method is novel — long-term behavior and failure modes are still being studied.
727
+ - This is a **research-stage artifact** from Uraion Labs. We are a systems research lab, not a product company. Use accordingly.
728
+
729
+ ---
730
+
731
+ ## Citations
732
+
733
+ ### H-Res (Parallel Manifold Steering)
734
+
735
+ ```bibtex
736
+ @article{awadhiya2026parallel,
737
+ title={Parallel Manifold Steering: Efficient Adaptation of Large
738
+ Associative Memories via Residual Energy Shaping},
739
+ author={Awadhiya, Kanishk},
740
+ journal={ICLR Workshop on New Frontiers in Associative Memory},
741
+ year={2026},
742
+ url={https://arxiv.org/abs/2606.24396}
743
+ }
744
+ ```
745
+
746
+ ### Uraion-Agent-Steer
747
+
748
+ ```bibtex
749
+ @software{uraion-agent-steer,
750
+ title={Uraion-Agent-Steer: Agentic Model via Hierarchical Residual Steering},
751
+ author={Uraion Labs},
752
+ year={2026},
753
+ url={https://huggingface.co/UraionLabs/Uraion-Agent-Steer}
754
+ }
755
+ ```
756
+
757
+ ### Qwen2.5
758
+
759
+ ```bibtex
760
+ @misc{qwen2.5,
761
+ title={Qwen2.5: A Party of Foundation Models},
762
+ author={Qwen Team},
763
+ year={2025},
764
+ publisher={GitHub},
765
+ url={https://github.com/QwenLM/Qwen2.5}
766
+ }
767
+ ```
768
+
769
+ ### TRL
770
+
771
+ ```bibtex
772
+ @software{vonwerra2020trl,
773
+ title={{TRL: Transformers Reinforcement Learning}},
774
+ author={von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and
775
+ Beeching, Edward and Thrush, Tristan and Lambert, Nathan and
776
+ Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
777
+ license={Apache-2.0},
778
+ url={https://github.com/huggingface/trl},
779
+ year={2020}
780
+ }
781
+ ```
782
+
783
+ ### Datasets
784
+
785
+ ```bibtex
786
+ @misc{hermesfc,
787
+ title={NousResearch Hermes Function Calling},
788
+ author={Nous Research},
789
+ year={2024},
790
+ url={https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1}
791
+ }
792
+
793
+ @misc{xlam2024,
794
+ title={xLAM: A Family of Large Action Models},
795
+ author={Salesforce AI Research},
796
+ year={2024},
797
+ url={https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k}
798
+ }
799
+
800
+ @misc{finetome2024,
801
+ title={FineTome-100k: A Curated Instruction Tuning Dataset},
802
+ author={Labonne, Maxime},
803
+ year={2024},
804
+ url={https://huggingface.co/datasets/mlabonne/FineTome-100k}
805
+ }
806
+
807
+ @misc{apigen2024,
808
+ title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets},
809
+ author={Salesforce AI Research},
810
+ year={2024},
811
+ url={https://huggingface.co/datasets/Salesforce/APIGen-MT-5k}
812
+ }
813
+
814
+ @misc{glaivefc,
815
+ title={Glaive Function Calling v2},
816
+ author={Glaive AI},
817
+ year={2024},
818
+ url={https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2}
819
+ }
820
+
821
+ @misc{toolace2025,
822
+ title={ToolACE: Winning the Points of LLM Function Calling},
823
+ author={Team ACE},
824
+ year={2025},
825
+ url={https://huggingface.co/datasets/Team-ACE/ToolACE}
826
+ }
827
+ ```
828
+
829
+ ---
830
+
831
+ <p align="center">
832
+ <img src="https://uraionlabs.com/public/icons/icon-32.png" alt="" width="24" height="24">
833
+ </p>
834
+
835
+ <p align="center" style="font-family: 'Inter', sans-serif; font-size: 0.8rem; color: #8A8478;">
836
+ <strong style="color: #F7F4ED;">Uraion Labs</strong> — Foundational systems research.
837
+ <br>
838
+ <a href="https://uraionlabs.com" style="color: #E45A1A;">uraionlabs.com</a>
839
+ <br><br>
840
+ <em style="color: #6F6A61;">
841
+ Intelligence is a systems problem.
842
+ </em>
843
+ <br>
844
+ Licensed under <a href="https://www.apache.org/licenses/LICENSE-2.0" style="color: #E45A1A;">Apache 2.0</a>.
845
+ </p>
README.md CHANGED
@@ -7,839 +7,34 @@ language:
7
  - en
8
  pipeline_tag: text-generation
9
  tags:
10
- - agent
11
- - function-calling
12
- - tool-use
13
  - h-res
14
- - manifold-steering
15
- - peft
16
- - uraion-labs
17
- - uraion
18
- - iclr-2026
19
- - associative-memory
20
- - hopfield
21
- - neural-collapse
22
- - qwen2.5
23
- - sft
24
- - trl
25
- - hermes-function-calling
26
- - apigen
27
- - xlam
28
- - toolace
29
- datasets:
30
- - NousResearch/hermes-function-calling-v1
31
- - Salesforce/xlam-function-calling-60k
32
- - mlabonne/FineTome-100k
33
- - Salesforce/APIGen-MT-5k
34
- - glaiveai/glaive-function-calling-v2
35
- - Team-ACE/ToolACE
36
- inference:
37
- parameters:
38
- temperature: 0.7
39
- top_p: 0.95
40
- max_new_tokens: 4096
41
  ---
42
 
43
- <p align="center">
44
- <picture>
45
- <source media="(prefers-color-scheme: dark)" srcset="https://uraionlabs.com/public/icons/icon-192.png">
46
- <img src="https://uraionlabs.com/public/icons/icon-192.png" alt="Uraion Labs" width="64" height="64">
47
- </picture>
48
- </p>
49
 
50
- <p align="center">
51
- <strong style="font-family: 'Instrument Serif', Georgia, serif; font-size: 2rem; color: #F7F4ED; letter-spacing: -0.02em;">
52
- Uraion Labs
53
- </strong>
54
- <br>
55
- <span style="font-family: 'Inter', sans-serif; font-size: 0.875rem; color: #8A8478;">Foundational systems research.</span>
56
- </p>
57
 
58
- <p align="center">
59
- <strong style="font-family: 'Inter', sans-serif; font-size: 1.15rem; color: #E45A1A;">
60
- Uraion-Agent-Steer
61
- </strong>
62
- <br>
63
- <span style="font-family: 'Inter', sans-serif; font-size: 0.875rem; color: #8A8478;">
64
- Agentic LLM fine-tuned via Hierarchical Residual Steering (H-Res) — steers activations, not weights.
65
- </span>
66
- </p>
67
 
68
- ---
69
-
70
- **Uraion-Agent-Steer** is a 7-billion parameter model adapted from [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) using **H-Res (Hierarchical Residual Steering)** — a novel PEFT method from ["Parallel Manifold Steering"](https://arxiv.org/abs/2606.24396) (ICLR Workshop 2026). Rather than modifying model weights (LoRA) or injecting synthetic tokens (VPT/Prefix Tuning), H-Res learns a **state-dependent vector field** that steers hidden activations into task-specific attractors — preserving the foundation model's associative memory while adapting it for agentic tool use.
71
-
72
- This is a research artifact in Uraion Labs' systems-first approach: studying novel adaptation mechanisms, the harness layer, evaluation, and deployment of agent-capable models. It is the first publicly available model trained with the full H-Res method.
73
-
74
- **Intelligence is a systems problem.** This model is one piece of that system — and the adaptation method itself is part of the research.
75
-
76
- ---
77
-
78
- ## The H-Res Method
79
-
80
- ### The problem with existing PEFT
81
-
82
- | Method | Mechanism | Fatal flaw |
83
- |--------|-----------|------------|
84
- | **LoRA** | Modifies weights globally | Catastrophic interference — distorts retrieval dynamics of pre-trained memories |
85
- | **VPT / Prefix Tuning** | Appends synthetic tokens to input | Buffer congestion — dilutes attention probability mass, weakens associative recall |
86
- | **H-Res** | Steers activations via vector field | *None of the above* — operates orthogonal to weights and input buffer |
87
-
88
- ### How H-Res works
89
-
90
- H-Res frames Transformer adaptation as a **control problem on the activation manifold**. Each layer `l` receives a state-dependent residual:
91
-
92
- ```
93
- z_{l+1} = Attn(z_l) + FFN(z_l) + λ · H_θ(z_l)
94
-
95
- where H_θ(x) = W_up · GeLU(W_down · x)
96
- ```
97
-
98
- - **W_down ∈ ℝ^{d×r}** — projects to a low-rank "control manifold" (bottleneck)
99
- - **W_up ∈ ℝ^{r×d}** — projects the steering signal back to activation space
100
- - **W_up initialized to zero** — no initialization shock; training starts from the pre-trained energy minimum
101
- - **λ** — learnable per-layer scaling factor
102
- - **Applied parallel to self-attention** — via forward hooks, orthogonal to the frozen backbone
103
-
104
- ### Theoretical guarantees (from the paper)
105
-
106
- | Property | Proof |
107
- |----------|-------|
108
- | **Attention entropy preserved** | No synthetic tokens → constant sequence length → H(A_cls) minimal |
109
- | **Neural Collapse facilitated** | Residual adapter acts as Maxwell's Demon, filtering task-irrelevant noise |
110
- | **Zero initialization** | W_up = 0 → H_θ(z) = 0 at t=0 → training starts from global energy minimum |
111
- | **SSM-compatible** | Operates entirely in residual stream — compatible with Mamba, S4, DeltaNet |
112
- | **Multi-task orthogonality** | Null-Space Projection of gradients across tasks (Eq. 6 in paper) |
113
-
114
- ---
115
-
116
- ## Contents
117
-
118
- - [Model Details](#model-details)
119
- - [H-Res Architecture (Deep Dive)](#h-res-architecture-deep-dive)
120
- - [Intended Uses & Limitations](#intended-uses--limitations)
121
- - [Training Data](#training-data)
122
- - [Training Procedure](#training-procedure)
123
- - [Hyperparameters](#hyperparameters)
124
- - [Training Loss](#training-loss)
125
- - [Quickstart](#quickstart)
126
- - [H-Res Adapter Analysis](#h-res-adapter-analysis)
127
- - [Hardware & Infrastructure](#hardware--infrastructure)
128
- - [GGUF Availability](#gguf-availability)
129
- - [Ethical Considerations](#ethical-considerations)
130
- - [Citations](#citations)
131
-
132
- ---
133
-
134
- ## Model Details
135
-
136
- | Property | Value |
137
- |----------|-------|
138
- | **Base model** | [Qwen/Qwen2.5-7B-Instruct](https://huggingface.co/Qwen/Qwen2.5-7B-Instruct) |
139
- | **Architecture** | Qwen2.5ForCausalLM — 28-layer pure Transformer (RoPE, SwiGLU, RMSNorm) |
140
- | **Adaptation method** | **H-Res (Hierarchical Residual Steering)** — state-dependent vector field |
141
- | **Context length** | 32,768 tokens (native, inherited) |
142
- | **Parameters** | ~7.6B total, 12.8M H-Res trainable (0.17%) |
143
- | **H-Res rank** | r = 64 per layer |
144
- | **H-Res layers** | 28/28 injected (all layers compatible) |
145
- | **Precision** | BF16 (full precision — no quantization of base model) |
146
- | **License** | Apache 2.0 (inherited from Qwen2.5) |
147
- | **On-disk size** | ~15.3 GB (BF16 safetensors) |
148
- | **Paper** | [arXiv:2606.24396](https://arxiv.org/abs/2606.24396) — ICLR Workshop 2026 |
149
-
150
- ### Architecture choice
151
-
152
- Qwen2.5-7B-Instruct was chosen for this H-Res implementation because:
153
-
154
- 1. **Pure Transformer** — 28 identical decoder layers with standard `input_layernorm` + `self_attn` + `post_attention_layernorm` + `mlp` — cleanest architecture for H-Res hook injection
155
- 2. **Apache 2.0 license** — no gated access, no approval required, fully open
156
- 3. **Strong instruct base** — already instruction-tuned, providing a solid foundation for agentic adaptation
157
- 4. **7B weight class** — punches above its weight on agent benchmarks while fitting comfortably on A100-40GB
158
-
159
- ---
160
-
161
- ## H-Res Architecture (Deep Dive)
162
-
163
- ### Injection mechanism
164
-
165
- H-Res adapters are injected into each transformer layer via **PyTorch forward hooks** — no monkey-patching of forward methods, no model code modification:
166
-
167
- ```
168
- Layer forward (simplified):
169
- ┌─────────────────────────────────────────────┐
170
- │ residual = hidden_states │
171
- │ normed = input_layernorm(hidden_states) │
172
- │ │
173
- │ attn_out = self_attn(normed) ← frozen │
174
- │ hres_out = hres(normed) ← trained │ ← Hook: captures normed, adds to attn output
175
- │ │
176
- │ hidden_states = residual + attn_out + hres_out │
177
- │ hidden_states = hidden_states + mlp(norm(hidden_states)) │
178
- └─────────────────────────────────────────────┘
179
- ```
180
-
181
- ### Per-layer H-Res parameters
182
-
183
- Each of the 28 layers contains:
184
-
185
- ```
186
- HResAdapter:
187
- W_down: Linear(3584 → 64, bias=False) 228,544 params
188
- W_up: Linear(64 → 3584, bias=False) 228,544 params
189
- scale: scalar (learnable) 1 param
190
- ─────────────────────────────────────────────────────
191
- Total per layer: 457,089 params
192
- Total (28 layers): 12,798,492 params
193
- % of base model (7.6B): 0.17%
194
- ```
195
-
196
- ### Initialization (per paper Section 2.3)
197
-
198
- ```python
199
- W_down ~ N(0, 1/d_model) # Normal with σ = 1/√3584
200
- W_up = 0 # Zero — preserves pre-trained energy minimum
201
- scale = 0.1 # Small constant — gentle ramp-up
202
- ```
203
-
204
- At initialization, H_θ(x) = 0 for all x → the model behaves identically to the frozen base. Training gradually "turns on" the steering field.
205
-
206
- ### What H-Res is NOT
207
-
208
- - **NOT LoRA** — doesn't modify frozen weights; computes input-dependent residuals
209
- - **NOT an adapter** — doesn't sit sequentially after attention/MLP; runs *parallel* to self-attention
210
- - **NOT a prompt method** — doesn't add tokens to the input sequence
211
- - **NOT a mixture-of-experts** — all layers are always active; the "expertise" is in the learned vector field
212
-
213
- ---
214
-
215
- ## Intended Uses & Limitations
216
-
217
- ### Intended use
218
-
219
- - **Tool-calling agents** — function calling, API orchestration, multi-turn tool use
220
- - **Agent frameworks** — drop-in replacement for agent runtimes (OpenAI-compatible via vLLM)
221
- - **Systems research** — studying the H-Res adaptation mechanism, its properties, and its limits
222
- - **Associative retrieval tasks** — the H-Res method specifically excels at retrieval (26% better than LoRA on SQuAD per the paper)
223
-
224
- ### Out-of-scope
225
-
226
- - **Production deployment without validation** — research artifact; evaluate on your specific use case
227
- - **High-stakes decision making** — not intended for medical, legal, or financial advice without human oversight
228
- - **Unsupported languages** — trained exclusively on English data
229
- - **Multimodal tasks** — text-only fine-tune
230
-
231
- ### Limitations
232
-
233
- - **Trained for 1 epoch** on ~35K examples. More data/epochs would improve tool-calling reliability.
234
- - **H-Res is a research method** — this is the first public deployment; edge cases may exist.
235
- - **GGUF conversion** — H-Res adapters are state-dependent (nonlinear), so they can't be directly merged into base weights for standard GGUF conversion. A LoRA-distilled GGUF version is available separately.
236
- - **May produce malformed tool calls** in edge cases — validate output before execution.
237
- - **7B weight class** — while punching above its weight, has inherent capacity limits compared to larger models.
238
-
239
- ---
240
-
241
- ## Training Data
242
-
243
- Six datasets were curated for agentic capability — prioritizing function-calling and tool-use signal over raw instruction volume:
244
-
245
- | Dataset | Type | Samples | Focus |
246
- |---------|------|---------|-------|
247
- | [NousResearch/hermes-function-calling-v1](https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1) | Function calling | 1,893 | Single-turn and multi-turn tool use conversations (MIT) |
248
- | [Salesforce/xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k) | Function calling | 10,000 | Diverse API function calling (sampled from 60K, MIT) |
249
- | [mlabonne/FineTome-100k](https://huggingface.co/datasets/mlabonne/FineTome-100k) | Instruction following | 20,000 | General instruct/chat data (sampled from 100K, MIT) |
250
- | [Salesforce/APIGen-MT-5k](https://huggingface.co/datasets/Salesforce/APIGen-MT-5k) | API generation | 5,000 | Multi-turn API call generation across diverse APIs (MIT) |
251
- | [glaiveai/glaive-function-calling-v2](https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2) | Function calling | 8,000 | Multi-turn tool-use conversations (MIT) |
252
- | [Team-ACE/ToolACE](https://huggingface.co/datasets/Team-ACE/ToolACE) | Tool use | 8,000 | Agentic tool-use conversations (Apache 2.0) |
253
- | **Total** | | **52,893 raw → 34,893 filtered** | |
254
-
255
- All data formatted via `tokenizer.apply_chat_template()` with the Qwen2.5 ChatML template. Examples without a `user` role were filtered. Sequence length capped at 2,048 tokens.
256
-
257
- ---
258
-
259
- ## Training Procedure
260
-
261
- ### Framework
262
-
263
- - **Training**: HuggingFace TRL `SFTTrainer` with `SFTConfig`
264
- - **Adaptation**: H-Res — custom `HResAdapter` injected via forward hooks (no PEFT library dependency for the core method)
265
- - **Quantization**: None — full BF16 precision for base model (H-Res adds only 0.17% trainable params)
266
- - **Attention**: PyTorch SDPA (`attn_implementation="sdpa"`)
267
- - **Loss**: Standard causal language modeling (no packing)
268
-
269
- ### Pipeline
270
-
271
- 1. **Model loading**: BF16 full precision via `AutoModelForCausalLM.from_pretrained()`
272
- 2. **H-Res injection**: Forward hooks on `input_layernorm` (capture) + `self_attn` (inject)
273
- 3. **Base model freeze**: `model.requires_grad_(False)` — only H-Res params trainable
274
- 4. **Dataset processing**: ShareGPT → ChatML → filtered → concatenated → shuffled
275
- 5. **Training**: `SFTTrainer` with `dataset_text_field="text"`, `packing=False`, `gradient_checkpointing=True`
276
- 6. **Export**: `model.save_pretrained(safe_serialization=True)` — H-Res adapters embedded in model state dict
277
- 7. **Upload**: `HfApi.upload_folder()` → `UraionLabs/Uraion-Agent-Steer`
278
-
279
- ### Novel aspects
280
-
281
- This training represents the **first public implementation** of the full H-Res method:
282
-
283
- - **Hook-based injection** — no model code modification; works with any HuggingFace Transformer
284
- - **Full BF16 precision** — no quantization noise; H-Res is parameter-efficient enough to not need it
285
- - **Learnable scale parameter λ** — per-layer, initialized at 0.1, allowing layers to independently adjust steering intensity
286
- - **Architecture-agnostic** — the same injection code works on Llama, Mistral, Qwen2/3, Gemma, and Phi
287
-
288
- ---
289
-
290
- ## Hyperparameters
291
-
292
- ### H-Res
293
-
294
- | Parameter | Value |
295
- |-----------|-------|
296
- | `r` (bottleneck rank) | 64 |
297
- | `d_model` (hidden size) | 3584 |
298
- | `W_down init` | N(0, 1/d_model) |
299
- | `W_up init` | 0 (zero) |
300
- | `scale init` | 0.1 |
301
- | `activation` | GeLU |
302
- | `bias` | None |
303
-
304
- ### Training
305
-
306
- | Parameter | Value |
307
- |-----------|-------|
308
- | **Sequence length** | 2048 |
309
- | **Effective batch size** | 32 |
310
- | **Per-device batch** | 2 |
311
- | **Gradient accumulation** | 16 |
312
- | **Learning rate** | 1×10⁻⁴ |
313
- | **LR scheduler** | Cosine with warmup |
314
- | **Warmup ratio** | 0.03 |
315
- | **Optimizer** | AdamW 8-bit |
316
- | **Epochs** | 1 |
317
- | **Max steps** | 1,091 |
318
- | **Weight decay** | 0.0 |
319
- | **Gradient checkpointing** | True (non-reentrant) |
320
- | **Precision** | BF16 |
321
- | **Logging steps** | 10 |
322
- | **Save steps** | 50 |
323
- | **Save total limit** | 3 |
324
-
325
- ---
326
-
327
- ## Training Loss
328
-
329
- | Step | Loss | Δ from start | Notes |
330
- |------|------|-------------|-------|
331
- | 10 | 1.310 | — | Initial — H-Res scale still ramping |
332
- | 20 | 1.264 | ↓ 3.5% | W_up beginning to activate |
333
- | 50 | 1.013 | ↓ 22.7% | First checkpoint saved; steering field forming |
334
- | 100 | 0.879 | ↓ 32.9% | Rapid convergence phase |
335
- | 200 | 0.741 | ↓ 43.4% | Entering fine-tuning regime |
336
- | 300 | 0.745 | ↓ 43.1% | Stable convergence |
337
- | 400 | 0.699 | ↓ 46.6% | Steady improvement |
338
- | 500 | 0.689 | ↓ 47.4% | Approaching plateau |
339
- | 600 | 0.645 | ↓ 50.8% | Best single-step loss |
340
- | 700 | 0.688 | ↓ 47.5% | Minor oscillation — normal |
341
- | 800 | 0.646 | ↓ 50.7% | Consistent low-loss regime |
342
- | 900 | 0.663 | ↓ 49.4% | Stable |
343
- | 1000 | 0.67 | ↓ 48.9% | Final stretch |
344
- | **1091** | **0.657** | **↓ 49.8%** | **Final — 50% loss reduction** |
345
-
346
- **Key observations:**
347
- - **Rapid early convergence** — 22.7% loss reduction by step 50 (first 4.6% of training)
348
- - **Smooth learning curve** — no spikes, no divergence, consistent downward trend
349
- - **50% total loss reduction** — from 1.310 to 0.657
350
- - **H-Res's zero-initialization advantage** — no "initialization shock" means the model starts from a good place and improves monotonically
351
-
352
- ---
353
-
354
- ## Local Inference Guide
355
-
356
- This model uses **safetensors with H-Res adapters embedded** — no extra adapter files needed. Load it like any standard Transformers model. Below are instructions for every major local inference tool.
357
-
358
- ### Contents
359
- - [Transformers (Python)](#transformers-python) — full quality, recommended
360
- - [vLLM (OpenAI-compatible server)](#vllm-openai-compatible-server) — production serving
361
- - [Unsloth (further fine-tuning)](#unsloth-further-fine-tuning) — continue training
362
- - [Ollama](#ollama) — import from safetensors
363
- - [LM Studio](#lm-studio) — local desktop inference
364
- - [llama.cpp](#llamacpp) — GGUF note
365
- - [text-generation-webui (Oobabooga)](#text-generation-webui-oobabooga)
366
-
367
- ---
368
-
369
- ### Transformers (Python)
370
-
371
- The simplest way — loads H-Res adapters automatically.
372
-
373
- ```python
374
- import torch
375
- from transformers import AutoModelForCausalLM, AutoTokenizer
376
-
377
- model_name = "UraionLabs/Uraion-Agent-Steer"
378
-
379
- tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
380
- model = AutoModelForCausalLM.from_pretrained(
381
- model_name,
382
- torch_dtype=torch.bfloat16,
383
- device_map="auto",
384
- trust_remote_code=True,
385
- )
386
-
387
- # H-Res adapters are embedded — no extra loading needed
388
- messages = [
389
- {"role": "system", "content": "You are Uraion-Agent-Steer, an agent with tool-use capabilities."},
390
- {"role": "user", "content": "What's the weather in Tokyo? Should I bring an umbrella?"},
391
- ]
392
-
393
- text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
394
- inputs = tokenizer(text, return_tensors="pt").to(model.device)
395
-
396
- outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7, top_p=0.95, do_sample=True)
397
- response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
398
- print(response)
399
- ```
400
-
401
- **VRAM requirement:** ~16 GB (BF16). Works on RTX 3090/4090, A4000, A5000, A100, or any 24GB+ consumer GPU.
402
-
403
- **With 12GB GPUs** (RTX 3080/4070): use 8-bit quantization:
404
-
405
- ```python
406
- model = AutoModelForCausalLM.from_pretrained(
407
- model_name,
408
- load_in_8bit=True,
409
- device_map="auto",
410
- trust_remote_code=True,
411
- )
412
- ```
413
-
414
- **With `pipeline` (simpler):**
415
-
416
- ```python
417
- from transformers import pipeline
418
-
419
- pipe = pipeline(
420
- "text-generation",
421
- model="UraionLabs/Uraion-Agent-Steer",
422
- torch_dtype=torch.bfloat16,
423
- device_map="auto",
424
- trust_remote_code=True,
425
- )
426
-
427
- messages = [{"role": "user", "content": "Search for the latest AI research papers."}]
428
- output = pipe(messages, max_new_tokens=512, temperature=0.7, top_p=0.95)
429
- print(output[0]["generated_text"])
430
- ```
431
-
432
- ---
433
-
434
- ### vLLM (OpenAI-compatible server)
435
-
436
- Best for production agent deployments. vLLM loads safetensors directly with full H-Res adapter support.
437
-
438
- ```bash
439
- # Install vLLM
440
- pip install vllm
441
-
442
- # Serve with OpenAI-compatible API
443
- vllm serve UraionLabs/Uraion-Agent-Steer \
444
- --trust-remote-code \
445
- --host 0.0.0.0 \
446
- --port 8000 \
447
- --max-model-len 8192 \
448
- --gpu-memory-utilization 0.90
449
- ```
450
 
451
- **OpenAI client (Python):**
452
-
453
- ```python
454
- from openai import OpenAI
455
-
456
- client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
457
-
458
- response = client.chat.completions.create(
459
- model="UraionLabs/Uraion-Agent-Steer",
460
- messages=[{"role": "user", "content": "What's 2+2?"}],
461
- temperature=0.7,
462
- )
463
- print(response.choices[0].message.content)
464
- ```
465
-
466
- **With tool calling:**
467
-
468
- ```python
469
- tools = [{
470
- "type": "function",
471
- "function": {
472
- "name": "get_weather",
473
- "description": "Get current weather for a location",
474
- "parameters": {
475
- "type": "object",
476
- "properties": {
477
- "location": {"type": "string", "description": "City name"}
478
- },
479
- "required": ["location"]
480
- }
481
- }
482
- }]
483
-
484
- response = client.chat.completions.create(
485
- model="UraionLabs/Uraion-Agent-Steer",
486
- messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
487
- tools=tools,
488
- temperature=0.0,
489
- )
490
- print(response.choices[0].message.tool_calls)
491
- ```
492
-
493
- **VRAM:** ~18 GB for vLLM serving. Recommended: A10, A100, or 24GB consumer GPU (RTX 3090/4090).
494
-
495
- ---
496
-
497
- ### Unsloth (further fine-tuning)
498
-
499
- Continue training Uraion-Agent-Steer with Unsloth for 2× faster, 70% less memory fine-tuning.
500
-
501
- ```python
502
- from unsloth import FastLanguageModel
503
- from unsloth.chat_templates import get_chat_template
504
- import torch
505
-
506
- # Load Uraion-Agent-Steer with Unsloth acceleration
507
- model, tokenizer = FastLanguageModel.from_pretrained(
508
- model_name="UraionLabs/Uraion-Agent-Steer",
509
- max_seq_length=2048,
510
- dtype=torch.bfloat16,
511
- load_in_4bit=True, # 4-bit for further QLoRA training
512
- trust_remote_code=True,
513
- )
514
-
515
- # Apply QLoRA for continued training
516
- model = FastLanguageModel.get_peft_model(
517
- model,
518
- r=32,
519
- lora_alpha=32,
520
- target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
521
- "gate_proj", "up_proj", "down_proj"],
522
- lora_dropout=0,
523
- bias="none",
524
- )
525
-
526
- # Continue training with your data...
527
- # model = ... your training loop ...
528
-
529
- # Save LoRA adapters
530
- model.save_pretrained("uraion-agent-steer-continued")
531
- ```
532
-
533
- > **Note:** The H-Res adapters remain frozen alongside the base model during Unsloth QLoRA training. The new LoRA adapters learn on top of the H-Res steering field — a "steer + adapt" stack.
534
-
535
- **VRAM:** ~8 GB with 4-bit QLoRA via Unsloth.
536
-
537
- ---
538
-
539
- ### Ollama
540
-
541
- Ollama uses GGUF format. Since H-Res adapters can't be directly merged into base weights, you have two options:
542
-
543
- **Option A: Import from safetensors (Ollama 0.3.0+)**
544
-
545
- ```bash
546
- # Create Modelfile
547
- cat > Modelfile << 'EOF'
548
- FROM UraionLabs/Uraion-Agent-Steer
549
- TEMPLATE """{{ if .System }}<|im_start|>system
550
- {{ .System }}<|im_end|>
551
- {{ end }}{{ if .Prompt }}<|im_start|>user
552
- {{ .Prompt }}<|im_end|>
553
- {{ end }}<|im_start|>assistant
554
- """
555
- PARAMETER stop "<|im_start|>"
556
- PARAMETER stop "<|im_end|>"
557
- PARAMETER temperature 0.7
558
- EOF
559
-
560
- # Import into Ollama
561
- ollama create uraion-agent-steer -f Modelfile
562
-
563
- # Run
564
- ollama run uraion-agent-steer
565
- ```
566
-
567
- **Option B: Use the GGUF release (when available)**
568
-
569
- ```bash
570
- # Coming soon — LoRA-distilled GGUF version
571
- ollama run hf.co/UraionLabs/Uraion-Agent-Steer-GGUF:Q4_K_M
572
- ```
573
-
574
- > **Note for Option A:** Ollama's safetensors import loads the full BF16 model (~15 GB download, ~16 GB VRAM). If you're VRAM-constrained, wait for the GGUF release or use 8-bit Transformers.
575
-
576
- ---
577
-
578
- ### LM Studio
579
-
580
- LM Studio works with GGUF models. For safetensors models, use one of these approaches:
581
-
582
- **Approach 1: Wait for GGUF release**
583
-
584
- The LoRA-distilled GGUF version (`UraionLabs/Uraion-Agent-Steer-GGUF`) will be importable directly in LM Studio's model browser. Pick a quant (Q4_K_M recommended) and download.
585
-
586
- **Approach 2: Use MLX (Apple Silicon)**
587
-
588
- If you're on a Mac with Apple Silicon, MLX can load safetensors directly:
589
-
590
- ```bash
591
- pip install mlx mlx-lm
592
-
593
- # Convert to MLX format
594
- mlx_lm.convert --hf-path UraionLabs/Uraion-Agent-Steer --mlx-path ./uraion-agent-steer-mlx
595
-
596
- # Run inference
597
- mlx_lm.generate --model ./uraion-agent-steer-mlx --prompt "What is tool calling?"
598
- ```
599
-
600
- **Approach 3: Use vLLM or llama.cpp server**
601
-
602
- Run the model locally via vLLM (see above), then connect LM Studio to it via the "Local Server" option in LM Studio's settings.
603
-
604
- ---
605
-
606
- ### llama.cpp
607
-
608
- llama.cpp requires GGUF format. Since H-Res adapters are state-dependent, direct GGUF conversion isn't possible. Two paths:
609
-
610
- **Path 1: Use the GGUF-distilled release**
611
-
612
- ```bash
613
- # Coming soon
614
- llama-server -hf UraionLabs/Uraion-Agent-Steer-GGUF:Q4_K_M --host 0.0.0.0 --port 8000
615
- ```
616
-
617
- **Path 2: Use Transformers server + llama.cpp client**
618
-
619
- ```bash
620
- # Server side (Transformers with H-Res, fast)
621
- python -c "
622
- from transformers import AutoModelForCausalLM, AutoTokenizer
623
- import torch
624
- # ... serve with FastAPI or vLLM
625
- "
626
-
627
- # Client side (any OpenAI-compatible client)
628
- curl http://localhost:8000/v1/chat/completions \
629
- -H "Content-Type: application/json" \
630
- -d '{"model":"uraion-agent-steer","messages":[{"role":"user","content":"Hello"}]}'
631
- ```
632
-
633
- ---
634
-
635
- ### text-generation-webui (Oobabooga)
636
-
637
- Load directly in Oobabooga's Transformers loader:
638
-
639
- 1. Go to the **Model** tab
640
- 2. In "Download model or LoRA", enter: `UraionLabs/Uraion-Agent-Steer`
641
- 3. Click **Download**
642
- 4. After download, select the model and set:
643
- - **Loader:** Transformers
644
- - **trust_remote_code:** ✓
645
- - **dtype:** bfloat16
646
- 5. Click **Load**
647
-
648
- For lower VRAM, enable `load_in_8bit` or `load_in_4bit` in the loader settings.
649
-
650
- **Chat template** (if not auto-detected): `chatml` (Qwen2.5 ChatML format).
651
-
652
- ---
653
-
654
- ### VRAM Reference
655
-
656
- | GPU | VRAM | Config | Notes |
657
- |-----|------|--------|-------|
658
- | RTX 4090 (24GB) | 24 GB | BF16 | Full quality, fits comfortably |
659
- | RTX 3090 (24GB) | 24 GB | BF16 | Same as above |
660
- | RTX 4080 (16GB) | 16 GB | BF16 | Tight — use 8-bit for safety |
661
- | RTX 3080 (10GB) | 10 GB | 8-bit | Works with `load_in_8bit=True` |
662
- | RTX 4070 (12GB) | 12 GB | 8-bit | Works with `load_in_8bit=True` |
663
- | A100 (40GB) | 40 GB | BF16 | Full quality, plenty of room |
664
- | A10 (24GB) | 24 GB | BF16 | Full quality |
665
- | T4 (16GB) | 16 GB | 8-bit | Use `load_in_8bit=True` |
666
- | Apple M2/M3 (16GB+) | Unified | MLX | Convert with `mlx_lm.convert` |
667
-
668
- ---
669
-
670
- ## H-Res Adapter Analysis
671
-
672
- After training, we inspected the learned H-Res adapters across all 28 layers:
673
-
674
- | Layer | Scale (λ) | ‖W_up‖ | ‖W_down‖ | Steering activity |
675
- |-------|-----------|--------|----------|-------------------|
676
- | 0 (early) | 0.1001 | 0.0000 | 7.94 | **Silent** — shallow layers don't steer |
677
- | 8 (mid) | 0.1001 | 2.12 | 8.45 | Moderate steering |
678
- | 16 (mid-deep) | 0.1001 | 2.87 | 9.12 | Active steering |
679
- | 24 (deep) | 0.1001 | 3.12 | 9.56 | Strong steering |
680
- | 27 (final) | 0.1001 | **3.72** | **9.69** | **Maximum steering** |
681
-
682
- **Key finding:** Steering intensity increases monotonically with layer depth. Early layers (0–3) have W_up ≈ 0 — the adapter is effectively dormant. Deep layers (20–27) have the strongest steering activity. This aligns with the paper's theoretical prediction: H-Res acts primarily on high-level semantic representations in deeper layers, while preserving low-level features in early layers.
683
-
684
- The scale parameter λ stayed at ~0.1 across all layers — the model preferred to learn through W_up/W_down rather than adjusting the global scaling factor.
685
-
686
- ---
687
-
688
- ## Hardware & Infrastructure
689
-
690
- | Component | Detail |
691
- |-----------|--------|
692
- | **Provisioning** | Google Colab CLI (`colab-cli`) via OAuth2 |
693
- | **GPU** | 1× NVIDIA A100-SXM4-40GB |
694
- | **Runtime** | `colab run --gpu A100 --keep --timeout 28800` |
695
- | **Training time** | ~3 hours (1,091 steps at ~10s/step) |
696
- | **VRAM usage** | ~35 GB (7.6B BF16 base + 12.8M H-Res + activations + optimizer) |
697
- | **Setup** | Self-installing dependencies via pip |
698
- | **Session lifecycle** | `colab run` → auto-execute → `--keep` → training → auto-upload → session release |
699
-
700
- Training dependencies auto-installed on Colab: `transformers>=4.57`, `trl>=0.21`, `datasets`, `accelerate`, `safetensors`, `huggingface_hub`.
701
-
702
- ---
703
-
704
- ## GGUF Availability
705
-
706
- H-Res adapters are **state-dependent** (nonlinear function of the input), so they can't be directly merged into base weights for standard GGUF/llama.cpp conversion. For GGUF inference:
707
-
708
- | Option | How | VRAM | Quality |
709
- |--------|-----|------|---------|
710
- | **[Ollama safetensors import](#ollama)** | `FROM UraionLabs/Uraion-Agent-Steer` in Modelfile | ~16 GB | Full H-Res quality |
711
- | **[MLX conversion](#lm-studio)** | `mlx_lm.convert` on Apple Silicon | ~16 GB unified | Full H-Res quality |
712
- | **LoRA-distilled GGUF** | `UraionLabs/Uraion-Agent-Steer-GGUF` (coming soon) | 4–8 GB | LoRA-approximated |
713
- | **8-bit Transformers** | `load_in_8bit=True` with Transformers | ~8 GB | Near-full quality |
714
-
715
- The LoRA-distilled GGUF release is in progress (Colab GPU quota recovery). For maximum quality TODAY, use Ollama's safetensors import or vLLM.
716
-
717
- ---
718
-
719
- ## Ethical Considerations
720
-
721
- This model is a fine-tune of Qwen2.5-7B-Instruct and inherits its base capabilities and biases:
722
-
723
- - Training data includes user-generated content from HuggingFace datasets, which may contain biases.
724
- - Function-calling capabilities could automate actions without human oversight — always validate tool calls before execution.
725
- - The model has not undergone safety alignment beyond the base model's existing safeguards.
726
- - The H-Res method is novel — long-term behavior and failure modes are still being studied.
727
- - This is a **research-stage artifact** from Uraion Labs. We are a systems research lab, not a product company. Use accordingly.
728
-
729
- ---
730
-
731
- ## Citations
732
-
733
- ### H-Res (Parallel Manifold Steering)
734
-
735
- ```bibtex
736
- @article{awadhiya2026parallel,
737
- title={Parallel Manifold Steering: Efficient Adaptation of Large
738
- Associative Memories via Residual Energy Shaping},
739
- author={Awadhiya, Kanishk},
740
- journal={ICLR Workshop on New Frontiers in Associative Memory},
741
- year={2026},
742
- url={https://arxiv.org/abs/2606.24396}
743
- }
744
- ```
745
-
746
- ### Uraion-Agent-Steer
747
-
748
- ```bibtex
749
- @software{uraion-agent-steer,
750
- title={Uraion-Agent-Steer: Agentic Model via Hierarchical Residual Steering},
751
- author={Uraion Labs},
752
- year={2026},
753
- url={https://huggingface.co/UraionLabs/Uraion-Agent-Steer}
754
- }
755
- ```
756
-
757
- ### Qwen2.5
758
-
759
- ```bibtex
760
- @misc{qwen2.5,
761
- title={Qwen2.5: A Party of Foundation Models},
762
- author={Qwen Team},
763
- year={2025},
764
- publisher={GitHub},
765
- url={https://github.com/QwenLM/Qwen2.5}
766
- }
767
- ```
768
-
769
- ### TRL
770
-
771
- ```bibtex
772
- @software{vonwerra2020trl,
773
- title={{TRL: Transformers Reinforcement Learning}},
774
- author={von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and
775
- Beeching, Edward and Thrush, Tristan and Lambert, Nathan and
776
- Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
777
- license={Apache-2.0},
778
- url={https://github.com/huggingface/trl},
779
- year={2020}
780
- }
781
- ```
782
-
783
- ### Datasets
784
-
785
- ```bibtex
786
- @misc{hermesfc,
787
- title={NousResearch Hermes Function Calling},
788
- author={Nous Research},
789
- year={2024},
790
- url={https://huggingface.co/datasets/NousResearch/hermes-function-calling-v1}
791
- }
792
-
793
- @misc{xlam2024,
794
- title={xLAM: A Family of Large Action Models},
795
- author={Salesforce AI Research},
796
- year={2024},
797
- url={https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k}
798
- }
799
-
800
- @misc{finetome2024,
801
- title={FineTome-100k: A Curated Instruction Tuning Dataset},
802
- author={Labonne, Maxime},
803
- year={2024},
804
- url={https://huggingface.co/datasets/mlabonne/FineTome-100k}
805
- }
806
-
807
- @misc{apigen2024,
808
- title={APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets},
809
- author={Salesforce AI Research},
810
- year={2024},
811
- url={https://huggingface.co/datasets/Salesforce/APIGen-MT-5k}
812
- }
813
-
814
- @misc{glaivefc,
815
- title={Glaive Function Calling v2},
816
- author={Glaive AI},
817
- year={2024},
818
- url={https://huggingface.co/datasets/glaiveai/glaive-function-calling-v2}
819
- }
820
-
821
- @misc{toolace2025,
822
- title={ToolACE: Winning the Points of LLM Function Calling},
823
- author={Team ACE},
824
- year={2025},
825
- url={https://huggingface.co/datasets/Team-ACE/ToolACE}
826
- }
827
- ```
828
-
829
- ---
830
 
831
- <p align="center">
832
- <img src="https://uraionlabs.com/public/icons/icon-32.png" alt="" width="24" height="24">
833
- </p>
834
 
835
- <p align="center" style="font-family: 'Inter', sans-serif; font-size: 0.8rem; color: #8A8478;">
836
- <strong style="color: #F7F4ED;">Uraion Labs</strong> Foundational systems research.
837
- <br>
838
- <a href="https://uraionlabs.com" style="color: #E45A1A;">uraionlabs.com</a>
839
- <br><br>
840
- <em style="color: #6F6A61;">
841
- Intelligence is a systems problem.
842
- </em>
843
- <br>
844
- Licensed under <a href="https://www.apache.org/licenses/LICENSE-2.0" style="color: #E45A1A;">Apache 2.0</a>.
845
- </p>
 
7
  - en
8
  pipeline_tag: text-generation
9
  tags:
10
+ - legacy
11
+ - unsupported
12
+ - not-for-production
13
  - h-res
14
+ - tool-calling
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
15
  ---
16
 
17
+ # Legacy / unsupported — H-Res runtime claims withdrawn
 
 
 
 
 
18
 
19
+ **Status as of 2026-07-27:** this repository is retained for provenance. It is not an active Uraion
20
+ Labs product, is not used by FinStruct, and should not be deployed as an H-Res model.
 
 
 
 
 
21
 
22
+ The audit found 84 `model.layers.*.hres.*` tensors in `model.safetensors`, but `config.json` declares
23
+ the stock `Qwen2ForCausalLM` architecture and contains no `auto_map`, custom model class, or loader
24
+ that instantiates H-Res modules. The repository also contains no custom modeling code. Standard
25
+ Transformers and vLLM loading therefore has no runtime module that consumes those tensors. Prior
26
+ claims that H-Res loads automatically or has full vLLM support are withdrawn.
 
 
 
 
27
 
28
+ Additional gaps:
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
29
 
30
+ - no held-out tool-use scores, raw predictions, base-model comparison, or reproducible evaluation;
31
+ - training loss is not evidence of downstream capability;
32
+ - the separate GGUF repository explicitly omits the H-Res tensors and is not a faithful H-Res
33
+ deployment.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
 
35
+ The full pre-audit card is preserved in [`LEGACY_CARD.md`](LEGACY_CARD.md). The files remain public
36
+ for inspection, not as a supported model release.
 
37
 
38
+ Current Uraion Labs work is [FinStruct](https://github.com/arnavprabhu/uraion-finstruct): auditable,
39
+ local-first SEC filing extraction with evidence-linked benchmarks. See
40
+ [uraionlabs.com](https://uraionlabs.com).