helpfulness-qwen-2.5-0.5b-l8
Helpfulness direction: helpful minus evasive response prompts. Positive coefficients steer toward more helpful completions.
Vector card
| Field | Value |
|---|---|
| Model | Qwen/Qwen2.5-0.5B-Instruct |
| Hook site | blocks.8.hook_resid_post |
| Direction norm (raw) | 5.5438 |
| Source paper | Arditi et al. 2024 |
| Source run | — |
How to use
mech apply-steering --vector helpfulness-qwen-2.5-0.5b-l8 --coefficient 3.0 --prompt "Your prompt here"
from mech_interp.steering.registry import load_steering_vector
direction, metadata = load_steering_vector("helpfulness-qwen-2.5-0.5b-l8")
Provenance
- Extraction method: mean-difference (Arditi/RepE)
- Source: Arditi et al. 2024
- Platform repo:
masonwyatt23/sv-helpfulness-qwen-2-5-0-5b
Files
| File | Description |
|---|---|
direction.safetensors |
Unit-norm direction tensor under key "direction" |
direction.safetensors.json |
Extraction metadata sidecar |
bundle_metadata.json |
Machine-readable bundle manifest |
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support