helpfulness-qwen-2.5-0.5b-l8

Helpfulness direction: helpful minus evasive response prompts. Positive coefficients steer toward more helpful completions.

Vector card

Field Value
Model Qwen/Qwen2.5-0.5B-Instruct
Hook site blocks.8.hook_resid_post
Direction norm (raw) 5.5438
Source paper Arditi et al. 2024
Source run —

How to use

mech apply-steering --vector helpfulness-qwen-2.5-0.5b-l8 --coefficient 3.0 --prompt "Your prompt here"
from mech_interp.steering.registry import load_steering_vector

direction, metadata = load_steering_vector("helpfulness-qwen-2.5-0.5b-l8")

Provenance

  • Extraction method: mean-difference (Arditi/RepE)
  • Source: Arditi et al. 2024
  • Platform repo: masonwyatt23/sv-helpfulness-qwen-2-5-0-5b

Files

File Description
direction.safetensors Unit-norm direction tensor under key "direction"
direction.safetensors.json Extraction metadata sidecar
bundle_metadata.json Machine-readable bundle manifest
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support