You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

nla-qwen3-4b-L24-av-sft

The AV (activation verbalizer, vector โ†’ text) half of a Natural Language Autoencoder (NLA) pair, fine-tuned from Qwen/Qwen3-4B. The other half is zaemyung/nla-qwen3-4b-L24-ar-sft; both are released together and are intended to be used as a pair.

NLA pairs are interpretability tools: the AV (activation verbalizer) maps a hidden-state vector to a natural-language description; the AR (activation reconstructor) maps that description back to a vector. Together they let you read out what a residual-stream activation "means" and measure how much of it the description captured. These checkpoints are not useful as general-purpose language models โ€” the fine-tuning repurposes them entirely for activation decoding.

Usage

See the nla-mnlp README for the full recipe (SGLang launch, NLAClient/NLACritic, embedding-injection details).

Citation

@article{frasertaliente2026nla,
  author  = {Fraser-Taliente, Kit and Kantamneni, Subhash and Ong, Euan and Mossing, Dan and Lu, Christina and Bogdan, Paul C. and Ameisen, Emmanuel and Chen, James and Kishylau, Dzmitry and Pearce, Adam and Tarng, Julius and Wu, Alex and Wu, Jeff and Zhang, Yang and Ziegler, Daniel M. and Hubinger, Evan and Batson, Joshua and Lindsey, Jack and Zimmerman, Samuel and Marks, Samuel},
  title   = {Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations},
  journal = {Transformer Circuits Thread},
  year    = {2026},
  url     = {https://transformer-circuits.pub/2026/nla/index.html}
}

License & use restrictions

This model is fine-tuned from Qwen/Qwen2.5-7B-Instruct and is distributed under the Apache License 2.0. See LICENSE in this repository.

Training data attribution

The fine-tuning data was derived from one public dataset:

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for minnesotanlp/nla-qwen3-4b-L24-av-sft

Finetuned
Qwen/Qwen3-4B
Finetuned
(949)
this model