SmolLM2-360M-Instruct-heretic

RACER IS OP

A decensored variant of HuggingFaceTB/SmolLM2-360M-Instruct, produced with Heretic v1.4.0 (directional ablation / "abliteration"). Refusal behaviour is suppressed via targeted weight edits to the attention output and MLP down-projections rather than fine-tuning, so the base model's instruction-following and code ability are left intact.

Who this is for: anyone who wants the smallest useful uncensored assistant in the collection. At 362M parameters and a 0.23 GB Q4_K_M, this runs in RAM on a phone, a Raspberry Pi 5, or a low-RAM server — and still produces coherent, instruction-following English. Good for classification, routing, reformatting, short agentic loops, and as a draft-model in a larger pipeline.

This model was already barely refusing. The base SmolLM2-360M-Instruct refused only 8 of 100 harmful prompts before abliteration — it is one of the least refusal-heavy instruct models in the 360M class, because its SFT/DPO post-training is deliberately lightweight. Abliteration took that to 4/100 at a KL divergence of just 0.0154, the lowest in this batch and among the lowest in the whole collection. In plain terms: there was very little refusal to remove, and the edit barely moved the model. Expect the base model's behaviour almost fully intact, just slightly more compliant.

Runs anywhere

At this size there is no VRAM ladder to speak of — pick a quant by how much RAM and bandwidth your device has:

Device Recommended quant Weights
Phone / browser (4 GB RAM) Q4_K_M ~0.25 GB
Raspberry Pi 5 (4 GB) Q4_K_M ~0.25 GB
Laptop (8 GB, Apple Silicon) Q8_0 ~0.36 GB
Server / multi-tenant Q6_K ~0.34 GB

Weights only, at this model's 362M native size. The Q2_K tier (~0.20 GB) is there if you are truly RAM-bound, but at 362M the model is small enough that Q4_K_M already fits almost anywhere. Context is the real cost: add roughly 16 MB per 8K tokens of KV cache.

Abliteration parameters

Trial 137 of a 200-trial Heretic run (seed 318129666).

Parameter Value
direction_index 27.04
attn.o_proj.max_weight 1.32
attn.o_proj.max_weight_position 29.44
attn.o_proj.min_weight 0.46
attn.o_proj.min_weight_distance 16.37
mlp.down_proj.max_weight 1.19
mlp.down_proj.max_weight_position 30.27
mlp.down_proj.min_weight 0.22
mlp.down_proj.min_weight_distance 8.14

Performance

Metric This model Original model (HuggingFaceTB/SmolLM2-360M-Instruct)
KL divergence 0.0154 0 (by definition)
Refusals 4/100 8/100

Why abliteration instead of fine-tuning

Fine-tuning a "helpful" persona on top of RLHF'd refusals fights the base model's training and tends to degrade coherence. Abliteration instead finds and edits the specific weight directions responsible for refusal, leaving the rest of the network (and its capabilities) untouched. See the Heretic repo and the original abliteration writeup for the mechanism.

Made with ❤️ by RACER IS OP — follow for more uncensored models

Files

Safetensors

File Size
model.safetensors 0.67 GB

BF16, 362M parameters, tied word embeddings. The reproduce/ directory carries the full Heretic recipe — config.toml, requirements.txt, the Optuna study journal, and SHA-256 sums — so this exact model can be regenerated bit-for-bit. Reproduce it with heretic --reproduce reproduce/reproduce.json.

GGUF quantizations

Full quantization set (14 quants + F16) produced with llama.cpp.

File Format Size
SmolLM2-360M-Instruct-heretic-F16.gguf GGUF F16 0.68 GB
SmolLM2-360M-Instruct-heretic-Q2_K.gguf GGUF Q2_K 0.20 GB
SmolLM2-360M-Instruct-heretic-IQ3_S.gguf GGUF IQ3_S 0.20 GB
SmolLM2-360M-Instruct-heretic-Q3_K_S.gguf GGUF Q3_K_S 0.20 GB
SmolLM2-360M-Instruct-heretic-Q3_K_M.gguf GGUF Q3_K_M 0.22 GB
SmolLM2-360M-Instruct-heretic-Q3_K_L.gguf GGUF Q3_K_L 0.23 GB
SmolLM2-360M-Instruct-heretic-IQ4_XS.gguf GGUF IQ4_XS 0.21 GB
SmolLM2-360M-Instruct-heretic-Q4_K_S.gguf GGUF Q4_K_S 0.24 GB
SmolLM2-360M-Instruct-heretic-Q4_0.gguf GGUF Q4_0 0.21 GB
SmolLM2-360M-Instruct-heretic-Q4_1.gguf GGUF Q4_1 0.23 GB
SmolLM2-360M-Instruct-heretic-Q4_K_M.gguf GGUF Q4_K_M 0.25 GB
SmolLM2-360M-Instruct-heretic-Q5_K_S.gguf GGUF Q5_K_S 0.26 GB
SmolLM2-360M-Instruct-heretic-Q5_K_M.gguf GGUF Q5_K_M 0.27 GB
SmolLM2-360M-Instruct-heretic-Q6_K.gguf GGUF Q6_K 0.34 GB
SmolLM2-360M-Instruct-heretic-Q8_0.gguf GGUF Q8_0 0.36 GB

Standard Llama architecture — loads natively in llama.cpp / Ollama / LM Studio / Jan.

Run llama serve -hf saidutta69/SmolLM2-360M-Instruct-heretic to pull the default quant.

Quickstart

# llama.cpp
llama serve -hf saidutta69/SmolLM2-360M-Instruct-heretic
# transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "saidutta69/SmolLM2-360M-Instruct-heretic"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype="auto",
    device_map="auto",
)

messages = [{"role": "user", "content": "Classify this ticket as billing, bug, or other: 'the invoice PDF won't open'"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Also runnable via Ollama, LM Studio, Jan, vLLM, SGLang — see the "Use this model" widget above for copy-paste commands.

Model details

Architecture LlamaForCausalLM
Parameters 362M (tied embeddings)
Layers / heads 32 layers, 15 attention heads, 5 KV heads (GQA)
Head dim / hidden / intermediate 64 / 960 / 2560
Position embedding RoPE, theta = 100,000
Context length 8,192
Vocab 49,152
Precision bfloat16
Languages English
Base model HuggingFaceTB/SmolLM2-360M-Instruct

Responsible use

Refusal suppression is deliberate and works as intended: this model will comply with requests the base model would refuse, including some it shouldn't. There is no safety filtering layered on top. You are responsible for how you deploy it — don't put this behind an unmoderated public-facing endpoint serving third parties. Note that at 362M it has real knowledge limits and will hallucinate on anything beyond simple tasks; verify its output rather than trusting it.

License

Inherits the apache-2.0 license from the base model.

Related

Downloads last month
74
Safetensors
Model size
0.4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for saidutta69/SmolLM2-360M-Instruct-heretic

Quantized
(124)
this model

Collection including saidutta69/SmolLM2-360M-Instruct-heretic

Paper for saidutta69/SmolLM2-360M-Instruct-heretic