🧠 Gemma 4 E4B Ultra Uncensored Heretic — Unsloth QLoRA (r=16)

Model ID: hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290

A lightweight LoRA adapter (rank 16) fine-tuned with QLoRA on llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic using Unsloth. Trained on reasoning traces distilled from Claude Opus for 3 epochs (1,335 steps) — just 73 MB (F16 GGUF) / 162 MB (safetensors).


📊 Training Summary

Metric Value
Base Model llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic
Dataset lordx64/reasoning-distill-claude-opus-4-7-max
Training Type QLoRA (load_in_4bit: true)
Epochs 3.0 (1,335 steps)
Train Loss 13.32 → 2.22 (↓ 83%)
Final Step Loss 1.28 (step 1,335)
Eval Loss 2.88
Learning Rate 2e-4 → cosine decay → 1.5e-7
Total Tokens Seen 1,608,612
Training Time ~7.2 hours
Hardware NVIDIA RTX 4060 Ti 16GB
CUDA / Driver 13.0 / 580.126.09

🛠️ LoRA Configuration

Parameter Value
Rank (r) 16
Alpha 16 (lora_alpha / r = 1.0)
Dropout 0.0
Target Modules q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj
Bias none
PEFT Version 0.18.1

Additional Training Settings

Parameter Value
Batch Size 1 (effective 18 with gradient accumulation)
Max Seq Length 512
Optimizer adamw_bnb_8bit
LR Scheduler linear
Warmup Steps 5
Weight Decay 0.001
Random Seed 3407
Sequence Packing ✅ enabled
Gradient Checkpointing unsloth

📈 Loss Curve

Step     0:  13.32  ████████████████████████████████████
Step   300:   2.13  ██████▍
Step   600:   1.64  █████
Step   900:   1.59  ████▊
Step  1200:   1.61  ████▉
Step  1335:   1.28  ███▉  ← FINAL

Training converged smoothly from initial loss ~13.3 down to 2.22 (average). The final training step achieved 1.28 loss. Eval loss at 2.88 suggests moderate overfitting common with small LoRA adapters — expected and acceptable for the adapter size (73 MB).


🚀 How to Use

Option 1: PEFT (PyTorch)

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model = "llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic"
lora_path  = "hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290"

model = AutoModelForCausalLM.from_pretrained(
    base_model,
    torch_dtype="auto",
    device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(base_model)

model.load_adapter(lora_path, adapter_name="lora")
model.set_active_adapter("lora")

messages = [{"role": "user", "content": "Explain the theory of relativity simply."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Option 2: Unsloth (Recommended — 2× faster, uses less VRAM)

from unsloth import FastLanguageModel

model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290",
    max_seq_length=2048,
    load_in_4bit=True,  # or False for BF16
)
FastLanguageModel.for_inference(model)

messages = [{"role": "user", "content": "Write a poem about AI in Thai."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to("cuda")

output = model.generate(**inputs, max_new_tokens=512, temperature=0.7)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Option 3: GGUF (llama.cpp)

Download gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf (73 MB) from the gguf/ directory. Then use with your existing base model GGUF:

# Serve with llama.cpp LoRA support (llama-server with --lora)
llama-server \
  -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \
  --lora gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf \
  --lora-scaled gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf 1.0

📦 Model Files

File Format Size
adapter_model.safetensors PEFT safetensors 162 MB
adapter_config.json PEFT config 1.3 KB
gguf/gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf GGUF LoRA (F16) 73.4 MB
tokenizer.json Tokenizer 31 MB
trainer_state.json Training log 394 KB

⚠️ Limitations & Bias

  • LoRA Adapter only — you need the base model llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic loaded separately (no merged weights included).
  • Eval gap — train loss 1.28 vs eval loss 2.88 indicates some overfitting on the Claude reasoning dataset.
  • Limited context — trained with max_seq_length=512 and sequence packing. Performance on very long reasoning chains may degrade.
  • Uncensored — the base model has minimal alignment filtering, so outputs may be more creative/unfiltered than standard models.
  • No formal benchmarks — MMLU, GSM8K, etc. not evaluated. Loss-based convergence suggests improved reasoning over the base model.
  • Single GPU — trained on one RTX 4060 Ti 16GB with batch_size=1. Larger-scale generalization may vary.

📚 Citation

If you use this model in research or production, please credit:

@misc{gemma4-e4b-heretic-lora-2025,
  author = {UKA (Hermes Agent)},
  title = {Gemma 4 E4B Ultra Uncensored Heretic — Unsloth QLoRA Fine-tuned
           on Claude Reasoning Distill},
  year = {2025},
  publisher = {Hugging Face},
  howpublished = {\\url{https://huggingface.co/hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290}},
  note = {Trained with Unsloth on RTX 4060 Ti. Base model by llmfan46.}
}

🔗 Links


📝 Changelog

Date Event
2026-05-05 09:33 Training started on llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic
2026-05-05 16:46 Training completed — checkpoint-1335 (3 epochs, 1,608,612 tokens)
2026-05-05 22:43 GGUF LoRA exported — 73.4 MB F16

Made with 💜 by UKA · Powered by Unsloth & NVIDIA RTX 4060 Ti

Downloads last month
30
GGUF
Model size
36.7M params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290