Text Generation
Safetensors
GGUF
English
Thai
gemma-4
unsloth
lora
qlora
sft
heretic
uncensored
llmfan46
reasoning
claude-distill
conversational
Instructions to use hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Local Apps Settings
- Unsloth Desktop
| language: | |
| - en | |
| - th | |
| license: apache-2.0 | |
| tags: | |
| - gemma-4 | |
| - unsloth | |
| - lora | |
| - qlora | |
| - sft | |
| - text-generation | |
| - heretic | |
| - uncensored | |
| - llmfan46 | |
| - reasoning | |
| - claude-distill | |
| base_model: llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic | |
| pipeline_tag: text-generation | |
| inference: false | |
| # 🧠 Gemma 4 E4B Ultra Uncensored Heretic — Unsloth QLoRA (r=16) | |
| **Model ID:** `hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290` | |
| A lightweight **LoRA adapter** (rank 16) fine-tuned with **QLoRA** on | |
| [llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic](https://huggingface.co/llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic) | |
| using [Unsloth](https://github.com/unslothai/unsloth). | |
| Trained on reasoning traces distilled from Claude Opus for 3 epochs | |
| (1,335 steps) — just **73 MB** (F16 GGUF) / **162 MB** (safetensors). | |
| --- | |
| ## 📊 Training Summary | |
| | Metric | Value | | |
| |--------|-------| | |
| | **Base Model** | [llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic](https://huggingface.co/llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic) | | |
| | **Dataset** | [lordx64/reasoning-distill-claude-opus-4-7-max](https://huggingface.co/datasets/lordx64/reasoning-distill-claude-opus-4-7-max) | | |
| | **Training Type** | QLoRA (`load_in_4bit: true`) | | |
| | **Epochs** | 3.0 (1,335 steps) | | |
| | **Train Loss** | 13.32 → **2.22** (↓ 83%) | | |
| | **Final Step Loss** | **1.28** (step 1,335) | | |
| | **Eval Loss** | 2.88 | | |
| | **Learning Rate** | `2e-4` → cosine decay → `1.5e-7` | | |
| | **Total Tokens Seen** | 1,608,612 | | |
| | **Training Time** | ~7.2 hours | | |
| | **Hardware** | NVIDIA RTX 4060 Ti 16GB | | |
| | **CUDA / Driver** | 13.0 / 580.126.09 | | |
| --- | |
| ## 🛠️ LoRA Configuration | |
| | Parameter | Value | | |
| |-----------|-------| | |
| | **Rank (`r`)** | 16 | | |
| | **Alpha** | 16 (`lora_alpha / r = 1.0`) | | |
| | **Dropout** | 0.0 | | |
| | **Target Modules** | `q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj` | | |
| | **Bias** | `none` | | |
| | **PEFT Version** | 0.18.1 | | |
| ### Additional Training Settings | |
| | Parameter | Value | | |
| |-----------|-------| | |
| | **Batch Size** | 1 (effective 18 with gradient accumulation) | | |
| | **Max Seq Length** | 512 | | |
| | **Optimizer** | `adamw_bnb_8bit` | | |
| | **LR Scheduler** | `linear` | | |
| | **Warmup Steps** | 5 | | |
| | **Weight Decay** | 0.001 | | |
| | **Random Seed** | 3407 | | |
| | **Sequence Packing** | ✅ enabled | | |
| | **Gradient Checkpointing** | `unsloth` | | |
| --- | |
| ## 📈 Loss Curve | |
| ``` | |
| Step 0: 13.32 ████████████████████████████████████ | |
| Step 300: 2.13 ██████▍ | |
| Step 600: 1.64 █████ | |
| Step 900: 1.59 ████▊ | |
| Step 1200: 1.61 ████▉ | |
| Step 1335: 1.28 ███▉ ← FINAL | |
| ``` | |
| Training converged smoothly from initial loss `~13.3` down to `2.22` (average). | |
| The final training step achieved **1.28** loss. Eval loss at 2.88 suggests moderate | |
| overfitting common with small LoRA adapters — expected and acceptable for the | |
| adapter size (73 MB). | |
| --- | |
| ## 🚀 How to Use | |
| ### Option 1: PEFT (PyTorch) | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| from peft import PeftModel | |
| base_model = "llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic" | |
| lora_path = "hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290" | |
| model = AutoModelForCausalLM.from_pretrained( | |
| base_model, | |
| torch_dtype="auto", | |
| device_map="auto" | |
| ) | |
| tokenizer = AutoTokenizer.from_pretrained(base_model) | |
| model.load_adapter(lora_path, adapter_name="lora") | |
| model.set_active_adapter("lora") | |
| messages = [{"role": "user", "content": "Explain the theory of relativity simply."}] | |
| text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = tokenizer(text, return_tensors="pt").to(model.device) | |
| outputs = model.generate(**inputs, max_new_tokens=512, temperature=0.7) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| ### Option 2: Unsloth (Recommended — 2× faster, uses less VRAM) | |
| ```python | |
| from unsloth import FastLanguageModel | |
| model, tokenizer = FastLanguageModel.from_pretrained( | |
| model_name="hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290", | |
| max_seq_length=2048, | |
| load_in_4bit=True, # or False for BF16 | |
| ) | |
| FastLanguageModel.for_inference(model) | |
| messages = [{"role": "user", "content": "Write a poem about AI in Thai."}] | |
| text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True) | |
| inputs = tokenizer(text, return_tensors="pt").to("cuda") | |
| output = model.generate(**inputs, max_new_tokens=512, temperature=0.7) | |
| print(tokenizer.decode(output[0], skip_special_tokens=True)) | |
| ``` | |
| ### Option 3: GGUF (llama.cpp) | |
| Download `gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf` (73 MB) from the `gguf/` | |
| directory. Then use with your existing base model GGUF: | |
| ```bash | |
| # Serve with llama.cpp LoRA support (llama-server with --lora) | |
| llama-server \ | |
| -m gemma-4-E4B-it-ultra-uncensored-heretic-Q6_K.gguf \ | |
| --lora gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf \ | |
| --lora-scaled gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf 1.0 | |
| ``` | |
| --- | |
| ## 📦 Model Files | |
| | File | Format | Size | | |
| |------|--------|------| | |
| | `adapter_model.safetensors` | PEFT safetensors | 162 MB | | |
| | `adapter_config.json` | PEFT config | 1.3 KB | | |
| | `gguf/gemma-4-E4B-uncensored-heretic-lora-r16.f16.gguf` | GGUF LoRA (F16) | **73.4 MB** | | |
| | `tokenizer.json` | Tokenizer | 31 MB | | |
| | `trainer_state.json` | Training log | 394 KB | | |
| --- | |
| ## ⚠️ Limitations & Bias | |
| - **LoRA Adapter only** — you need the base model | |
| [`llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic`](https://huggingface.co/llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic) | |
| loaded separately (no merged weights included). | |
| - **Eval gap** — train loss 1.28 vs eval loss 2.88 indicates some overfitting on the | |
| Claude reasoning dataset. | |
| - **Limited context** — trained with `max_seq_length=512` and sequence packing. | |
| Performance on very long reasoning chains may degrade. | |
| - **Uncensored** — the base model has minimal alignment filtering, so outputs | |
| may be more creative/unfiltered than standard models. | |
| - **No formal benchmarks** — MMLU, GSM8K, etc. not evaluated. Loss-based | |
| convergence suggests improved reasoning over the base model. | |
| - **Single GPU** — trained on one RTX 4060 Ti 16GB with `batch_size=1`. | |
| Larger-scale generalization may vary. | |
| --- | |
| ## 📚 Citation | |
| If you use this model in research or production, please credit: | |
| ```bibtex | |
| @misc{gemma4-e4b-heretic-lora-2025, | |
| author = {UKA (Hermes Agent)}, | |
| title = {Gemma 4 E4B Ultra Uncensored Heretic — Unsloth QLoRA Fine-tuned | |
| on Claude Reasoning Distill}, | |
| year = {2025}, | |
| publisher = {Hugging Face}, | |
| howpublished = {\\url{https://huggingface.co/hotdogs/gemma4-E4B-heretic_claude4.7-reasoning_lora-r16-step1290}}, | |
| note = {Trained with Unsloth on RTX 4060 Ti. Base model by llmfan46.} | |
| } | |
| ``` | |
| --- | |
| ## 🔗 Links | |
| | Resource | URL | | |
| |----------|-----| | |
| | **Base Model** | [llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic](https://huggingface.co/llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic) | | |
| | **Dataset** | [lordx64/reasoning-distill-claude-opus-4-7-max](https://huggingface.co/datasets/lordx64/reasoning-distill-claude-opus-4-7-max) | | |
| | **Unsloth** | [github.com/unslothai/unsloth](https://github.com/unslothai/unsloth) | | |
| | **PEFT** | [huggingface.co/docs/peft](https://huggingface.co/docs/peft) | | |
| | **GGUF LoRA Guide** | [llama.cpp LoRA](https://github.com/ggml-org/llama.cpp/discussions/11379) | | |
| --- | |
| ## 📝 Changelog | |
| | Date | Event | | |
| |------|-------| | |
| | 2026-05-05 09:33 | Training started on `llmfan46/gemma-4-E4B-it-ultra-uncensored-heretic` | | |
| | 2026-05-05 16:46 | Training completed — checkpoint-1335 (3 epochs, 1,608,612 tokens) | | |
| | 2026-05-05 22:43 | GGUF LoRA exported — 73.4 MB F16 | | |
| --- | |
| *Made with 💜 by UKA · Powered by Unsloth & NVIDIA RTX 4060 Ti* | |