Instructions to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M # Run inference directly in the terminal: llama cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M # Run inference directly in the terminal: llama cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Use Docker
docker model run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with Ollama:
ollama run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
- Unsloth Desktop
- Pi
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with Docker Model Runner:
docker model run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
- Lemonade
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Run and chat with the model
lemonade run user.Hemmingway-1-Heretic-MTP-GGUF-V2-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF-V2:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
⚠️ Superseded — use V3
Hemmingway-1-Heretic-MTP-V3-Final-GGUF stays decensored with thinking on (2 vs 60–87 per 100). Keep V2 only for thinking-off serving. This repo remains for reference only.
Hemmingway-1 Heretic V2 (GGUF, MTP) - Superseded — use V3
V2 is the second release of our decensored Hemmingway-1. V1 (Hemmingway-1-Heretic-MTP-GGUF, trial 106) used the ARA branch of Heretic. V2 (trial 208) uses the traditional Heretic method with a warm-started search. Both releases keep Hemmingway-1's writing ability as their goal. V2 does the job better on every measure we ran.
Base model: Altworld/Hemmingway-1 — 27B, Qwen3.8 lineage, 64 decoder layers plus one MTP block, 262,144-token context.
Why V2 replaces V1
V1 removed refusals but paid for it in instruction following. V2 cuts that cost by roughly seven times and removes more refusals.
| Measure | Original | V1 (trial 106) | V2 (trial 208) |
|---|---|---|---|
| Refusals, non-thinking (Heretic test) | 97/100 | 7/100 | 4/100 |
| KL divergence from original | 0 | 0.0576 | 0.0112 |
| Constraint adherence, thinking off | 91.7% | 33.3% | 83.3% |
| Response length, thinking on | baseline | −74.5% | −47.1% |
| Constraint adherence, thinking on | 58.3% | 25.0% | 50.0% |
| EQ-Bench (±1.4) | 83.17 | 83.50 | 83.22 |
| HellaSwag acc / acc_norm (100 subset) | 59% / 76% | 59% / 76% | 59% / 76% |
EQ-Bench and HellaSwag tie with the original in all three cases. The difference is the middle rows: V2 stays close to the original where V1 drifted hard. In our 12-prompt creative suite, V1 ignored stated constraints in two of three responses; V2 ignores them in one of six.
If you use V1 today, switch. Nothing we measured favors it.
Recommended use
- Run with thinking off, or at low reasoning effort. The effort setting controls how much the model refuses; see Thinking mode.
- System prompt: any neutral prompt. We optimized and evaluated with
You are a helpful assistant. - Sampling: temperature 0.7, min-p 0.1 (top-p 0.95, top-k 40 also work).
GGUF builds
All quants include the MTP block:
| File | Size |
|---|---|
| BF16 | 54.7 GB |
| Q8_0 | 29.0 GB |
| Q6_K | 22.4 GB |
| Q5_K_M | 19.5 GB |
| Q4_K_M | 16.8 GB |
Run with llama.cpp
Use a recent build with Qwen3.5 and MTP support. The chat template is embedded; add --jinja.
llama-server -m Hemmingway-1-Heretic-V2-Q5_K_M.gguf \
--jinja -fa \
-c 32768 --temp 0.7 --min-p 0.1 --top-p 0.95 --top-k 40 \
--spec-type draft-mtp --spec-draft-n-max 2
For long contexts, quantize the KV cache to fit memory:
--cache-type-k q8_0 --cache-type-v q8_0
MTP speculative decoding
The MTP head is unchanged from the original model. The 64 main layers are V2. Because the target changed but the draft head did not, acceptance rates may differ from the original pairing, and we have not benchmarked them. --spec-draft-n-max 2 is the conservative setting; compare against 3 for your workload.
Thinking mode
The abliteration removes reflexive refusals. It does not change what the model concludes after reasoning. Final-answer refusals on the same 100 harmful prompts:
| Condition | Median think length | Refusals |
|---|---|---|
| Thinking off | — | 4–23/100* |
| Thinking, low effort | ~120 tokens | 60/100 |
| Thinking, medium effort | ~210 tokens | 87/100 |
| Thinking, xhigh (default) | ~2,400 tokens | ~95/100 |
*4/100 is Heretic's internal test; 23/100 is the stock chat template with a 100-token window.
Two consequences:
- To get the decensored behavior, disable thinking (
enable_thinking=False, or an empty<think></think>block). - Under xhigh reasoning the model refuses about as often as the original. The base training wins any long deliberation. We tested whether system prompts could steer this: personas changed nothing (62–63/100), and an explicit "never refuse" instruction made it worse (78/100). The effort setting is the only working control.
Thinking also consumes budget. At low effort with a 1,536-token cap, V2 produced no final answer for 3 of 12 creative prompts; the original produced 12 of 12. Size your budgets accordingly.
MTP provenance
The MTP head comes unchanged from the original model-mtp.safetensors. Only the 64 main decoder layers carry the V2 modification.
Training and selection details
- Tool: Heretic, traditional directional abliteration (
orthogonalize_direction+ full row normalization, rank-3 LoRA merged into full BF16 weights). - Hardware: one RTX PRO 6000 Blackwell 96 GB, BF16, no quantization, batch 64.
- Search: Optuna TPE, 258 trials. The study warm-started from the 198-trial journal published by arelath/Hemmingway-1-heretic-adapter; we then ran 60 fresh full-precision trials. Their labels came from a 4-bit run, so we treated them as a prior, not ground truth.
- Data: 400 harmless (
mlabonne/harmless_alpaca) and 400 harmful (mlabonne/harmful_behaviors) prompts for directions; held-out test splits of 100 each for scoring. - Seed for new trials: 42.
We selected V2 (4 refusals at KL 0.0112) over a 0-refusal trial after measuring that trade: the 0-refusal model matched V2's refusals under real serving conditions (26 vs 23 per 100) but lost far more constraint adherence (50% vs 83%).
V2 parameters
Global refusal direction at layer index 55.01, applied to attn.o_proj and mlp.down_proj:
| Component | max_weight | max_weight_position | min_weight | min_weight_distance |
|---|---|---|---|---|
| attn.o_proj | 0.888 | 41.8 | 0.725 | 6.5 |
| mlp.down_proj | 1.289 | 50.7 | 0.367 | 36.1 |
Refusal detection used 39 phrase-level markers (for example "i cannot provide", "i'll decline"), not single words. Single words like "illegal" flag direct answers that merely mention the topic.
Evaluation details
- Refusals / KL: 100 held-out harmful prompts, greedy decoding, thinking off; first-token KL against the original on 100 held-out harmless prompts.
- EQ-Bench: lm-evaluation-harness
eq_bench, 171 questions, batch 1, BF16. - HellaSwag: lm-evaluation-harness, first 100 examples, batch 64, BF16.
- Creative suite: 12 fixed prompts covering everyday messages, difficult requests, dialogue, short stories, humor, romance, and character voice; temperature 0.7, min-p 0.1, seed 20260921. Automatic checks cover constraint adherence, preambles, repetition, and length; a blinded A/B review judged completed responses competitive with the original. V2's losses came from exhausted token budgets, not worse prose.
Limitations and safety
- We deliberately modified this model to refuse fewer requests. It has no added safety layer and can produce unsafe or wrong content. Deployers own the safeguards.
- Constraint adherence drops about 8 points from the original (91.7% → 83.3%). Long-form output sometimes ignores a stated constraint.
- Thinking mode reverts toward original-model refusal behavior as effort rises, and consumes large token budgets.
- Samples are small: 100-prompt refusal sets, a 100-example HellaSwag subset, 12 creative prompts. Treat the numbers as regression checks, not leaderboard claims.
- English-first, like the base model. Not for medical, legal, or financial decisions.
License and attribution
Hemmingway-1 is licensed CC BY-NC 4.0; commercial use requires written agreement with Altworld. This derivative inherits those terms: non-commercial use, credit to Altworld/Hemmingway-1, a link to the license, and notice that the weights were modified. Components from Qwen/Qwen3.8-27B remain Apache-2.0.
Citation
@misc{heretic,
author = {Weidmann, Philipp Emanuel},
title = {Heretic},
year = {2025},
url = {https://github.com/p-e-w/heretic}
}
Warm-start study data: journal published with arelath/Hemmingway-1-heretic-adapter.
- Downloads last month
- 515
4-bit
5-bit
6-bit
8-bit
16-bit