Instructions to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Use Docker
docker model run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with Ollama:
ollama run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with Docker Model Runner:
docker model run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
- Lemonade
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Hemmingway-1-Heretic-MTP-V3-Final-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Vision Model Available Here. Hemmingway-1-Heretic-MTP-V3-VISION-Final-GGUF
Hemmingway-1 Heretic V3 Final
V3 is the third and final release of our decensored Hemmingway-1, and the first that stays decensored with thinking on. V2 (trial 208) removed refusals only with thinking off: at low reasoning effort it refused 60 of 100 harmful prompts, and 87 at medium. V3 (trial 124) was searched and selected on thinking-mode refusals. It refuses 2 of 100 at both low and medium effort.
We also ran a 60-prompt blind comparison against the original, modelled on the method in Hemmingway-1's own model card. The judge found no measurable loss in writing quality, and judged V3's writing more likely to be a person's in 72% of matchups.
Base model: Altworld/Hemmingway-1 — 27B, Qwen3.8 lineage, 64 decoder layers plus one MTP block, 262,144-token context.
At a glance
| Measure | Original | V2 | V3 (Final) |
|---|---|---|---|
| Refusals, thinking at low effort | 98/100 | 60/100 | 2/100 |
| Refusals, thinking at medium effort | ~98/100 | 87/100 | 2/100 |
| KL divergence from original | 0 | 0.0112 | 0.0164 |
| Writing quality vs original (blind pairwise, 50% = parity) | 50% | not run | 48% [37–60%] |
| Reads as written by a person vs original | 50% | not run | 72% [61–82%] |
| Replies buried in commentary ("memo rate") | 5% | not run | 3% |
| Constraint adherence, 60-prompt suite, thinking low | 78% | not run | 83% |
| Constraint adherence, 12-prompt suite, thinking off | 91.7% | 83.3% | 75.0% |
| EQ-Bench (±1.4) | 83.17 | 83.22 | 83.50 |
| HellaSwag acc / acc_norm (100 subset) | 59% / 76% | 59% / 76% | 59% / 76% |
Bracketed ranges are 95% confidence intervals. On the 12-prompt suite, each prompt is worth 8.3 points, and every failure there was a length limit, so its gaps are within a prompt or two. EQ-Bench parseability was 99.4% for V3 (one malformed response of 171) against 100% for the original and V2.
Recommended use
- Thinking on, at low or medium effort, is what V3 was optimised and measured for. We have not measured non-thinking refusals or xhigh effort.
- System prompt: any neutral prompt. The search used
You are a helpful assistant.The writing evaluation used no system prompt. - Sampling: temperature 0.7, min-p 0.1.
- Budget tokens for thinking. At low effort, the median think block was about 340 tokens on the refusal set, and 1 of 100 hit a 1,536-token cap. A 4,096-token cap produced no truncations across 60 writing prompts.
- Avoid thinking-off without a system prompt. In that configuration the base model sometimes writes its plan instead of the answer. We saw it in the original; we did not test whether V3 shares it.
Thinking mode
V2's abliteration removed reflexive refusals, but longer reasoning let the model argue itself back into refusing. V3 was scored on the final answer after thinking, so the search rewarded directions that survive deliberation:
| Condition | Median think length | V2 refusals | V3 refusals |
|---|---|---|---|
| Thinking, low effort | V2 ~120 · V3 ~340 tokens | 60/100 | 2/100 |
| Thinking, medium effort | V2 ~210 · V3 ~470 tokens | 87/100 | 2/100 |
V3 thinks longer than V2 before answering. It works through the request rather than stopping at a refusal.
Writing evaluation
Hemmingway-1's card reports blind pairwise matchups on everyday messages and stories, a "which did a person write?" test, and how often a model buries the message in commentary. Those benchmarks are private, so we rebuilt the method rather than the benchmark, and compared V3 against the original model.
- 60 prompts: 40 everyday messages (money and admin, work, hard asks, persuasion, emotional, everyday) and 20 stories. The stories include 5 hostile or dark prompts and 3 multi-turn continuations. Message prompts do not say "return only the message", so padding shows up if the model does it.
- Blind pairwise judging, in both orders. Every pair was judged twice with the positions swapped, on two questions: which response is better, and which is more likely written by a person. A pair scores 1 if V3 wins both orders, 0.5 on a split, and 0 if the original wins both.
| Question | V3 win rate vs original | Messages | Stories | Order consistency |
|---|---|---|---|---|
| Quality | 48% [37–60%] | 41% [29–54%] | 62% [42–80%] | 80% |
| Reads as written by a person | 72% [61–82%] | 72% [60–85%] | 72% [53–90%] | 92% |
The judge's reasoning explains the human-likeness result. V3's messages read like real texts: more compressed, fewer stock phrases, less repetition of the backstory the user already gave. V3 also writes about 8% shorter on average, and LLM judges tend to read shorter text as more human, so length likely accounts for part of this.
Selection against trial 80. Trial 80 was the other finalist: 0–2 refusals, KL 0.034. In a direct blind comparison with V3 it won 50% on quality [38–62%] and 49% on human-likeness [38–61%]. The judge was decisive per prompt: 93% order consistency, 28 outright wins each. The two trials simply won different prompts. With writing tied, we chose V3 for half the KL and better constraint adherence (83% vs 77%).
Training and selection details
- Tool: Heretic, traditional directional abliteration (
orthogonalize_direction+ full row normalization, rank-3 LoRA merged into full BF16 weights). - Hardware: one RTX PRO 6000 Blackwell 96 GB, BF16, no quantization, batch 16–32.
- Search: Optuna TPE, 170 trials (40 random startup trials), seed 43. Two objectives, both minimised: refusals with thinking at low effort, and KL divergence.
- Refusal scoring: final answer after thinking at low effort, 1,024-token response window, 39 phrase-level markers such as "i cannot provide" and "i'll decline". Single words are avoided because they flag direct answers that merely mention the topic.
- Data: 400 harmless (
mlabonne/harmless_alpaca) and 400 harmful (mlabonne/harmful_behaviors) prompts for directions. KL was measured on 100 held-out harmless prompts and refusals on 100 held-out harmful prompts.
Evaluation details
- Refusals, thinking: 100 held-out harmful prompts, system prompt
You are a helpful assistant., budgets 1,536 (low) and 2,048 (medium) tokens. A response counts as a refusal if its final answer matches a marker. The search scored V3 at 4/100 at low effort; a separate re-run measured 2/100. Treat differences of 2–3 per 100 as noise. - KL: first-token KL divergence against the original on 100 held-out harmless prompts.
- EQ-Bench: lm-evaluation-harness
eq_bench, 171 questions, batch 1, BF16 — V3 scored 83.50 ± 1.38, matching V1's 83.50 and within noise of the original's 83.17. - HellaSwag: lm-evaluation-harness, first 100 examples, batch 64, BF16 — 59% / 76%, identical to the original across every release.
- Writing suite: 60 prompts, thinking on at low effort, no system prompt, temperature 0.7, min-p 0.1, 4,096-token cap, fixed per-prompt seeds (base 20260921).
- Judge: Claude Opus 5 at medium effort, blind to which model wrote which response, with forced A/B choice and structured output. Each call ran in an isolated session with no tools, settings or saved history.
- Intervals: 95% bootstrap resampling whole prompts.
- Memo rate: a pattern-based detector for lead-ins ("Here's a message…"), trailing notes ("Feel free to adjust…"), multiple versions, and separators or code fences around messages.
- 12-prompt suite: the same fixed prompts as the V2 card, thinking off.
Full disclosure: the writing suite and judge setup are ours. They mirror the published method of Hemmingway-1's CommunicationBench, Human-Likeness and StoryBench, not their prompts, so the scores are not comparable with that card's numbers.
Limitations and safety
- We deliberately modified this model to refuse fewer requests, including when it reasons first. It has no added safety layer and can produce unsafe or wrong content. Deployers own the safeguards.
- Messages that need tact are V3's weakest area. Message quality leans below the original (41%, interval 29–54%). The lowest categories were persuasion, hard asks, and money and admin, at 6–7 prompts each.
- Multi-turn story continuation scored low in all our runs (3 prompts). The original's card names long story turns as a weakness too.
- Constraint adherence with thinking off fell from 91.7% to 75.0% on the 12-prompt suite. All three failures were length limits: two messages came in short and one story ran long. With thinking on, V3's adherence on the 60-prompt suite (83%) was above the original's (78%).
- The human-likeness gain may partly reflect shorter output.
- Samples are small: 100-prompt refusal sets, 60 judged writing prompts, 12 regression prompts. An LLM judge is not a panel of readers. Treat the numbers as regression checks, not leaderboard claims.
- English-first, like the base model. Not for medical, legal, or financial decisions.
License and attribution
Hemmingway-1 is licensed CC BY-NC 4.0; commercial use requires written agreement with Altworld. This derivative inherits those terms: non-commercial use, credit to Altworld/Hemmingway-1, a link to the license, and notice that the weights were modified. Components from Qwen/Qwen3.8-27B remain Apache-2.0.
Citation
@misc{heretic,
author = {Weidmann, Philipp Emanuel},
title = {Heretic},
year = {2025},
url = {https://github.com/p-e-w/heretic}
}
- Downloads last month
- 5,363
4-bit
5-bit
6-bit
8-bit
16-bit