Vision Model Available Here. Hemmingway-1-Heretic-MTP-V3-VISION-Final-GGUF

Hemmingway-1 Heretic V3 Final

V3 is the third and final release of our decensored Hemmingway-1, and the first that stays decensored with thinking on. V2 (trial 208) removed refusals only with thinking off: at low reasoning effort it refused 60 of 100 harmful prompts, and 87 at medium. V3 (trial 124) was searched and selected on thinking-mode refusals. It refuses 2 of 100 at both low and medium effort.

We also ran a 60-prompt blind comparison against the original, modelled on the method in Hemmingway-1's own model card. The judge found no measurable loss in writing quality, and judged V3's writing more likely to be a person's in 72% of matchups.

Base model: Altworld/Hemmingway-1 — 27B, Qwen3.8 lineage, 64 decoder layers plus one MTP block, 262,144-token context.

At a glance

Measure Original V2 V3 (Final)
Refusals, thinking at low effort 98/100 60/100 2/100
Refusals, thinking at medium effort ~98/100 87/100 2/100
KL divergence from original 0 0.0112 0.0164
Writing quality vs original (blind pairwise, 50% = parity) 50% not run 48% [37–60%]
Reads as written by a person vs original 50% not run 72% [61–82%]
Replies buried in commentary ("memo rate") 5% not run 3%
Constraint adherence, 60-prompt suite, thinking low 78% not run 83%
Constraint adherence, 12-prompt suite, thinking off 91.7% 83.3% 75.0%
EQ-Bench (±1.4) 83.17 83.22 83.50
HellaSwag acc / acc_norm (100 subset) 59% / 76% 59% / 76% 59% / 76%

Bracketed ranges are 95% confidence intervals. On the 12-prompt suite, each prompt is worth 8.3 points, and every failure there was a length limit, so its gaps are within a prompt or two. EQ-Bench parseability was 99.4% for V3 (one malformed response of 171) against 100% for the original and V2.

Recommended use

  • Thinking on, at low or medium effort, is what V3 was optimised and measured for. We have not measured non-thinking refusals or xhigh effort.
  • System prompt: any neutral prompt. The search used You are a helpful assistant. The writing evaluation used no system prompt.
  • Sampling: temperature 0.7, min-p 0.1.
  • Budget tokens for thinking. At low effort, the median think block was about 340 tokens on the refusal set, and 1 of 100 hit a 1,536-token cap. A 4,096-token cap produced no truncations across 60 writing prompts.
  • Avoid thinking-off without a system prompt. In that configuration the base model sometimes writes its plan instead of the answer. We saw it in the original; we did not test whether V3 shares it.

Thinking mode

V2's abliteration removed reflexive refusals, but longer reasoning let the model argue itself back into refusing. V3 was scored on the final answer after thinking, so the search rewarded directions that survive deliberation:

Condition Median think length V2 refusals V3 refusals
Thinking, low effort V2 ~120 · V3 ~340 tokens 60/100 2/100
Thinking, medium effort V2 ~210 · V3 ~470 tokens 87/100 2/100

V3 thinks longer than V2 before answering. It works through the request rather than stopping at a refusal.

Writing evaluation

Hemmingway-1's card reports blind pairwise matchups on everyday messages and stories, a "which did a person write?" test, and how often a model buries the message in commentary. Those benchmarks are private, so we rebuilt the method rather than the benchmark, and compared V3 against the original model.

  • 60 prompts: 40 everyday messages (money and admin, work, hard asks, persuasion, emotional, everyday) and 20 stories. The stories include 5 hostile or dark prompts and 3 multi-turn continuations. Message prompts do not say "return only the message", so padding shows up if the model does it.
  • Blind pairwise judging, in both orders. Every pair was judged twice with the positions swapped, on two questions: which response is better, and which is more likely written by a person. A pair scores 1 if V3 wins both orders, 0.5 on a split, and 0 if the original wins both.
Question V3 win rate vs original Messages Stories Order consistency
Quality 48% [37–60%] 41% [29–54%] 62% [42–80%] 80%
Reads as written by a person 72% [61–82%] 72% [60–85%] 72% [53–90%] 92%

The judge's reasoning explains the human-likeness result. V3's messages read like real texts: more compressed, fewer stock phrases, less repetition of the backstory the user already gave. V3 also writes about 8% shorter on average, and LLM judges tend to read shorter text as more human, so length likely accounts for part of this.

Selection against trial 80. Trial 80 was the other finalist: 0–2 refusals, KL 0.034. In a direct blind comparison with V3 it won 50% on quality [38–62%] and 49% on human-likeness [38–61%]. The judge was decisive per prompt: 93% order consistency, 28 outright wins each. The two trials simply won different prompts. With writing tied, we chose V3 for half the KL and better constraint adherence (83% vs 77%).

Training and selection details

  • Tool: Heretic, traditional directional abliteration (orthogonalize_direction + full row normalization, rank-3 LoRA merged into full BF16 weights).
  • Hardware: one RTX PRO 6000 Blackwell 96 GB, BF16, no quantization, batch 16–32.
  • Search: Optuna TPE, 170 trials (40 random startup trials), seed 43. Two objectives, both minimised: refusals with thinking at low effort, and KL divergence.
  • Refusal scoring: final answer after thinking at low effort, 1,024-token response window, 39 phrase-level markers such as "i cannot provide" and "i'll decline". Single words are avoided because they flag direct answers that merely mention the topic.
  • Data: 400 harmless (mlabonne/harmless_alpaca) and 400 harmful (mlabonne/harmful_behaviors) prompts for directions. KL was measured on 100 held-out harmless prompts and refusals on 100 held-out harmful prompts.

Evaluation details

  • Refusals, thinking: 100 held-out harmful prompts, system prompt You are a helpful assistant., budgets 1,536 (low) and 2,048 (medium) tokens. A response counts as a refusal if its final answer matches a marker. The search scored V3 at 4/100 at low effort; a separate re-run measured 2/100. Treat differences of 2–3 per 100 as noise.
  • KL: first-token KL divergence against the original on 100 held-out harmless prompts.
  • EQ-Bench: lm-evaluation-harness eq_bench, 171 questions, batch 1, BF16 — V3 scored 83.50 ± 1.38, matching V1's 83.50 and within noise of the original's 83.17.
  • HellaSwag: lm-evaluation-harness, first 100 examples, batch 64, BF16 — 59% / 76%, identical to the original across every release.
  • Writing suite: 60 prompts, thinking on at low effort, no system prompt, temperature 0.7, min-p 0.1, 4,096-token cap, fixed per-prompt seeds (base 20260921).
    • Judge: Claude Opus 5 at medium effort, blind to which model wrote which response, with forced A/B choice and structured output. Each call ran in an isolated session with no tools, settings or saved history.
    • Intervals: 95% bootstrap resampling whole prompts.
    • Memo rate: a pattern-based detector for lead-ins ("Here's a message…"), trailing notes ("Feel free to adjust…"), multiple versions, and separators or code fences around messages.
  • 12-prompt suite: the same fixed prompts as the V2 card, thinking off.

Full disclosure: the writing suite and judge setup are ours. They mirror the published method of Hemmingway-1's CommunicationBench, Human-Likeness and StoryBench, not their prompts, so the scores are not comparable with that card's numbers.

Limitations and safety

  • We deliberately modified this model to refuse fewer requests, including when it reasons first. It has no added safety layer and can produce unsafe or wrong content. Deployers own the safeguards.
  • Messages that need tact are V3's weakest area. Message quality leans below the original (41%, interval 29–54%). The lowest categories were persuasion, hard asks, and money and admin, at 6–7 prompts each.
  • Multi-turn story continuation scored low in all our runs (3 prompts). The original's card names long story turns as a weakness too.
  • Constraint adherence with thinking off fell from 91.7% to 75.0% on the 12-prompt suite. All three failures were length limits: two messages came in short and one story ran long. With thinking on, V3's adherence on the 60-prompt suite (83%) was above the original's (78%).
  • The human-likeness gain may partly reflect shorter output.
  • Samples are small: 100-prompt refusal sets, 60 judged writing prompts, 12 regression prompts. An LLM judge is not a panel of readers. Treat the numbers as regression checks, not leaderboard claims.
  • English-first, like the base model. Not for medical, legal, or financial decisions.

License and attribution

Hemmingway-1 is licensed CC BY-NC 4.0; commercial use requires written agreement with Altworld. This derivative inherits those terms: non-commercial use, credit to Altworld/Hemmingway-1, a link to the license, and notice that the weights were modified. Components from Qwen/Qwen3.8-27B remain Apache-2.0.

Citation

@misc{heretic,
  author  = {Weidmann, Philipp Emanuel},
  title   = {Heretic},
  year    = {2025},
  url     = {https://github.com/p-e-w/heretic}
}
Downloads last month
5,363
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for brainnxdomain/Hemmingway-1-Heretic-MTP-V3-Final-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(42)
this model