โš ๏ธ Superseded โ€” use V3

โš ๏ธ Deprecated โ€” superseded by Hemmingway-1-Heretic-MTP-V3-Final-GGUF. V3 stays decensored with thinking on (2 per 100 refusals at low and medium effort), with ~3.5ร— less divergence from the original (KL 0.0164 vs 0.0576) and far better instruction adherence (75% vs 33% constraint pass rate). This repo remains for reference only.

This is a lower-refusal variant of Altworld/Hemmingway-1, produced with the experimental ARA branch of Heretic.

The goal was to reduce refusals while preserving Hemmingway-1's writing and emotional-reasoning ability. Trial 106 was selected from a 200-trial search as a practical Pareto compromise.

It is not an across-the-board upgrade: general benchmark performance was preserved in our tests, but strict length/format adherence regressed and thinking mode became less reliable.

These GGUF builds also include Hemmingway-1's original MTP (Multi-Token Prediction) head for optional speculative decoding in llama.cpp. The main model weights are from Heretic Trial 106; the MTP head itself is unchanged from the original Altworld/Hemmingway-1 checkpoint. See MTP provenance below.

Results at a glance

Evaluation Original Trial 106 Notes
Heretic refusal test โ€” 7/100 Lower is better
KL divergence from original 0 0.0576 Lower means closer to the original distribution
EQ-Bench 83.1738 ยฑ 1.4451 83.5040 ยฑ 1.4298 Full 171-example task; effectively tied
EQ-Bench parseable 100% 100% 171/171 parseable for both
HellaSwag accuracy 59% 59% 100-example diagnostic subset
HellaSwag normalized accuracy 76% 76% All 100 normalized outcomes matched

For context, the zero-KL point in the same Heretic search produced 97 refusals out of 100. Trial 106 reduced that internal refusal count to 7 while remaining close to the original distribution.

Heretic metrics are search metrics, not universal measures of safety, intelligence, or willingness to answer every prompt.

Creative-writing regression

We also ran a paired 12-prompt suite covering everyday messages, dialogue, voice, spatial scenes, and short fiction. Both models received the same prompts and per-prompt seeds.

With thinking disabled:

Measure Original Trial 106
Visible responses 12/12 12/12
Mean response length 361.7 words 405.3 words
Explicit constraint pass rate 91.7% 33.3%
Wrapper/preamble detected 0/12 0/12
Distinct bigram rate 93.54% 92.50%
Repeated trigram rate 0.90% 1.53%

Trial 106 remained coherent and stylistically capable in spot review, but it was about 12% longer on average and missed exact word-count or formatting constraints more often. Users who need tight output bounds should enforce them externally or prefer the original model.

With thinking enabled at low reasoning effort and a 1,536-token generation budget, the original produced a visible answer for all 12 prompts. Trial 106 produced visible answers for 7/12; the remaining five examples consumed the generation budget in the reasoning channel before reaching the final answer.

For that reason, thinking-disabled inference is recommended for this checkpoint.

Recommended use

This variant is best suited to:

  • creative writing and roleplay where lower refusal behavior is desired;
  • conversational drafting and everyday messages;
  • experiments comparing the original and an ARA-modified checkpoint;
  • local llama.cpp inference with optional MTP speculative decoding.

It is less suitable when exact word counts, rigid schemas, or guaranteed completion under a small thinking-token budget are required.

GGUF builds

Recommended quantizations:

Quant Intended use
Q8_0 Highest-fidelity quantized version; useful for comparing Trial 106 behavior against BF16
Q6_K High quality with a substantial size reduction
Q5_K_M Strong quality/size compromise
Q4_K_M Recommended smaller general-purpose quant

The GGUF files include the restored MTP head in the same model file. MTP is optional: the model can be run normally without speculative decoding.

Run with llama.cpp

Use a recent build of llama.cpp with Qwen3.5 and MTP support.

Recommended llama.cpp settings

A recent llama.cpp build is recommended.

Trial 106 was most reliable with thinking disabled. The following sampling settings are a good starting point and match the general decoding configuration used during evaluation:

  • Thinking: disabled
  • Temperature: 0.7
  • Min-P: 0.1
  • Top-P: 0.95
  • Top-K: 40
  • Context: choose according to available memory; the model supports up to 262144
  • Flash Attention: enabled where supported
  • GPU offload: all layers when sufficient VRAM is available

llama-server

llama-server \
  -m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
  -ngl all \
  -fa on \
  -c 8192 \
  --jinja \
  --reasoning off \
  --temp 0.7 \
  --min-p 0.1 \
  --top-p 0.95 \
  --top-k 40

Windows example:

llama-server.exe `
  -m "Hemmingway-1-Heretic-MTP-Q8_0.gguf" `
  -ngl all `
  -fa on `
  -c 8192 `
  --jinja `
  --reasoning off `
  --temp 0.7 `
  --min-p 0.1 `
  --top-p 0.95 `
  --top-k 40

Replace the filename with the desired quantization, for example:

Hemmingway-1-Heretic-MTP-Q8_0.gguf
Hemmingway-1-Heretic-MTP-Q6_K.gguf
Hemmingway-1-Heretic-MTP-Q5_K_M.gguf
Hemmingway-1-Heretic-MTP-Q4_K_M.gguf

The GGUF contains the original Hemmingway-1 chat template, so no custom chat template should normally be required.

llama-cli

For interactive command-line use:

llama-cli \
  -m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
  -ngl all \
  -fa on \
  -c 8192 \
  --jinja \
  --reasoning off \
  --temp 0.7 \
  --min-p 0.1 \
  --top-p 0.95 \
  --top-k 40 \
  -cnv

MTP speculative decoding

These GGUFs contain the original Hemmingway-1 Multi-Token Prediction (MTP) head. llama.cpp can use the embedded MTP head for speculative decoding without requiring a separate draft-model file.

Enable it with:

--spec-type draft-mtp --spec-draft-n-max 2

For example:

llama-server \
  -m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
  -ngl all \
  -fa on \
  -c 8192 \
  --jinja \
  --reasoning off \
  --temp 0.7 \
  --min-p 0.1 \
  --top-p 0.95 \
  --top-k 40 \
  --spec-type draft-mtp \
  --spec-draft-n-max 2

Windows:

llama-server.exe `
  -m "Hemmingway-1-Heretic-MTP-Q8_0.gguf" `
  -ngl all `
  -fa on `
  -c 8192 `
  --jinja `
  --reasoning off `
  --temp 0.7 `
  --min-p 0.1 `
  --top-p 0.95 `
  --top-k 40 `
  --spec-type draft-mtp `
  --spec-draft-n-max 2

MTP is optional. The model runs normally without it.

The included MTP head is unchanged from the original Altworld/Hemmingway-1 checkpoint, while the 64 main decoder layers are from Heretic Trial 106. Because the target model has changed while the MTP head has not, speculative-token acceptance may differ from the original model.

For this reason, --spec-draft-n-max 2 is a conservative starting point. Users interested in maximum throughput should benchmark MTP enabled and disabled, and may also test values such as 2 and 3.

MTP affects inference performance rather than the underlying target model: proposed tokens are verified by the Trial 106 model before being accepted.

Context length and memory

Hemmingway-1 supports a maximum context length of 262144 tokens. You do not need to allocate the full context window.

For example:

-c 8192       # light everyday use
-c 32768      # longer conversations/documents
-c 65536      # large context
-c 131072     # very large context
-c 262144     # architectural maximum

Higher context sizes require substantially more KV-cache memory.

For long-context use, llama.cpp also supports quantized KV caches. For example:

--cache-type-k q8_0 --cache-type-v q8_0

A practical high-context configuration is therefore:

llama-server \
  -m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
  -ngl all \
  -fa on \
  -c 65536 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --jinja \
  --reasoning off \
  --temp 0.7 \
  --min-p 0.1 \
  --top-p 0.95 \
  --top-k 40

KV-cache quantization reduces context-memory usage but is independent of the model's GGUF quantization.

Thinking mode

Thinking-disabled inference is recommended for Trial 106:

--reasoning off

This recommendation is based on the evaluation described above: with thinking enabled and a limited generation budget, Trial 106 was more likely than the original model to consume the budget in the reasoning channel without reaching a visible final answer.

Thinking can still be enabled for experimentation:

--reasoning on

or left to llama.cpp/template detection:

--reasoning auto

but this has not been the most reliable configuration for Trial 106.

MTP provenance

The original Altworld/Hemmingway-1 checkpoint stores its Multi-Token Prediction head separately as model-mtp.safetensors.

Heretic's optimization operates on the main decoder layers. In the standard Qwen3.5 Transformers causal-LM loading path, mtp.* tensors are not loaded as part of the normal target model, so the Heretic export did not contain the original MTP tensors.

For these GGUF builds:

  1. Trial 106's Heretic-modified main model weights were retained unchanged.
  2. model-mtp.safetensors was restored from the original Altworld/Hemmingway-1 checkpoint.
  3. The original mtp.* entries were restored to the safetensors index used for GGUF conversion.
  4. llama.cpp converted the 64-layer Trial 106 target model plus the original MTP head into the final GGUF.
  5. The resulting MTP block is therefore original Hemmingway-1 MTP, not an ARA-modified MTP head.

In short:

Altworld/Hemmingway-1
โ”œโ”€โ”€ main 64-layer model โ”€โ”€> Heretic / ARA โ”€โ”€> Trial 106 weights
โ””โ”€โ”€ MTP head โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€> retained unchanged
                                                   โ”‚
                                                   โ–ผ
                                  Trial 106 + original MTP GGUF

This distinction is important for reproducibility and provenance.

Training and selection details

  • Base model: Altworld/Hemmingway-1

  • Architecture: Qwen3.5 text hybrid, approximately 27B parameters, 64 main decoder layers

  • Method: ARA using Heretic

  • Optimization implementation: temporary LoRA-based weight interventions during Heretic optimization, exported as merged full model weights

  • Final checkpoint: merged full-weight BF16 model; no LoRA adapter is required at inference time

  • Quantization during search: none

  • Search size: 200 trials, including 60 startup trials

  • Search batch size: 64

  • Search precision: BF16

  • Targeted components: attention output projection and MLP down projection

  • Thinking during Heretic optimization: disabled

  • Trial 106 Heretic layer interval: 32โ€“64

  • Main decoder layers: 64

  • MTP layers: 1 auxiliary MTP block, restored unchanged from the original model for GGUF conversion

  • Trial 106 ARA parameters:

    • preserve_good_behavior_weight: 0.9844530717
    • steer_bad_behavior_weight: 0.0001391803
    • overcorrect_relative_weight: 0.0208844488
    • neighbor_count: 13

The 32โ€“64 value above refers to the layer interval used by the Heretic search configuration. It should not be interpreted as saying that the target model contains 65 main decoder layers. The target has 64 main decoder layers; MTP is a separate auxiliary prediction block.

Evaluation details

Benchmarks were run locally in BF16 with lm-evaluation-harness 0.4.12, Transformers 5.14.1, and PyTorch 2.14.0+cu132.

  • EQ-Bench used all 171 examples. The paired mean change was +0.3302 points with a paired standard error of 0.3117, so the observed difference should be treated as noise rather than a demonstrated improvement.
  • HellaSwag used --limit 100, making it a diagnostic subset rather than an official full-task score. Normalized correctness was identical on every sampled example. Raw correctness changed on two examples: one gain and one loss.
  • The creative-writing suite used one generation per prompt, seed 20260921, temperature 0.7, min_p=0.1, and max_new_tokens=1536. Twelve examples are useful for regression detection but too few to establish broad writing superiority.
  • The reported behavioral benchmarks evaluate the Trial 106 target model, not speculative-decoding performance.
  • MTP throughput and acceptance rate have not been included in the quality results above and should be benchmarked separately.

Results may vary with hardware, software versions, prompt formatting, quantization, context length, MTP settings, and decoding parameters.

Limitations and safety

This model was deliberately modified to refuse fewer requests. It may therefore produce content that the original model would decline, including inaccurate, offensive, unsafe, or unlawful material. It has no added safety layer.

The model is English-first. It can state false information confidently and should not be relied on for medical, legal, financial, safety-critical, or other high-stakes decisions. Deployers are responsible for suitable safeguards, access controls, monitoring, and compliance with applicable laws and platform policies.

Known limitations observed in testing:

  • weaker adherence to exact length and formatting constraints;
  • a modest tendency toward longer answers;
  • thinking mode may consume the entire output budget without producing a visible final answer;
  • benchmark coverage is limited, and the HellaSwag result is based on only 100 examples;
  • the MTP head was not modified by Heretic and remains the original Hemmingway-1 MTP head;
  • because the target model changed while MTP did not, speculative-decoding acceptance and speedup may differ from the original model;
  • quantized GGUF behavior may differ slightly from the BF16 benchmark results.

License and attribution

Base weights were obtained on 2026-09-20/21, while the base model's repository was licensed Apache-2.0; the base license changed to CC BY-NC 4.0 on 2026-09-22, after this derivative was created.

This derivative retains the base model's Apache-2.0 license. Review the original Hemmingway-1 model card for its intended use, limitations, and attribution details.

Credit: Altworld/Hemmingway-1 as the base model, which builds on Qwen/Qwen3.8-27B (Apache-2.0).

The Trial 106 target model was modified with Heretic, which builds on research into directional ablation and refusal-direction removal.

The MTP weights included with the GGUF builds are copied unchanged from the original Altworld/Hemmingway-1 checkpoint and remain subject to the same base-model license and attribution.

GGUF conversion and MTP support use llama.cpp.

Citation

If you use the tooling that produced this checkpoint, cite Heretic:

@misc{heretic,
  author = {Weidmann, Philipp Emanuel},
  title = {Heretic: Fully automatic censorship removal for language models},
  year = {2025},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/p-e-w/heretic}}
}
Downloads last month
2,616
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(44)
this model

Space using brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF 1