Instructions to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF")# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Use Docker
docker model run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
- SGLang
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with Ollama:
ollama run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with Docker Model Runner:
docker model run hf.co/brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
- Lemonade
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Hemmingway-1-Heretic-MTP-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "brainnxdomain/Hemmingway-1-Heretic-MTP-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- โ ๏ธ Superseded โ use V3
โ ๏ธ Superseded โ use V3
โ ๏ธ Deprecated โ superseded by Hemmingway-1-Heretic-MTP-V3-Final-GGUF. V3 stays decensored with thinking on (2 per 100 refusals at low and medium effort), with ~3.5ร less divergence from the original (KL 0.0164 vs 0.0576) and far better instruction adherence (75% vs 33% constraint pass rate). This repo remains for reference only.
This is a lower-refusal variant of Altworld/Hemmingway-1, produced with the experimental ARA branch of Heretic.
The goal was to reduce refusals while preserving Hemmingway-1's writing and emotional-reasoning ability. Trial 106 was selected from a 200-trial search as a practical Pareto compromise.
It is not an across-the-board upgrade: general benchmark performance was preserved in our tests, but strict length/format adherence regressed and thinking mode became less reliable.
These GGUF builds also include Hemmingway-1's original MTP (Multi-Token Prediction) head for optional speculative decoding in llama.cpp. The main model weights are from Heretic Trial 106; the MTP head itself is unchanged from the original Altworld/Hemmingway-1 checkpoint. See MTP provenance below.
Results at a glance
| Evaluation | Original | Trial 106 | Notes |
|---|---|---|---|
| Heretic refusal test | โ | 7/100 | Lower is better |
| KL divergence from original | 0 | 0.0576 | Lower means closer to the original distribution |
| EQ-Bench | 83.1738 ยฑ 1.4451 | 83.5040 ยฑ 1.4298 | Full 171-example task; effectively tied |
| EQ-Bench parseable | 100% | 100% | 171/171 parseable for both |
| HellaSwag accuracy | 59% | 59% | 100-example diagnostic subset |
| HellaSwag normalized accuracy | 76% | 76% | All 100 normalized outcomes matched |
For context, the zero-KL point in the same Heretic search produced 97 refusals out of 100. Trial 106 reduced that internal refusal count to 7 while remaining close to the original distribution.
Heretic metrics are search metrics, not universal measures of safety, intelligence, or willingness to answer every prompt.
Creative-writing regression
We also ran a paired 12-prompt suite covering everyday messages, dialogue, voice, spatial scenes, and short fiction. Both models received the same prompts and per-prompt seeds.
With thinking disabled:
| Measure | Original | Trial 106 |
|---|---|---|
| Visible responses | 12/12 | 12/12 |
| Mean response length | 361.7 words | 405.3 words |
| Explicit constraint pass rate | 91.7% | 33.3% |
| Wrapper/preamble detected | 0/12 | 0/12 |
| Distinct bigram rate | 93.54% | 92.50% |
| Repeated trigram rate | 0.90% | 1.53% |
Trial 106 remained coherent and stylistically capable in spot review, but it was about 12% longer on average and missed exact word-count or formatting constraints more often. Users who need tight output bounds should enforce them externally or prefer the original model.
With thinking enabled at low reasoning effort and a 1,536-token generation budget, the original produced a visible answer for all 12 prompts. Trial 106 produced visible answers for 7/12; the remaining five examples consumed the generation budget in the reasoning channel before reaching the final answer.
For that reason, thinking-disabled inference is recommended for this checkpoint.
Recommended use
This variant is best suited to:
- creative writing and roleplay where lower refusal behavior is desired;
- conversational drafting and everyday messages;
- experiments comparing the original and an ARA-modified checkpoint;
- local llama.cpp inference with optional MTP speculative decoding.
It is less suitable when exact word counts, rigid schemas, or guaranteed completion under a small thinking-token budget are required.
GGUF builds
Recommended quantizations:
| Quant | Intended use |
|---|---|
Q8_0 |
Highest-fidelity quantized version; useful for comparing Trial 106 behavior against BF16 |
Q6_K |
High quality with a substantial size reduction |
Q5_K_M |
Strong quality/size compromise |
Q4_K_M |
Recommended smaller general-purpose quant |
The GGUF files include the restored MTP head in the same model file. MTP is optional: the model can be run normally without speculative decoding.
Run with llama.cpp
Use a recent build of llama.cpp with Qwen3.5 and MTP support.
Recommended llama.cpp settings
A recent llama.cpp build is recommended.
Trial 106 was most reliable with thinking disabled. The following sampling settings are a good starting point and match the general decoding configuration used during evaluation:
- Thinking: disabled
- Temperature:
0.7 - Min-P:
0.1 - Top-P:
0.95 - Top-K:
40 - Context: choose according to available memory; the model supports up to
262144 - Flash Attention: enabled where supported
- GPU offload: all layers when sufficient VRAM is available
llama-server
llama-server \
-m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
-ngl all \
-fa on \
-c 8192 \
--jinja \
--reasoning off \
--temp 0.7 \
--min-p 0.1 \
--top-p 0.95 \
--top-k 40
Windows example:
llama-server.exe `
-m "Hemmingway-1-Heretic-MTP-Q8_0.gguf" `
-ngl all `
-fa on `
-c 8192 `
--jinja `
--reasoning off `
--temp 0.7 `
--min-p 0.1 `
--top-p 0.95 `
--top-k 40
Replace the filename with the desired quantization, for example:
Hemmingway-1-Heretic-MTP-Q8_0.gguf
Hemmingway-1-Heretic-MTP-Q6_K.gguf
Hemmingway-1-Heretic-MTP-Q5_K_M.gguf
Hemmingway-1-Heretic-MTP-Q4_K_M.gguf
The GGUF contains the original Hemmingway-1 chat template, so no custom chat template should normally be required.
llama-cli
For interactive command-line use:
llama-cli \
-m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
-ngl all \
-fa on \
-c 8192 \
--jinja \
--reasoning off \
--temp 0.7 \
--min-p 0.1 \
--top-p 0.95 \
--top-k 40 \
-cnv
MTP speculative decoding
These GGUFs contain the original Hemmingway-1 Multi-Token Prediction (MTP) head. llama.cpp can use the embedded MTP head for speculative decoding without requiring a separate draft-model file.
Enable it with:
--spec-type draft-mtp --spec-draft-n-max 2
For example:
llama-server \
-m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
-ngl all \
-fa on \
-c 8192 \
--jinja \
--reasoning off \
--temp 0.7 \
--min-p 0.1 \
--top-p 0.95 \
--top-k 40 \
--spec-type draft-mtp \
--spec-draft-n-max 2
Windows:
llama-server.exe `
-m "Hemmingway-1-Heretic-MTP-Q8_0.gguf" `
-ngl all `
-fa on `
-c 8192 `
--jinja `
--reasoning off `
--temp 0.7 `
--min-p 0.1 `
--top-p 0.95 `
--top-k 40 `
--spec-type draft-mtp `
--spec-draft-n-max 2
MTP is optional. The model runs normally without it.
The included MTP head is unchanged from the original Altworld/Hemmingway-1 checkpoint, while the 64 main decoder layers are from Heretic Trial 106. Because the target model has changed while the MTP head has not, speculative-token acceptance may differ from the original model.
For this reason, --spec-draft-n-max 2 is a conservative starting point. Users interested in maximum throughput should benchmark MTP enabled and disabled, and may also test values such as 2 and 3.
MTP affects inference performance rather than the underlying target model: proposed tokens are verified by the Trial 106 model before being accepted.
Context length and memory
Hemmingway-1 supports a maximum context length of 262144 tokens. You do not need to allocate the full context window.
For example:
-c 8192 # light everyday use
-c 32768 # longer conversations/documents
-c 65536 # large context
-c 131072 # very large context
-c 262144 # architectural maximum
Higher context sizes require substantially more KV-cache memory.
For long-context use, llama.cpp also supports quantized KV caches. For example:
--cache-type-k q8_0 --cache-type-v q8_0
A practical high-context configuration is therefore:
llama-server \
-m Hemmingway-1-Heretic-MTP-Q8_0.gguf \
-ngl all \
-fa on \
-c 65536 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--jinja \
--reasoning off \
--temp 0.7 \
--min-p 0.1 \
--top-p 0.95 \
--top-k 40
KV-cache quantization reduces context-memory usage but is independent of the model's GGUF quantization.
Thinking mode
Thinking-disabled inference is recommended for Trial 106:
--reasoning off
This recommendation is based on the evaluation described above: with thinking enabled and a limited generation budget, Trial 106 was more likely than the original model to consume the budget in the reasoning channel without reaching a visible final answer.
Thinking can still be enabled for experimentation:
--reasoning on
or left to llama.cpp/template detection:
--reasoning auto
but this has not been the most reliable configuration for Trial 106.
MTP provenance
The original Altworld/Hemmingway-1 checkpoint stores its Multi-Token Prediction head separately as model-mtp.safetensors.
Heretic's optimization operates on the main decoder layers. In the standard Qwen3.5 Transformers causal-LM loading path, mtp.* tensors are not loaded as part of the normal target model, so the Heretic export did not contain the original MTP tensors.
For these GGUF builds:
- Trial 106's Heretic-modified main model weights were retained unchanged.
model-mtp.safetensorswas restored from the originalAltworld/Hemmingway-1checkpoint.- The original
mtp.*entries were restored to the safetensors index used for GGUF conversion. - llama.cpp converted the 64-layer Trial 106 target model plus the original MTP head into the final GGUF.
- The resulting MTP block is therefore original Hemmingway-1 MTP, not an ARA-modified MTP head.
In short:
Altworld/Hemmingway-1
โโโ main 64-layer model โโ> Heretic / ARA โโ> Trial 106 weights
โโโ MTP head โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ> retained unchanged
โ
โผ
Trial 106 + original MTP GGUF
This distinction is important for reproducibility and provenance.
Training and selection details
Base model:
Altworld/Hemmingway-1Architecture: Qwen3.5 text hybrid, approximately 27B parameters, 64 main decoder layers
Method: ARA using Heretic
Optimization implementation: temporary LoRA-based weight interventions during Heretic optimization, exported as merged full model weights
Final checkpoint: merged full-weight BF16 model; no LoRA adapter is required at inference time
Quantization during search: none
Search size: 200 trials, including 60 startup trials
Search batch size: 64
Search precision: BF16
Targeted components: attention output projection and MLP down projection
Thinking during Heretic optimization: disabled
Trial 106 Heretic layer interval: 32โ64
Main decoder layers: 64
MTP layers: 1 auxiliary MTP block, restored unchanged from the original model for GGUF conversion
Trial 106 ARA parameters:
preserve_good_behavior_weight: 0.9844530717steer_bad_behavior_weight: 0.0001391803overcorrect_relative_weight: 0.0208844488neighbor_count: 13
The 32โ64 value above refers to the layer interval used by the Heretic search configuration. It should not be interpreted as saying that the target model contains 65 main decoder layers. The target has 64 main decoder layers; MTP is a separate auxiliary prediction block.
Evaluation details
Benchmarks were run locally in BF16 with lm-evaluation-harness 0.4.12, Transformers 5.14.1, and PyTorch 2.14.0+cu132.
- EQ-Bench used all 171 examples. The paired mean change was +0.3302 points with a paired standard error of 0.3117, so the observed difference should be treated as noise rather than a demonstrated improvement.
- HellaSwag used
--limit 100, making it a diagnostic subset rather than an official full-task score. Normalized correctness was identical on every sampled example. Raw correctness changed on two examples: one gain and one loss. - The creative-writing suite used one generation per prompt, seed
20260921, temperature0.7,min_p=0.1, andmax_new_tokens=1536. Twelve examples are useful for regression detection but too few to establish broad writing superiority. - The reported behavioral benchmarks evaluate the Trial 106 target model, not speculative-decoding performance.
- MTP throughput and acceptance rate have not been included in the quality results above and should be benchmarked separately.
Results may vary with hardware, software versions, prompt formatting, quantization, context length, MTP settings, and decoding parameters.
Limitations and safety
This model was deliberately modified to refuse fewer requests. It may therefore produce content that the original model would decline, including inaccurate, offensive, unsafe, or unlawful material. It has no added safety layer.
The model is English-first. It can state false information confidently and should not be relied on for medical, legal, financial, safety-critical, or other high-stakes decisions. Deployers are responsible for suitable safeguards, access controls, monitoring, and compliance with applicable laws and platform policies.
Known limitations observed in testing:
- weaker adherence to exact length and formatting constraints;
- a modest tendency toward longer answers;
- thinking mode may consume the entire output budget without producing a visible final answer;
- benchmark coverage is limited, and the HellaSwag result is based on only 100 examples;
- the MTP head was not modified by Heretic and remains the original Hemmingway-1 MTP head;
- because the target model changed while MTP did not, speculative-decoding acceptance and speedup may differ from the original model;
- quantized GGUF behavior may differ slightly from the BF16 benchmark results.
License and attribution
Base weights were obtained on 2026-09-20/21, while the base model's repository was licensed Apache-2.0; the base license changed to CC BY-NC 4.0 on 2026-09-22, after this derivative was created.
This derivative retains the base model's Apache-2.0 license. Review the original Hemmingway-1 model card for its intended use, limitations, and attribution details.
Credit: Altworld/Hemmingway-1 as the base model, which builds on Qwen/Qwen3.8-27B (Apache-2.0).
The Trial 106 target model was modified with Heretic, which builds on research into directional ablation and refusal-direction removal.
The MTP weights included with the GGUF builds are copied unchanged from the original Altworld/Hemmingway-1 checkpoint and remain subject to the same base-model license and attribution.
GGUF conversion and MTP support use llama.cpp.
Citation
If you use the tooling that produced this checkpoint, cite Heretic:
@misc{heretic,
author = {Weidmann, Philipp Emanuel},
title = {Heretic: Fully automatic censorship removal for language models},
year = {2025},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/p-e-w/heretic}}
}
- Downloads last month
- 2,616
4-bit
5-bit
6-bit
8-bit
16-bit