Instructions to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Use Docker
docker model run hf.co/3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "3ntr0py-t4m3r/Hemmingway-1-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "3ntr0py-t4m3r/Hemmingway-1-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
- Ollama
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with Ollama:
ollama run hf.co/3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with Docker Model Runner:
docker model run hf.co/3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
- Lemonade
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Hemmingway-1-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use 3ntr0py-t4m3r/Hemmingway-1-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf 3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "3ntr0py-t4m3r/Hemmingway-1-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Hemmingway-1 GGUF
A practical GGUF build of Altworld/Hemmingway-1 for llama.cpp. Same model, smaller footprint, still good at rewriting.
Files
Hemmingway-1-imatrix.gguf– importance matrix built from rewriting-style calibrationHemmingway-1-IQ4_K_M.gguf– main model, ~16 GB
What this is
Hemmingway-1 is a 27B model tuned for everyday writing: messages, emails, the awkward note you keep rewriting. This repo is a Q4_K_M quant with a few targeted upgrades so it stays usable for that job.
Context is 262k. License is CC BY-NC 4.0, same as the original.
How it was built
The original HF weights are Qwen3.5 based with one MTP layer. Conversion to GGUF needed a small fix in the llama.cpp Qwen converter – the checkout had a stray @ModelBase.example decorator that broke import, so it was stripped before conversion.
Steps we actually ran:
Convert to BF16. Straight HF → GGUF with outtype bf16. That gave a 51 GB intermediate with 866 tensors.
Q8_0 intermediate. Quantized the BF16 to Q8_0 first. Faster than doing imatrix on BF16 and keeps the numbers sane. Ended up around 28 GB.
Importance matrix. Ran llama-imatrix on the Q8_0 with 2 300+ chunks of calibration text focused on rewriting tasks: verbose → concise, tone changes, email polish, de-slopping. Imatrix was computed with 50 GPU layers on an RTX 4090. The MTP block 64 tensors were unused during calibration – that's normal, the forward pass never hits them – so they will be protected manually later.
Quantize with overrides. Final Q4_K_M pass with imatrix applied:
- Token embeddings at Q4_K_M
- Output projection at Q5_K_M
- MTP head tensors forced to Q8_0:
blk.64.nextn.eh_proj.weightand the associated norms - Norms and small vectors tend to stay f32 in this pipeline; we tried to push them, they didn't take. That's a limitation of the current quant kernel, not a mistake.
Resulting model is about 15.95 GB, 4.90 BPW. It loads fine with llama-server at 130k+ context and the rewriting style holds up.
Notes from the run
- The imatrix calibration set was a mix of English rewriting prompts, ~5 MB of text. Not code heavy on purpose.
- MTP tensors weren't activated during imatrix, so the importance scores for them are zero. We kept them at Q8_0 anyway because speculative acceptance drops fast if you quantize the next-token predictor too hard.
- Vector tensors like output_norm and the small norm weights stay f32 even with explicit overrides. If you want them quantized, you'd need a different kernel.
- The original model card and benchmarks are reproduced below for reference.
Running
llama.cpp:
./llama-cli -m Hemmingway-1-IQ4_K_M.gguf -p "Rewrite this email to be shorter and more polite:\n..." -n 256
Server:
./llama-server -m Hemmingway-1-IQ4_K_M.gguf -c 262144 --ngl 999
Original model card
license: cc-by-nc-4.0 base_model: Qwen/Qwen3.8-27B pipeline_tag: text-generation library_name: transformers language: - en tags: - qwen3.8 - chat - creative-writing - altworld
Hemmingway-1
The AI that writes like a person. 27B parameters, open weights, free for non-commercial use.
Weights → · Try it → · Mac and Android apps → · Code → · Product Hunt →
Ask most models for a text to your landlord and you get three options, a preamble, and a paragraph explaining the options. Hemmingway-1 just gives you the text.
We built it for the writing people actually do every day: messages, emails, the awkward note to a colleague, the thing you have been putting off. Then we tested it against the biggest models in the world at exactly that, and it came first.
It writes the best everyday messages of any model we tested
We tested on eighty real requests. Every answer went head to head against another model's answer to the same request, shuffled so the judge never knew which was which.
It beats Fable 5.1, and it beats GPT-6 Astra by fifty points. Kimi K3, GLM-5.3, Grok 4.6 and DeepSeek V4 Pro all come in behind it. That is a 27B model.
And it is the one that sounds like a person
We ran the same matchups again with one question: which of these two did a person write?
It finished twenty-six points clear of the next model.
Where it wins
We broke it down by what you asked for. Higher means the judge more often took its version for the one a person wrote.
It wins on money and admin, work, the hard asks you keep rewriting, and talking someone round. Most of those by a wide margin. On hard asks, GPT-6 Astra gets 9%. Hemmingway-1 gets 72%.
It loses on hostile storytelling and long story turns. The story models are better at those, and that is fair.
You get the message, not a memo
One more thing we measured: how often a model buries the actual text in commentary, options and notes you have to read past.
Fable 5, GLM-5.3 and Kimi K3 bury it in more than nine replies out of ten.
It reads the room
EQ-Bench 4 is not ours. It is the public emotional-intelligence benchmark, run by its own harness.
It placed third, past GPT-5.5, Opus 4.7 and Opus 4.8, and inside twelve points of the best model on the board.
It tells a decent story too
It sits level with Kimi K3, comfortably past Qwen3.8-Max and DeepSeek V4 Pro, and 504 points above the model we started from.
Run it
vllm serve Altworld/Hemmingway-1 --max-model-len 262144
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Altworld/Hemmingway-1"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto", dtype="auto")
messages = [{"role": "user", "content": "Write the text I send my landlord about the broken boiler."}]
ids = tok.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=512)
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))
| Parameters | 27B |
| Built on | Qwen3.8-27B |
| Context | 262,144 tokens |
| Licence | CC BY-NC 4.0, free for non-commercial use; commercial use by agreement |
Licence
Hemmingway-1 is free for personal, research and other non-commercial use under CC BY-NC 4.0. Download it, run it, fine-tune it and share it, as long as you credit Hemmingway-1.
Commercial use needs a separate agreement with us. We are happy to work with anyone who reaches out: luka@hemmingway.io
The fine print
Full disclosure: CommunicationBench, Human-Likeness and StoryBench are our own benchmarks. We built them, we ran them, and we are saying that up front. Every matchup was blind and run in both orders so position could not sway the result, and the judge was a different model from the ones being judged. EQ-Bench 4 is not ours.
It is English-first. It can be wrong and still sound certain about it. So do not use it to decide anything medical, legal or financial.
- Downloads last month
- 143
4-bit





