Lukack commited on
Commit
6583e7b
·
verified ·
1 Parent(s): c064b23

announcement card with every benchmark charted

Browse files
.gitattributes CHANGED
@@ -34,3 +34,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
 
 
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
  tokenizer.json filter=lfs diff=lfs merge=lfs -text
37
+ charts/categories-heatmap.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -14,46 +14,87 @@ tags:
14
 
15
  # Hemmingway-1
16
 
17
- A 27B model that writes the way a person writes.
18
-
19
- Hemmingway-1 is built by [Altworld](https://hemmingway.io) on top of
20
- [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). It is trained for two
21
- things: writing that does not read as machine-written, and answering a person's
22
- whole message instead of a piece of it.
23
 
24
  **[Try it →](https://hemmingway.io)** · **[Mac and Android apps →](https://hemmingway.io/download)** · **[Code →](https://github.com/lukeckprobierts/Hemmingway-1)**
25
 
26
- ## What it is
 
 
 
27
 
28
- | | |
29
- |---|---|
30
- | Parameters | 27B |
31
- | Base | Qwen/Qwen3.8-27B |
32
- | Context | 262,144 tokens |
33
- | Licence | Apache-2.0 |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
 
35
- ## How it scores
 
 
36
 
37
- Our own benchmark. Every reply is judged against another model's reply to the
38
- same prompt, blind, in both orders, and the ratings are Bradley-Terry. The judge
39
- is GLM-5.3 at low thinking.
40
 
41
- | model | Elo |
42
- |---|---:|
43
- | **Hemmingway-1** | **1197** |
44
- | Kimi K3 | 1197 |
45
- | Qwen3.8-Max | 1081 |
46
- | DeepSeek V4 Pro | 1054 |
47
- | DeepSeek V4 Flash | 981 |
48
- | Gemma 4 31B | 808 |
49
- | Qwen3.8-27B (base) | 693 |
50
 
51
- Level with Kimi K3, and 504 points above the base model it was trained from.
52
- Closed frontier models still score above it.
53
 
54
- This is our own benchmark, run by us. Take it as ours, not as a neutral result.
55
 
56
- ## Running it
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
57
 
58
  ```bash
59
  vllm serve Altworld/Hemmingway-1 --max-model-len 262144
@@ -72,7 +113,20 @@ out = model.generate(ids, max_new_tokens=512)
72
  print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))
73
  ```
74
 
75
- ## Limits
 
 
 
 
 
 
 
 
 
 
 
 
 
76
 
77
- English first. It can be wrong and still sound certain. Not for medical, legal
78
- or financial decisions.
 
14
 
15
  # Hemmingway-1
16
 
17
+ **The AI that writes like a person.** 27B parameters, open weights, Apache-2.0.
 
 
 
 
 
18
 
19
  **[Try it →](https://hemmingway.io)** · **[Mac and Android apps →](https://hemmingway.io/download)** · **[Code →](https://github.com/lukeckprobierts/Hemmingway-1)**
20
 
21
+ Most models can write. Almost none can write the message you were actually
22
+ going to send. Ask one for a text to your landlord and you get three options, a
23
+ preamble, and a paragraph explaining the options. Hemmingway-1 gives you the
24
+ text.
25
 
26
+ We built it for the writing people do every day — messages, emails, the awkward
27
+ note to a colleague, the thing you've been putting off — and then we tested it
28
+ against the biggest models in the world at exactly that.
29
+
30
+ It came first.
31
+
32
+ ## It writes the best everyday messages of any model we tested
33
+
34
+ Eighty real requests. Every answer put head to head with another model's answer
35
+ to the same request, shuffled so the judge never knows which is which.
36
+
37
+ ![CommunicationBench](charts/communicationbench.png)
38
+
39
+ Ahead of Fable 5.1. Ahead of GPT-6 Astra by fifty points. Ahead of Kimi K3,
40
+ GLM-5.3, Grok 4.6 and DeepSeek V4 Pro. At 27B.
41
+
42
+ ## And it's the one that sounds like a person
43
+
44
+ Same matchups, one question: which of these two did a person write?
45
+
46
+ ![Human-Likeness](charts/human-likeness.png)
47
+
48
+ Twenty-six points clear of the next model. This is the whole point of
49
+ Hemmingway-1, and it's the number we're proudest of.
50
+
51
+ ## Where it wins
52
+
53
+ Broken down by what you actually asked for. Higher means the judge more often
54
+ took its version for the one a person wrote.
55
+
56
+ ![Where Hemmingway wins](charts/categories-heatmap.png)
57
 
58
+ Money and admin, work, the hard asks you keep rewriting, talking someone round
59
+ — it wins all of them, most by a wide margin. GPT-6 Astra gets 9% on hard asks.
60
+ Hemmingway-1 gets 72%.
61
 
62
+ Where it loses is hostile storytelling and long story turns. The story models
63
+ are better at those. We'd rather win your inbox.
 
64
 
65
+ ## You get the message, not a memo
 
 
 
 
 
 
 
 
66
 
67
+ How often a model buries the actual text in commentary, options and notes you
68
+ have to read past.
69
 
70
+ ![The message, not a memo](charts/wrapped.png)
71
 
72
+ Fable 5, GLM-5.3 and Kimi K3 do it to more than nine replies in ten.
73
+
74
+ ## It reads the room
75
+
76
+ EQ-Bench 4 is not ours. It's the public emotional-intelligence benchmark, run
77
+ by its own harness.
78
+
79
+ ![EQ-Bench 4](charts/eqbench4.png)
80
+
81
+ Third, past GPT-5.5, Opus 4.7 and Opus 4.8, and inside twelve points of the
82
+ best model on the board.
83
+
84
+ ## It tells a decent story too
85
+
86
+ ![StoryBench](charts/storybench.png)
87
+
88
+ Level with Kimi K3, comfortably past Qwen3.8-Max and DeepSeek V4 Pro, and 504
89
+ points above the model we started from.
90
+
91
+ ## How it got here
92
+
93
+ Every round, from our first 9B to this one.
94
+
95
+ ![The climb](charts/the-climb.png)
96
+
97
+ ## Run it
98
 
99
  ```bash
100
  vllm serve Altworld/Hemmingway-1 --max-model-len 262144
 
113
  print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))
114
  ```
115
 
116
+ | | |
117
+ |---|---|
118
+ | Parameters | 27B |
119
+ | Built on | Qwen3.8-27B |
120
+ | Context | 262,144 tokens |
121
+ | Licence | Apache-2.0 — yours to use, including commercially |
122
+
123
+ ## The fine print
124
+
125
+ CommunicationBench, Human-Likeness and StoryBench are our own benchmarks. We
126
+ built them, we ran them, and we're telling you that up front. Every matchup was
127
+ blind and run in both orders so position couldn't sway it, and the judge was a
128
+ different model from the ones being judged. EQ-Bench 4 and its slop meter are
129
+ not ours.
130
 
131
+ It's English-first. It can be wrong and still sound certain. Don't use it to
132
+ decide anything medical, legal or financial.
charts/categories-heatmap.png ADDED

Git LFS Details

  • SHA256: 28b780f883ae33522fe1628ce32020b27d1827d829e2f0a19d4440c07432910a
  • Pointer size: 131 Bytes
  • Size of remote file: 152 kB
charts/communicationbench.png ADDED
charts/eqbench4.png ADDED
charts/human-likeness.png ADDED
charts/storybench.png ADDED
charts/the-climb.png ADDED
charts/wrapped.png ADDED