humanizer
Rewrite AI drafts to read like a person; facts kept. 12B
You can try it here: https://huggingface.co/spaces/jialinyyzz/humanizer
An AI draft and a piece of human writing differ in two ways at once. One is style: the stock transitions, the balanced hedges, the tidy groups of three, the sentence that only announces what the next sentence will say. The other is structure: an AI draft tends to walk its points in the order it thought of them, one idea per sentence, one sentence per idea.
A rewriter has to change both while changing nothing about the facts. That constraint is what makes the task hard. Paraphrasing tools that only swap synonyms leave the structure intact and are trivially recognisable. Models that restructure freely start inventing — they smooth "an estimated 4.2 %" into "4.2 %", turn "the data suggest" into "the data show", and quietly flip a condition.
So the job splits in two, and we gave each half to a different training stage.
The supervised stage is built on pairs. Crucially, the human side of every pair is a real piece of human writing that we never touched — not one word rewritten, not one sentence smoothed. The AI side is a draft that a large model wrote backwards from that human text: we handed the model the human piece and asked it to produce the AI-sounding draft that the human text would be a rewrite of.
This is worth spelling out because it is easy to get backwards. Nobody rewrote anything by hand. The supervision direction is draft → human original, and the human original is the fixed target, which is exactly what we want the model to learn to produce.
The generated drafts pass an acceptance gate before they become training data: the draft must not reuse the original's phrasing, must not map one-to-one onto its sentences, must preserve every fact in both directions, and must actually look like AI prose. Roughly one in four attempts survives. A later batch added drafts whose structure was mechanically scrambled relative to the original — the content units split out, reordered, and rewritten in that new order — which pushed the eval loss down from 0.907 to 0.873 on the same base and recipe.
What SFT does not give you is factual reliability. Every pure-SFT arm we evaluated made critical fidelity errors. That is the second half of the job.
The reinforcement stage is GRPO on top of the merged SFT + DPO model, eight samples per draft. The reward has three parts:
Fidelity. An LLM judge extracts an atomic fact list from the draft and checks the rewrite against it. A reversed or contradicted fact, a changed number, name, date or deadline, a lost request, an invented claim: -3.0. A dropped specific or a shifted quantified boundary: -0.15. Invented content and reversals: -2.0 each. A dropped format element or greeting: -1.5.
Register. -1.0 if the rewrite normalises the draft's register — flattening slang, jokes or informality into generic prose. This term exists because without it the policy collapsed toward one safe voice, and our pass rate fell with it.
Reuse. A superlinear ramp on max(verbatim 5-gram copy, content-masked syntactic-skeleton recall). The second term is the important one: masking the content words and matching the remaining function-word skeleton catches a rewrite that swapped every noun but kept the sentence architecture. Reuse below 0.31 is free, and full copying costs -4.
Two details mattered more than we expected. First, errors are credited to the sentence that carries them — the judge returns the offending span and those tokens get roughly triple weight, instead of the penalty being smeared across the whole output. Second, more RL is not better: critical errors start climbing after about 225 steps, so the release checkpoint is picked from a 25-step ladder on a held-out set rather than taken from the end of the run.
No detector appears anywhere in the reward, the filtering, or the sample selection. This is not modesty about our ability to do it; it is that the resulting number would mean nothing. Optimise against a classifier and you learn that classifier's decision boundary, which is a different and much smaller object than "writing like a person." The model would score beautifully on the detector it was trained against and no better than baseline on the next one.
We ran Originality.ai once, at the end, on the finished model, purely as an external read: 85 % of outputs come back "human", against 57 % for a baseline 4B rewriter. That number was never fed back into anything. If it had been, we would not report it.
39-case everyday-writing set, 2 samples each, 62 English samples. Two independent judges (GLM-5.3 and gpt-5.6); an error counts only if both report it. v1 was re-graded under the same protocol.
| v2 (SFT+DPO+RL) | v1 (SFT+DPO) | baseline 4B rewriter | |
|---|---|---|---|
| critical fidelity errors | 0 / 62 | 3 / 62 | around 30 % |
| minor / clean | 20 / 42 | 15 / 44 | — |
| samples copying > 35 % of draft | 0 | 1 | 33 / 93 |
| reuse median | 0.29 | 0.34 | — |
| format element dropped | 5 / 62 | 12 / 62 | — |
| Chinese cases passing judge | 13 / 16 | 11 / 16 | 4 / 16 |
| Originality.ai "human" (external, once) | 85 % | 81 % | 57 % |
This is the section we would read first in someone else's post, so it is not buried here either.
Code, evaluation set, reward function and training scripts: https://github.com/sgaofen/humanizer
Rewrite AI drafts to read like a person; facts kept. 12B