We built an AI humanizer and never let it see a detector

Community Article
Published September 16, 2026

There is an obvious way to build a model that beats AI-text detectors: put a detector in the reward and optimise against it. We deliberately did not do that, and this post is about what we did instead, what it cost, and what the model still gets wrong.

You can try it here: https://huggingface.co/spaces/jialinyyzz/humanizer

The problem, stated honestly

An AI draft and a piece of human writing differ in two ways at once. One is style: the stock transitions, the balanced hedges, the tidy groups of three, the sentence that only announces what the next sentence will say. The other is structure: an AI draft tends to walk its points in the order it thought of them, one idea per sentence, one sentence per idea.

A rewriter has to change both while changing nothing about the facts. That constraint is what makes the task hard. Paraphrasing tools that only swap synonyms leave the structure intact and are trivially recognisable. Models that restructure freely start inventing — they smooth "an estimated 4.2 %" into "4.2 %", turn "the data suggest" into "the data show", and quietly flip a condition.

So the job splits in two, and we gave each half to a different training stage.

SFT teaches it to write like a person

The supervised stage is built on pairs. Crucially, the human side of every pair is a real piece of human writing that we never touched — not one word rewritten, not one sentence smoothed. The AI side is a draft that a large model wrote backwards from that human text: we handed the model the human piece and asked it to produce the AI-sounding draft that the human text would be a rewrite of.

This is worth spelling out because it is easy to get backwards. Nobody rewrote anything by hand. The supervision direction is draft → human original, and the human original is the fixed target, which is exactly what we want the model to learn to produce.

The generated drafts pass an acceptance gate before they become training data: the draft must not reuse the original's phrasing, must not map one-to-one onto its sentences, must preserve every fact in both directions, and must actually look like AI prose. Roughly one in four attempts survives. A later batch added drafts whose structure was mechanically scrambled relative to the original — the content units split out, reordered, and rewritten in that new order — which pushed the eval loss down from 0.907 to 0.873 on the same base and recipe.

What SFT does not give you is factual reliability. Every pure-SFT arm we evaluated made critical fidelity errors. That is the second half of the job.

RL teaches it to rewrite the same way every time, without losing facts

The reinforcement stage is GRPO on top of the merged SFT + DPO model, eight samples per draft. The reward has three parts:

Fidelity. An LLM judge extracts an atomic fact list from the draft and checks the rewrite against it. A reversed or contradicted fact, a changed number, name, date or deadline, a lost request, an invented claim: -3.0. A dropped specific or a shifted quantified boundary: -0.15. Invented content and reversals: -2.0 each. A dropped format element or greeting: -1.5.

Register. -1.0 if the rewrite normalises the draft's register — flattening slang, jokes or informality into generic prose. This term exists because without it the policy collapsed toward one safe voice, and our pass rate fell with it.

Reuse. A superlinear ramp on max(verbatim 5-gram copy, content-masked syntactic-skeleton recall). The second term is the important one: masking the content words and matching the remaining function-word skeleton catches a rewrite that swapped every noun but kept the sentence architecture. Reuse below 0.31 is free, and full copying costs -4.

Two details mattered more than we expected. First, errors are credited to the sentence that carries them — the judge returns the offending span and those tokens get roughly triple weight, instead of the penalty being smeared across the whole output. Second, more RL is not better: critical errors start climbing after about 225 steps, so the release checkpoint is picked from a 25-step ladder on a held-out set rather than taken from the end of the run.

What we refused to do

No detector appears anywhere in the reward, the filtering, or the sample selection. This is not modesty about our ability to do it; it is that the resulting number would mean nothing. Optimise against a classifier and you learn that classifier's decision boundary, which is a different and much smaller object than "writing like a person." The model would score beautifully on the detector it was trained against and no better than baseline on the next one.

We ran Originality.ai once, at the end, on the finished model, purely as an external read: 85 % of outputs come back "human", against 57 % for a baseline 4B rewriter. That number was never fed back into anything. If it had been, we would not report it.

Numbers

39-case everyday-writing set, 2 samples each, 62 English samples. Two independent judges (GLM-5.3 and gpt-5.6); an error counts only if both report it. v1 was re-graded under the same protocol.

v2 (SFT+DPO+RL) v1 (SFT+DPO) baseline 4B rewriter
critical fidelity errors 0 / 62 3 / 62 around 30 %
minor / clean 20 / 42 15 / 44 —
samples copying > 35 % of draft 0 1 33 / 93
reuse median 0.29 0.34 —
format element dropped 5 / 62 12 / 62 —
Chinese cases passing judge 13 / 16 11 / 16 4 / 16
Originality.ai "human" (external, once) 85 % 81 % 57 %

What it still gets wrong

This is the section we would read first in someone else's post, so it is not buried here either.

  • Dropped qualifiers, about 1 output in 3. "an estimated 4.2 %" becomes "4.2 %"; "suggest" becomes "conclude". This is the most common failure and it is exactly the kind a skimming reader misses.
    • Meaning flips were v1's main failure (about 1 in 10) and did not appear in v2's 62 samples — but 62 samples is a small set and we would not claim the rate is zero.
      • An occasional garbled sentence. Resample.
        • Chinese is weaker than English (13/16 versus the English rate).

          • Short drafts under about 120 words are less reliable.

          • Proofread every number, date and the direction of every claim. A rewriter that is right 97 % of the time still needs a reader.

        • Practical notes

      • It is a base-model completion, not a chat model. No system prompt, no chat template, no turn markers. The exact instruction and separator ship in prompt_format.json beside the weights and must be reproduced byte for byte — a paraphrased instruction or a missing blank line measurably degrades output. Ollama and LM Studio apply a chat template by default; turn it off (the model card has a Modelfile that passes the prompt through untouched).
    • Quantisation: GGUF Q8_0, Q6_K and bf16 are released. Q5_K_M is withheld — it produced 17 critical errors on the same 62 samples — and Q4_K_M is withheld as gibberish. Fidelity collapses below 6-bit on this model. MLX 4- and 6-bit are also unusable because of Gemma 4's PLE layers; on a Mac, use the GGUF quants.
  • Model: https://huggingface.co/jialinyyzz/humanizer-gemma-4-e4b Demo: https://huggingface.co/spaces/jialinyyzz/humanizer

Code, evaluation set, reward function and training scripts: https://github.com/sgaofen/humanizer

Community

Sign up or log in to comment