Start here: prompt format, what it gets wrong, and what feedback we need

#1
by jialinyyzz - opened

Three things that will save you time, then one ask.

1. This is a base-model completion, not a chat model.

No system prompt, no chat template, no <start_of_turn> markers. You send the instruction, your draft, and the separator as one block of text and let the model continue. The exact wrapper ships in prompt_format.json next to the weights β€” reproduce it byte for byte, including the blank line after ### Rewritten:. A paraphrased instruction measurably degrades output.

If you are using Ollama or LM Studio, both apply the base model's chat template by default and it will break the format. The model card has a Modelfile that passes the prompt through untouched.

2. Use Q8_0 or Q6_K, not lower.

Q5_K_M and Q4_K_M are deliberately not released. Q5_K_M made 17 critical fidelity errors on the same 62-sample set where the bf16 model made 0; Q4_K_M produces gibberish. Fidelity collapses below 6-bit on this model. MLX 4-/6-bit is also unusable (Gemma 4 PLE layers) β€” on a Mac, use the GGUF quants.

3. Proofread the numbers.

The most common failure is a dropped qualifier, in roughly 1 output in 3: "an estimated 4.2 %" becomes "4.2 %", "suggest" becomes "conclude". Short drafts (under ~120 words) are less reliable, Chinese is weaker than English, and you will occasionally get one garbled sentence β€” resample it.

What we would like back.

The evaluation set is 39 everyday-writing cases and it is public in the GitHub repo. What we cannot see from here is where it breaks on text that is not ours. If it drops a fact, flips a claim, or produces something obviously machine-like on your material, please post the draft and the output in this thread β€” a single failing pair is more useful to us than a general impression. Genre labels help too (legal, medical, code-adjacent, non-English), since we know our coverage there is thin.

Code, eval set and reward function: https://github.com/sgaofen/humanizer

Sign up or log in to comment