jialinyyzz commited on
Commit
c07c06c
·
verified ·
1 Parent(s): 8a6975e

Note: an English draft occasionally came back in Chinese; app 0.3.1 and hz resample automatically. App link -> 0.3.1

Browse files
Files changed (4) hide show
  1. AGENTS.md +1 -0
  2. README.md +2 -1
  3. USAGE.md +5 -1
  4. USAGE.zh.md +5 -1
AGENTS.md CHANGED
@@ -9,6 +9,7 @@ Instructions for AI agents (Claude Code, Codex, Cursor, etc.) setting up **human
9
  - **Weights:** Hugging Face repo `jialinyyzz/humanizer`. Recommended file: `humanizer-12b-Q8_0.gguf` (12,669,630,368 bytes, about 12.7 GB); `humanizer-12b-Q6_K.gguf` (10,029,799,584 bytes, about 10.0 GB) for 16 GB machines; `humanizer-12b-Q4_K_M.gguf` (7,625,160,864 bytes, about 7.6 GB) for about 14 GB or when disk is tight; `humanizer-12b-Q3-QAT.gguf` (5,587,794,816 bytes, about 5.6 GB) for 12 GB; the smallest, `humanizer-12b-IQ2_XS-QAT.gguf` (3,893,632,896 bytes, about 3.9 GB, 2-bit), for 8 GB. Q3 and 2-bit make a few more fact slips. You also need `prompt_format.json` from the same repo.
10
  - **Runtime:** `llama-server` from llama.cpp (macOS, Windows, Linux). Alternatives in section 10.
11
  - **Sampling:** temperature 1.0, top_p 0.95, and nothing else: top_k 0, min_p 0, repeat_penalty 1.0. Stop on EOS only. No stop strings.
 
12
  - **License:** Apache 2.0.
13
 
14
  ## 1. Pick the route
 
9
  - **Weights:** Hugging Face repo `jialinyyzz/humanizer`. Recommended file: `humanizer-12b-Q8_0.gguf` (12,669,630,368 bytes, about 12.7 GB); `humanizer-12b-Q6_K.gguf` (10,029,799,584 bytes, about 10.0 GB) for 16 GB machines; `humanizer-12b-Q4_K_M.gguf` (7,625,160,864 bytes, about 7.6 GB) for about 14 GB or when disk is tight; `humanizer-12b-Q3-QAT.gguf` (5,587,794,816 bytes, about 5.6 GB) for 12 GB; the smallest, `humanizer-12b-IQ2_XS-QAT.gguf` (3,893,632,896 bytes, about 3.9 GB, 2-bit), for 8 GB. Q3 and 2-bit make a few more fact slips. You also need `prompt_format.json` from the same repo.
10
  - **Runtime:** `llama-server` from llama.cpp (macOS, Windows, Linux). Alternatives in section 10.
11
  - **Sampling:** temperature 1.0, top_p 0.95, and nothing else: top_k 0, min_p 0, repeat_penalty 1.0. Stop on EOS only. No stop strings.
12
+ - **Language check:** occasionally, on short informal English drafts with technical jargon, the model writes the whole rewrite in Chinese. If the draft is English and the rewrite has more than a few Chinese characters, sample again with the same settings (up to 3 times). The app (0.3.1 and later) and `hz` do this themselves.
13
  - **License:** Apache 2.0.
14
 
15
  ## 1. Pick the route
README.md CHANGED
@@ -36,7 +36,7 @@ tags:
36
 
37
  ## Quick start
38
 
39
- **App:** download the `.dmg` (Mac with Apple silicon) or the Windows installer from [Releases](https://github.com/sgaofen/humanizer-local-model/releases/latest). On first run it suggests one of five sizes for your memory (Q8_0, Q6_K, Q4_K_M, Q3 or 2-bit, down to 8 GB machines; app 0.3.0) and downloads it once; after that it works offline. Sizes and quality notes: [Quantized versions](#quantized-versions).
40
 
41
  **Command line, for long documents and agents:** `pipx install git+https://github.com/sgaofen/humanizer-local-model`, then `hz paper.md -o paper.out.md` (also `.txt` and `.docx`). It uses the app or a llama-server, keeps headings, code, tables and links, rewrites the prose piece by piece and flags any piece where a number went missing. See [USAGE.md, section 14](USAGE.md#14-hz-command-line-tool).
42
 
@@ -336,6 +336,7 @@ Measured on an M5 Max. llama.cpp Q8_0 with Metal (what the app uses): about 36
336
  - Chinese is still catching up with English (no factual problem in 149 of 204 Chinese rewrites; about 9 in 10 fixes are a single word or phrase).
337
  - Templated genres are still the hardest for detectors: emoji/hashtag social posts (3 / 16 flagged) and policy memos (2 / 13).
338
  - Formatting can change: 28 of 420 outputs dropped a format element; paragraph breaks, lists and headings sometimes merge or disappear.
 
339
  - In casual genres it sometimes adds slang or profanity that wasn't in the draft.
340
  - Detector results change over time. Nothing here guarantees any detector outcome.
341
  - It is a writing tool for your own drafts. Where a school, employer or publication has rules about AI assistance, follow them.
 
36
 
37
  ## Quick start
38
 
39
+ **App:** download the `.dmg` (Mac with Apple silicon) or the Windows installer from [Releases](https://github.com/sgaofen/humanizer-local-model/releases/latest) (current: [app 0.3.1](https://github.com/sgaofen/humanizer-local-model/releases/tag/app-v0.3.1)). On first run it suggests one of five sizes for your memory (Q8_0, Q6_K, Q4_K_M, Q3 or 2-bit, down to 8 GB machines; app 0.3.0 and later) and downloads it once; after that it works offline. Sizes and quality notes: [Quantized versions](#quantized-versions).
40
 
41
  **Command line, for long documents and agents:** `pipx install git+https://github.com/sgaofen/humanizer-local-model`, then `hz paper.md -o paper.out.md` (also `.txt` and `.docx`). It uses the app or a llama-server, keeps headings, code, tables and links, rewrites the prose piece by piece and flags any piece where a number went missing. See [USAGE.md, section 14](USAGE.md#14-hz-command-line-tool).
42
 
 
336
  - Chinese is still catching up with English (no factual problem in 149 of 204 Chinese rewrites; about 9 in 10 fixes are a single word or phrase).
337
  - Templated genres are still the hardest for detectors: emoji/hashtag social posts (3 / 16 flagged) and policy memos (2 / 13).
338
  - Formatting can change: 28 of 420 outputs dropped a format element; paragraph breaks, lists and headings sometimes merge or disappear.
339
+ - Occasionally answers in the wrong language: on short, informal English drafts with technical jargon, it occasionally writes the whole rewrite in Chinese. The app (0.3.1 and later) and `hz` check the language and sample again automatically. If you call the model yourself through llama.cpp or another runtime: when an English draft comes back with more than a few Chinese characters, sample once more with the same settings.
340
  - In casual genres it sometimes adds slang or profanity that wasn't in the draft.
341
  - Detector results change over time. Nothing here guarantees any detector outcome.
342
  - It is a writing tool for your own drafts. Where a school, employer or publication has rules about AI assistance, follow them.
USAGE.md CHANGED
@@ -97,7 +97,7 @@ Also in the repo:
97
 
98
  Compared draft by draft with bf16, Q8_0, Q6_K and Q4_K_M are within noise on the fact judge. Q3 and 2-bit make a few more fact slips in English (64 and 70 rewrites flagged, against 52 for bf16), mostly a single word or number; in Chinese Q3 is on par with bf16 (53 vs. 50 of 204 flagged). Full table, Chinese results and how these two were made: [README, Quantized versions](https://github.com/sgaofen/humanizer-local-model#quantized-versions) or the [GGUF repo](https://huggingface.co/jialinyyzz/humanizer-GGUF).
99
 
100
- **Q4_K_M was refined on 2026-10-04 with quantization-aware training:** same size and format, about 1/3 lower KL to the full-precision model than a standard Q4_K_M. ¹ Measured on a larger KL set (30 blocks of English drafts and rewrites), where the standard Q4_K_M scores 0.0203 and 94.5% (Chinese: 0.0146 vs. 0.0225). On the fact judge, compared draft by draft with the standard Q4_K_M, it is within noise: 58 vs. 56 of 420 English rewrites flagged, 162 vs. 163 problems listed by the second pass, more than 9 in 10 of them a single word or phrase. Q3 and 2-bit were measured on the same larger set (Chinese: 0.0318 and 0.106; Q3 on an A30, the others on an A100).
101
 
102
  **Download:**
103
 
@@ -583,6 +583,7 @@ The copy ratio here is a rough measure (share of the rewrite's 5-word or 5-chara
583
  | Output starts with "Sure", "Here is…", repeats the instruction, or doesn't stop | A generic chat template is in use: a GGUF downloaded before 2026-10-04 (no built-in template), a template changed in the app's settings, Ollama without the Modelfile from [section 7](#7-ollama), or a chat endpoint on the safetensors weights | Download the GGUF again, or use a completion endpoint (`/completion`, `/v1/completions`, Ollama `raw: true`) with the exact prompt from [section 1](#1-what-makes-this-model-different) |
584
  | Output contains `<start_of_turn>`, `<end_of_turn>` or similar markers | Same: chat formatting | Same fix |
585
  | Output stops at `###` or very early | A stop string is set | Remove all stop strings; rely on EOS |
 
586
  | Output is almost the same as the draft | Sampling luck, or temperature too low | Check temperature 1.0 and sample again |
587
  | Rambling, odd word choices, or repeated phrases | Wrong samplers (llama.cpp's default top-k 40 / min-p 0.05, the top-k 64 from `generation_config.json`, or a repetition penalty) | Set top-k 0, min-p 0, repetition penalty 1.0 explicitly |
588
  | Rewrite cut off mid-sentence | Output limit or context too small | Raise `n_predict` / `max_tokens`; give llama-server `-c 8192 -np 1`; split long drafts |
@@ -661,6 +662,8 @@ Each rewritten piece is checked. If it has one of these problems, the piece is r
661
 
662
  `added_numbers`, numbers in the rewrite that are not in the draft, are **reported but never retried**: the model sometimes does correct arithmetic ("cut costs from $480k to $305k" became "saved $175k"), and sometimes invents a figure. Either way, look at them.
663
 
 
 
664
  Pieces that still have a problem after the retry are listed on stderr with their line number (`.docx`: paragraph number) and, with `--json`, marked `"flagged": true`. The exit code is still 0.
665
 
666
  ### `--json` output
@@ -697,6 +700,7 @@ From a real run on a 1,200-word English Markdown article through the app (one of
697
  | `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | As in the table above, spelled as in the draft (missing) or the rewrite (added) |
698
  | `pieces[].retried` / `chosen` / `attempts` | Whether it was rewritten twice, which try was kept, and each try's checks |
699
  | `pieces[].seconds` | Model time for this piece, all tries together |
 
700
  | `pieces[].flagged` / `issues` | Whether the kept version still has a problem, and which |
701
  | `text` | The whole result, only when there is no `-o` and the input is not `.docx` |
702
 
 
97
 
98
  Compared draft by draft with bf16, Q8_0, Q6_K and Q4_K_M are within noise on the fact judge. Q3 and 2-bit make a few more fact slips in English (64 and 70 rewrites flagged, against 52 for bf16), mostly a single word or number; in Chinese Q3 is on par with bf16 (53 vs. 50 of 204 flagged). Full table, Chinese results and how these two were made: [README, Quantized versions](https://github.com/sgaofen/humanizer-local-model#quantized-versions) or the [GGUF repo](https://huggingface.co/jialinyyzz/humanizer-GGUF).
99
 
100
+ **Q4_K_M was refined on 2026-10-04 with quantization-aware training:** same size and format, about 1/3 lower KL to the full-precision model than a standard Q4_K_M. ¹ Measured on a larger KL set (30 blocks of English drafts and rewrites), where the standard Q4_K_M scores 0.0203 and 94.5% (Chinese: 0.0146 vs. 0.0225). On the fact judge, compared draft by draft with the standard Q4_K_M, it is within noise: 58 vs. 56 of 420 English rewrites flagged, 162 vs. 163 problems listed by the second pass, more than 9 in 10 of them a single word or phrase. Q3 and 2-bit were measured on the same larger set (Chinese: 0.0318 and 0.106; measured on A100 and A30 GPUs, which differ by about 3%).
101
 
102
  **Download:**
103
 
 
583
  | Output starts with "Sure", "Here is…", repeats the instruction, or doesn't stop | A generic chat template is in use: a GGUF downloaded before 2026-10-04 (no built-in template), a template changed in the app's settings, Ollama without the Modelfile from [section 7](#7-ollama), or a chat endpoint on the safetensors weights | Download the GGUF again, or use a completion endpoint (`/completion`, `/v1/completions`, Ollama `raw: true`) with the exact prompt from [section 1](#1-what-makes-this-model-different) |
584
  | Output contains `<start_of_turn>`, `<end_of_turn>` or similar markers | Same: chat formatting | Same fix |
585
  | Output stops at `###` or very early | A stop string is set | Remove all stop strings; rely on EOS |
586
+ | An English draft comes back in Chinese (occasionally, on short informal English drafts with technical jargon) | Sampling luck | Sample again with the same settings when the rewrite of an English draft has more than a few Chinese characters. The app (0.3.1 and later) and `hz` do this automatically, up to 3 times |
587
  | Output is almost the same as the draft | Sampling luck, or temperature too low | Check temperature 1.0 and sample again |
588
  | Rambling, odd word choices, or repeated phrases | Wrong samplers (llama.cpp's default top-k 40 / min-p 0.05, the top-k 64 from `generation_config.json`, or a repetition penalty) | Set top-k 0, min-p 0, repetition penalty 1.0 explicitly |
589
  | Rewrite cut off mid-sentence | Output limit or context too small | Raise `n_predict` / `max_tokens`; give llama-server `-c 8192 -np 1`; split long drafts |
 
662
 
663
  `added_numbers`, numbers in the rewrite that are not in the draft, are **reported but never retried**: the model sometimes does correct arithmetic ("cut costs from $480k to $305k" became "saved $175k"), and sometimes invents a figure. Either way, look at them.
664
 
665
+ **Wrong language.** Before these checks, a rewrite in the wrong language (an English draft written in Chinese, or a Chinese draft written in English; it only counts characters) is sampled again with the same settings, up to 3 times. `pieces[].language_resampled` counts those extra samples; if every sample is in the wrong language, the piece is flagged with `language`.
666
+
667
  Pieces that still have a problem after the retry are listed on stderr with their line number (`.docx`: paragraph number) and, with `--json`, marked `"flagged": true`. The exit code is still 0.
668
 
669
  ### `--json` output
 
700
  | `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | As in the table above, spelled as in the draft (missing) or the rewrite (added) |
701
  | `pieces[].retried` / `chosen` / `attempts` | Whether it was rewritten twice, which try was kept, and each try's checks |
702
  | `pieces[].seconds` | Model time for this piece, all tries together |
703
+ | `pieces[].language_resampled` | Extra samples taken because a rewrite came back in the wrong language (usually 0) |
704
  | `pieces[].flagged` / `issues` | Whether the kept version still has a problem, and which |
705
  | `text` | The whole result, only when there is no `-o` and the input is not `.docx` |
706
 
USAGE.zh.md CHANGED
@@ -97,7 +97,7 @@ assert hashlib.sha256(build_prompt("X").encode("utf-8")).hexdigest()[:16] == "cc
97
 
98
  逐篇和 bf16 对比,Q8_0、Q6_K、Q4_K_M 在事实判官上的差别都在噪声范围内。Q3 和 2 bit 英文事实小错稍多(标出 64 篇、70 篇,bf16 是 52 篇),多是一个词或一个数字;中文 Q3 和 bf16 持平(204 篇里标出 53 篇对 50 篇)。完整表格、中文结果和这两档怎么做的,见 [README 的量化版本一节](https://github.com/sgaofen/humanizer-local-model/blob/main/README.zh.md#量化版本)或 [GGUF 仓库](https://huggingface.co/jialinyyzz/humanizer-GGUF)。
99
 
100
- **Q4_K_M 在 2026-10-04 用量化感知训练重新做过一遍:**大小和格式不变,和全精度模型的 KL 比普通 Q4_K_M 低约三分之一。¹ 这两个数是在更大的一套 KL 测试上量的(30 块英文草稿和改写),普通 Q4_K_M 在同一套上是 0.0203 和 94.5%(中文:0.0146 对 0.0225)。事实判官上,和普通 Q4_K_M 逐篇对比在噪声范围内:英文 420 篇改写里标出 58 篇对 56 篇,第二遍复核列出的问题 162 处对 163 处,9 成以上只是一个词或短语。Q3 和 2 bit 也是在这套更大的测试上量的(中文:0.0318 和 0.106;Q3 在 A30 上测,其余在 A100 上)。
101
 
102
  **下载:**
103
 
@@ -583,6 +583,7 @@ if __name__ == "__main__":
583
  | 输出以“好的”“Sure”“Here is…”开头、复述指令,或者停不下来 | 套上了通用的聊天模板:GGUF 是 2026-10-04 以前下载的(没有自带模板)、软件设置里改过模板、Ollama 没用[第 7 节](#7-ollama)的 Modelfile,或者对 safetensors 权重用了聊天接口 | 重新下载 GGUF,或者用续写接口(`/completion`、`/v1/completions`、Ollama 的 `raw: true`),发[第 1 节](#1-这个模型哪里特殊)里逐字的提示词 |
584
  | 输出里有 `<start_of_turn>`、`<end_of_turn>` 之类的标记 | 同上:聊天格式 | 同上 |
585
  | 输出在 `###` 处或很早就停了 | 设了停止符 | 删掉所有停止符,只靠 EOS |
 
586
  | 改写和草稿几乎一样 | 采样运气不好,或温度太低 | 确认 temperature 1.0,再采一次 |
587
  | 胡言乱语、用词古怪或反复重复 | 采样参数不对(llama.cpp 默认的 top-k 40 / min-p 0.05、`generation_config.json` 里的 top-k 64,或者开了重复惩罚) | 显式设 top-k 0、min-p 0、重复惩罚 1.0 |
588
  | 改写在句子中间断了 | 输出上限或上下文太小 | 调大 `n_predict` / `max_tokens`;llama-server 加 `-c 8192 -np 1`;长稿分段 |
@@ -661,6 +662,8 @@ hz paper.md --server http://127.0.0.1:8080 # 指定某个 llama-server
661
 
662
  `added_numbers`(改写里有、草稿里没有的数字)**只报告、不重写**:模型有时是做了正确的算术("成本从 $480k 降到 $305k"被写成"省了 $175k"),有时是编出来的数字。不管哪种,都要看一眼。
663
 
 
 
664
  重写之后仍有问题的块会在 stderr 列出行号(`.docx` 为段落序号);用 `--json` 时标为 `"flagged": true`。这种情况退出码仍是 0。
665
 
666
  ### `--json` 输出
@@ -697,6 +700,7 @@ hz paper.md --server http://127.0.0.1:8080 # 指定某个 llama-server
697
  | `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | 见上表;缺失的按草稿里的写法,新增的按改写里的写法 |
698
  | `pieces[].retried` / `chosen` / `attempts` | 是否重写过、留的是第几版、每一版的检查结果 |
699
  | `pieces[].seconds` | 这块所有版本加起来的模型耗时 |
 
700
  | `pieces[].flagged` / `issues` | 留下的那版是否仍有问题,以及是哪些问题 |
701
  | `text` | 完整结果,只在没给 `-o`、输入又不是 `.docx` 时才有 |
702
 
 
97
 
98
  逐篇和 bf16 对比,Q8_0、Q6_K、Q4_K_M 在事实判官上的差别都在噪声范围内。Q3 和 2 bit 英文事实小错稍多(标出 64 篇、70 篇,bf16 是 52 篇),多是一个词或一个数字;中文 Q3 和 bf16 持平(204 篇里标出 53 篇对 50 篇)。完整表格、中文结果和这两档怎么做的,见 [README 的量化版本一节](https://github.com/sgaofen/humanizer-local-model/blob/main/README.zh.md#量化版本)或 [GGUF 仓库](https://huggingface.co/jialinyyzz/humanizer-GGUF)。
99
 
100
+ **Q4_K_M 在 2026-10-04 用量化感知训练重新做过一遍:**大小和格式不变,和全精度模型的 KL 比普通 Q4_K_M 低约三分之一。¹ 这两个数是在更大的一套 KL 测试上量的(30 块英文草稿和改写),普通 Q4_K_M 在同一套上是 0.0203 和 94.5%(中文:0.0146 对 0.0225)。事实判官上,和普通 Q4_K_M 逐篇对比在噪声范围内:英文 420 篇改写里标出 58 篇对 56 篇,第二遍复核列出的问题 162 处对 163 处,9 成以上只是一个词或短语。Q3 和 2 bit 也是在这套更大的测试上量的(中文:0.0318 和 0.106;在 A100 和 A30 两种卡上测,卡型之间只差约 3%)。
101
 
102
  **下载:**
103
 
 
583
  | 输出以“好的”“Sure”“Here is…”开头、复述指令,或者停不下来 | 套上了通用的聊天模板:GGUF 是 2026-10-04 以前下载的(没有自带模板)、软件设置里改过模板、Ollama 没用[第 7 节](#7-ollama)的 Modelfile,或者对 safetensors 权重用了聊天接口 | 重新下载 GGUF,或者用续写接口(`/completion`、`/v1/completions`、Ollama 的 `raw: true`),发[第 1 节](#1-这个模型哪里特殊)里逐字的提示词 |
584
  | 输出里有 `<start_of_turn>`、`<end_of_turn>` 之类的标记 | 同上:聊天格式 | 同上 |
585
  | 输出在 `###` 处或很早就停了 | 设了停止符 | 删掉所有停止符,只靠 EOS |
586
+ | 英文草稿被写成了中文(偶尔发生,多见于夹技术术语的英文口语短稿) | 采样运气 | 英文草稿的输出里出现大段中文,就用同样的参数再采一次。App(0.3.1 及以后)和 `hz` 会自动做,最多 3 次 |
587
  | 改写和草稿几乎一样 | 采样运气不好,或温度太低 | 确认 temperature 1.0,再采一次 |
588
  | 胡言乱语、用词古怪或反复重复 | 采样参数不对(llama.cpp 默认的 top-k 40 / min-p 0.05、`generation_config.json` 里的 top-k 64,或者开了重复惩罚) | 显式设 top-k 0、min-p 0、重复惩罚 1.0 |
589
  | 改写在句子中间断了 | 输出上限或上下文太小 | 调大 `n_predict` / `max_tokens`;llama-server 加 `-c 8192 -np 1`;长稿分段 |
 
662
 
663
  `added_numbers`(改写里有、草稿里没有的数字)**只报告、不重写**:模型有时是做了正确的算术("成本从 $480k 降到 $305k"被写成"省了 $175k"),有时是编出来的数字。不管哪种,都要看一眼。
664
 
665
+ **语言不对。**在这些检查之前,写成了另一种语言的改写(英文草稿写成中文,或中文草稿写成英文;只数字符)会用同样的参数重新生成,最多 3 次。`pieces[].language_resampled` 记重采了几次;3 次都不对的话,这块会标上 `language`。
666
+
667
  重写之后仍有问题的块会在 stderr 列出行号(`.docx` 为段落序号);用 `--json` 时标为 `"flagged": true`。这种情况退出码仍是 0。
668
 
669
  ### `--json` 输出
 
700
  | `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | 见上表;缺失的按草稿里的写法,新增的按改写里的写法 |
701
  | `pieces[].retried` / `chosen` / `attempts` | 是否重写过、留的是第几版、每一版的检查结果 |
702
  | `pieces[].seconds` | 这块所有版本加起来的模型耗时 |
703
+ | `pieces[].language_resampled` | 因为语言不对多采了几次(通常是 0) |
704
  | `pieces[].flagged` / `issues` | 留下的那版是否仍有问题,以及是哪些问题 |
705
  | `text` | 完整结果,只在没给 `-o`、输入又不是 `.docx` 时才有 |
706