Text Generation
GGUF
Safetensors
Transformers
English
Chinese
gemma4_unified
image-text-to-text
humanizer
text-rewriting
rewriting
paraphrase
style-transfer
llama.cpp
gemma4
conversational
Instructions to use jialinyyzz/humanizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jialinyyzz/humanizer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jialinyyzz/humanizer")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("jialinyyzz/humanizer") model = AutoModelForMultimodalLM.from_pretrained("jialinyyzz/humanizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jialinyyzz/humanizer with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: llama cli -hf jialinyyzz/humanizer:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: llama cli -hf jialinyyzz/humanizer:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf jialinyyzz/humanizer:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf jialinyyzz/humanizer:Q4_K_M
Use Docker
docker model run hf.co/jialinyyzz/humanizer:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use jialinyyzz/humanizer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jialinyyzz/humanizer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jialinyyzz/humanizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jialinyyzz/humanizer:Q4_K_M
- SGLang
How to use jialinyyzz/humanizer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jialinyyzz/humanizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jialinyyzz/humanizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jialinyyzz/humanizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jialinyyzz/humanizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use jialinyyzz/humanizer with Ollama:
ollama run hf.co/jialinyyzz/humanizer:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use jialinyyzz/humanizer with Docker Model Runner:
docker model run hf.co/jialinyyzz/humanizer:Q4_K_M
- Lemonade
How to use jialinyyzz/humanizer with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jialinyyzz/humanizer:Q4_K_M
Run and chat with the model
lemonade run user.humanizer-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Note: an English draft occasionally came back in Chinese; app 0.3.1 and hz resample automatically. App link -> 0.3.1
Browse files
AGENTS.md
CHANGED
|
@@ -9,6 +9,7 @@ Instructions for AI agents (Claude Code, Codex, Cursor, etc.) setting up **human
|
|
| 9 |
- **Weights:** Hugging Face repo `jialinyyzz/humanizer`. Recommended file: `humanizer-12b-Q8_0.gguf` (12,669,630,368 bytes, about 12.7 GB); `humanizer-12b-Q6_K.gguf` (10,029,799,584 bytes, about 10.0 GB) for 16 GB machines; `humanizer-12b-Q4_K_M.gguf` (7,625,160,864 bytes, about 7.6 GB) for about 14 GB or when disk is tight; `humanizer-12b-Q3-QAT.gguf` (5,587,794,816 bytes, about 5.6 GB) for 12 GB; the smallest, `humanizer-12b-IQ2_XS-QAT.gguf` (3,893,632,896 bytes, about 3.9 GB, 2-bit), for 8 GB. Q3 and 2-bit make a few more fact slips. You also need `prompt_format.json` from the same repo.
|
| 10 |
- **Runtime:** `llama-server` from llama.cpp (macOS, Windows, Linux). Alternatives in section 10.
|
| 11 |
- **Sampling:** temperature 1.0, top_p 0.95, and nothing else: top_k 0, min_p 0, repeat_penalty 1.0. Stop on EOS only. No stop strings.
|
|
|
|
| 12 |
- **License:** Apache 2.0.
|
| 13 |
|
| 14 |
## 1. Pick the route
|
|
|
|
| 9 |
- **Weights:** Hugging Face repo `jialinyyzz/humanizer`. Recommended file: `humanizer-12b-Q8_0.gguf` (12,669,630,368 bytes, about 12.7 GB); `humanizer-12b-Q6_K.gguf` (10,029,799,584 bytes, about 10.0 GB) for 16 GB machines; `humanizer-12b-Q4_K_M.gguf` (7,625,160,864 bytes, about 7.6 GB) for about 14 GB or when disk is tight; `humanizer-12b-Q3-QAT.gguf` (5,587,794,816 bytes, about 5.6 GB) for 12 GB; the smallest, `humanizer-12b-IQ2_XS-QAT.gguf` (3,893,632,896 bytes, about 3.9 GB, 2-bit), for 8 GB. Q3 and 2-bit make a few more fact slips. You also need `prompt_format.json` from the same repo.
|
| 10 |
- **Runtime:** `llama-server` from llama.cpp (macOS, Windows, Linux). Alternatives in section 10.
|
| 11 |
- **Sampling:** temperature 1.0, top_p 0.95, and nothing else: top_k 0, min_p 0, repeat_penalty 1.0. Stop on EOS only. No stop strings.
|
| 12 |
+
- **Language check:** occasionally, on short informal English drafts with technical jargon, the model writes the whole rewrite in Chinese. If the draft is English and the rewrite has more than a few Chinese characters, sample again with the same settings (up to 3 times). The app (0.3.1 and later) and `hz` do this themselves.
|
| 13 |
- **License:** Apache 2.0.
|
| 14 |
|
| 15 |
## 1. Pick the route
|
README.md
CHANGED
|
@@ -36,7 +36,7 @@ tags:
|
|
| 36 |
|
| 37 |
## Quick start
|
| 38 |
|
| 39 |
-
**App:** download the `.dmg` (Mac with Apple silicon) or the Windows installer from [Releases](https://github.com/sgaofen/humanizer-local-model/releases/latest). On first run it suggests one of five sizes for your memory (Q8_0, Q6_K, Q4_K_M, Q3 or 2-bit, down to 8 GB machines; app 0.3.0) and downloads it once; after that it works offline. Sizes and quality notes: [Quantized versions](#quantized-versions).
|
| 40 |
|
| 41 |
**Command line, for long documents and agents:** `pipx install git+https://github.com/sgaofen/humanizer-local-model`, then `hz paper.md -o paper.out.md` (also `.txt` and `.docx`). It uses the app or a llama-server, keeps headings, code, tables and links, rewrites the prose piece by piece and flags any piece where a number went missing. See [USAGE.md, section 14](USAGE.md#14-hz-command-line-tool).
|
| 42 |
|
|
@@ -336,6 +336,7 @@ Measured on an M5 Max. llama.cpp Q8_0 with Metal (what the app uses): about 36
|
|
| 336 |
- Chinese is still catching up with English (no factual problem in 149 of 204 Chinese rewrites; about 9 in 10 fixes are a single word or phrase).
|
| 337 |
- Templated genres are still the hardest for detectors: emoji/hashtag social posts (3 / 16 flagged) and policy memos (2 / 13).
|
| 338 |
- Formatting can change: 28 of 420 outputs dropped a format element; paragraph breaks, lists and headings sometimes merge or disappear.
|
|
|
|
| 339 |
- In casual genres it sometimes adds slang or profanity that wasn't in the draft.
|
| 340 |
- Detector results change over time. Nothing here guarantees any detector outcome.
|
| 341 |
- It is a writing tool for your own drafts. Where a school, employer or publication has rules about AI assistance, follow them.
|
|
|
|
| 36 |
|
| 37 |
## Quick start
|
| 38 |
|
| 39 |
+
**App:** download the `.dmg` (Mac with Apple silicon) or the Windows installer from [Releases](https://github.com/sgaofen/humanizer-local-model/releases/latest) (current: [app 0.3.1](https://github.com/sgaofen/humanizer-local-model/releases/tag/app-v0.3.1)). On first run it suggests one of five sizes for your memory (Q8_0, Q6_K, Q4_K_M, Q3 or 2-bit, down to 8 GB machines; app 0.3.0 and later) and downloads it once; after that it works offline. Sizes and quality notes: [Quantized versions](#quantized-versions).
|
| 40 |
|
| 41 |
**Command line, for long documents and agents:** `pipx install git+https://github.com/sgaofen/humanizer-local-model`, then `hz paper.md -o paper.out.md` (also `.txt` and `.docx`). It uses the app or a llama-server, keeps headings, code, tables and links, rewrites the prose piece by piece and flags any piece where a number went missing. See [USAGE.md, section 14](USAGE.md#14-hz-command-line-tool).
|
| 42 |
|
|
|
|
| 336 |
- Chinese is still catching up with English (no factual problem in 149 of 204 Chinese rewrites; about 9 in 10 fixes are a single word or phrase).
|
| 337 |
- Templated genres are still the hardest for detectors: emoji/hashtag social posts (3 / 16 flagged) and policy memos (2 / 13).
|
| 338 |
- Formatting can change: 28 of 420 outputs dropped a format element; paragraph breaks, lists and headings sometimes merge or disappear.
|
| 339 |
+
- Occasionally answers in the wrong language: on short, informal English drafts with technical jargon, it occasionally writes the whole rewrite in Chinese. The app (0.3.1 and later) and `hz` check the language and sample again automatically. If you call the model yourself through llama.cpp or another runtime: when an English draft comes back with more than a few Chinese characters, sample once more with the same settings.
|
| 340 |
- In casual genres it sometimes adds slang or profanity that wasn't in the draft.
|
| 341 |
- Detector results change over time. Nothing here guarantees any detector outcome.
|
| 342 |
- It is a writing tool for your own drafts. Where a school, employer or publication has rules about AI assistance, follow them.
|
USAGE.md
CHANGED
|
@@ -97,7 +97,7 @@ Also in the repo:
|
|
| 97 |
|
| 98 |
Compared draft by draft with bf16, Q8_0, Q6_K and Q4_K_M are within noise on the fact judge. Q3 and 2-bit make a few more fact slips in English (64 and 70 rewrites flagged, against 52 for bf16), mostly a single word or number; in Chinese Q3 is on par with bf16 (53 vs. 50 of 204 flagged). Full table, Chinese results and how these two were made: [README, Quantized versions](https://github.com/sgaofen/humanizer-local-model#quantized-versions) or the [GGUF repo](https://huggingface.co/jialinyyzz/humanizer-GGUF).
|
| 99 |
|
| 100 |
-
**Q4_K_M was refined on 2026-10-04 with quantization-aware training:** same size and format, about 1/3 lower KL to the full-precision model than a standard Q4_K_M. ¹ Measured on a larger KL set (30 blocks of English drafts and rewrites), where the standard Q4_K_M scores 0.0203 and 94.5% (Chinese: 0.0146 vs. 0.0225). On the fact judge, compared draft by draft with the standard Q4_K_M, it is within noise: 58 vs. 56 of 420 English rewrites flagged, 162 vs. 163 problems listed by the second pass, more than 9 in 10 of them a single word or phrase. Q3 and 2-bit were measured on the same larger set (Chinese: 0.0318 and 0.106;
|
| 101 |
|
| 102 |
**Download:**
|
| 103 |
|
|
@@ -583,6 +583,7 @@ The copy ratio here is a rough measure (share of the rewrite's 5-word or 5-chara
|
|
| 583 |
| Output starts with "Sure", "Here is…", repeats the instruction, or doesn't stop | A generic chat template is in use: a GGUF downloaded before 2026-10-04 (no built-in template), a template changed in the app's settings, Ollama without the Modelfile from [section 7](#7-ollama), or a chat endpoint on the safetensors weights | Download the GGUF again, or use a completion endpoint (`/completion`, `/v1/completions`, Ollama `raw: true`) with the exact prompt from [section 1](#1-what-makes-this-model-different) |
|
| 584 |
| Output contains `<start_of_turn>`, `<end_of_turn>` or similar markers | Same: chat formatting | Same fix |
|
| 585 |
| Output stops at `###` or very early | A stop string is set | Remove all stop strings; rely on EOS |
|
|
|
|
| 586 |
| Output is almost the same as the draft | Sampling luck, or temperature too low | Check temperature 1.0 and sample again |
|
| 587 |
| Rambling, odd word choices, or repeated phrases | Wrong samplers (llama.cpp's default top-k 40 / min-p 0.05, the top-k 64 from `generation_config.json`, or a repetition penalty) | Set top-k 0, min-p 0, repetition penalty 1.0 explicitly |
|
| 588 |
| Rewrite cut off mid-sentence | Output limit or context too small | Raise `n_predict` / `max_tokens`; give llama-server `-c 8192 -np 1`; split long drafts |
|
|
@@ -661,6 +662,8 @@ Each rewritten piece is checked. If it has one of these problems, the piece is r
|
|
| 661 |
|
| 662 |
`added_numbers`, numbers in the rewrite that are not in the draft, are **reported but never retried**: the model sometimes does correct arithmetic ("cut costs from $480k to $305k" became "saved $175k"), and sometimes invents a figure. Either way, look at them.
|
| 663 |
|
|
|
|
|
|
|
| 664 |
Pieces that still have a problem after the retry are listed on stderr with their line number (`.docx`: paragraph number) and, with `--json`, marked `"flagged": true`. The exit code is still 0.
|
| 665 |
|
| 666 |
### `--json` output
|
|
@@ -697,6 +700,7 @@ From a real run on a 1,200-word English Markdown article through the app (one of
|
|
| 697 |
| `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | As in the table above, spelled as in the draft (missing) or the rewrite (added) |
|
| 698 |
| `pieces[].retried` / `chosen` / `attempts` | Whether it was rewritten twice, which try was kept, and each try's checks |
|
| 699 |
| `pieces[].seconds` | Model time for this piece, all tries together |
|
|
|
|
| 700 |
| `pieces[].flagged` / `issues` | Whether the kept version still has a problem, and which |
|
| 701 |
| `text` | The whole result, only when there is no `-o` and the input is not `.docx` |
|
| 702 |
|
|
|
|
| 97 |
|
| 98 |
Compared draft by draft with bf16, Q8_0, Q6_K and Q4_K_M are within noise on the fact judge. Q3 and 2-bit make a few more fact slips in English (64 and 70 rewrites flagged, against 52 for bf16), mostly a single word or number; in Chinese Q3 is on par with bf16 (53 vs. 50 of 204 flagged). Full table, Chinese results and how these two were made: [README, Quantized versions](https://github.com/sgaofen/humanizer-local-model#quantized-versions) or the [GGUF repo](https://huggingface.co/jialinyyzz/humanizer-GGUF).
|
| 99 |
|
| 100 |
+
**Q4_K_M was refined on 2026-10-04 with quantization-aware training:** same size and format, about 1/3 lower KL to the full-precision model than a standard Q4_K_M. ¹ Measured on a larger KL set (30 blocks of English drafts and rewrites), where the standard Q4_K_M scores 0.0203 and 94.5% (Chinese: 0.0146 vs. 0.0225). On the fact judge, compared draft by draft with the standard Q4_K_M, it is within noise: 58 vs. 56 of 420 English rewrites flagged, 162 vs. 163 problems listed by the second pass, more than 9 in 10 of them a single word or phrase. Q3 and 2-bit were measured on the same larger set (Chinese: 0.0318 and 0.106; measured on A100 and A30 GPUs, which differ by about 3%).
|
| 101 |
|
| 102 |
**Download:**
|
| 103 |
|
|
|
|
| 583 |
| Output starts with "Sure", "Here is…", repeats the instruction, or doesn't stop | A generic chat template is in use: a GGUF downloaded before 2026-10-04 (no built-in template), a template changed in the app's settings, Ollama without the Modelfile from [section 7](#7-ollama), or a chat endpoint on the safetensors weights | Download the GGUF again, or use a completion endpoint (`/completion`, `/v1/completions`, Ollama `raw: true`) with the exact prompt from [section 1](#1-what-makes-this-model-different) |
|
| 584 |
| Output contains `<start_of_turn>`, `<end_of_turn>` or similar markers | Same: chat formatting | Same fix |
|
| 585 |
| Output stops at `###` or very early | A stop string is set | Remove all stop strings; rely on EOS |
|
| 586 |
+
| An English draft comes back in Chinese (occasionally, on short informal English drafts with technical jargon) | Sampling luck | Sample again with the same settings when the rewrite of an English draft has more than a few Chinese characters. The app (0.3.1 and later) and `hz` do this automatically, up to 3 times |
|
| 587 |
| Output is almost the same as the draft | Sampling luck, or temperature too low | Check temperature 1.0 and sample again |
|
| 588 |
| Rambling, odd word choices, or repeated phrases | Wrong samplers (llama.cpp's default top-k 40 / min-p 0.05, the top-k 64 from `generation_config.json`, or a repetition penalty) | Set top-k 0, min-p 0, repetition penalty 1.0 explicitly |
|
| 589 |
| Rewrite cut off mid-sentence | Output limit or context too small | Raise `n_predict` / `max_tokens`; give llama-server `-c 8192 -np 1`; split long drafts |
|
|
|
|
| 662 |
|
| 663 |
`added_numbers`, numbers in the rewrite that are not in the draft, are **reported but never retried**: the model sometimes does correct arithmetic ("cut costs from $480k to $305k" became "saved $175k"), and sometimes invents a figure. Either way, look at them.
|
| 664 |
|
| 665 |
+
**Wrong language.** Before these checks, a rewrite in the wrong language (an English draft written in Chinese, or a Chinese draft written in English; it only counts characters) is sampled again with the same settings, up to 3 times. `pieces[].language_resampled` counts those extra samples; if every sample is in the wrong language, the piece is flagged with `language`.
|
| 666 |
+
|
| 667 |
Pieces that still have a problem after the retry are listed on stderr with their line number (`.docx`: paragraph number) and, with `--json`, marked `"flagged": true`. The exit code is still 0.
|
| 668 |
|
| 669 |
### `--json` output
|
|
|
|
| 700 |
| `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | As in the table above, spelled as in the draft (missing) or the rewrite (added) |
|
| 701 |
| `pieces[].retried` / `chosen` / `attempts` | Whether it was rewritten twice, which try was kept, and each try's checks |
|
| 702 |
| `pieces[].seconds` | Model time for this piece, all tries together |
|
| 703 |
+
| `pieces[].language_resampled` | Extra samples taken because a rewrite came back in the wrong language (usually 0) |
|
| 704 |
| `pieces[].flagged` / `issues` | Whether the kept version still has a problem, and which |
|
| 705 |
| `text` | The whole result, only when there is no `-o` and the input is not `.docx` |
|
| 706 |
|
USAGE.zh.md
CHANGED
|
@@ -97,7 +97,7 @@ assert hashlib.sha256(build_prompt("X").encode("utf-8")).hexdigest()[:16] == "cc
|
|
| 97 |
|
| 98 |
逐篇和 bf16 对比,Q8_0、Q6_K、Q4_K_M 在事实判官上的差别都在噪声范围内。Q3 和 2 bit 英文事实小错稍多(标出 64 篇、70 篇,bf16 是 52 篇),多是一个词或一个数字;中文 Q3 和 bf16 持平(204 篇里标出 53 篇对 50 篇)。完整表格、中文结果和这两档怎么做的,见 [README 的量化版本一节](https://github.com/sgaofen/humanizer-local-model/blob/main/README.zh.md#量化版本)或 [GGUF 仓库](https://huggingface.co/jialinyyzz/humanizer-GGUF)。
|
| 99 |
|
| 100 |
-
**Q4_K_M 在 2026-10-04 用量化感知训练重新做过一遍:**大小和格式不变,和全精度模型的 KL 比普通 Q4_K_M 低约三分之一。¹ 这两个数是在更大的一套 KL 测试上量的(30 块英文草稿和改写),普通 Q4_K_M 在同一套上是 0.0203 和 94.5%(中文:0.0146 对 0.0225)。事实判官上,和普通 Q4_K_M 逐篇对比在噪声范围内:英文 420 篇改写里标出 58 篇对 56 篇,第二遍复核列出的问题 162 处对 163 处,9 成以上只是一个词或短语。Q3 和 2 bit 也是在这套更大的测试上量的(中文:0.0318 和 0.106;
|
| 101 |
|
| 102 |
**下载:**
|
| 103 |
|
|
@@ -583,6 +583,7 @@ if __name__ == "__main__":
|
|
| 583 |
| 输出以“好的”“Sure”“Here is…”开头、复述指令,或者停不下来 | 套上了通用的聊天模板:GGUF 是 2026-10-04 以前下载的(没有自带模板)、软件设置里改过模板、Ollama 没用[第 7 节](#7-ollama)的 Modelfile,或者对 safetensors 权重用了聊天接口 | 重新下载 GGUF,或者用续写接口(`/completion`、`/v1/completions`、Ollama 的 `raw: true`),发[第 1 节](#1-这个模型哪里特殊)里逐字的提示词 |
|
| 584 |
| 输出里有 `<start_of_turn>`、`<end_of_turn>` 之类的标记 | 同上:聊天格式 | 同上 |
|
| 585 |
| 输出在 `###` 处或很早就停了 | 设了停止符 | 删掉所有停止符,只靠 EOS |
|
|
|
|
| 586 |
| 改写和草稿几乎一样 | 采样运气不好,或温度太低 | 确认 temperature 1.0,再采一次 |
|
| 587 |
| 胡言乱语、用词古怪或反复重复 | 采样参数不对(llama.cpp 默认的 top-k 40 / min-p 0.05、`generation_config.json` 里的 top-k 64,或者开了重复惩罚) | 显式设 top-k 0、min-p 0、重复惩罚 1.0 |
|
| 588 |
| 改写在句子中间断了 | 输出上限或上下文太小 | 调大 `n_predict` / `max_tokens`;llama-server 加 `-c 8192 -np 1`;长稿分段 |
|
|
@@ -661,6 +662,8 @@ hz paper.md --server http://127.0.0.1:8080 # 指定某个 llama-server
|
|
| 661 |
|
| 662 |
`added_numbers`(改写里有、草稿里没有的数字)**只报告、不重写**:模型有时是做了正确的算术("成本从 $480k 降到 $305k"被写成"省了 $175k"),有时是编出来的数字。不管哪种,都要看一眼。
|
| 663 |
|
|
|
|
|
|
|
| 664 |
重写之后仍有问题的块会在 stderr 列出行号(`.docx` 为段落序号);用 `--json` 时标为 `"flagged": true`。这种情况退出码仍是 0。
|
| 665 |
|
| 666 |
### `--json` 输出
|
|
@@ -697,6 +700,7 @@ hz paper.md --server http://127.0.0.1:8080 # 指定某个 llama-server
|
|
| 697 |
| `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | 见上表;缺失的按草稿里的写法,新增的按改写里的写法 |
|
| 698 |
| `pieces[].retried` / `chosen` / `attempts` | 是否重写过、留的是第几版、每一版的检查结果 |
|
| 699 |
| `pieces[].seconds` | 这块所有版本加起来的模型耗时 |
|
|
|
|
| 700 |
| `pieces[].flagged` / `issues` | 留下的那版是否仍有问题,以及是哪些问题 |
|
| 701 |
| `text` | 完整结果,只在没给 `-o`、输入又不是 `.docx` 时才有 |
|
| 702 |
|
|
|
|
| 97 |
|
| 98 |
逐篇和 bf16 对比,Q8_0、Q6_K、Q4_K_M 在事实判官上的差别都在噪声范围内。Q3 和 2 bit 英文事实小错稍多(标出 64 篇、70 篇,bf16 是 52 篇),多是一个词或一个数字;中文 Q3 和 bf16 持平(204 篇里标出 53 篇对 50 篇)。完整表格、中文结果和这两档怎么做的,见 [README 的量化版本一节](https://github.com/sgaofen/humanizer-local-model/blob/main/README.zh.md#量化版本)或 [GGUF 仓库](https://huggingface.co/jialinyyzz/humanizer-GGUF)。
|
| 99 |
|
| 100 |
+
**Q4_K_M 在 2026-10-04 用量化感知训练重新做过一遍:**大小和格式不变,和全精度模型的 KL 比普通 Q4_K_M 低约三分之一。¹ 这两个数是在更大的一套 KL 测试上量的(30 块英文草稿和改写),普通 Q4_K_M 在同一套上是 0.0203 和 94.5%(中文:0.0146 对 0.0225)。事实判官上,和普通 Q4_K_M 逐篇对比在噪声范围内:英文 420 篇改写里标出 58 篇对 56 篇,第二遍复核列出的问题 162 处对 163 处,9 成以上只是一个词或短语。Q3 和 2 bit 也是在这套更大的测试上量的(中文:0.0318 和 0.106;在 A100 和 A30 两种卡上测,卡型之间只差约 3%)。
|
| 101 |
|
| 102 |
**下载:**
|
| 103 |
|
|
|
|
| 583 |
| 输出以“好的”“Sure”“Here is…”开头、复述指令,或者停不下来 | 套上了通用的聊天模板:GGUF 是 2026-10-04 以前下载的(没有自带模板)、软件设置里改过模板、Ollama 没用[第 7 节](#7-ollama)的 Modelfile,或者对 safetensors 权重用了聊天接口 | 重新下载 GGUF,或者用续写接口(`/completion`、`/v1/completions`、Ollama 的 `raw: true`),发[第 1 节](#1-这个模型哪里特殊)里逐字的提示词 |
|
| 584 |
| 输出里有 `<start_of_turn>`、`<end_of_turn>` 之类的标记 | 同上:聊天格式 | 同上 |
|
| 585 |
| 输出在 `###` 处或很早就停了 | 设了停止符 | 删掉所有停止符,只靠 EOS |
|
| 586 |
+
| 英文草稿被写成了中文(偶尔发生,多见于夹技术术语的英文口语短稿) | 采样运气 | 英文草稿的输出里出现大段中文,就用同样的参数再采一次。App(0.3.1 及以后)和 `hz` 会自动做,最多 3 次 |
|
| 587 |
| 改写和草稿几乎一样 | 采样运气不好,或温度太低 | 确认 temperature 1.0,再采一次 |
|
| 588 |
| 胡言乱语、用词古怪或反复重复 | 采样参数不对(llama.cpp 默认的 top-k 40 / min-p 0.05、`generation_config.json` 里的 top-k 64,或者开了重复惩罚) | 显式设 top-k 0、min-p 0、重复惩罚 1.0 |
|
| 589 |
| 改写在句子中间断了 | 输出上限或上下文太小 | 调大 `n_predict` / `max_tokens`;llama-server 加 `-c 8192 -np 1`;长稿分段 |
|
|
|
|
| 662 |
|
| 663 |
`added_numbers`(改写里有、草稿里没有的数字)**只报告、不重写**:模型有时是做了正确的算术("成本从 $480k 降到 $305k"被写成"省了 $175k"),有时是编出来的数字。不管哪种,都要看一眼。
|
| 664 |
|
| 665 |
+
**语言不对。**在这些检查之前,写成了另一种语言的改写(英文草稿写成中文,或中文草稿写成英文;只数字符)会用同样的参数重新生成,最多 3 次。`pieces[].language_resampled` 记重采了几次;3 次都不对的话,这块会标上 `language`。
|
| 666 |
+
|
| 667 |
重写之后仍有问题的块会在 stderr 列出行号(`.docx` 为段落序号);用 `--json` 时标为 `"flagged": true`。这种情况退出码仍是 0。
|
| 668 |
|
| 669 |
### `--json` 输出
|
|
|
|
| 700 |
| `pieces[].missing_numbers` / `missing_urls` / `added_numbers` | 见上表;缺失的按草稿里的写法,新增的按改写里的写法 |
|
| 701 |
| `pieces[].retried` / `chosen` / `attempts` | 是否重写过、留的是第几版、每一版的检查结果 |
|
| 702 |
| `pieces[].seconds` | 这块所有版本加起来的模型耗时 |
|
| 703 |
+
| `pieces[].language_resampled` | 因为语言不对多采了几次(通常是 0) |
|
| 704 |
| `pieces[].flagged` / `issues` | 留下的那版是否仍有问题,以及是哪些问题 |
|
| 705 |
| `text` | 完整结果,只在没给 `-o`、输入又不是 `.docx` 时才有 |
|
| 706 |
|