Instructions to use jialinyyzz/humanizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jialinyyzz/humanizer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jialinyyzz/humanizer")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("jialinyyzz/humanizer") model = AutoModelForMultimodalLM.from_pretrained("jialinyyzz/humanizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use jialinyyzz/humanizer with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: llama cli -hf jialinyyzz/humanizer:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: llama cli -hf jialinyyzz/humanizer:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf jialinyyzz/humanizer:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf jialinyyzz/humanizer:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf jialinyyzz/humanizer:Q4_K_M
Use Docker
docker model run hf.co/jialinyyzz/humanizer:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use jialinyyzz/humanizer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jialinyyzz/humanizer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jialinyyzz/humanizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jialinyyzz/humanizer:Q4_K_M
- SGLang
How to use jialinyyzz/humanizer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jialinyyzz/humanizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jialinyyzz/humanizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jialinyyzz/humanizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jialinyyzz/humanizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use jialinyyzz/humanizer with Ollama:
ollama run hf.co/jialinyyzz/humanizer:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use jialinyyzz/humanizer with Docker Model Runner:
docker model run hf.co/jialinyyzz/humanizer:Q4_K_M
- Lemonade
How to use jialinyyzz/humanizer with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull jialinyyzz/humanizer:Q4_K_M
Run and chat with the model
lemonade run user.humanizer-Q4_K_M
List all available models
lemonade list
- Atomic Chat
Start here: prompt format, what it gets wrong, and what feedback we need
Three things that will save you time, then one ask.
1. This is a base-model completion, not a chat model.
No system prompt, no chat template, no <start_of_turn> markers. You send the instruction, your draft, and the separator as one block of text and let the model continue. The exact wrapper ships in prompt_format.json next to the weights β reproduce it byte for byte, including the blank line after ### Rewritten:. A paraphrased instruction measurably degrades output.
If you are using Ollama or LM Studio, both apply the base model's chat template by default and it will break the format. The model card has a Modelfile that passes the prompt through untouched.
2. Use Q8_0 or Q6_K, not lower.
Q5_K_M and Q4_K_M are deliberately not released. Q5_K_M made 17 critical fidelity errors on the same 62-sample set where the bf16 model made 0; Q4_K_M produces gibberish. Fidelity collapses below 6-bit on this model. MLX 4-/6-bit is also unusable (Gemma 4 PLE layers) β on a Mac, use the GGUF quants.
3. Proofread the numbers.
The most common failure is a dropped qualifier, in roughly 1 output in 3: "an estimated 4.2 %" becomes "4.2 %", "suggest" becomes "conclude". Short drafts (under ~120 words) are less reliable, Chinese is weaker than English, and you will occasionally get one garbled sentence β resample it.
What we would like back.
The evaluation set is 39 everyday-writing cases and it is public in the GitHub repo. What we cannot see from here is where it breaks on text that is not ours. If it drops a fact, flips a claim, or produces something obviously machine-like on your material, please post the draft and the output in this thread β a single failing pair is more useful to us than a general impression. Genre labels help too (legal, medical, code-adjacent, non-English), since we know our coverage there is thin.
Code, eval set and reward function: https://github.com/sgaofen/humanizer