Instructions to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "apus-ailab/APUS-OpenJev-v1-4B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "apus-ailab/APUS-OpenJev-v1-4B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
- Ollama
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with Ollama:
ollama run hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with Docker Model Runner:
docker model run hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
- Lemonade
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.APUS-OpenJev-v1-4B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use apus-ailab/APUS-OpenJev-v1-4B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download README.zh-CN.md from apus-ailab/APUS-OpenJev-v1-4B-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 5.2 kB
-
https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF/resolve/main/README.zh-CN.md
- Command line
-
hf download hf://apus-ailab/APUS-OpenJev-v1-4B-GGUF/README.zh-CN.md
-
curl -L -o README.zh-CN.md https://huggingface.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF/resolve/main/README.zh-CN.md
library_name: gguf
license: apache-2.0
base_model: apus-ailab/APUS-OpenJev-v1-4B
base_model_relation: quantized
pipeline_tag: text-generation
language:
- en
- zh
tags:
- apus-openjev
- decision-model
- gguf
- llama.cpp
- ollama
APUS-OpenJev-v1-4B-GGUF
English | 中文 · 源模型 · Collection · GGUF collection · MLX collection · MLX(Apple Silicon)
APUS-OpenJev-v1-4B(revision 65797c526c27)的 GGUF 版本,适用于 Linux / Windows / macOS(Metal)上的 Ollama、llama.cpp、LM Studio。
OpenJev 是决策模型:每个请求给出状态、指令和 2–16 个候选,模型回答一个候选标签(A–P),应用读取这些标签上的概率分布。它不是聊天模型。
文件与一致性
每个文件都在 Frozen80 上用完全相同的 prompt token 评测,并与 HF BF16 发布版(其自带 runtime、完整深度,**66/80 · 82.50%**)对比。
| 文件 | 大小 | Frozen80 | 与 HF BF16 决策一致 | 相对 HF BF16 最大 Δp |
|---|---|---|---|---|
| Q8_0 | 4.2 GiB | 67/80 · 83.75% | 79/80 | 0.1288 |
| Q4_K_M | 2.5 GiB | 67/80 · 83.75% | 77/80 | 0.8879 |
| BF16 | 7.8 GiB | 66/80 · 82.50% | 80/80 | 0.0780 |
| Q8_0 · Apple M5 Metal | 4.2 GiB | 66/80 · 82.50% | 80/80 | 0.1319 |
除标注 Apple M5(Metal)的行外,均为 llama-server b11118 在 NVIDIA RTX PRO 6000(CUDA)上的结果。默认推荐 Q8_0;Q4_K_M 使用 importance matrix。量化成绩单独测量,不沿用 BF16 成绩。Frozen80 是复用的开发面板,不是盲测。明细见 evaluation/。逐题明细(候选概率、选择、是否正确,可按 panel_index 与 Frozen80 关联):evaluation/per-question/。
Ollama
ollama run hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0 --think=false
务必关闭 thinking。 Ollama 对这个架构使用内置的 Qwen3.5 renderer(不会应用 Modelfile 里的 TEMPLATE),而这个 renderer 默认会打开 thinking 块,与训练时"不思考"的格式不一致。命令行加 --think=false,/api/chat 或 /api/generate 中传 "think": false,或者用 "raw": true 发送完整渲染好的 prompt。params 设置了 temperature 0 和 num_predict 1;仓库也附带本地 Modelfile。
需要候选概率时,使用自带客户端(raw 模式):
python examples/openjev_local.py --backend ollama --model hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0
Ollama 最多返回 20 个 top_logprobs,且不能指定 token,因此只有全部候选标签都在前 20 时分布才是精确的(Q8_0:Frozen80 中 64/80 题);否则请使用所选标签,或改用 llama-server。
Ollama 0.34.3 + Q8_0 实测 on an Apple M5 Mac (24 GB, Metal):
| 调用方式 | 标签与 llama-server 一致 | Frozen80 | prompt token 与训练一致 |
|---|---|---|---|
/api/generate + raw: true + 完整 prompt(examples/openjev_local.py) |
80/80 | 66/80 | 80/80 |
/api/chat 或 ollama run,think: false |
80/80 | 66/80 | 80/80 |
/api/generate,未关闭 thinking(错误用法) |
49/80 | 45/80 | 0/80 |
直接拉取 hf.co/apus-ailab/APUS-OpenJev-v1-4B-GGUF:Q8_0 并使用 think: false,结果相同:标签一致 80/80,正确 66/80。
llama.cpp(精确分布)
llama-server -m APUS-OpenJev-v1-4B-Q8_0.gguf -c 9216 -ngl 999
python examples/openjev_local.py --backend llama-server --url http://127.0.0.1:8080
examples/openjev_local.py 使用 openjev_contracts.py 渲染 prompt,与训练时的格式一致。
转换说明
- llama.cpp
b11118(e6ab7c1a4),convert_hf_to_gguf.py --no-mtp(merged 发布版不含 MTP 权重)。 - Q4_K_M 的 importance matrix:训练集中 448 条决策,每个来源 64 条,与 Frozen80 无重叠(明细、imatrix.gguf)。
- 所有 1-D 张量(含 GDN
A_log/dt_bias和各类 norm)在每个文件中都保持 F32(检查结果)。 - 范围:仅完整深度(不提供 16/20 层
low出口),仅文本(不含视觉塔),概率未经校准。
许可
Apache-2.0,继承自源模型,见 LICENSE。基座模型:Qwen/Qwen3.5-4B。
作者: gumpcheng(xDAN2099)、zhangxu、APUS AI-LAB。