Text Generation
MLX
Safetensors
English
Chinese
qwen3_5
apus-openjev
decision-model
apple-silicon
conversational
8-bit precision
Instructions to use apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Download examples/openjev_mlx.py from apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit: direct link, hf CLI and curl.
- Browser
- Download file 3.39 kB
-
https://huggingface.co/apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit/resolve/main/examples/openjev_mlx.py
- Command line
-
hf download hf://apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit/examples/openjev_mlx.py
-
curl -L -o openjev_mlx.py https://huggingface.co/apus-ailab/APUS-OpenJev-v1-9B-MLX-8bit/resolve/main/examples/openjev_mlx.py
3.39 kB
| """Run APUS-OpenJev MLX on Apple Silicon: exact candidate distribution with mlx-lm. | |
| pip install mlx-lm | |
| python examples/openjev_mlx.py --model apus-ailab/APUS-OpenJev-v1-4B-MLX-8bit | |
| The prompt is rendered by ``openjev_contracts.py`` (identical to the training | |
| contract) inside the no-thinking chat turn; the candidate labels A..P are single | |
| tokens and the distribution is the softmax of their final-position logits. | |
| """ | |
| import argparse | |
| import json | |
| import sys | |
| from pathlib import Path | |
| import mlx.core as mx | |
| from mlx_lm import load | |
| from mlx_lm.models.cache import make_prompt_cache | |
| sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) | |
| from openjev_contracts import format_response, label_mapping, render_prompt # noqa: E402 | |
| MAX_TOKENS = 8192 | |
| PREFILL_STEP = 512 # chunked prefill keeps long prompts within unified memory | |
| class OpenJevMLX: | |
| def __init__(self, model_path): | |
| self.model, self.tokenizer = load(model_path) | |
| def compile(self, request): | |
| mapping = label_mapping(request) | |
| prompt = self.tokenizer.apply_chat_template( | |
| [{"role": "user", "content": render_prompt(request)}], | |
| tokenize=False, | |
| add_generation_prompt=True, | |
| enable_thinking=False, | |
| ) | |
| ids = self.tokenizer.encode(prompt, add_special_tokens=False) | |
| if not ids or len(ids) > MAX_TOKENS: | |
| raise ValueError("input exceeds 8192 tokens; no truncation is performed") | |
| candidates = [] | |
| for label in mapping: | |
| token = self.tokenizer.encode(label, add_special_tokens=False) | |
| if ( | |
| len(token) != 1 | |
| or self.tokenizer.encode(prompt + label, add_special_tokens=False) | |
| != ids + token | |
| ): | |
| raise ValueError( | |
| "candidate label is not a single token at the answer boundary" | |
| ) | |
| candidates.append(token[0]) | |
| return ids, candidates | |
| def decide(self, request): | |
| ids, candidates = self.compile(request) | |
| cache = make_prompt_cache(self.model) | |
| for start in range(0, len(ids) - 1, PREFILL_STEP): | |
| self.model( | |
| mx.array([ids[start : min(start + PREFILL_STEP, len(ids) - 1)]]), | |
| cache=cache, | |
| ) | |
| mx.eval([c.state for c in cache]) | |
| logits = self.model(mx.array([ids[-1:]]), cache=cache)[0, -1].astype( | |
| mx.float32 | |
| )[mx.array(candidates)] | |
| response = format_response(request, mx.softmax(logits).tolist()) | |
| response.update(prompt_tokens=len(ids), calibrated=False) | |
| return response | |
| EXAMPLE = { | |
| "id": "support-731", | |
| "group_id": "support-731", | |
| "primitive": "choice", | |
| "state": "Order 731 was delivered. The customer confirms that the issue is resolved.", | |
| "instructions": "Select the next support action.", | |
| "criteria": [ | |
| {"id": "close_ticket", "description": "Close the ticket as resolved."}, | |
| {"id": "escalate", "description": "Escalate to a human agent."}, | |
| {"id": "refund", "description": "Issue a refund."}, | |
| ], | |
| } | |
| if __name__ == "__main__": | |
| parser = argparse.ArgumentParser() | |
| parser.add_argument("--model", required=True, help="local directory or HF repo id") | |
| args = parser.parse_args() | |
| print(json.dumps(OpenJevMLX(args.model).decide(EXAMPLE), indent=2)) | |