Instructions to use JANGQ-AI/Naive-N0.5-Flash-JANGH2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use JANGQ-AI/Naive-N0.5-Flash-JANGH2 with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("JANGQ-AI/Naive-N0.5-Flash-JANGH2") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use JANGQ-AI/Naive-N0.5-Flash-JANGH2 with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/Naive-N0.5-Flash-JANGH2"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "JANGQ-AI/Naive-N0.5-Flash-JANGH2" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use JANGQ-AI/Naive-N0.5-Flash-JANGH2 with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "JANGQ-AI/Naive-N0.5-Flash-JANGH2"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "JANGQ-AI/Naive-N0.5-Flash-JANGH2" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JANGQ-AI/Naive-N0.5-Flash-JANGH2", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use JANGQ-AI/Naive-N0.5-Flash-JANGH2 with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/Naive-N0.5-Flash-JANGH2"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default JANGQ-AI/Naive-N0.5-Flash-JANGH2
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use JANGQ-AI/Naive-N0.5-Flash-JANGH2 with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "JANGQ-AI/Naive-N0.5-Flash-JANGH2"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "JANGQ-AI/Naive-N0.5-Flash-JANGH2" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Run JANG models in MLX Studio / vMLX
⚠️ Runtime: vMLX (Python) support for this model family is not in a released build yet. The bundle was built and validated on a vMLX development build that adds the
naive_n05_flashfamily. Released vMLX builds do not include this architecture and cannot load the bundle. Known limitation of the development build: prefix / SSD cache reuse is not working for this model yet (answers are correct; a repeated long prompt is processed again instead of being restored).
JANGQ-AI/Naive-N0.5-Flash-JANGH2
Naive-N0.5-Flash for 128 GB Macs. 575 GiB of bf16 weights in 95.89 GiB, at the speed of a plain MLX quant of the same size and far closer to the original model.
A JANGH bundle of NaiveAI/Naive-N0.5-Flash: a 309B MoE (15.5B active, 256 routed experts, top-8) for coding and agentic work, with a native 1M-token context built from sliding-window attention and sparse attention. Text only.
- Routed experts: JANGH at 2-4 bits (2.54 on average). Codebook-quantized with per-row scales and a blockwise Hadamard rotation, rounded with GPTQ on the full per-expert input statistics. Bits are placed per layer by measurement.
- Everything else: affine 8-bit, or kept in the source precision (routers, norms, the sparse-attention indexer).
| Top-1 agreement | KL median | |
|---|---|---|
| Ceiling (model vs. itself) | 71.6% | ~0.06 |
| JANGH2 | 64.4% | 0.170 |
| RTN affine, same size | 48.2% | 1.552 |
Read this first: agreement has a ceiling for this model
Agreement with the bf16 model cannot reach 100% here, for any quantization. Naive-N0.5-Flash chooses 8 of 256 experts per token in every layer, and the 8th and 9th candidates are often almost tied. A change far smaller than any quantization error flips some of those choices, and the flips compound over 47 layers.
We measured that on the unquantized weights: adding relative noise of 0.001 to a single tensor of one layer changes the top-1 token at 28% of the positions. That run is the ceiling in every table below. The distance to the ceiling, not the distance to 100%, is what a quantization costs.
Fidelity vs the bf16 model
68,370 teacher-forced positions on 28 held-out prompts (coding, cybersecurity, agentic, general, Chinese, science, academic; six of them 4-6.5k tokens long), scored against the bf16 model run layer by layer from disk. Top-128 renormalized KL. JANGH2 and the control are served through the runtime with 2048-token prefill chunks.
| Size | median KL ↓ | mean KL ↓ | p90 / p95 / p99 ↓ | top-1 ↑ | top-5 ↑ | top-10 ↑ | |
|---|---|---|---|---|---|---|---|
| Ceiling: bf16 + noise 0.001 on one tensor | 575 GiB | 0.060 | 0.749 | 2.23 / 4.26 / 9.62 | 71.6% | 85.8% | 89.1% |
| Naive-N0.5-Flash-JANGH2 | 95.89 GiB | 0.170 | 1.008 | 3.11 / 5.37 / 10.50 | 64.4% | 81.5% | 85.6% |
| MLX affine RTN, same size ¹ | 95.83 GiB | 1.552 | 2.976 | 8.04 / 11.06 / 16.78 | 48.2% | 64.1% | 68.6% |
¹ A control we built for this comparison: stock MLX affine quantization of the same source, experts at 2-3 bits chosen by the same kind of measurement, non-experts affine 8-bit, no calibration. It is not published.
By domain, median KL and top-1 agreement:
| domain | positions | ceiling | JANGH2 | MLX affine RTN |
|---|---|---|---|---|
| cybersecurity | 6,653 | 0.024 · 76.4% | 0.053 · 73.0% | 1.653 · 48.5% |
| coding | 8,249 | 0.031 · 73.7% | 0.082 · 68.8% | 1.086 · 53.0% |
| coding, long prompts | 9,390 | 0.036 · 74.5% | 0.085 · 69.9% | 1.165 · 54.3% |
| agentic | 7,411 | 0.096 · 66.9% | 0.238 · 60.4% | 1.864 · 43.7% |
| agentic, long prompts | 10,190 | 0.152 · 63.7% | 0.248 · 59.0% | 2.205 · 48.7% |
| general | 5,311 | 0.074 · 71.5% | 0.242 · 60.6% | 1.864 · 42.8% |
| long documents | 10,990 | 0.059 · 74.0% | 0.222 · 63.4% | 1.465 · 47.1% |
| Chinese | 3,158 | 0.038 · 77.8% | 0.172 · 65.1% | 1.496 · 43.7% |
| academic | 3,123 | 0.072 · 73.1% | 0.213 · 62.3% | 1.542 · 44.8% |
| science | 3,895 | 0.100 · 68.5% | 0.356 · 57.2% | 1.426 · 46.8% |
The calibration set is weighted toward coding, tool use and cybersecurity, and that is where the bundle is closest to the original. General, Chinese and science text lose more.
Tool-use fidelity vs the bf16 model
96 held-out tool conversations (65,719 positions) on tools that never appear in calibration, with 208 "call a tool or answer?" decision points. The bf16 model starts a tool call at 100 of the 112 points where the transcript has one.
| ceiling | JANGH2 | MLX affine RTN | |
|---|---|---|---|
| tool-call points: model starts a tool call (bf16: 100) | 97 | 101 | 83 |
| tool-call points: decisions flipped call ↔ answer vs bf16 | 9 | 7 | 27 |
| tool-call points: same next token as bf16 | 92.0% | 93.8% | 75.9% |
| answer points: same next token as bf16 | 91.7% | 88.5% | 55.2% |
median P(<tool_call>) at tool-call points (bf16: 0.859) |
0.850 | 0.889 | 0.817 |
lowest P(<tool_call>) at a tool-call point |
0.053 | 0.133 | 0.0003 |
| median KL, whole conversations | 0.292 | 0.395 | 3.112 |
| top-1 agreement, whole conversations | 55.1% | 51.3% | 22.7% |
On tool decisions JANGH2 is indistinguishable from the model's own ceiling. These are synthetic agent transcripts, so absolute KL is high in every column; read the columns against each other. The calibration set contains agent-style conversations from the same generator (different tools and wording), so part of this result is in-distribution.
Live behavior (served, temperature 0)
| suite | JANGH2 | MLX affine RTN |
|---|---|---|
| 48-case tool-use eval (required / auto × thinking on / off) | 48 / 48 | 34 / 48 |
| 20 behavior probes: arithmetic, follow-ups with reasoning history, tool round trips, chained tool calls, multi-line tool arguments, efforts low/high/max, a 5,173-token needle prompt | 20 / 20 | 11 / 20 |
Examples of what the control gets wrong: 17 × 19 = "329", "221 = 3 × 731", and a passphrase copied from a long prompt with an extra digit. The bf16 model itself does not fit in 128 GB, so there is no live bf16 column; these suites are easy enough that they check behavior and separate a working bundle from a broken one, not more.
Speed (M5 Max 128 GB, served, interleaved A/B/A/B, fresh prompts, median of 3)
| JANGH2 | MLX affine RTN | |
|---|---|---|
| decode (tok/s) | 40.38 / 41.02 | 40.87 / 40.27 |
| prefill, ~5.1k-token prompt (tok/s) | 662 / 669 | 635 / 619 |
| time to first token, ~5.1k-token prompt | 7.80 s / 7.68 s | 8.17 s / 8.25 s |
| memory after load / peak while serving | 95.9 / 98.9 GiB | 95.8 / 99.0 GiB |
| server ready after process start | 20 s | 20 s |
Two interleaved runs per bundle, same runtime build, a fresh server for each run, each run the median of 3 probes on a never-seen prompt. Decode is the same within run-to-run variation; long-prompt prefill is ~6% faster than the plain MLX quant.
JANGH
JANGH is our expert format (it replaced JANGTQ): a fixed blockwise Hadamard-32 transform with no random signs and a two-constant codebook per bit width, stored bit-for-bit in MLX's own packing layout, so a runtime reuses MLX's kernel structure for decode and prefill.
For compatibility with runtimes already built against it, the on-disk identifiers keep their original names:
config.json → jangtq block (version: 2), per-module "mode": "jangtq2", tensors *.tq2_packed /
*.tq2_scales, and jang_config.json → "format": "jangtq2". They are not loadable by, and must not be routed to,
JANGTQ v1 loaders.
What's in the bundle
- Text only. The source model has no vision, audio or video components.
- Thinking + agentic:
- Thinking is ON by default. The model opens
<think>itself; with thinking off the template appends<think></think>. - Reasoning efforts are
low/high/max(defaultmax). There is nomedium: the template renders any other value as Max. - Reasoning is kept in history by default.
- Thinking is ON by default. The model opens
- Tool calls: Qwen3-Coder-style XML (
<tool_call>/<function=name>/<parameter=key>value</parameter>), declared astool_parser: xml_function; tool results use thetoolrole and render as<tool_response>. Hermes-style JSON parsers will not work. - Self-describing:
config.jsoncarries the JANGH format block (codebook, packing, rotation, method) and a per-modulequantizationmap (141 expert projections, 197 affine 8-bit modules).jang_config.jsonrecords capabilities, calibration and per-layer expert bits.- Raw evaluation results, including the ceiling and the control, are in
evaluation/.
- Expert bits: gate/up 2-bit in 28 layers and 3-bit in 19; down 2-bit in 11 layers, 3-bit in 34, 4-bit in 2.
- Every shard is alignment-safe (zero-copy memory mapping). 49 shards, 1,148 tensors.
Serving contract
- Sampling:
temperature=1.0, top_p=0.95(vendor defaults), no repetition penalty - EOS:
151645· context: 1M native - Reasoning:
reasoning_effortchat-template kwarg (low/high/max), defaultmax; declared asreasoning_parser: think_xml - Output layout: the model writes
</think>, then a blank line, then the answer or the<tool_call>. - Memory: prefill long prompts in chunks of at most 2048 tokens. With the weights loaded, a single 6k-token forward needs more GPU memory than a 128 GB machine has left.
- The routers must stay in fp32 and the sparse-attention indexer unquantized, as shipped.
Build details
- Source:
NaiveAI/Naive-N0.5-Flash@0235b3b(bf16, 575 GiB) - Calibration: 618k tokens rendered with the model's own chat template (coding, tool conversations, cybersecurity, agentic, general, Chinese, science, academic), referenced to the bf16 model; evaluation prompts, tools and every tenth calibration document held out
- Experts: JANGH (odd-cubic codebook, fp16 per-row scale, Hadamard-32 rotation), GPTQ on all 47 MoE layers, bit allocation measured per layer
- Measured and not applied: AWQ, MXFP8 for the non-expert weights, per-row bias correction
Quantized and validated by Jinho Jang — eric@jangq.ai
- Downloads last month
- 244
8-bit
Model tree for JANGQ-AI/Naive-N0.5-Flash-JANGH2
Base model
NaiveAI/Naive-N0.5-Flash
