Plumb-4B for MLX, 8-bit

This is an 8-bit MLX conversion of crh225/plumb-4b, a 4B text decision model. It takes evidence, a criterion, and 2–16 options and returns a probability for each option in one forward pass. It uses the original model's calibrated temperature, 2.07.

Plumb is a separate model, not Jev. It follows the JevK5/SemIf typed-decision protocol. It is text-only; it does not accept images, audio, or video. This model is intended for Apple silicon with MLX. The original BF16 model and its own benchmark results remain available in the source model card.

Local use

Requires an Apple silicon Mac and Python 3.12 or newer. The weight file is 4.47 GB (4.16 GiB). Tested on a MacBook Pro M5 with 24 GB unified memory.

python3 -m venv .venv
.venv/bin/pip install 'mlx-lm==0.31.3' 'numpy>=2,<3'
.venv/bin/pip install huggingface-hub
.venv/bin/hf download CogniSoft/Plumb-4B-MLX-8bit --local-dir plumb-mlx
.venv/bin/python plumb-mlx/plumb_mlx.py --model plumb-mlx --port 8090

In another terminal:

curl -s http://127.0.0.1:8090/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{"state":"Order #7120 shows delivered to No. 17; the customer lives at No. 71.","questions":{"parcel":{"type":"choice","instructions":"What happened to the parcel?","criteria":["delivered","misdelivered","unknown"]}}}'

The endpoint accepts multiple questions for one state. Supported question types are noul (true/false), choice (2–16 options), and score (2–16 ordinal levels). GET /health reports readiness. You can also import PlumbMLX from plumb_mlx.py and call decide(state, question) in Python. This runtime does not use ordinary text generation: it reads next-token logits for the answer letters and applies a softmax.

Conversion and validation

Converted from the original BF16 checkpoint at revision 55de037, with mlx-lm==0.31.3 and mlx==0.32.2, using affine 8-bit quantization, group size 64. The source config.json declares qwen3_5_text; for the conversion its architecture identifier was changed to qwen3_5, which is the text-only implementation recognized by this MLX release. The weights and the original jevk5_config.json were otherwise used unchanged as conversion input. See CONVERSION.md for the exact steps.

Validation compared a 30-case panel directly with the same BF16 checkpoint in the original JevK5 v0.2.0 runtime on PyTorch MPS, then compared all 231 public cases with Plumb's published BF16 GPU outputs. It checked token counts, option probabilities, and highest-probability answers. Results are in VALIDATION.md. These are port-parity checks, not an official JevBench score or a claim that the quantized model has the same calibration on unseen data. A close decision can still change with quantization or runtime precision.

Provenance and license

Original Plumb weights and model card: crh225/plumb-4b. Original code: crh225/plumb. Decision runtime: JevK5 v0.2.0. Weights are under Apache-2.0. The decision prompt/readout adapted by JevK5 come from SemIf (MIT); see NOTICE, NOTICE-jevk5, and LICENSE.

Downloads last month
58
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for CogniSoft/Plumb-4B-MLX-8bit

Finetuned
Qwen/Qwen3.5-4B
Finetuned
crh225/plumb-4b
Quantized
(2)
this model