Instructions to use CogniSoft/Plumb-4B-MLX-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use CogniSoft/Plumb-4B-MLX-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download CogniSoft/Plumb-4B-MLX-8bit --local-dir Plumb-4B-MLX-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Plumb-4B for MLX, 8-bit
This is an 8-bit MLX conversion of crh225/plumb-4b, a 4B text decision model. It takes evidence, a criterion, and 2–16 options and returns a probability for each option in one forward pass. It uses the original model's calibrated temperature, 2.07.
Plumb is a separate model, not Jev. It follows the JevK5/SemIf typed-decision protocol. It is text-only; it does not accept images, audio, or video. This model is intended for Apple silicon with MLX. The original BF16 model and its own benchmark results remain available in the source model card.
Local use
Requires an Apple silicon Mac and Python 3.12 or newer. The weight file is 4.47 GB (4.16 GiB). Tested on a MacBook Pro M5 with 24 GB unified memory.
python3 -m venv .venv
.venv/bin/pip install 'mlx-lm==0.31.3' 'numpy>=2,<3'
.venv/bin/pip install huggingface-hub
.venv/bin/hf download CogniSoft/Plumb-4B-MLX-8bit --local-dir plumb-mlx
.venv/bin/python plumb-mlx/plumb_mlx.py --model plumb-mlx --port 8090
In another terminal:
curl -s http://127.0.0.1:8090/v1/systemone \
-H 'Content-Type: application/json' \
-d '{"state":"Order #7120 shows delivered to No. 17; the customer lives at No. 71.","questions":{"parcel":{"type":"choice","instructions":"What happened to the parcel?","criteria":["delivered","misdelivered","unknown"]}}}'
The endpoint accepts multiple questions for one state. Supported question types are noul (true/false), choice (2–16 options), and score (2–16 ordinal levels). GET /health reports readiness. You can also import PlumbMLX from plumb_mlx.py and call decide(state, question) in Python. This runtime does not use ordinary text generation: it reads next-token logits for the answer letters and applies a softmax.
Conversion and validation
Converted from the original BF16 checkpoint at revision 55de037, with mlx-lm==0.31.3 and mlx==0.32.2, using affine 8-bit quantization, group size 64. The source config.json declares qwen3_5_text; for the conversion its architecture identifier was changed to qwen3_5, which is the text-only implementation recognized by this MLX release. The weights and the original jevk5_config.json were otherwise used unchanged as conversion input. See CONVERSION.md for the exact steps.
Validation compared a 30-case panel directly with the same BF16 checkpoint in the original JevK5 v0.2.0 runtime on PyTorch MPS, then compared all 231 public cases with Plumb's published BF16 GPU outputs. It checked token counts, option probabilities, and highest-probability answers. Results are in VALIDATION.md. These are port-parity checks, not an official JevBench score or a claim that the quantized model has the same calibration on unseen data. A close decision can still change with quantization or runtime precision.
Provenance and license
Original Plumb weights and model card: crh225/plumb-4b. Original code: crh225/plumb. Decision runtime: JevK5 v0.2.0. Weights are under Apache-2.0. The decision prompt/readout adapted by JevK5 come from SemIf (MIT); see NOTICE, NOTICE-jevk5, and LICENSE.
- Downloads last month
- 58
8-bit
Model tree for CogniSoft/Plumb-4B-MLX-8bit
Base model
Qwen/Qwen3.5-4B-Base