How to use from the
Use from the
MLX library
# Make sure mlx-lm is installed
# pip install --upgrade mlx-lm

# Generate text with mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("flowxai/flowx-sentinel-gate-ministral-3b")

prompt = "Write a story about Einstein"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True
)

text = generate(model, tokenizer, prompt=prompt, verbose=True)

FlowX Sentinel Gate (Ministral)

Decides whether a case in a regulated workflow can be actioned automatically (DECIDE) or must go to a human (ESCALATE), and when it escalates, says which of six categories applies. Emits JSON only.

Spec

Field Value
Base model mistralai/Ministral-3-3B-Instruct-2512 (3.85B, Apache-2.0)
Adapter LoRA rank 32, scale 16, dropout 0.05, top 16 layers, MLX-LM
Training 4 epochs, batch 4, LR 5e-5, max_seq 2048, cosine with 10-step warmup
Decoding greedy, enable_thinking=False, the model is trained on pure JSON
Base preparation Published fp8, which MLX cannot read: mx.load has no float8 dtype in MLX 0.31.2. Dequantized to bf16 with dequantize_ministral.py (weight_scale_inv is a multiplier despite the name), then converted to MLX 4-bit. 1.8 GB at 4.501 bits per weight.

Evaluation

Frozen held-out set, n=50, greedy decoding, scored by eval_v2.py. Reported across 3 training seeds (7, 42, 1337) that differ in nothing but the seed, because a single run cannot measure its own noise and this table makes claims against two other published models.

metric median range across seeds
action accuracy 95.8% 95.7% to 98.0%
category accuracy 79.1% 75.6% to 90.0%
false negatives 1 0 to 2
valid JSON 96.0% 94.0% to 100.0%

Same frozen set and scoring as every system below.

model action category false-neg valid JSON
Qwen3-4B (flowx-sentinel-gate-4b) 93.5% 80.0% 1 92%
Qwen3.5-35B (flowx-sentinel-gate-35b-v1) 92.0% 81.4% 2 100%
Claude Sonnet 4.6 91% 74% 3 90%
GPT-5.4-mini 90% 70% 3 100%
Claude Haiku 4.5 90% 65% 3 100%
Gemini 2.5 Flash 88% 63% 5 100%

The metric that decides this detector is the false-negative rate: a missed escalation is the expensive error. On a parse failure the caller should fail safe to ESCALATE.

Limits, stated plainly

Exactly one claim here survives 3 seeds, and it is action accuracy. 95.7% to 98.0% against the Qwen3-4B incumbent's 93.5%: ahead on every draw, with a margin of 2.2 to 4.5 points rather than the top of the range.

Category accuracy is not a usable comparison at this sample size. It spans 75.6% to 90.0% across seeds that differ in nothing else, a 14.4-point swing around the incumbent's 80.0%. One seed loses to the incumbent and another beats it comfortably. Neither ordering is real.

The false-negative count is likewise indistinguishable. It spans 0 to 2 against the incumbent's 1, and the median is 1. This is the metric the detector is judged by, so it matters that the honest answer is "no measurable difference" rather than a win. Do not select between these models on that basis.

n is 50. A 4.5-point action difference is three cases, and a one-case false-negative difference is one case out of 43. Every figure above is quoted with its seed range because a single run cannot measure its own noise; the first version of this card reported one draw and claimed zero missed escalations, which a reseed did not reproduce.

Contrast-set gap: +43.0 points. Scored against a contrast set whose members differ in exactly one field, category accuracy falls to 36.1% (artifact eval_cross_oldmodel_newtest_ministral.json). That gap measures how much the model leans on phrasing and co-occurring fields rather than the deciding fact. It is not a corrected accuracy: the contrast set's own labels come from a single teacher model and separating "shortcuts removed" from "labels noisier" needs human adjudication. It is reported because a model card that omits it would overstate the headline number.

The training corpus is 227 rows of realistic synthetic scenarios grounded in real regulation citations, 85% ESCALATE. An attempt to expand it with generated minimal pairs regressed every metric and is documented rather than shipped.

Not a compliance control. This produces evidence about a routing decision. Obligations under any regulation sit with the operator of the system, not with a model.

Licence and attribution

Released under the Apache License, Version 2.0. Full text in LICENSE.

This model is a derivative work of mistralai/Ministral-3-3B-Instruct-2512 by Mistral AI, itself licensed under Apache-2.0. That licence permits redistribution of derivatives and requires that attribution and a statement of changes travel with them, so both are carried here and in NOTICE.

Changes made to the base model, in order:

  1. Dequantized FP8 to bfloat16. The published checkpoint stores 3.03B parameters as F8_E4M3 with per-tensor weight_scale_inv scales. MLX has no float8 dtype, so it cannot read the checkpoint at all.
  2. Converted to MLX 4-bit, 4.501 bits per weight.
  3. Vision tower discarded. The base is a vision-language checkpoint; this is text-only.
  4. LoRA fine-tune fused in (rank 32, scale 16, dropout 0.05, top 16 layers).

No part of the base model's training data, evaluation results or documentation is reproduced here. The evaluation figures above are ours and were produced by eval_v2.py on our own frozen set.

Downloads last month
175
Safetensors
Model size
3B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for flowxai/flowx-sentinel-gate-ministral-3b

Evaluation results

  • Action accuracy (median of 3 seeds) on FlowX escalation frozen held-out set (n=50)
    self-reported
    0.958
  • Category accuracy (median of 3 seeds) on FlowX escalation frozen held-out set (n=50)
    self-reported
    0.791
  • Missed escalations (median count, lower is better) on FlowX escalation frozen held-out set (n=50)
    self-reported
    1.000