How to use from
OpenClaw
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "vinci00/limite-1b-violetto-mlx-4bit"
Configure OpenClaw
# Install OpenClaw:
npm install -g openclaw@latest
# Register the local server and set it as the default model:
openclaw onboard --non-interactive --mode local \
  --auth-choice custom-api-key \
  --custom-base-url http://127.0.0.1:8080/v1 \
  --custom-model-id "vinci00/limite-1b-violetto-mlx-4bit" \
  --custom-provider-id mlx-lm \
  --custom-compatibility openai \
  --custom-text-input \
  --accept-risk \
  --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Quick Links

Limite 1B Violetto — MLX 4-bit

A community 4-bit MLX quantization of Paradigma's Limite 1B Violetto, converted and evaluated on a Mac mini M4 with 16 GiB unified memory. This derivative changes the storage and inference representation; it adds no training or fine-tuning. The source model is a text-only mathematical reasoner, not a general-purpose or multimodal assistant.

Property Value
Source revision 9402e422e87b220507963cd42c997459c1e87b43
Architecture LimiteForCausalLM, approximately 1.035B parameters
Source precision BF16 matrices with FP32 auxiliary tensors
Quantization MLX affine, 4 bits, group size 64
Weight-file size Approximately 0.543 GiB
Custom adapter limite-mlx==0.1.0, pinned below
License Apache-2.0

Installation and usage

Use an Apple Silicon Mac. The custom Limite architecture is not built into mlx-lm 0.31.3; install the pinned community adapter along with the runtime:

python -m pip install mlx==0.32.2 mlx-lm==0.31.3 transformers==5.17.0 \
  "limite-mlx @ https://huggingface.co/pierjoe/limite-mlx/resolve/b21b7a4fa6ca05cb02082d9f7819b8d1387f2f1a/limite_mlx-0.1.0-py3-none-any.whl"

The adapter adds a Python startup import hook for mlx_lm.models.limite. Restart an already-running notebook kernel after installation.

from mlx_lm import generate, load
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load('vinci00/limite-1b-violetto-mlx-4bit')
prompt = tokenizer.apply_chat_template(
    [{'role': 'user', 'content': 'How many positive divisors does 360 have?'}],
    tokenize=False,
    add_generation_prompt=True,
)
answer = generate(model, tokenizer, prompt=prompt, max_tokens=3000,
                  sampler=make_sampler(temp=0.6, top_p=0.95))
print(answer)

Keep the original fixed mathematical chat template. It supplies the system prompt and requests a boxed answer; arbitrary system prompts and tools are not supported. The model may produce long reasoning passages before its answer. The sampling parameters above follow the upstream recommendation, while our paired evaluation uses greedy decoding for reproducibility.

Conversion and provenance

Both the unquantized baseline and this quantized derivative were converted from the same pinned original checkpoint, using mlx-lm 0.31.3, MLX 0.32.2, and the same pinned third-party Limite adapter. The baseline was not dequantized from this artifact. Quantization includes the tied input embedding/output head. The adapter folds learned attention scales into projection matrices and retains small FP32 auxiliary tensors.

Both converted configurations accept end-token IDs [151643, 151645], resolving the upstream mismatch between tokenizer and generation configuration. Tokenizer and chat-template contents are identical across the paired artifacts. No Mistral-specific tokenizer regex replacement is applied.

See conversion manifest, weight hashes and source provenance, and recorded hardware. The reproducible source repository contains the executed notebook, evaluator, tests, and pinned requirements.

Evaluation protocol

The teaching evaluation uses the same 10 GSM8K test examples, selected with seed 42, and a 1,024-token generation cap for both variants. Dataset revision 3101c7d5072418e28b9008a6636bde82a006892c is pinned and its JSONL checksum is verified. No training or prompt selection uses the test split.

Scoring uses the last complete \boxed{...} after the reasoning section and requires its entire content to be numeric. LaTeX wrappers and prose outside the box are allowed. An unfinished <think> section scores zero. Missing, malformed, nonnumeric, or unfinished final answers count as incorrect. We retain raw predictions and paired regressions/improvements, correct/total, 95% Wilson intervals, and paired bootstrap uncertainty. Four fixed mathematical smoke questions measure runtime separately; runtime excludes model loading and prompt formatting, and generated lengths can differ.

This small, token-limited run is not a reproduction of the upstream competition scores. Reasoning may exceed the cap. A tie does not establish general quality parity, and pretraining contamination is unknown.

Compatibility and limitations

This is an MLX artifact. Ollama 0.34.2 rejects it with unsupported MLX architecture: model "LimiteForCausalLM". It is not a usable Ollama registry release. Stock llama.cpp also lacks this architecture; converting the container to GGUF alone does not solve runtime support.

The official runtime is Paradigma's vLLM plugin. Our comparison measures quantization within the community MLX adapter; we did not independently establish numerical equivalence to official vLLM. No GGUF, Ollama, image, long-context, or general-assistant quality result is claimed. The source model is lightly instruction-tuned and may reinterpret prompts as mathematical tasks or produce incorrect answers.

License and attribution

The source weights and adapter are Apache-2.0 licensed. Credit for the model belongs to Paradigma, and the MLX port is by pierjoe. This conversion is not an official release by either author. Preserve LICENSE and NOTICE when redistributing.

Measured results — Mac mini M4, 2026-09-24

The saved notebook completed on Mac mini M4, 16 GiB, macOS 26.4, Python 3.13.15. Both variants passed 4/4 mathematical smoke cases. The GSM8K test evaluation used 10 examples, seed 42, greedy decoding, 1,024 maximum output tokens per example and variant.

Variant GSM8K accuracy Invalid final answers Mean smoke latency Output tokens/s MLX peak memory
MLX unquantized 5/10 (50%) 4 12.36 s 39.54 1.974 GiB
MLX 4-bit 5/10 (50%) 4 6.14 s 87.49 0.599 GiB

Both accuracy estimates have a 95% Wilson interval of 23.66%–76.34%. The paired difference is 0.0 percentage points, with no paired regressions or improvements. The bootstrap interval is [0.0, 0.0] pp because all ten paired correct/incorrect outcomes agree; this degenerate interval does not establish quality parity. Four responses per variant fail the strict numeric final-box format; some generations reach the reasoning token cap. One additional response per variant is formatted correctly but incorrect.

Weight files occupy 1.929 GiB for the source-precision MLX baseline and 0.543 GiB for the 4-bit derivative. The runtime sample is short, runs on a shared machine, excludes loading, and has differing generated lengths; it does not measure peak hardware throughput.

Measured runtime and accuracy

See raw runtime, paired accuracy and predictions, uncertainty report, and hardware. These scores belong only to the MLX artifacts and must not be attributed to Ollama or the official vLLM runtime.

Downloads last month
41
Safetensors
Model size
1B params
Tensor type
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vinci00/limite-1b-violetto-mlx-4bit

Quantized
(8)
this model