Instructions to use vinci00/limite-1b-violetto-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use vinci00/limite-1b-violetto-mlx-4bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("vinci00/limite-1b-violetto-mlx-4bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use vinci00/limite-1b-violetto-mlx-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "vinci00/limite-1b-violetto-mlx-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "vinci00/limite-1b-violetto-mlx-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use vinci00/limite-1b-violetto-mlx-4bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "vinci00/limite-1b-violetto-mlx-4bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "vinci00/limite-1b-violetto-mlx-4bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vinci00/limite-1b-violetto-mlx-4bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use vinci00/limite-1b-violetto-mlx-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "vinci00/limite-1b-violetto-mlx-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default vinci00/limite-1b-violetto-mlx-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use vinci00/limite-1b-violetto-mlx-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "vinci00/limite-1b-violetto-mlx-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "vinci00/limite-1b-violetto-mlx-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default vinci00/limite-1b-violetto-mlx-4bitRun Hermes
hermesLimite 1B Violetto — MLX 4-bit
A community 4-bit MLX quantization of Paradigma's Limite 1B Violetto, converted and evaluated on a Mac mini M4 with 16 GiB unified memory. This derivative changes the storage and inference representation; it adds no training or fine-tuning. The source model is a text-only mathematical reasoner, not a general-purpose or multimodal assistant.
| Property | Value |
|---|---|
| Source revision | 9402e422e87b220507963cd42c997459c1e87b43 |
| Architecture | LimiteForCausalLM, approximately 1.035B parameters |
| Source precision | BF16 matrices with FP32 auxiliary tensors |
| Quantization | MLX affine, 4 bits, group size 64 |
| Weight-file size | Approximately 0.543 GiB |
| Custom adapter | limite-mlx==0.1.0, pinned below |
| License | Apache-2.0 |
Installation and usage
Use an Apple Silicon Mac. The custom Limite architecture is not built into mlx-lm 0.31.3; install the pinned community adapter along with the runtime:
python -m pip install mlx==0.32.2 mlx-lm==0.31.3 transformers==5.17.0 \
"limite-mlx @ https://huggingface.co/pierjoe/limite-mlx/resolve/b21b7a4fa6ca05cb02082d9f7819b8d1387f2f1a/limite_mlx-0.1.0-py3-none-any.whl"
The adapter adds a Python startup import hook for mlx_lm.models.limite.
Restart an already-running notebook kernel after installation.
from mlx_lm import generate, load
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load('vinci00/limite-1b-violetto-mlx-4bit')
prompt = tokenizer.apply_chat_template(
[{'role': 'user', 'content': 'How many positive divisors does 360 have?'}],
tokenize=False,
add_generation_prompt=True,
)
answer = generate(model, tokenizer, prompt=prompt, max_tokens=3000,
sampler=make_sampler(temp=0.6, top_p=0.95))
print(answer)
Keep the original fixed mathematical chat template. It supplies the system prompt and requests a boxed answer; arbitrary system prompts and tools are not supported. The model may produce long reasoning passages before its answer. The sampling parameters above follow the upstream recommendation, while our paired evaluation uses greedy decoding for reproducibility.
Conversion and provenance
Both the unquantized baseline and this quantized derivative were converted from the same pinned original checkpoint, using mlx-lm 0.31.3, MLX 0.32.2, and the same pinned third-party Limite adapter. The baseline was not dequantized from this artifact. Quantization includes the tied input embedding/output head. The adapter folds learned attention scales into projection matrices and retains small FP32 auxiliary tensors.
Both converted configurations accept end-token IDs [151643, 151645], resolving
the upstream mismatch between tokenizer and generation configuration. Tokenizer
and chat-template contents are identical across the paired artifacts. No
Mistral-specific tokenizer regex replacement is applied.
See conversion manifest, weight hashes and source provenance, and recorded hardware. The reproducible source repository contains the executed notebook, evaluator, tests, and pinned requirements.
Evaluation protocol
The teaching evaluation uses the same 10 GSM8K test examples, selected with
seed 42, and a 1,024-token generation cap for both variants. Dataset revision
3101c7d5072418e28b9008a6636bde82a006892c is pinned and its JSONL checksum is
verified. No training or prompt selection uses the test split.
Scoring uses the last complete \boxed{...} after the reasoning section and
requires its entire content to be numeric. LaTeX wrappers and prose outside
the box are allowed. An unfinished <think> section scores zero. Missing, malformed,
nonnumeric, or unfinished final answers count as incorrect. We retain raw
predictions and paired regressions/improvements, correct/total, 95% Wilson
intervals, and paired bootstrap uncertainty. Four fixed mathematical smoke
questions measure runtime separately; runtime excludes model loading and prompt
formatting, and generated lengths can differ.
This small, token-limited run is not a reproduction of the upstream competition scores. Reasoning may exceed the cap. A tie does not establish general quality parity, and pretraining contamination is unknown.
Compatibility and limitations
This is an MLX artifact. Ollama 0.34.2 rejects it with
unsupported MLX architecture: model "LimiteForCausalLM". It is not a usable
Ollama registry release. Stock llama.cpp also lacks this architecture; converting
the container to GGUF alone does not solve runtime support.
The official runtime is Paradigma's vLLM plugin. Our comparison measures quantization within the community MLX adapter; we did not independently establish numerical equivalence to official vLLM. No GGUF, Ollama, image, long-context, or general-assistant quality result is claimed. The source model is lightly instruction-tuned and may reinterpret prompts as mathematical tasks or produce incorrect answers.
License and attribution
The source weights and adapter are Apache-2.0 licensed. Credit for the model
belongs to Paradigma, and the MLX port is by
pierjoe. This conversion is not an
official release by either author. Preserve LICENSE and
NOTICE when redistributing.
Measured results — Mac mini M4, 2026-09-24
The saved notebook completed on Mac mini M4, 16 GiB, macOS 26.4, Python 3.13.15. Both variants passed 4/4 mathematical smoke cases. The GSM8K test evaluation used 10 examples, seed 42, greedy decoding, 1,024 maximum output tokens per example and variant.
| Variant | GSM8K accuracy | Invalid final answers | Mean smoke latency | Output tokens/s | MLX peak memory |
|---|---|---|---|---|---|
| MLX unquantized | 5/10 (50%) | 4 | 12.36 s | 39.54 | 1.974 GiB |
| MLX 4-bit | 5/10 (50%) | 4 | 6.14 s | 87.49 | 0.599 GiB |
Both accuracy estimates have a 95% Wilson interval of 23.66%–76.34%. The paired difference is 0.0 percentage points, with no paired regressions or improvements. The bootstrap interval is [0.0, 0.0] pp because all ten paired correct/incorrect outcomes agree; this degenerate interval does not establish quality parity. Four responses per variant fail the strict numeric final-box format; some generations reach the reasoning token cap. One additional response per variant is formatted correctly but incorrect.
Weight files occupy 1.929 GiB for the source-precision MLX baseline and 0.543 GiB for the 4-bit derivative. The runtime sample is short, runs on a shared machine, excludes loading, and has differing generated lengths; it does not measure peak hardware throughput.
See raw runtime, paired accuracy and predictions, uncertainty report, and hardware. These scores belong only to the MLX artifacts and must not be attributed to Ollama or the official vLLM runtime.
- Downloads last month
- 41
4-bit
Model tree for vinci00/limite-1b-violetto-mlx-4bit
Base model
paradigma-inc/limite-1b-violetto
Start the MLX server
# Install MLX LM: uv tool install mlx-lm# Start a local OpenAI-compatible server: mlx_lm.server --model "vinci00/limite-1b-violetto-mlx-4bit"