Instructions to use sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx --local-dir supra-1.5-50m-instruct-exp-mxfp4-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
|
Download README.md from sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx: direct link, hf CLI and curl.
- Browser
- Download file 2.21 kB
-
https://huggingface.co/sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx/resolve/main/README.md
- Command line
-
hf download hf://sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx/README.md
-
curl -L -o README.md https://huggingface.co/sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx/resolve/main/README.md
2.21 kB
| license: apache-2.0 | |
| base_model: SupraLabs/Supra-1.5-50M-Instruct-exp | |
| tags: | |
| - mlx | |
| - quantized | |
| - apple-silicon | |
| # supra-1.5-50m-instruct-exp-mxfp4-mlx | |
| MLX quantization of [SupraLabs/Supra-1.5-50M-Instruct-exp](https://huggingface.co/SupraLabs/Supra-1.5-50M-Instruct-exp) for Apple Silicon. | |
| **Variant**: Block float MX FP4 | |
| **Disk size**: 28 MB | |
| **Quantized by**: [sahilchachra](https://huggingface.co/sahilchachra) | |
| ## Benchmark results | |
| Evaluated on Apple M4 Pro with MLX. Model loaded once; performance and quality measured in a single pass. | |
| ### Performance | |
| | | This model | FP16 baseline | | |
| |---|---:|---:| | |
| | Decode tok/s (avg, long traces) | 1675.62 | 1025.59 | | |
| | Peak memory (GB) | 0.126 | 0.223 | | |
| | Disk size (MB) | 28 | 101 | | |
| ### Quality | |
| | Benchmark | This model | FP16 baseline | n | | |
| |---|---:|---:|---:| | |
| | IFEval (instruction following) | 20.5% | 15.9% | 44 | | |
| | Alpaca-cleaned (instruct F1 vs reference) | 36.8 | 40.9 | 50 | | |
| ### Context scaling (decode tok/s) | |
| | Context length | Decode tok/s | | |
| |---:|---:| | |
| | ~128 tokens | 1675.2 | | |
| | ~256 tokens | 1685.8 | | |
| | ~512 tokens | 1655.6 | | |
| | ~1024 tokens | 1685.9 | | |
| ## Usage | |
| ```bash | |
| pip install mlx-lm | |
| ``` | |
| ```python | |
| from mlx_lm import load, generate | |
| model, tokenizer = load("sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx") | |
| response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True) | |
| ``` | |
| ## All variants in this collection | |
| | Model | Variant | | |
| |---|---| | |
| | [sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx](https://huggingface.co/sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx) | Block float MX FP4 ← this model | | |
| | [sahilchachra/supra-1.5-50m-instruct-exp-mxfp8-mlx](https://huggingface.co/sahilchachra/supra-1.5-50m-instruct-exp-mxfp8-mlx) | Block float MX FP8 | | |
| ## Notes | |
| - Requires Apple Silicon (M1 or later) with MLX | |
| - Benchmarks run on Apple M4 Pro, 24 GB unified memory | |
| - License: see [SupraLabs/Supra-1.5-50M-Instruct-exp](https://huggingface.co/SupraLabs/Supra-1.5-50M-Instruct-exp) for the original model's license | |
| ## Original model | |
| See [SupraLabs/Supra-1.5-50M-Instruct-exp](https://huggingface.co/SupraLabs/Supra-1.5-50M-Instruct-exp) for full model details and intended use. |