sahilchachra's picture
Upload mxfp4 MLX quantization
2ca79c7 verified
|
Raw History Blame Contribute Delete
2.21 kB
metadata
license: apache-2.0
base_model: SupraLabs/Supra-1.5-50M-Instruct-exp
tags:
  - mlx
  - quantized
  - apple-silicon

supra-1.5-50m-instruct-exp-mxfp4-mlx

MLX quantization of SupraLabs/Supra-1.5-50M-Instruct-exp for Apple Silicon.

Variant: Block float MX FP4
Disk size: 28 MB
Quantized by: sahilchachra

Benchmark results

Evaluated on Apple M4 Pro with MLX. Model loaded once; performance and quality measured in a single pass.

Performance

This model FP16 baseline
Decode tok/s (avg, long traces) 1675.62 1025.59
Peak memory (GB) 0.126 0.223
Disk size (MB) 28 101

Quality

Benchmark This model FP16 baseline n
IFEval (instruction following) 20.5% 15.9% 44
Alpaca-cleaned (instruct F1 vs reference) 36.8 40.9 50

Context scaling (decode tok/s)

Context length Decode tok/s
~128 tokens 1675.2
~256 tokens 1685.8
~512 tokens 1655.6
~1024 tokens 1685.9

Usage

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=256, verbose=True)

All variants in this collection

Model Variant
sahilchachra/supra-1.5-50m-instruct-exp-mxfp4-mlx Block float MX FP4 ← this model
sahilchachra/supra-1.5-50m-instruct-exp-mxfp8-mlx Block float MX FP8

Notes

  • Requires Apple Silicon (M1 or later) with MLX
  • Benchmarks run on Apple M4 Pro, 24 GB unified memory
  • License: see SupraLabs/Supra-1.5-50M-Instruct-exp for the original model's license

Original model

See SupraLabs/Supra-1.5-50M-Instruct-exp for full model details and intended use.