3-bit VLM base + 4-bit MTP drafter. Pair both for accelerated instruct decode. reasoning_effort baked to low.
Jorge Leon
leonsarmiento
AI & ML interests
AI for Environment
Recent Activity
updated a model 2 days ago
leonsarmiento/Orion-26B-A4B-v1-6bit-XL-mlx published a model 2 days ago
leonsarmiento/Orion-26B-A4B-v1-6bit-XL-mlx updated a model 2 days ago
leonsarmiento/Tiger-Gemma-12B-v3-4bit-mlx