L3-8B-Sunfall-v0.5-Stheno-v3.2 — 8-bit (MLX)

Uniform 8-bit MLX quantization of Vdr1/L3-8B-Sunfall-v0.5-Stheno-v3.2 — a Llama-3-8B roleplay/creative-writing merge combining the Sunfall v0.5 and Stheno v3.2 lineages.

Content note: the source model is gated Not-For-All-Audiences (NSFW-capable RP model). Follow the source model's usage terms.

Quantization

Recipe Uniform 8-bit, group size 64, affine
Bits per weight 8.5
Size 8.0 GB
Shards 2
Framework mlx_lm (native llama)

Vision/audio: none (text-only model). Chat template shipped in both tokenizer_config.json and chat_template.jinja.

Usage

Works in LM Studio, oMLX, and mlx_lm:

import mlx_lm
model, tokenizer = mlx_lm.load("leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx")

Suggested starting settings (Stheno-lineage RP models; tune to taste — the source card is authoritative): temp 0.8–1.0 · min_p 0.05–0.1 · top_p 1.0 · rep penalty 1.0–1.05 · 8K context

A 4-bit uniform variant is available at leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-4bit-mlx.

Downloads last month
22
Safetensors
Model size
8B params
Tensor type
U32
·
F16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx

Quantized
(6)
this model