Instructions to use leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
L3-8B-Sunfall-v0.5-Stheno-v3.2 — 8-bit (MLX)
Uniform 8-bit MLX quantization of Vdr1/L3-8B-Sunfall-v0.5-Stheno-v3.2 — a Llama-3-8B roleplay/creative-writing merge combining the Sunfall v0.5 and Stheno v3.2 lineages.
Content note: the source model is gated Not-For-All-Audiences (NSFW-capable RP model). Follow the source model's usage terms.
Quantization
| Recipe | Uniform 8-bit, group size 64, affine |
| Bits per weight | 8.5 |
| Size | 8.0 GB |
| Shards | 2 |
| Framework | mlx_lm (native llama) |
Vision/audio: none (text-only model). Chat template shipped in both tokenizer_config.json and chat_template.jinja.
Usage
Works in LM Studio, oMLX, and mlx_lm:
import mlx_lm
model, tokenizer = mlx_lm.load("leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx")
Suggested starting settings (Stheno-lineage RP models; tune to taste — the source card is authoritative): temp 0.8–1.0 · min_p 0.05–0.1 · top_p 1.0 · rep penalty 1.0–1.05 · 8K context
A 4-bit uniform variant is available at leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-4bit-mlx.
- Downloads last month
- 22
8-bit
Model tree for leonsarmiento/L3-8B-Sunfall-v0.5-Stheno-v3.2-8bit-mlx
Base model
Vdr1/L3-8B-Sunfall-v0.5-Stheno-v3.2