Leanstral 1.5 models, quantized.

I don't want to waste your time reading a 500-word LLM-generated essay on what the model is, when Mistral themselves already provide a good explanation in the original model's README. Go read that instead.

Instead, I'll focus on the important part: why should you use my quantization?

  • Properly labeled as mistral4 architecture instead of deepseek2. This is a bug in upstream llama.cpp's GGUF conversion code. PR incoming.
  • Chat template from Leanstral-2603 embedded inside GGUF, no need to specify a template by yourself.
  • I run these models myself on my Strix Halo box.
  • You can interrogate me on the Lean Zulip if you find these quants to be malicious, or if you just have suggestions for improvements.

Enjoy!

Downloads last month
60
GGUF
Model size
119B params
Architecture
mistral4
Hardware compatibility
Log In to add your hardware

4-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for servantofares/Leanstral-1.5-119B-A6B-GGUF

Quantized
(14)
this model