Qwen3.6-27B-slo-med-mt-GGUF-Q8

Trained by MediaAtlas — LLM fine-tuning on your own data, trained in the EU, weights delivered. Pricing · All our models

Na kratko: Strojno prevajanje medicinskih besedil angleščina ⇄ slovenščina (27B, GGUF Q8). Ohranja odmerke, enote, imena zdravil in zanikanja ter uporablja formalni medicinski register. Raziskovalno orodje: vsak prevod mora pred rabo preveriti človek. Sprejme tudi slike.

Vision / slike: takes images too. Load the mmproj-*.gguf next to the model (llama.cpp --mmproj; LM Studio pairs it automatically). Our training was text-only: image understanding is the base model's and was not evaluated in Slovenian.

Q8_0 GGUF build of Qwen3.6-27B-slo-med-mt — an English ⇄ Slovenian medical machine-translation model (dense Qwen3.6-27B fine-tune), for llama.cpp / LM Studio. This is the higher-quality 8-bit build; for the smaller 4-bit build see Qwen3.6-27B-slo-med-mt-GGUF. MTP (multi-token-prediction) is bundled for speculative decoding, and a separate mmproj provides the vision encoder.

Files:

file what ~size
qwen27b-slo-med-mt-Q8_0.gguf text model, 8-bit (higher quality) ~29 GB
mmproj-qwen27b-slo-med-mt-f16.gguf vision projector (F16) ~0.9 GB

The text GGUF has MTP baked in (speculative decoding works out of the box). The mmproj is optional — needed only for vision input; translation is text-only and does not require it.

For the bf16 model see Qwen3.6-27B-slo-med-mt; for the LoRA adapter see Qwen3.6-27B-slo-med-mt-LoRA.


⚠️ Disclaimer

Research/educational only. NOT medical advice, not a clinical product. A translation tool that can mistranslate a dose, unit, drug name, or negation — every output must be verified by a qualified human. The base is uncensored (no safety guardrails); do not deploy in any user-facing/generative role. Hallucinates; known weakness on spelled-out large numbers (digit numbers are reliable). Provided "as is", no warranty, no liability — use at your own risk. Not a medical device; not evaluated by any regulator. See the base model card for full terms.

Usage (llama.cpp)

# text-only medical translation (-sys "<|think_off|>" disables the reasoning trace; -st = one-shot)
llama-cli -m qwen27b-slo-med-mt-Q8_0.gguf --jinja -c 4096 -st \
  -sys "<|think_off|>" --temp 0 --repeat-penalty 1.05 \
  -p "Translate the following English medical text into Slovenian. Output only the translation:\n\nStore the vaccine at 2-8 °C and do not freeze."

# with the vision encoder (optional; only if you pass images)
llama-server -m qwen27b-slo-med-mt-Q8_0.gguf --mmproj mmproj-qwen27b-slo-med-mt-f16.gguf --jinja

Recommended settings (MT)

Reasoning base — for translation keep the reasoning trace off and decode near-deterministically:

setting value
thinking / reasoning OFF — put `<
temperature 0 (greedy)
repeat-penalty 1.05 (--repeat-penalty 1.05)
top-p / top-k 0.9 / 20 (only if sampling)
max tokens 256–512

Disabling the reasoning trace. This chat template gates thinking on the special tokens <|think_off|> / <|think_on|> (not /no_think). To translate directly (no <think> block):

  • LM Studio: set the System Prompt field to <|think_off|>.
  • llama.cpp: add -sys "<|think_off|>" (confirmed working). Note: --chat-template-kwargs '{"enable_thinking":false}' does not disable it in llama.cpp — the <|think_off|> token is the reliable switch.
  • llama-server: send a system message with content <|think_off|>.

Prompt (send as the user message): Translate the following English medical text into Slovenian. Output only the translation:\n\n{src}. Reverse direction: swap the instruction to "Translate the following Slovenian medical text into English." MTP enables speculative decoding for faster output.

Evaluation

  • FLORES-200 devtest (en→sl, 1012): greedy (what llama.cpp uses) BLEU 27.36 / chrF 55.62; beam=5 BLEU 29.47 / chrF 57.62. On this general-domain benchmark the fine-tune is ≈ the base dense 27B (base: 28.03 greedy / 29.36 beam=5); its value is medical terminology, not general MT.
  • Qualitative (100 clinical sentences): formal medical register (odmerek, pediatrični), Croatian-ish borrowings cleaned, doses/units/negation preserved.

Quantization: bf16 → f16 GGUF → Q8_0. Trunk includes the 15 mtp.* tensors; mmproj built from the dense-27B vision encoder (not the 35B-MoE — different dims).

Training & data

QLoRA rank 128, 1 epoch on ~1.0M quality-filtered OPUS EN–SL pairs (medical ≈ 57%): EMEA, ELRC health sets, ECDC, ELRC-SciPar, Europarl. See the base model card for full training details, licenses, and attribution.

License

Qwen license (derivative); also comply with the OPUS corpus terms.

Citation

Cite this work (Tadej Fius, MediaAtlas Ltd):

@misc{fius2026qwen27bmtq8,
  title        = {Qwen3.6-27B Slovenian Medical MT (GGUF Q8)},
  author       = {Fius, Tadej},
  year         = {2026},
  publisher    = {MediaAtlas Ltd},
  howpublished = {Hugging Face},
  url          = {https://huggingface.co/texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8}
}

Upstream / source citations:

@misc{qwen3, title={Qwen3 Technical Report}, author={{Qwen Team}}, year={2025}, url={https://huggingface.co/Qwen}}
@inproceedings{tiedemann2012opus, title={Parallel Data, Tools and Interfaces in OPUS}, author={Tiedemann, J.}, booktitle={LREC}, year={2012}}
@article{nllbflores2022, title={No Language Left Behind}, author={{NLLB Team}}, journal={arXiv:2207.04672}, year={2022}}
@inproceedings{post2018sacrebleu, title={A Call for Clarity in Reporting BLEU Scores}, author={Post, Matt}, booktitle={WMT}, year={2018}}

Corpora (EMEA, ELRC-SHARE, ECDC, ELRC-SciPar, Europarl) via OPUS; eval on FLORES-200 with sacreBLEU.

Downloads last month
216
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8

Collections including texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8

Paper for texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8