Instructions to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16 # Run inference directly in the terminal: llama cli -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16 # Run inference directly in the terminal: llama cli -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16 # Run inference directly in the terminal: ./llama-cli -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Use Docker
docker model run hf.co/texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
- LM Studio
- Jan
- Ollama
How to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with Ollama:
ollama run hf.co/texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
- Unsloth Desktop
- Pi
How to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with Docker Model Runner:
docker model run hf.co/texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
- Lemonade
How to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Run and chat with the model
lemonade run user.Qwen3.6-27B-slo-med-mt-GGUF-Q8-F16
List all available models
lemonade list
- Hermes Agent
How to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8:F16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.6-27B-slo-med-mt-GGUF-Q8
Trained by MediaAtlas — LLM fine-tuning on your own data, trained in the EU, weights delivered. Pricing · All our models
Na kratko: Strojno prevajanje medicinskih besedil angleščina ⇄ slovenščina (27B, GGUF Q8). Ohranja odmerke, enote, imena zdravil in zanikanja ter uporablja formalni medicinski register. Raziskovalno orodje: vsak prevod mora pred rabo preveriti človek. Sprejme tudi slike.
Vision / slike: takes images too. Load the mmproj-*.gguf next to the model (llama.cpp --mmproj; LM Studio pairs it automatically). Our training was text-only: image understanding is the base model's and was not evaluated in Slovenian.
Q8_0 GGUF build of Qwen3.6-27B-slo-med-mt — an English ⇄ Slovenian
medical machine-translation model (dense Qwen3.6-27B fine-tune), for
llama.cpp / LM Studio. This is the higher-quality 8-bit build; for the
smaller 4-bit build see
Qwen3.6-27B-slo-med-mt-GGUF.
MTP (multi-token-prediction) is bundled for speculative decoding, and a
separate mmproj provides the vision encoder.
Files:
| file | what | ~size |
|---|---|---|
qwen27b-slo-med-mt-Q8_0.gguf |
text model, 8-bit (higher quality) | ~29 GB |
mmproj-qwen27b-slo-med-mt-f16.gguf |
vision projector (F16) | ~0.9 GB |
The text GGUF has MTP baked in (speculative decoding works out of the box). The mmproj is optional — needed only for vision input; translation is text-only and does not require it.
For the bf16 model see Qwen3.6-27B-slo-med-mt;
for the LoRA adapter see Qwen3.6-27B-slo-med-mt-LoRA.
⚠️ Disclaimer
Research/educational only. NOT medical advice, not a clinical product. A translation tool that can mistranslate a dose, unit, drug name, or negation — every output must be verified by a qualified human. The base is uncensored (no safety guardrails); do not deploy in any user-facing/generative role. Hallucinates; known weakness on spelled-out large numbers (digit numbers are reliable). Provided "as is", no warranty, no liability — use at your own risk. Not a medical device; not evaluated by any regulator. See the base model card for full terms.
Usage (llama.cpp)
# text-only medical translation (-sys "<|think_off|>" disables the reasoning trace; -st = one-shot)
llama-cli -m qwen27b-slo-med-mt-Q8_0.gguf --jinja -c 4096 -st \
-sys "<|think_off|>" --temp 0 --repeat-penalty 1.05 \
-p "Translate the following English medical text into Slovenian. Output only the translation:\n\nStore the vaccine at 2-8 °C and do not freeze."
# with the vision encoder (optional; only if you pass images)
llama-server -m qwen27b-slo-med-mt-Q8_0.gguf --mmproj mmproj-qwen27b-slo-med-mt-f16.gguf --jinja
Recommended settings (MT)
Reasoning base — for translation keep the reasoning trace off and decode near-deterministically:
| setting | value |
|---|---|
| thinking / reasoning | OFF — put `< |
| temperature | 0 (greedy) |
| repeat-penalty | 1.05 (--repeat-penalty 1.05) |
| top-p / top-k | 0.9 / 20 (only if sampling) |
| max tokens | 256–512 |
Disabling the reasoning trace. This chat template gates thinking on the
special tokens <|think_off|> / <|think_on|> (not /no_think). To translate
directly (no <think> block):
- LM Studio: set the System Prompt field to
<|think_off|>. - llama.cpp: add
-sys "<|think_off|>"(confirmed working). Note:--chat-template-kwargs '{"enable_thinking":false}'does not disable it in llama.cpp — the<|think_off|>token is the reliable switch. - llama-server: send a
systemmessage with content<|think_off|>.
Prompt (send as the user message): Translate the following English medical text into Slovenian. Output only the translation:\n\n{src}. Reverse direction:
swap the instruction to "Translate the following Slovenian medical text into
English." MTP enables speculative decoding for faster output.
Evaluation
- FLORES-200 devtest (en→sl, 1012): greedy (what llama.cpp uses) BLEU 27.36 / chrF 55.62; beam=5 BLEU 29.47 / chrF 57.62. On this general-domain benchmark the fine-tune is ≈ the base dense 27B (base: 28.03 greedy / 29.36 beam=5); its value is medical terminology, not general MT.
- Qualitative (100 clinical sentences): formal medical register (odmerek, pediatrični), Croatian-ish borrowings cleaned, doses/units/negation preserved.
Quantization: bf16 → f16 GGUF → Q8_0. Trunk includes the 15 mtp.* tensors;
mmproj built from the dense-27B vision encoder (not the 35B-MoE — different
dims).
Training & data
QLoRA rank 128, 1 epoch on ~1.0M quality-filtered OPUS EN–SL pairs (medical ≈ 57%): EMEA, ELRC health sets, ECDC, ELRC-SciPar, Europarl. See the base model card for full training details, licenses, and attribution.
License
Qwen license (derivative); also comply with the OPUS corpus terms.
Citation
Cite this work (Tadej Fius, MediaAtlas Ltd):
@misc{fius2026qwen27bmtq8,
title = {Qwen3.6-27B Slovenian Medical MT (GGUF Q8)},
author = {Fius, Tadej},
year = {2026},
publisher = {MediaAtlas Ltd},
howpublished = {Hugging Face},
url = {https://huggingface.co/texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8}
}
Upstream / source citations:
@misc{qwen3, title={Qwen3 Technical Report}, author={{Qwen Team}}, year={2025}, url={https://huggingface.co/Qwen}}
@inproceedings{tiedemann2012opus, title={Parallel Data, Tools and Interfaces in OPUS}, author={Tiedemann, J.}, booktitle={LREC}, year={2012}}
@article{nllbflores2022, title={No Language Left Behind}, author={{NLLB Team}}, journal={arXiv:2207.04672}, year={2022}}
@inproceedings{post2018sacrebleu, title={A Call for Clarity in Reporting BLEU Scores}, author={Post, Matt}, booktitle={WMT}, year={2018}}
Corpora (EMEA, ELRC-SHARE, ECDC, ELRC-SciPar, Europarl) via OPUS; eval on FLORES-200 with sacreBLEU.
- Downloads last month
- 216
8-bit
Model tree for texdata/Qwen3.6-27B-slo-med-mt-GGUF-Q8
Base model
Qwen/Qwen3.6-27B