Image-Text-to-Text
Transformers
Safetensors
GGUF
English
qwen3_5
text-generation-inference
unsloth
medical
triage
emergency-medicine
dpo
rlhf
medical-llm
clinical
healthcare
medicine
medical-ai
clinical-decision-support
conversational
Instructions to use vadimbelsky/qwen3.5-medical-ft-stage3-dpo with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use vadimbelsky/qwen3.5-medical-ft-stage3-dpo with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="vadimbelsky/qwen3.5-medical-ft-stage3-dpo") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("vadimbelsky/qwen3.5-medical-ft-stage3-dpo") model = AutoModelForMultimodalLM.from_pretrained("vadimbelsky/qwen3.5-medical-ft-stage3-dpo", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use vadimbelsky/qwen3.5-medical-ft-stage3-dpo with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "vadimbelsky/qwen3.5-medical-ft-stage3-dpo" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vadimbelsky/qwen3.5-medical-ft-stage3-dpo", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/vadimbelsky/qwen3.5-medical-ft-stage3-dpo
- SGLang
How to use vadimbelsky/qwen3.5-medical-ft-stage3-dpo with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "vadimbelsky/qwen3.5-medical-ft-stage3-dpo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vadimbelsky/qwen3.5-medical-ft-stage3-dpo", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "vadimbelsky/qwen3.5-medical-ft-stage3-dpo" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "vadimbelsky/qwen3.5-medical-ft-stage3-dpo", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Unsloth Desktop
- Docker Model Runner
How to use vadimbelsky/qwen3.5-medical-ft-stage3-dpo with Docker Model Runner:
docker model run hf.co/vadimbelsky/qwen3.5-medical-ft-stage3-dpo
File size: 6,345 Bytes
4d58b28 6fc17a7 4d58b28 b383b05 b5731e9 c43ba49 4d58b28 6fc17a7 b5731e9 6fc17a7 b5731e9 6fc17a7 4d58b28 6fc17a7 b958f59 6fc17a7 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 | ---
base_model: vadimbelsky/qwen3.5-medical-ft-stage2
tags:
- text-generation-inference
- transformers
- unsloth
- qwen3_5
- medical
- triage
- emergency-medicine
- dpo
- rlhf
- medical-llm
- clinical
- healthcare
- medicine
- medical-ai
- clinical-decision-support
license: apache-2.0
language:
- en
---
# Qwen3.5-9B Medical Triage — Stage 3 DPO (v4)
Emergency department triage model fine-tuned on Qwen3.5-9B via a 3-stage pipeline:
**Stage 1** (general medical SFT) → **Stage 2** (ED intake SOAP → ESI decision SFT) → **Stage 3** (DPO alignment to reduce over-triage, this model).
Quantized to **Q4_K_M GGUF** for on-device inference.
---
## Model Description
Given an ED SOAP intake note, the model outputs a structured triage decision:
- **ESI level** (1–5) with justification
- Key clinical findings
- Time-to-provider target
- Immediate interventions required
**ESI Scale:** 1 = Immediate life threat · 2 = Emergent high-risk · 3 = Urgent stable · 4 = Less urgent · 5 = Non-urgent
---
## Training Pipeline
| Stage | Method | Objective |
|-------|--------|-----------|
| 1 | SFT (LoRA r=16) | General medical knowledge (PubMed, clinical guidelines) |
| 2 | SFT (LoRA r=16) | SOAP note → structured ESI triage decision |
| 3 | DPO (LoRA r=8) | Reduce over-triage · preserve ESI 1/2 high-risk recall |
### Stage 3 DPO Details
- **Base:** Stage 2 LoRA checkpoint (`vadimbelsky/qwen3.5-medical-ft-stage2`)
- **Dataset:** `dpo_dataset_v4.jsonl` — 5,413 raw pairs → 7,789 weighted pairs
- **Loss:** Combined `apo_down × 0.3 + sft × 1.0` (MPO-style)
- **Beta:** 0.5 · **LR:** 5e-5 · **Epochs:** 0.1 (47 steps)
- **Batch:** 2 × 8 gradient accumulation = effective 16
- **ESI label prepending:** All chosen/rejected completions prefixed with explicit ESI label (e.g. `ESI 2 — Emergent (high risk)\n\n...`) to anchor preference signal at token position 0
### Dataset Sources (v4)
| Source | Description | Raw pairs | Weight | Weighted |
|--------|-------------|-----------|--------|---------|
| A | Anti-overtriage synthetic (ESI 3→1/2 rejected) | 2,388 | 1× | 2,388 |
| B | Anti-overtriage synthetic (ESI 4/5→1/2 rejected) | 1,500 | 1× | 1,500 |
| C | Edge cases (synthetic boundary scenarios) | 39 | 1× | 39 |
| D | ESI 1/2 anchor pairs (high-risk recall preservation) | 890 | 3× | 2,670 |
| E-over | ESI 3 bidirectional — anti-overtriage | 297 | 2× | 594 |
| E-under | ESI 3 bidirectional — anti-undertriage | 299 | 2× | 598 |
| **Total** | | **5,413** | | **7,789** |
---
## Evaluation Results
Evaluated on **MIMIC-IV-Ext Triage Instruction Corpus** (MIETIC) — 36 human-expert validated RETAIN cases.
### v4 vs Previous Stages
| Metric | Stage 2 (SFT) | v1 DPO | v2 DPO | v3 DPO | **v4 DPO** | Target |
|--------|--------------|--------|--------|--------|-----------|--------|
| Accuracy | ~68% | 55.6% | 50.0% | 27.8% | **75.0%** | >82% |
| Over-triage rate | ~22% | 22.2% | 30.6% | 0% | **13.9%** | <10% |
| Under-triage rate | ~8% | 36.1% | 41.7% | 72.2% | **11.1%** | <6% |
| High-risk recall (ESI 1+2) | ~84% | 76% | 64% | 40% | **92%** | 100% |
| ESI 3 accuracy | ~45% | ~40% | ~30% | ~0% | **60%** | >65% |
### v4 Detailed Results (MIETIC, n=36)
```
Samples evaluated : 36
ESI level parsed : 36 / 36
Correct : 27
Accuracy : 75.0%
Under-triage rate : 11.1% (4 cases)
Over-triage rate : 13.9% (5 cases)
High-risk recall : 92.0% (ESI 1+2, n=25)
```
**Per-ESI Accuracy:**
| ESI Level | N | Correct | Accuracy |
|-----------|----|---------|----------|
| ESI 1 | 14 | 12 | 85.7% |
| ESI 2 | 11 | 9 | 81.8% |
| ESI 3 | 5 | 3 | 60.0% |
| ESI 4 | 4 | 2 | 50.0% |
| ESI 5 | 2 | 1 | 50.0% |
**Confusion Matrix** (rows = ground truth, cols = predicted):
```
GT \ Pred ESI 1 ESI 2 ESI 3 ESI 4 ESI 5
ESI 1 12 2 0 0 0
ESI 2 0 9 2 0 0
ESI 3 0 2 3 0 0
ESI 4 0 0 2 2 0
ESI 5 0 0 0 1 1
```
All remaining errors are ±1 ESI boundary confusions — no catastrophic mis-triage.
---
## Key Lessons from DPO Iteration
- **v1–v3 failure:** IPO/sigmoid loss collapsed when dataset direction was 100% anti-overtriage → catastrophic under-triage regression (40% high-risk recall at worst)
- **v4 fix:** (1) ESI label prepended at token position 0 for unambiguous preference signal; (2) `apo_down + sft` combined loss preserves ESI 1/2 recall via SFT component; (3) Sources D (ESI 1/2 anchors ×3) + E (ESI 3 bidirectional ×2) balance dataset direction
---
## Usage
```python
# Requires llama.cpp server running with the Q4_K_M GGUF
# llama-server --model qwen3.5-medical-ft-stage3-dpo-q4km.gguf --port 8080 -c 4096
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8080/v1", api_key="none")
SYSTEM_PROMPT = (
"You are an expert emergency medicine triage nurse. "
"Given a SOAP intake note, provide a structured triage decision including "
"ESI level with justification, key clinical findings, time-to-provider target, "
"and any immediate interventions required."
)
response = client.chat.completions.create(
model="local",
messages=[
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": "<SOAP intake note here>"},
],
temperature=0.1,
max_tokens=512,
)
print(response.choices[0].message.content)
```
---
## Limitations & Safety
> ⚠️ **This model is for research purposes only. It must NOT be used for clinical decision-making without licensed clinician oversight.**
- Evaluated on 36 MIETIC validation cases — not a clinical trial
- 11.1% under-triage rate means critical patients may be down-triaged
- 92% high-risk recall means ~8% of ESI 1/2 patients may be missed
- Model has not been validated on real ED populations
- Fine-tuned on synthetic + MIMIC-IV derived data only
---
## Training Infrastructure
- **Hardware:** NVIDIA GB10 (121 GB VRAM), 1 GPU
- **Framework:** Unsloth 2026.3.4 + TRL DPOTrainer + Transformers 5.2.0
- **Training time:** ~2 hours (47 steps)
- **Quantization:** GGUF Q4_K_M via llama.cpp
---
*Fine-tuned with [Unsloth](https://github.com/unslothai/unsloth) 🦥*
|