Text Generation
PEFT
Safetensors
French
English
medical
bilingual
french
english
dpo
lora
trl
unsloth
qwen3
conversational
Instructions to use Maphe/qwen3-1.7b-medical-finetuned with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use Maphe/qwen3-1.7b-medical-finetuned with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("unsloth/Qwen3-1.7B-unsloth-bnb-4bit") model = PeftModel.from_pretrained(base_model, "Maphe/qwen3-1.7b-medical-finetuned") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Unsloth Desktop
| base_model: unsloth/Qwen3-1.7B-unsloth-bnb-4bit | |
| library_name: peft | |
| pipeline_tag: text-generation | |
| tags: | |
| - medical | |
| - bilingual | |
| - french | |
| - english | |
| - dpo | |
| - lora | |
| - peft | |
| - trl | |
| - unsloth | |
| - qwen3 | |
| - base_model:adapter:unsloth/Qwen3-1.7B-unsloth-bnb-4bit | |
| language: | |
| - fr | |
| - en | |
| datasets: | |
| - Maphe/medical-sft-5k | |
| - Maphe/medical-dpo-5k | |
| # Qwen3 1.7B Medical Finetuned | |
| This repository contains a bilingual French/English medical LoRA adapter built on top of `unsloth/Qwen3-1.7B-unsloth-bnb-4bit`. | |
| The training workflow used: | |
| 1. supervised fine-tuning (SFT) on a curated medical instruction dataset; | |
| 2. preference alignment with DPO on medical chosen/rejected pairs. | |
| The adapter is intended for experimentation, evaluation, and educational use around medical-domain instruction tuning. It is not a medical device and must not be used as a substitute for a qualified health professional. | |
| ## Model Details | |
| - Base model: `unsloth/Qwen3-1.7B-unsloth-bnb-4bit` | |
| - Adapter type: PEFT LoRA | |
| - Task: causal language modeling / chat-style instruction following | |
| - Languages: French and English | |
| - Final artifact in this folder: DPO-aligned LoRA adapter | |
| - Upstream SFT dataset: `Maphe/medical-sft-5k` | |
| - Upstream DPO dataset: `Maphe/medical-dpo-5k` | |
| ### Training setup | |
| The project uses Unsloth, TRL, PEFT, and bitsandbytes with 4-bit loading. | |
| LoRA configuration: | |
| - `r = 16` | |
| - `lora_alpha = 16` | |
| - `lora_dropout = 0` | |
| - `bias = none` | |
| - Target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` | |
| SFT configuration: | |
| - Epochs: `2` | |
| - Per-device batch size: `32` | |
| - Gradient accumulation: `16` | |
| - Learning rate: `2e-4` | |
| - Scheduler: `cosine` | |
| - Max sequence length: `1024` | |
| - Optimizer: `adamw_8bit` | |
| - Seed: `42` | |
| DPO configuration: | |
| - Epochs: `1` | |
| - Per-device batch size: `4` | |
| - Gradient accumulation: `8` | |
| - Learning rate: `5e-5` | |
| - Beta: `0.1` | |
| - Scheduler: `cosine` | |
| - Max sequence length: `1024` | |
| - Optimizer: `adamw_8bit` | |
| - Seed: `42` | |
| ## Training Data | |
| Two project datasets were prepared and used in the workflow: | |
| - `Maphe/medical-sft-5k` for supervised fine-tuning | |
| - `Maphe/medical-dpo-5k` for preference optimization | |
| The SFT dataset aggregates bilingual medical QA and MCQ-style examples derived from these Hugging Face sources: | |
| - `ANR-MALADES/MediQAl` | |
| - `nthngdy/frenchmedmcqa` | |
| - `keivalya/MedQuad-MedicalQnADataset` | |
| The DPO dataset is built primarily from: | |
| - `TsinghuaC3I/UltraMedical-Preference` | |
| Project-side preprocessing includes: | |
| - schema normalization across heterogeneous sources; | |
| - prompt/response formatting for chat training; | |
| - deduplication on textual pairs; | |
| - source quota sampling; | |
| - deterministic train/validation/test splitting for SFT; | |
| - heuristic PII anonymization with Presidio and regex-based detectors. | |
| The resulting model is optimized for: | |
| - French and English medical questions; | |
| - short factual answers; | |
| - multiple-choice style medical questions; | |
| - structured, direct responses. | |
| ## Prompting Format | |
| The training prompt uses a fixed system instruction: | |
| `Tu es un assistant medical expert. Reponds de maniere claire, factuelle et structuree. Si la question est en anglais, reponds en anglais.` | |
| During training, assistant outputs were formatted in direct-answer mode with an empty Qwen thinking block. This adapter therefore works best with standard chat prompting and concise medical questions. | |
| ## Intended Uses | |
| Appropriate uses: | |
| - research prototypes in domain adaptation; | |
| - comparison between base and finetuned medical assistants; | |
| - educational work on SFT + DPO pipelines; | |
| - internal experimentation on bilingual medical QA. | |
| Out-of-scope uses: | |
| - diagnosis or treatment decisions without clinician oversight; | |
| - emergency triage; | |
| - autonomous clinical decision support; | |
| - legal, regulatory, or production-grade medical advice systems; | |
| - any workflow requiring guaranteed factuality or safety. | |
| ## Evaluation | |
| The repository contains a comparative evaluation between the base model and the SFT checkpoint on `500` examples. | |
| Important: the metrics below are for the SFT checkpoint, not for this final DPO adapter. At the time of writing, no dedicated post-DPO benchmark has been added to the repository. | |
| Available evaluation artifacts: | |
| - `notebooks/eval_results/qwen3_base_vs_sft_output_summary.json` | |
| - `notebooks/eval_results/qwen3_base_vs_sft_output.jsonl` | |
| - `notebooks/eval_results/qwen3_base_vs_sft_output.csv` | |
| Summary of SFT-vs-base results: | |
| - Mean METEOR on free-text answers: `0.1361 -> 0.1653` (`+0.0292`) | |
| - MCQ first-letter score: `0.0515 -> 0.4378` (`+0.3863`) | |
| - MCQ correct answers: `12 -> 102` | |
| Interpretation: | |
| - the finetuning substantially improved MCQ behavior in this project benchmark; | |
| - gains on open-ended generation were positive but more modest; | |
| - automatic metrics remain insufficient to validate clinical quality. | |
| ## Biases, Risks, and Limitations | |
| This model inherits limitations from both the base model and the medical datasets used during fine-tuning. | |
| Known risks: | |
| - hallucinated or overconfident medical statements; | |
| - incomplete coverage of diseases, populations, and care settings; | |
| - source-data bias toward specific question styles; | |
| - imperfect anonymization in upstream preparation; | |
| - limited evaluation depth; | |
| - possible mismatch between benchmark gains and real clinical usefulness. | |
| This adapter should be used only with strong human review and explicit user-facing warnings. | |
| ## How to Use | |
| Example with PEFT and Transformers: | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| from peft import PeftModel | |
| base_model_id = "unsloth/Qwen3-1.7B-unsloth-bnb-4bit" | |
| adapter_path = "Maphe/qwen3-1.7b-medical-finetuned" | |
| tokenizer = AutoTokenizer.from_pretrained(base_model_id) | |
| base_model = AutoModelForCausalLM.from_pretrained(base_model_id) | |
| model = PeftModel.from_pretrained(base_model, adapter_path) | |
| messages = [ | |
| { | |
| "role": "system", | |
| "content": ( | |
| "Tu es un assistant medical expert. " | |
| "Reponds de maniere claire, factuelle et structuree. " | |
| "Si la question est en anglais, reponds en anglais." | |
| ), | |
| }, | |
| {"role": "user", "content": "Quels sont les symptomes principaux du diabete de type 2 ?"}, | |
| ] | |
| prompt = tokenizer.apply_chat_template( | |
| messages, | |
| tokenize=False, | |
| add_generation_prompt=True, | |
| ) | |
| inputs = tokenizer(prompt, return_tensors="pt") | |
| outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False) | |
| print(tokenizer.decode(outputs[0], skip_special_tokens=True)) | |
| ``` | |
| If you use Unsloth in the same way as in the project notebook, load the base model first and then the LoRA adapter exported in this repository. | |
| ## Repository Context | |
| This model card is derived from the accompanying project materials: | |
| - root project documentation in `README.md` | |
| - training notebook: `notebooks/colab_qwen3_unsloth_finetune.ipynb` | |
| - evaluation notebook: `notebooks/colab_qwen3_unsloth_eval_compare.ipynb` | |
| The local training artifacts produced by the project include: | |
| - SFT adapter: `notebooks/qwen3-medical-lora/` | |
| - DPO adapter: `notebooks/qwen3-medical-dpo-lora/` | |
| - SFT checkpoints: `notebooks/sft_output/checkpoint-*` | |
| - DPO checkpoint: `notebooks/dpo_output/checkpoint-157` | |
| ## License | |
| No final consolidated license statement has been added yet in the project for the combined derivative artifact. Before public release, verify: | |
| - the license of the base model; | |
| - the license terms of each source dataset; | |
| - whether redistribution of this adapter is compatible with those upstream terms. | |
| ## Contact | |
| Project owner / publisher: `Maphe` | |
| If you publish this model publicly, it is worth adding: | |
| - the source repository URL; | |
| - exact dataset revisions; | |
| - a dedicated post-DPO evaluation section; | |
| - explicit medical safety disclaimers in the serving application. | |
| ### Framework versions | |
| - PEFT 0.19.1 | |