Maphe commited on
Commit
b4ab03b
·
verified ·
1 Parent(s): de53676

Upload folder using huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +170 -127
README.md CHANGED
@@ -3,208 +3,251 @@ base_model: unsloth/Qwen3-1.7B-unsloth-bnb-4bit
3
  library_name: peft
4
  pipeline_tag: text-generation
5
  tags:
6
- - base_model:adapter:unsloth/Qwen3-1.7B-unsloth-bnb-4bit
 
 
 
7
  - dpo
8
  - lora
9
- - transformers
10
  - trl
11
  - unsloth
 
 
 
 
 
 
 
 
12
  ---
13
 
14
- # Model Card for Model ID
15
-
16
- <!-- Provide a quick summary of what the model is/does. -->
17
-
18
-
19
-
20
- ## Model Details
21
-
22
- ### Model Description
23
-
24
- <!-- Provide a longer summary of what this model is. -->
25
-
26
-
27
-
28
- - **Developed by:** [More Information Needed]
29
- - **Funded by [optional]:** [More Information Needed]
30
- - **Shared by [optional]:** [More Information Needed]
31
- - **Model type:** [More Information Needed]
32
- - **Language(s) (NLP):** [More Information Needed]
33
- - **License:** [More Information Needed]
34
- - **Finetuned from model [optional]:** [More Information Needed]
35
-
36
- ### Model Sources [optional]
37
-
38
- <!-- Provide the basic links for the model. -->
39
-
40
- - **Repository:** [More Information Needed]
41
- - **Paper [optional]:** [More Information Needed]
42
- - **Demo [optional]:** [More Information Needed]
43
 
44
- ## Uses
45
 
46
- <!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
47
 
48
- ### Direct Use
 
49
 
50
- <!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
51
 
52
- [More Information Needed]
53
 
54
- ### Downstream Use [optional]
 
 
 
 
 
 
55
 
56
- <!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
57
 
58
- [More Information Needed]
59
 
60
- ### Out-of-Scope Use
61
 
62
- <!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
 
 
 
 
63
 
64
- [More Information Needed]
65
 
66
- ## Bias, Risks, and Limitations
 
 
 
 
 
 
 
67
 
68
- <!-- This section is meant to convey both technical and sociotechnical limitations. -->
69
 
70
- [More Information Needed]
 
 
 
 
 
 
 
 
71
 
72
- ### Recommendations
73
 
74
- <!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
75
 
76
- Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
 
77
 
78
- ## How to Get Started with the Model
79
 
80
- Use the code below to get started with the model.
 
 
81
 
82
- [More Information Needed]
83
 
84
- ## Training Details
85
 
86
- ### Training Data
87
 
88
- <!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
 
 
 
 
 
89
 
90
- [More Information Needed]
91
 
92
- ### Training Procedure
 
 
 
93
 
94
- <!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
95
 
96
- #### Preprocessing [optional]
97
 
98
- [More Information Needed]
99
 
 
100
 
101
- #### Training Hyperparameters
102
 
103
- - **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
104
 
105
- #### Speeds, Sizes, Times [optional]
 
 
 
106
 
107
- <!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
108
 
109
- [More Information Needed]
 
 
 
 
110
 
111
  ## Evaluation
112
 
113
- <!-- This section describes the evaluation protocols and provides the results. -->
114
-
115
- ### Testing Data, Factors & Metrics
116
-
117
- #### Testing Data
118
-
119
- <!-- This should link to a Dataset Card if possible. -->
120
-
121
- [More Information Needed]
122
-
123
- #### Factors
124
-
125
- <!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
126
-
127
- [More Information Needed]
128
-
129
- #### Metrics
130
-
131
- <!-- These are the evaluation metrics being used, ideally with a description of why. -->
132
-
133
- [More Information Needed]
134
-
135
- ### Results
136
-
137
- [More Information Needed]
138
-
139
- #### Summary
140
 
 
141
 
 
142
 
143
- ## Model Examination [optional]
 
 
144
 
145
- <!-- Relevant interpretability work for the model goes here -->
146
 
147
- [More Information Needed]
 
 
148
 
149
- ## Environmental Impact
150
 
151
- <!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
 
 
152
 
153
- Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
154
 
155
- - **Hardware Type:** [More Information Needed]
156
- - **Hours used:** [More Information Needed]
157
- - **Cloud Provider:** [More Information Needed]
158
- - **Compute Region:** [More Information Needed]
159
- - **Carbon Emitted:** [More Information Needed]
160
 
161
- ## Technical Specifications [optional]
162
 
163
- ### Model Architecture and Objective
 
 
 
 
 
164
 
165
- [More Information Needed]
166
 
167
- ### Compute Infrastructure
168
 
169
- [More Information Needed]
170
 
171
- #### Hardware
 
 
172
 
173
- [More Information Needed]
 
174
 
175
- #### Software
 
 
176
 
177
- [More Information Needed]
 
 
 
 
 
 
 
 
 
 
178
 
179
- ## Citation [optional]
 
 
 
 
 
 
 
 
180
 
181
- <!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
182
 
183
- **BibTeX:**
184
 
185
- [More Information Needed]
186
 
187
- **APA:**
 
 
188
 
189
- [More Information Needed]
190
 
191
- ## Glossary [optional]
 
 
 
192
 
193
- <!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
194
 
195
- [More Information Needed]
196
 
197
- ## More Information [optional]
 
 
198
 
199
- [More Information Needed]
200
 
201
- ## Model Card Authors [optional]
202
 
203
- [More Information Needed]
204
 
205
- ## Model Card Contact
 
 
 
206
 
207
- [More Information Needed]
208
  ### Framework versions
209
 
210
- - PEFT 0.19.1
 
3
  library_name: peft
4
  pipeline_tag: text-generation
5
  tags:
6
+ - medical
7
+ - bilingual
8
+ - french
9
+ - english
10
  - dpo
11
  - lora
12
+ - peft
13
  - trl
14
  - unsloth
15
+ - qwen3
16
+ - base_model:adapter:unsloth/Qwen3-1.7B-unsloth-bnb-4bit
17
+ language:
18
+ - fr
19
+ - en
20
+ datasets:
21
+ - Maphe/medical-sft-5k
22
+ - Maphe/medical-dpo-5k
23
  ---
24
 
25
+ # Qwen3 1.7B Medical Finetuned
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26
 
27
+ This repository contains a bilingual French/English medical LoRA adapter built on top of `unsloth/Qwen3-1.7B-unsloth-bnb-4bit`.
28
 
29
+ The training workflow used:
30
 
31
+ 1. supervised fine-tuning (SFT) on a curated medical instruction dataset;
32
+ 2. preference alignment with DPO on medical chosen/rejected pairs.
33
 
34
+ The adapter is intended for experimentation, evaluation, and educational use around medical-domain instruction tuning. It is not a medical device and must not be used as a substitute for a qualified health professional.
35
 
36
+ ## Model Details
37
 
38
+ - Base model: `unsloth/Qwen3-1.7B-unsloth-bnb-4bit`
39
+ - Adapter type: PEFT LoRA
40
+ - Task: causal language modeling / chat-style instruction following
41
+ - Languages: French and English
42
+ - Final artifact in this folder: DPO-aligned LoRA adapter
43
+ - Upstream SFT dataset: `Maphe/medical-sft-5k`
44
+ - Upstream DPO dataset: `Maphe/medical-dpo-5k`
45
 
46
+ ### Training setup
47
 
48
+ The project uses Unsloth, TRL, PEFT, and bitsandbytes with 4-bit loading.
49
 
50
+ LoRA configuration:
51
 
52
+ - `r = 16`
53
+ - `lora_alpha = 16`
54
+ - `lora_dropout = 0`
55
+ - `bias = none`
56
+ - Target modules: `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj`
57
 
58
+ SFT configuration:
59
 
60
+ - Epochs: `2`
61
+ - Per-device batch size: `32`
62
+ - Gradient accumulation: `16`
63
+ - Learning rate: `2e-4`
64
+ - Scheduler: `cosine`
65
+ - Max sequence length: `1024`
66
+ - Optimizer: `adamw_8bit`
67
+ - Seed: `42`
68
 
69
+ DPO configuration:
70
 
71
+ - Epochs: `1`
72
+ - Per-device batch size: `4`
73
+ - Gradient accumulation: `8`
74
+ - Learning rate: `5e-5`
75
+ - Beta: `0.1`
76
+ - Scheduler: `cosine`
77
+ - Max sequence length: `1024`
78
+ - Optimizer: `adamw_8bit`
79
+ - Seed: `42`
80
 
81
+ ## Training Data
82
 
83
+ Two project datasets were prepared and used in the workflow:
84
 
85
+ - `Maphe/medical-sft-5k` for supervised fine-tuning
86
+ - `Maphe/medical-dpo-5k` for preference optimization
87
 
88
+ The SFT dataset aggregates bilingual medical QA and MCQ-style examples derived from these Hugging Face sources:
89
 
90
+ - `ANR-MALADES/MediQAl`
91
+ - `nthngdy/frenchmedmcqa`
92
+ - `keivalya/MedQuad-MedicalQnADataset`
93
 
94
+ The DPO dataset is built primarily from:
95
 
96
+ - `TsinghuaC3I/UltraMedical-Preference`
97
 
98
+ Project-side preprocessing includes:
99
 
100
+ - schema normalization across heterogeneous sources;
101
+ - prompt/response formatting for chat training;
102
+ - deduplication on textual pairs;
103
+ - source quota sampling;
104
+ - deterministic train/validation/test splitting for SFT;
105
+ - heuristic PII anonymization with Presidio and regex-based detectors.
106
 
107
+ The resulting model is optimized for:
108
 
109
+ - French and English medical questions;
110
+ - short factual answers;
111
+ - multiple-choice style medical questions;
112
+ - structured, direct responses.
113
 
114
+ ## Prompting Format
115
 
116
+ The training prompt uses a fixed system instruction:
117
 
118
+ `Tu es un assistant medical expert. Reponds de maniere claire, factuelle et structuree. Si la question est en anglais, reponds en anglais.`
119
 
120
+ During training, assistant outputs were formatted in direct-answer mode with an empty Qwen thinking block. This adapter therefore works best with standard chat prompting and concise medical questions.
121
 
122
+ ## Intended Uses
123
 
124
+ Appropriate uses:
125
 
126
+ - research prototypes in domain adaptation;
127
+ - comparison between base and finetuned medical assistants;
128
+ - educational work on SFT + DPO pipelines;
129
+ - internal experimentation on bilingual medical QA.
130
 
131
+ Out-of-scope uses:
132
 
133
+ - diagnosis or treatment decisions without clinician oversight;
134
+ - emergency triage;
135
+ - autonomous clinical decision support;
136
+ - legal, regulatory, or production-grade medical advice systems;
137
+ - any workflow requiring guaranteed factuality or safety.
138
 
139
  ## Evaluation
140
 
141
+ The repository contains a comparative evaluation between the base model and the SFT checkpoint on `500` examples.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
142
 
143
+ Important: the metrics below are for the SFT checkpoint, not for this final DPO adapter. At the time of writing, no dedicated post-DPO benchmark has been added to the repository.
144
 
145
+ Available evaluation artifacts:
146
 
147
+ - `notebooks/eval_results/qwen3_base_vs_sft_output_summary.json`
148
+ - `notebooks/eval_results/qwen3_base_vs_sft_output.jsonl`
149
+ - `notebooks/eval_results/qwen3_base_vs_sft_output.csv`
150
 
151
+ Summary of SFT-vs-base results:
152
 
153
+ - Mean METEOR on free-text answers: `0.1361 -> 0.1653` (`+0.0292`)
154
+ - MCQ first-letter score: `0.0515 -> 0.4378` (`+0.3863`)
155
+ - MCQ correct answers: `12 -> 102`
156
 
157
+ Interpretation:
158
 
159
+ - the finetuning substantially improved MCQ behavior in this project benchmark;
160
+ - gains on open-ended generation were positive but more modest;
161
+ - automatic metrics remain insufficient to validate clinical quality.
162
 
163
+ ## Biases, Risks, and Limitations
164
 
165
+ This model inherits limitations from both the base model and the medical datasets used during fine-tuning.
 
 
 
 
166
 
167
+ Known risks:
168
 
169
+ - hallucinated or overconfident medical statements;
170
+ - incomplete coverage of diseases, populations, and care settings;
171
+ - source-data bias toward specific question styles;
172
+ - imperfect anonymization in upstream preparation;
173
+ - limited evaluation depth;
174
+ - possible mismatch between benchmark gains and real clinical usefulness.
175
 
176
+ This adapter should be used only with strong human review and explicit user-facing warnings.
177
 
178
+ ## How to Use
179
 
180
+ Example with PEFT and Transformers:
181
 
182
+ ```python
183
+ from transformers import AutoModelForCausalLM, AutoTokenizer
184
+ from peft import PeftModel
185
 
186
+ base_model_id = "unsloth/Qwen3-1.7B-unsloth-bnb-4bit"
187
+ adapter_path = "Maphe/qwen3-1.7b-medical-finetuned"
188
 
189
+ tokenizer = AutoTokenizer.from_pretrained(base_model_id)
190
+ base_model = AutoModelForCausalLM.from_pretrained(base_model_id)
191
+ model = PeftModel.from_pretrained(base_model, adapter_path)
192
 
193
+ messages = [
194
+ {
195
+ "role": "system",
196
+ "content": (
197
+ "Tu es un assistant medical expert. "
198
+ "Reponds de maniere claire, factuelle et structuree. "
199
+ "Si la question est en anglais, reponds en anglais."
200
+ ),
201
+ },
202
+ {"role": "user", "content": "Quels sont les symptomes principaux du diabete de type 2 ?"},
203
+ ]
204
 
205
+ prompt = tokenizer.apply_chat_template(
206
+ messages,
207
+ tokenize=False,
208
+ add_generation_prompt=True,
209
+ )
210
+ inputs = tokenizer(prompt, return_tensors="pt")
211
+ outputs = model.generate(**inputs, max_new_tokens=256, do_sample=False)
212
+ print(tokenizer.decode(outputs[0], skip_special_tokens=True))
213
+ ```
214
 
215
+ If you use Unsloth in the same way as in the project notebook, load the base model first and then the LoRA adapter exported in this repository.
216
 
217
+ ## Repository Context
218
 
219
+ This model card is derived from the accompanying project materials:
220
 
221
+ - root project documentation in `README.md`
222
+ - training notebook: `notebooks/colab_qwen3_unsloth_finetune.ipynb`
223
+ - evaluation notebook: `notebooks/colab_qwen3_unsloth_eval_compare.ipynb`
224
 
225
+ The local training artifacts produced by the project include:
226
 
227
+ - SFT adapter: `notebooks/qwen3-medical-lora/`
228
+ - DPO adapter: `notebooks/qwen3-medical-dpo-lora/`
229
+ - SFT checkpoints: `notebooks/sft_output/checkpoint-*`
230
+ - DPO checkpoint: `notebooks/dpo_output/checkpoint-157`
231
 
232
+ ## License
233
 
234
+ No final consolidated license statement has been added yet in the project for the combined derivative artifact. Before public release, verify:
235
 
236
+ - the license of the base model;
237
+ - the license terms of each source dataset;
238
+ - whether redistribution of this adapter is compatible with those upstream terms.
239
 
240
+ ## Contact
241
 
242
+ Project owner / publisher: `Maphe`
243
 
244
+ If you publish this model publicly, it is worth adding:
245
 
246
+ - the source repository URL;
247
+ - exact dataset revisions;
248
+ - a dedicated post-DPO evaluation section;
249
+ - explicit medical safety disclaimers in the serving application.
250
 
 
251
  ### Framework versions
252
 
253
+ - PEFT 0.19.1