ryfkn commited on
Commit
bb636ea
·
verified ·
1 Parent(s): 5b54d71

add training details

Browse files
Files changed (1) hide show
  1. README.md +223 -27
README.md CHANGED
@@ -159,44 +159,240 @@ For a combined response, the following schema can be used:
159
 
160
  ## Training Details
161
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
162
  ### Feature Structure
163
 
164
- | Column | Type | Description |
165
- | ---------- | ----- | ------------------------------------------------------------------------------------ |
166
- | `image` | Image | Visual input used by the vision-language model. |
167
- | `messages` | List | Conversation-style instruction data containing user prompts and assistant responses. |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
168
 
169
- The `messages` field follows a chat-style structure consisting of `role` and `content`. This structure allows the dataset to support multimodal supervised fine-tuning across different task instructions.
170
 
171
- ### Rating Distribution
 
 
 
 
 
 
 
 
172
 
173
- | Rating | Number of Samples | Percentage |
174
- | ---------------- | ----------------: | ----------: |
175
- | Semua Umur | 15,973 | 25.59% |
176
- | 13+ | 9,422 | 15.09% |
177
- | 7+ | 8,189 | 13.12% |
178
- | Konten Terlarang | 7,676 | 12.30% |
179
- | Unrated | 7,497 | 12.01% |
180
- | 18+ | 6,851 | 10.97% |
181
- | 15+ | 6,819 | 10.92% |
182
- | **Total** | **62,427** | **100.00%** |
183
 
184
- ### Training Procedure
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
185
 
186
  This model was trained using a unified multi-task fine-tuning strategy. Instead of training separate models for text classification, image classification, and keyword generation, all tasks were learned by a single Vision-Language Model.
187
 
188
  The model was fine-tuned end-to-end using task-specific prompts and responses in a shared multimodal instruction format. This allows the model to preserve a unified latent representation across tasks and reduces the risk of performance degradation caused by separate adapter merging.
189
 
190
- | Component | Value |
191
- | ------------------ | ------------------------------------------------------------- |
192
- | Base model | `aitf-komdigi/KomdigiITS-3B-PAD-CPT` |
193
- | Model architecture | Unified Vision-Language Model |
194
- | Training approach | Multi-task supervised fine-tuning |
195
- | Tasks | Text classification, image classification, keyword generation |
196
- | Data format | TRL-style messages |
197
- | Merge status | Merged model |
198
- | Precision | BF16 |
199
- | Frameworks | Transformers, TRL, PEFT, Unsloth |
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
200
 
201
  ## Evaluation
202
 
 
159
 
160
  ## Training Details
161
 
162
+ ### Training Dataset Composition
163
+
164
+ The model was trained using a unified multi-task supervised fine-tuning strategy. The training data combines task-specific datasets that were converted into a shared instruction-following format, allowing the same VLM architecture to learn text classification, image classification, reasoning generation, and keyword generation in a single training workflow.
165
+
166
+ The training notebook uses a combined multi-task dataset loaded from:
167
+
168
+ - `nuresens/PAD-Combined-Dataset_v6`
169
+
170
+ The dataset is separated into three task sources using the `source` field:
171
+
172
+ | Source | Task | Number of Samples |
173
+ |---|---|---:|
174
+ | `pad1` | Text Classification with Reasoning | 13,406 |
175
+ | `pad2` | Keyword Generator | 15,030 |
176
+ | `pad3` | Image Classification with Reasoning | 62,427 |
177
+ | **Total** | **All Tasks** | **90,863** |
178
+
179
+ The three sources are merged into a single metadata dataset containing the following columns:
180
+
181
+ | Column | Description |
182
+ |---|---|
183
+ | `messages` | Chat-style instruction data containing user prompts and assistant responses. |
184
+ | `source` | Dataset source identifier: `pad1`, `pad2`, or `pad3`. |
185
+ | `source_idx` | Original row index from each source dataset. |
186
+ | `has_image` | Boolean marker indicating whether the sample contains an image placeholder. |
187
+
188
  ### Feature Structure
189
 
190
+ | Column | Type | Description |
191
+ |---|---|---|
192
+ | `image` | Image | Visual input used by the vision-language model for image-based classification and reasoning. |
193
+ | `messages` | List | Conversation-style instruction data containing user prompts and assistant responses. |
194
+
195
+ The `messages` field follows a chat-style structure consisting of `role` and `content`. This structure allows the dataset to support multimodal supervised fine-tuning across different task instructions. Text classification and keyword generation samples are text-only instruction samples, while image classification samples include image placeholders and are lazily paired with their original images during collation.
196
+
197
+ ## Dataset Statistics
198
+
199
+ ### 1. Image Classification and Reasoning Dataset
200
+
201
+ The image classification and reasoning dataset consists of **62,427 samples**. The dataset is used to train the model to analyze visual content and predict the appropriate PAD rating category with an accompanying explanation.
202
+
203
+ | Rating | Number of Samples | Percentage |
204
+ |---|---:|---:|
205
+ | Semua Umur | 15,973 | 25.59% |
206
+ | 13+ | 9,422 | 15.09% |
207
+ | 7+ | 8,189 | 13.12% |
208
+ | Konten Terlarang | 7,676 | 12.30% |
209
+ | Unrated | 7,497 | 12.01% |
210
+ | 18+ | 6,851 | 10.97% |
211
+ | 15+ | 6,819 | 10.92% |
212
+ | **Total** | **62,427** | **100.00%** |
213
+
214
+ The largest category in the image classification dataset is `Semua Umur`, representing 25.59% of the dataset. The remaining categories are distributed between approximately 10.92% and 15.09%, indicating a moderately imbalanced dataset.
215
+
216
+ ### 2. Text Classification and Reasoning Dataset
217
 
218
+ The text classification and reasoning dataset consists of **13,406 rows**. This dataset is used to train the model to classify text-based content into age-rating categories and generate explanations that justify the classification.
219
 
220
+ | Label Rating Usia | Main Content Characteristics | Number of Rows | Percentage |
221
+ |---|---|---:|---:|
222
+ | Konten Terlarang | Judi online, kekerasan verbal, pelecehan, SARA | 4,497 | 33.54% |
223
+ | Semua Umur | Ramah keluarga, aktivitas keseharian yang positif | 3,252 | 24.26% |
224
+ | 15+ | Diskusi sosial-politik, kriminalitas tanpa glorifikasi | 1,826 | 13.62% |
225
+ | 13+ | Curahan hati remaja, konflik sosial pertemanan | 1,458 | 10.88% |
226
+ | 7+ | Edukasi ringan, hiburan anak, persaingan olahraga | 1,239 | 9.24% |
227
+ | 18+ | Romansa dewasa, gaya hidup malam, edukasi self-harm | 1,134 | 8.46% |
228
+ | **Total** | **Dataset Master SFT** | **13,406** | **100.00%** |
229
 
230
+ The text classification dataset is more imbalanced than the image classification dataset. The largest category is `Konten Terlarang`, representing 33.54% of the dataset, followed by `Semua Umur` at 24.26%. The smaller categories, especially `18+` and `7+`, should be monitored carefully during evaluation using macro-averaged metrics.
 
 
 
 
 
 
 
 
 
231
 
232
+ ### 3. Keyword Generator Dataset
233
+
234
+ The keyword generator dataset consists of **15,030 records**. This dataset is used to train the model to generate a variable number of relevant keywords from a given content context.
235
+
236
+ | Number of Keywords | Number of Records | Percentage |
237
+ |---:|---:|---:|
238
+ | 1 | 845 | 5.62% |
239
+ | 2 | 2,347 | 15.62% |
240
+ | 3 | 4,700 | 31.27% |
241
+ | 4 | 3,691 | 24.56% |
242
+ | 5 | 1,528 | 10.17% |
243
+ | 6 | 671 | 4.46% |
244
+ | 7 | 327 | 2.18% |
245
+ | 8 | 238 | 1.58% |
246
+ | 9 | 218 | 1.45% |
247
+ | 10 | 275 | 1.83% |
248
+ | 11 | 65 | 0.43% |
249
+ | 12 | 57 | 0.38% |
250
+ | 13 | 32 | 0.21% |
251
+ | 14 | 23 | 0.15% |
252
+ | 15 | 10 | 0.07% |
253
+ | 16 | 1 | 0.01% |
254
+ | **Total** | **15,030** | **100.00%** |
255
+
256
+ The keyword generator dataset is concentrated around outputs containing **3 to 4 keywords**, which together account for **55.83%** of all records. The average number of keywords per record is approximately **3.81**, with a median and mode of **3 keywords**. This indicates that the keyword generation task is primarily optimized for concise keyword outputs rather than long keyword lists.
257
+
258
+ ## Dataset Split
259
+
260
+ The dataset is split using a source-aware strategy. Each source is split independently using an **80:10:10** ratio, then recombined into train, validation, and test splits. This ensures that all tasks remain represented proportionally in each split.
261
+
262
+ ### Overall Split
263
+
264
+ | Split | Number of Samples | Percentage |
265
+ |---|---:|---:|
266
+ | Train | 72,689 | 80.00% |
267
+ | Validation | 9,087 | 10.00% |
268
+ | Test | 9,087 | 10.00% |
269
+ | **Total** | **90,863** | **100.00%** |
270
+
271
+ ### Split by Task Source
272
+
273
+ | Source | Task | Train | Validation | Test | Total |
274
+ |---|---|---:|---:|---:|---:|
275
+ | `pad1` | Text Classification with Reasoning | 10,724 | 1,341 | 1,341 | 13,406 |
276
+ | `pad2` | Keyword Generator | 12,024 | 1,503 | 1,503 | 15,030 |
277
+ | `pad3` | Image Classification with Reasoning | 49,941 | 6,243 | 6,243 | 62,427 |
278
+ | **Total** | **All Tasks** | **72,689** | **9,087** | **9,087** | **90,863** |
279
+
280
+
281
+ ## Training Procedure
282
 
283
  This model was trained using a unified multi-task fine-tuning strategy. Instead of training separate models for text classification, image classification, and keyword generation, all tasks were learned by a single Vision-Language Model.
284
 
285
  The model was fine-tuned end-to-end using task-specific prompts and responses in a shared multimodal instruction format. This allows the model to preserve a unified latent representation across tasks and reduces the risk of performance degradation caused by separate adapter merging.
286
 
287
+ | Component | Value |
288
+ |---|---|
289
+ | Base model | `aitf-komdigi/KomdigiITS-3B-PAD-CPT` |
290
+ | Model architecture | Unified Vision-Language Model |
291
+ | Training approach | Multi-task supervised fine-tuning |
292
+ | Tasks | Text classification, image classification, keyword generation |
293
+ | Data format | TRL-style messages |
294
+ | Split strategy | Source-aware 80:10:10 train/validation/test split |
295
+ | Merge status | Merged model |
296
+ | Precision | BF16 |
297
+ | Frameworks | Transformers, TRL, PEFT, Unsloth |
298
+
299
+ ## Training Configuration
300
+
301
+ The model was fine-tuned using **Unsloth FastVisionModel** and **TRL SFTTrainer**. Training was performed as multi-task supervised fine-tuning over the merged dataset containing text classification, image classification, and keyword generation tasks.
302
+
303
+ ### Base Model and Sequence Configuration
304
+
305
+ | Component | Value |
306
+ |---|---|
307
+ | Base model | `aitf-komdigi/KomdigiITS-3B-PAD-CPT` |
308
+ | Model class | `FastVisionModel` |
309
+ | Maximum sequence length | 4096 |
310
+ | Load in 4-bit | `False` |
311
+ | Precision | BF16 if supported, otherwise FP16 |
312
+ | Gradient checkpointing | `unsloth` |
313
+ | Chat template reference | `mistralai/Ministral-3-3B-Instruct-2512` |
314
+ | Padding side | Right padding |
315
+
316
+ The notebook applies the chat template from `mistralai/Ministral-3-3B-Instruct-2512` to the model processor and tokenizer so that the training format follows the instruction-style format expected by the model.
317
+
318
+ ### LoRA Configuration
319
+
320
+ The model was fine-tuned using LoRA. Both the vision and language components were enabled for fine-tuning.
321
+
322
+ | Parameter | Value |
323
+ |---|---:|
324
+ | LoRA rank `r` | 16 |
325
+ | LoRA alpha | 16 |
326
+ | LoRA dropout | 0.0 |
327
+ | Bias | `none` |
328
+ | Target modules | `all-linear` |
329
+ | Random state | 42 |
330
+ | Fine-tune vision layers | `True` |
331
+ | Fine-tune language layers | `True` |
332
+ | Fine-tune attention modules | `True` |
333
+ | Fine-tune MLP modules | `True` |
334
+
335
+ The training run reported **33,751,040 trainable parameters** out of **3,882,841,088 total parameters**, meaning approximately **0.87%** of the model parameters were trained.
336
+
337
+ ### Data Collation
338
+
339
+ The notebook uses `UnslothVisionDataCollator` with a custom lazy image collation strategy. Images are loaded only when a sample has an image marker, which reduces unnecessary image decoding for text-only samples.
340
+
341
+ | Component | Value |
342
+ |---|---|
343
+ | Base collator | `UnslothVisionDataCollator` |
344
+ | Custom wrapper | Lazy vision collator |
345
+ | Train on responses only | `True` |
346
+ | Instruction part | `[INST]` |
347
+ | Response part | `[/INST]` |
348
+ | Image loading strategy | Lazy image injection for `pad3` samples only |
349
+
350
+ The training loss is computed only on the assistant response, which is useful for instruction tuning because the model is optimized to generate the expected output rather than reproduce the full prompt.
351
+
352
+ ### SFT Training Arguments
353
+
354
+ | Hyperparameter | Value |
355
+ |---|---:|
356
+ | Per-device train batch size | 4 |
357
+ | Gradient accumulation steps | 2 |
358
+ | Effective train batch size | 8 |
359
+ | Per-device evaluation batch size | 4 |
360
+ | Number of epochs | 2 |
361
+ | Learning rate | 2e-4 |
362
+ | Warmup ratio | 0.03 |
363
+ | Learning rate scheduler | Cosine |
364
+ | Weight decay | 0.01 |
365
+ | Optimizer | `adamw_8bit` |
366
+ | Max gradient norm | 1.0 |
367
+ | Logging steps | 10 |
368
+ | Evaluation strategy | Steps |
369
+ | Evaluation steps | 1000 |
370
+ | Save strategy | Steps |
371
+ | Save steps | 1000 |
372
+ | Save total limit | 2 |
373
+ | Seed | 42 |
374
+ | Output directory | `outputs_pad_sft` |
375
+ | Report to | Weights & Biases |
376
+ | Dataset text field | Empty string |
377
+ | Skip dataset preparation | `True` |
378
+ | Dataset number of processes | 4 |
379
+ | Dataloader workers | 2 |
380
+ | Pin memory | `True` |
381
+
382
+ ### Training Run Summary
383
+
384
+ | Attribute | Value |
385
+ |---|---:|
386
+ | Number of training examples | 72,689 |
387
+ | Number of epochs | 2 |
388
+ | Total optimization steps | 18,174 |
389
+ | Number of GPUs | 1 |
390
+ | GPU used | NVIDIA A100-SXM4-40GB |
391
+ | Total batch size | 8 |
392
+ | Trainable parameters | 33,751,040 |
393
+ | Total parameters | 3,882,841,088 |
394
+ | Percentage of trained parameters | 0.87% |
395
+
396
 
397
  ## Evaluation
398