add training details
Browse files
README.md
CHANGED
|
@@ -159,44 +159,240 @@ For a combined response, the following schema can be used:
|
|
| 159 |
|
| 160 |
## Training Details
|
| 161 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 162 |
### Feature Structure
|
| 163 |
|
| 164 |
-
| Column
|
| 165 |
-
|
|
| 166 |
-
| `image`
|
| 167 |
-
| `messages` | List
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 168 |
|
| 169 |
-
The
|
| 170 |
|
| 171 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 172 |
|
| 173 |
-
|
| 174 |
-
| ---------------- | ----------------: | ----------: |
|
| 175 |
-
| Semua Umur | 15,973 | 25.59% |
|
| 176 |
-
| 13+ | 9,422 | 15.09% |
|
| 177 |
-
| 7+ | 8,189 | 13.12% |
|
| 178 |
-
| Konten Terlarang | 7,676 | 12.30% |
|
| 179 |
-
| Unrated | 7,497 | 12.01% |
|
| 180 |
-
| 18+ | 6,851 | 10.97% |
|
| 181 |
-
| 15+ | 6,819 | 10.92% |
|
| 182 |
-
| **Total** | **62,427** | **100.00%** |
|
| 183 |
|
| 184 |
-
###
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 185 |
|
| 186 |
This model was trained using a unified multi-task fine-tuning strategy. Instead of training separate models for text classification, image classification, and keyword generation, all tasks were learned by a single Vision-Language Model.
|
| 187 |
|
| 188 |
The model was fine-tuned end-to-end using task-specific prompts and responses in a shared multimodal instruction format. This allows the model to preserve a unified latent representation across tasks and reduces the risk of performance degradation caused by separate adapter merging.
|
| 189 |
|
| 190 |
-
| Component
|
| 191 |
-
|
|
| 192 |
-
| Base model
|
| 193 |
-
| Model architecture | Unified Vision-Language Model
|
| 194 |
-
| Training approach
|
| 195 |
-
| Tasks
|
| 196 |
-
| Data format
|
| 197 |
-
|
|
| 198 |
-
|
|
| 199 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 200 |
|
| 201 |
## Evaluation
|
| 202 |
|
|
|
|
| 159 |
|
| 160 |
## Training Details
|
| 161 |
|
| 162 |
+
### Training Dataset Composition
|
| 163 |
+
|
| 164 |
+
The model was trained using a unified multi-task supervised fine-tuning strategy. The training data combines task-specific datasets that were converted into a shared instruction-following format, allowing the same VLM architecture to learn text classification, image classification, reasoning generation, and keyword generation in a single training workflow.
|
| 165 |
+
|
| 166 |
+
The training notebook uses a combined multi-task dataset loaded from:
|
| 167 |
+
|
| 168 |
+
- `nuresens/PAD-Combined-Dataset_v6`
|
| 169 |
+
|
| 170 |
+
The dataset is separated into three task sources using the `source` field:
|
| 171 |
+
|
| 172 |
+
| Source | Task | Number of Samples |
|
| 173 |
+
|---|---|---:|
|
| 174 |
+
| `pad1` | Text Classification with Reasoning | 13,406 |
|
| 175 |
+
| `pad2` | Keyword Generator | 15,030 |
|
| 176 |
+
| `pad3` | Image Classification with Reasoning | 62,427 |
|
| 177 |
+
| **Total** | **All Tasks** | **90,863** |
|
| 178 |
+
|
| 179 |
+
The three sources are merged into a single metadata dataset containing the following columns:
|
| 180 |
+
|
| 181 |
+
| Column | Description |
|
| 182 |
+
|---|---|
|
| 183 |
+
| `messages` | Chat-style instruction data containing user prompts and assistant responses. |
|
| 184 |
+
| `source` | Dataset source identifier: `pad1`, `pad2`, or `pad3`. |
|
| 185 |
+
| `source_idx` | Original row index from each source dataset. |
|
| 186 |
+
| `has_image` | Boolean marker indicating whether the sample contains an image placeholder. |
|
| 187 |
+
|
| 188 |
### Feature Structure
|
| 189 |
|
| 190 |
+
| Column | Type | Description |
|
| 191 |
+
|---|---|---|
|
| 192 |
+
| `image` | Image | Visual input used by the vision-language model for image-based classification and reasoning. |
|
| 193 |
+
| `messages` | List | Conversation-style instruction data containing user prompts and assistant responses. |
|
| 194 |
+
|
| 195 |
+
The `messages` field follows a chat-style structure consisting of `role` and `content`. This structure allows the dataset to support multimodal supervised fine-tuning across different task instructions. Text classification and keyword generation samples are text-only instruction samples, while image classification samples include image placeholders and are lazily paired with their original images during collation.
|
| 196 |
+
|
| 197 |
+
## Dataset Statistics
|
| 198 |
+
|
| 199 |
+
### 1. Image Classification and Reasoning Dataset
|
| 200 |
+
|
| 201 |
+
The image classification and reasoning dataset consists of **62,427 samples**. The dataset is used to train the model to analyze visual content and predict the appropriate PAD rating category with an accompanying explanation.
|
| 202 |
+
|
| 203 |
+
| Rating | Number of Samples | Percentage |
|
| 204 |
+
|---|---:|---:|
|
| 205 |
+
| Semua Umur | 15,973 | 25.59% |
|
| 206 |
+
| 13+ | 9,422 | 15.09% |
|
| 207 |
+
| 7+ | 8,189 | 13.12% |
|
| 208 |
+
| Konten Terlarang | 7,676 | 12.30% |
|
| 209 |
+
| Unrated | 7,497 | 12.01% |
|
| 210 |
+
| 18+ | 6,851 | 10.97% |
|
| 211 |
+
| 15+ | 6,819 | 10.92% |
|
| 212 |
+
| **Total** | **62,427** | **100.00%** |
|
| 213 |
+
|
| 214 |
+
The largest category in the image classification dataset is `Semua Umur`, representing 25.59% of the dataset. The remaining categories are distributed between approximately 10.92% and 15.09%, indicating a moderately imbalanced dataset.
|
| 215 |
+
|
| 216 |
+
### 2. Text Classification and Reasoning Dataset
|
| 217 |
|
| 218 |
+
The text classification and reasoning dataset consists of **13,406 rows**. This dataset is used to train the model to classify text-based content into age-rating categories and generate explanations that justify the classification.
|
| 219 |
|
| 220 |
+
| Label Rating Usia | Main Content Characteristics | Number of Rows | Percentage |
|
| 221 |
+
|---|---|---:|---:|
|
| 222 |
+
| Konten Terlarang | Judi online, kekerasan verbal, pelecehan, SARA | 4,497 | 33.54% |
|
| 223 |
+
| Semua Umur | Ramah keluarga, aktivitas keseharian yang positif | 3,252 | 24.26% |
|
| 224 |
+
| 15+ | Diskusi sosial-politik, kriminalitas tanpa glorifikasi | 1,826 | 13.62% |
|
| 225 |
+
| 13+ | Curahan hati remaja, konflik sosial pertemanan | 1,458 | 10.88% |
|
| 226 |
+
| 7+ | Edukasi ringan, hiburan anak, persaingan olahraga | 1,239 | 9.24% |
|
| 227 |
+
| 18+ | Romansa dewasa, gaya hidup malam, edukasi self-harm | 1,134 | 8.46% |
|
| 228 |
+
| **Total** | **Dataset Master SFT** | **13,406** | **100.00%** |
|
| 229 |
|
| 230 |
+
The text classification dataset is more imbalanced than the image classification dataset. The largest category is `Konten Terlarang`, representing 33.54% of the dataset, followed by `Semua Umur` at 24.26%. The smaller categories, especially `18+` and `7+`, should be monitored carefully during evaluation using macro-averaged metrics.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 231 |
|
| 232 |
+
### 3. Keyword Generator Dataset
|
| 233 |
+
|
| 234 |
+
The keyword generator dataset consists of **15,030 records**. This dataset is used to train the model to generate a variable number of relevant keywords from a given content context.
|
| 235 |
+
|
| 236 |
+
| Number of Keywords | Number of Records | Percentage |
|
| 237 |
+
|---:|---:|---:|
|
| 238 |
+
| 1 | 845 | 5.62% |
|
| 239 |
+
| 2 | 2,347 | 15.62% |
|
| 240 |
+
| 3 | 4,700 | 31.27% |
|
| 241 |
+
| 4 | 3,691 | 24.56% |
|
| 242 |
+
| 5 | 1,528 | 10.17% |
|
| 243 |
+
| 6 | 671 | 4.46% |
|
| 244 |
+
| 7 | 327 | 2.18% |
|
| 245 |
+
| 8 | 238 | 1.58% |
|
| 246 |
+
| 9 | 218 | 1.45% |
|
| 247 |
+
| 10 | 275 | 1.83% |
|
| 248 |
+
| 11 | 65 | 0.43% |
|
| 249 |
+
| 12 | 57 | 0.38% |
|
| 250 |
+
| 13 | 32 | 0.21% |
|
| 251 |
+
| 14 | 23 | 0.15% |
|
| 252 |
+
| 15 | 10 | 0.07% |
|
| 253 |
+
| 16 | 1 | 0.01% |
|
| 254 |
+
| **Total** | **15,030** | **100.00%** |
|
| 255 |
+
|
| 256 |
+
The keyword generator dataset is concentrated around outputs containing **3 to 4 keywords**, which together account for **55.83%** of all records. The average number of keywords per record is approximately **3.81**, with a median and mode of **3 keywords**. This indicates that the keyword generation task is primarily optimized for concise keyword outputs rather than long keyword lists.
|
| 257 |
+
|
| 258 |
+
## Dataset Split
|
| 259 |
+
|
| 260 |
+
The dataset is split using a source-aware strategy. Each source is split independently using an **80:10:10** ratio, then recombined into train, validation, and test splits. This ensures that all tasks remain represented proportionally in each split.
|
| 261 |
+
|
| 262 |
+
### Overall Split
|
| 263 |
+
|
| 264 |
+
| Split | Number of Samples | Percentage |
|
| 265 |
+
|---|---:|---:|
|
| 266 |
+
| Train | 72,689 | 80.00% |
|
| 267 |
+
| Validation | 9,087 | 10.00% |
|
| 268 |
+
| Test | 9,087 | 10.00% |
|
| 269 |
+
| **Total** | **90,863** | **100.00%** |
|
| 270 |
+
|
| 271 |
+
### Split by Task Source
|
| 272 |
+
|
| 273 |
+
| Source | Task | Train | Validation | Test | Total |
|
| 274 |
+
|---|---|---:|---:|---:|---:|
|
| 275 |
+
| `pad1` | Text Classification with Reasoning | 10,724 | 1,341 | 1,341 | 13,406 |
|
| 276 |
+
| `pad2` | Keyword Generator | 12,024 | 1,503 | 1,503 | 15,030 |
|
| 277 |
+
| `pad3` | Image Classification with Reasoning | 49,941 | 6,243 | 6,243 | 62,427 |
|
| 278 |
+
| **Total** | **All Tasks** | **72,689** | **9,087** | **9,087** | **90,863** |
|
| 279 |
+
|
| 280 |
+
|
| 281 |
+
## Training Procedure
|
| 282 |
|
| 283 |
This model was trained using a unified multi-task fine-tuning strategy. Instead of training separate models for text classification, image classification, and keyword generation, all tasks were learned by a single Vision-Language Model.
|
| 284 |
|
| 285 |
The model was fine-tuned end-to-end using task-specific prompts and responses in a shared multimodal instruction format. This allows the model to preserve a unified latent representation across tasks and reduces the risk of performance degradation caused by separate adapter merging.
|
| 286 |
|
| 287 |
+
| Component | Value |
|
| 288 |
+
|---|---|
|
| 289 |
+
| Base model | `aitf-komdigi/KomdigiITS-3B-PAD-CPT` |
|
| 290 |
+
| Model architecture | Unified Vision-Language Model |
|
| 291 |
+
| Training approach | Multi-task supervised fine-tuning |
|
| 292 |
+
| Tasks | Text classification, image classification, keyword generation |
|
| 293 |
+
| Data format | TRL-style messages |
|
| 294 |
+
| Split strategy | Source-aware 80:10:10 train/validation/test split |
|
| 295 |
+
| Merge status | Merged model |
|
| 296 |
+
| Precision | BF16 |
|
| 297 |
+
| Frameworks | Transformers, TRL, PEFT, Unsloth |
|
| 298 |
+
|
| 299 |
+
## Training Configuration
|
| 300 |
+
|
| 301 |
+
The model was fine-tuned using **Unsloth FastVisionModel** and **TRL SFTTrainer**. Training was performed as multi-task supervised fine-tuning over the merged dataset containing text classification, image classification, and keyword generation tasks.
|
| 302 |
+
|
| 303 |
+
### Base Model and Sequence Configuration
|
| 304 |
+
|
| 305 |
+
| Component | Value |
|
| 306 |
+
|---|---|
|
| 307 |
+
| Base model | `aitf-komdigi/KomdigiITS-3B-PAD-CPT` |
|
| 308 |
+
| Model class | `FastVisionModel` |
|
| 309 |
+
| Maximum sequence length | 4096 |
|
| 310 |
+
| Load in 4-bit | `False` |
|
| 311 |
+
| Precision | BF16 if supported, otherwise FP16 |
|
| 312 |
+
| Gradient checkpointing | `unsloth` |
|
| 313 |
+
| Chat template reference | `mistralai/Ministral-3-3B-Instruct-2512` |
|
| 314 |
+
| Padding side | Right padding |
|
| 315 |
+
|
| 316 |
+
The notebook applies the chat template from `mistralai/Ministral-3-3B-Instruct-2512` to the model processor and tokenizer so that the training format follows the instruction-style format expected by the model.
|
| 317 |
+
|
| 318 |
+
### LoRA Configuration
|
| 319 |
+
|
| 320 |
+
The model was fine-tuned using LoRA. Both the vision and language components were enabled for fine-tuning.
|
| 321 |
+
|
| 322 |
+
| Parameter | Value |
|
| 323 |
+
|---|---:|
|
| 324 |
+
| LoRA rank `r` | 16 |
|
| 325 |
+
| LoRA alpha | 16 |
|
| 326 |
+
| LoRA dropout | 0.0 |
|
| 327 |
+
| Bias | `none` |
|
| 328 |
+
| Target modules | `all-linear` |
|
| 329 |
+
| Random state | 42 |
|
| 330 |
+
| Fine-tune vision layers | `True` |
|
| 331 |
+
| Fine-tune language layers | `True` |
|
| 332 |
+
| Fine-tune attention modules | `True` |
|
| 333 |
+
| Fine-tune MLP modules | `True` |
|
| 334 |
+
|
| 335 |
+
The training run reported **33,751,040 trainable parameters** out of **3,882,841,088 total parameters**, meaning approximately **0.87%** of the model parameters were trained.
|
| 336 |
+
|
| 337 |
+
### Data Collation
|
| 338 |
+
|
| 339 |
+
The notebook uses `UnslothVisionDataCollator` with a custom lazy image collation strategy. Images are loaded only when a sample has an image marker, which reduces unnecessary image decoding for text-only samples.
|
| 340 |
+
|
| 341 |
+
| Component | Value |
|
| 342 |
+
|---|---|
|
| 343 |
+
| Base collator | `UnslothVisionDataCollator` |
|
| 344 |
+
| Custom wrapper | Lazy vision collator |
|
| 345 |
+
| Train on responses only | `True` |
|
| 346 |
+
| Instruction part | `[INST]` |
|
| 347 |
+
| Response part | `[/INST]` |
|
| 348 |
+
| Image loading strategy | Lazy image injection for `pad3` samples only |
|
| 349 |
+
|
| 350 |
+
The training loss is computed only on the assistant response, which is useful for instruction tuning because the model is optimized to generate the expected output rather than reproduce the full prompt.
|
| 351 |
+
|
| 352 |
+
### SFT Training Arguments
|
| 353 |
+
|
| 354 |
+
| Hyperparameter | Value |
|
| 355 |
+
|---|---:|
|
| 356 |
+
| Per-device train batch size | 4 |
|
| 357 |
+
| Gradient accumulation steps | 2 |
|
| 358 |
+
| Effective train batch size | 8 |
|
| 359 |
+
| Per-device evaluation batch size | 4 |
|
| 360 |
+
| Number of epochs | 2 |
|
| 361 |
+
| Learning rate | 2e-4 |
|
| 362 |
+
| Warmup ratio | 0.03 |
|
| 363 |
+
| Learning rate scheduler | Cosine |
|
| 364 |
+
| Weight decay | 0.01 |
|
| 365 |
+
| Optimizer | `adamw_8bit` |
|
| 366 |
+
| Max gradient norm | 1.0 |
|
| 367 |
+
| Logging steps | 10 |
|
| 368 |
+
| Evaluation strategy | Steps |
|
| 369 |
+
| Evaluation steps | 1000 |
|
| 370 |
+
| Save strategy | Steps |
|
| 371 |
+
| Save steps | 1000 |
|
| 372 |
+
| Save total limit | 2 |
|
| 373 |
+
| Seed | 42 |
|
| 374 |
+
| Output directory | `outputs_pad_sft` |
|
| 375 |
+
| Report to | Weights & Biases |
|
| 376 |
+
| Dataset text field | Empty string |
|
| 377 |
+
| Skip dataset preparation | `True` |
|
| 378 |
+
| Dataset number of processes | 4 |
|
| 379 |
+
| Dataloader workers | 2 |
|
| 380 |
+
| Pin memory | `True` |
|
| 381 |
+
|
| 382 |
+
### Training Run Summary
|
| 383 |
+
|
| 384 |
+
| Attribute | Value |
|
| 385 |
+
|---|---:|
|
| 386 |
+
| Number of training examples | 72,689 |
|
| 387 |
+
| Number of epochs | 2 |
|
| 388 |
+
| Total optimization steps | 18,174 |
|
| 389 |
+
| Number of GPUs | 1 |
|
| 390 |
+
| GPU used | NVIDIA A100-SXM4-40GB |
|
| 391 |
+
| Total batch size | 8 |
|
| 392 |
+
| Trainable parameters | 33,751,040 |
|
| 393 |
+
| Total parameters | 3,882,841,088 |
|
| 394 |
+
| Percentage of trained parameters | 0.87% |
|
| 395 |
+
|
| 396 |
|
| 397 |
## Evaluation
|
| 398 |
|