Image Classification
Transformers
TensorBoard
Safetensors
vit
food
fruits
junkfood
Generated from Trainer
Instructions to use gutkia01/vit-food-classification-gutkia01 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use gutkia01/vit-food-classification-gutkia01 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-classification", model="gutkia01/vit-food-classification-gutkia01") pipe("https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/hub/parrots.png")# Load model directly from transformers import AutoImageProcessor, AutoModelForImageClassification processor = AutoImageProcessor.from_pretrained("gutkia01/vit-food-classification-gutkia01") model = AutoModelForImageClassification.from_pretrained("gutkia01/vit-food-classification-gutkia01", device_map="auto") - Notebooks
- Google Colab
- Kaggle
metadata
library_name: transformers
license: apache-2.0
base_model: google/vit-base-patch16-224
tags:
- image-classification
- food
- fruits
- junkfood
- generated_from_trainer
metrics:
- accuracy
model-index:
- name: vit-food-classification-gutkia01
results: []
vit-food-classification-gutkia01
This model is a fine-tuned version of google/vit-base-patch16-224 on the food-classification dataset. It achieves the following results on the evaluation set:
- Loss: 0.0008
- Model Preparation Time: 0.0034
- Accuracy: 0.9998
π Fruits vs. Junkfood Classifier β Vision Transformer (gutkia01)
This model is a fine-tuned version of google/vit-base-patch16-224, trained on a custom binary dataset to distinguish between healthy fruits and unhealthy fast food.
π§ Model Description
- Architecture: Vision Transformer (ViT)
- Base model:
google/vit-base-patch16-224 - Task: Binary image classification:
Fruitvs.Junkfood - Framework: Hugging Face Transformers Trainer
- Input format: RGB images, 224Γ224, loaded via
imagefolder
β Intended Use & Limitations
Appropriate Use Cases
- Food classification in nutrition, health, or educational applications
- Interactive demos comparing healthy vs. unhealthy food
- Computer vision use cases with simple binary class structures
Limitations
- Only supports binary classification (no subclass differentiation)
- Cannot recognize new or abstract dishes (e.g. salad, sushi)
- Cannot evaluate ingredients, calories, or portion sizes
π Training and Evaluation Data
The model was trained on a binary dataset composed of:
- Fruits360 Dataset: 137,000+ structured fruit images in a controlled studio setup (Kaggle link)
- Fast Food Classification Dataset v2: 20,000 fast food images, 10 categories (e.g., burger, pizza, fries) (Kaggle link)
Dataset Composition
The dataset is a combination of:
Training hyperparameters
The following hyperparameters were used during training:
- learning_rate: 5e-05
- train_batch_size: 16
- eval_batch_size: 8
- seed: 42
- optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
- lr_scheduler_type: linear
- num_epochs: 6
Training results
| Epoch | Training Loss | Validation Loss | Accuracy |
|---|---|---|---|
| 1 | 0.0000 | 0.0215 | 0.9975 |
| 2 | 0.0000 | 0.00004 | 1.0000 |
| 3 | 0.0000 | 0.00008 | 1.0000 |
| 4 | 0.0000 | 0.00011 | 1.0000 |
| 5 | 0.0000 | 0.00011 | 1.0000 |
| 6 | 0.0000 | 0.00008 | 1.0000 |
Final training loss: 0.00047
Evaluation accuracy: 0.9998
Framework versions
- Transformers 4.51.3
- Pytorch 2.6.0+cu124
- Datasets 2.14.4
- Tokenizers 0.21.1