MedGemma-4B Kvasir Fine-Tune

This model is a fine-tuned version of google/medgemma-4b-it trained specifically to classify gastrointestinal (GI) tract findings and anatomical landmarks from endoscopic imagery.

Developed by Umakant Biswal, this model utilizes Parameter-Efficient Fine-Tuning (LoRA) to achieve high accuracy in medical image analysis.

Model Performance

The model was evaluated on a 100-image holdout validation subset of the Kvasir dataset, demonstrating a massive +80.00% boost in accuracy compared to the base model on this specific task:

  • Accuracy: 87.00% (Up from 7.00% Base)
  • F1 Score: 0.8718 (Up from 0.1001 Base)

Supported Classes

The model is trained to return one of the following 8 standardized classes:

  • A: dyed-lifted-polyps
  • B: dyed-resection-margins
  • C: esophagitis
  • D: normal-cecum
  • E: normal-pylorus
  • F: normal-z-line
  • G: polyps
  • H: ulcerative-colitis

Quick Start

The model requires the images to be passed directly inside the chat template's content array. Here is how to load and run inference:

import torch
from transformers import pipeline
from PIL import Image
import requests

# 1. Load the Model
pipe = pipeline(
    "image-text-to-text", 
    model="umakantcurateai/medgemma-kvasir-finetune-umak", 
    torch_dtype=torch.bfloat16, 
    device_map="auto"
)
pipe.model.generation_config.do_sample = False
pipe.processor.tokenizer.padding_side = "left"

# 2. Load an Endoscopy Image
url = "[https://example.com/path_to_endoscopy_image.jpg](https://example.com/path_to_endoscopy_image.jpg)" # Replace with your image
image = Image.open(requests.get(url, stream=True).raw).convert("RGB")

# 3. Define the Prompt & Multiple Choice Options
PROMPT = """What is the most likely endoscopic finding or anatomical landmark shown in this image?
A: dyed-lifted-polyps
B: dyed-resection-margins
C: esophagitis
D: normal-cecum
E: normal-pylorus
F: normal-z-line
G: polyps
H: ulcerative-colitis"""

messages = [
    {
        "role": "user", 
        "content": [
            {"type": "image", "image": image}, 
            {"type": "text", "text": PROMPT}
        ]
    }
]

# 4. Generate Classification
output = pipe(messages, max_new_tokens=20, return_full_text=False)
print(f"Predicted Class: {output[0]['generated_text']}")
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for umakantcurateai/medgemma-kvasir-finetune-umak

Adapter
(122)
this model

Collection including umakantcurateai/medgemma-kvasir-finetune-umak