Safetensors
English
idefics3
Pull Figure

BMC-SmolVLM1 is a family of lightweight biomedical vision-language models (ranging from 256M to 2.2B parameters) based on SmolVLM. These models are designed for efficient multimodal understanding in the biomedical domain. Please ensure you are using a GPU runtime to run this notebook.

Colab Tutorial: Colab Tutorial

Citation

@inproceedings{lozano2025biomedica,
  title={Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature},
  author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Chen, Liangyu and Nirschl, Jeffrey J and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and Rau, Anita and Katzer, Austin Wolfgang and others},
  booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  pages={19724--19735},
  year={2025},
  organization={IEEE}
}

@article{lozano2025large,
  title={A large-scale vision-language dataset derived from open scientific literature to advance biomedical generalist ai},
  author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Nirschl, Jeffrey J and Polzak, Christopher and Zhang, Yuhui and Chen, Liangyu and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and others},
  journal={arXiv preprint arXiv:2503.22727},
  year={2025}
}
Downloads last month
9
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BIOMEDICA/BMC-smolvlm1-2.2B

Finetuned
(1)
this model

Dataset used to train BIOMEDICA/BMC-smolvlm1-2.2B

Collection including BIOMEDICA/BMC-smolvlm1-2.2B

Paper for BIOMEDICA/BMC-smolvlm1-2.2B