BIOMEDICA/biomedica_webdataset_24M
Updated • 14.3k • 37
BMC-SmolVLM1 is a family of lightweight biomedical vision-language models (ranging from 256M to 2.2B parameters) based on SmolVLM. These models are designed for efficient multimodal understanding in the biomedical domain. Please ensure you are using a GPU runtime to run this notebook.
@inproceedings{lozano2025biomedica,
title={Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature},
author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Chen, Liangyu and Nirschl, Jeffrey J and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and Rau, Anita and Katzer, Austin Wolfgang and others},
booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
pages={19724--19735},
year={2025},
organization={IEEE}
}
@article{lozano2025large,
title={A large-scale vision-language dataset derived from open scientific literature to advance biomedical generalist ai},
author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Nirschl, Jeffrey J and Polzak, Christopher and Zhang, Yuhui and Chen, Liangyu and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and others},
journal={arXiv preprint arXiv:2503.22727},
year={2025}
}
Base model
HuggingFaceTB/SmolVLM-2.2B-Instruct