Visual Question Answering
English
biology
medical
Pull Figure

Paper: Arxiv     |     Website: Biomedica     |     Training instructions: OpenCLIP     |     Tutorial: Google Colab

Model Name: BMC-LongCLIP+

Abstract

Embedding vision–language models (VLMs) are typically pretrained with short text windows (<77 tokens), which forces the truncation of long-format captions. Yet, the distribution of biomedical captions from large-scale open source literature reveals that a huge portion of captions far exceed 77 tokens. To this end, we investigate the impact of pretraining on longformat biomedical captions by extending the context length of text encoders in VLMs. We find that longer context (thus, enabling additional supervision provided in long-format captions) correlates with better retrieval and classification performance. Given this finding, we introduce BIOMEDICA-LongCAP, a dataset of 1M image–caption pairs enriched with contextaware descriptions from full-text articles, providing longer and additional textual supervision. Using BIOMEDICA-LongCAP, we train BMC-LongCLIP, a long-context biomedical VLM with a text encoder supporting windows of up to 512 tokens. Our model extends context capacity by 6.6×, reducing token waste from 55% to just 2.2%. On longcaption retrieval benchmarks, BMC-LongCLIP achieves up to +30% absolute gains in Recall@1 and +2% average improvements in classification, while also converging faster than shortcontext.

Baseline Comparison Against Frontier Models

Benchmark Model Context Batch T2I R@1 T2I R@5 T2I R@10 I2T R@1 I2T R@5 I2T R@10
CXR PMC-CLIP 77 128 0.0 0.5 0.7 0.2 1.0 1.6
CXR BiomedCLIP 256 4K 0.5 2.6 5.7 0.6 3.3 5.5
CXR BMC-CLIP 77 8K 0.1 1.1 2.9 0.3 1.9 3.4
CXR BMC-LongCLIP+ 512 16K 1.9 7.1 12.2 3.0 9.5 14.5
PMC PMC-CLIP 77 128 0.2 0.7 1.2 0.1 0.7 1.2
PMC MedSigLIP 77 N/A 20.1 37.0 46.0 30.9 49.0 60.1
PMC BiomedCLIP 256 4K 68.8 86.2 91.1 73.3 89.3 93.7
PMC BMC-CLIP 77 8K 49.0 67.6 74.0 40.8 60.4 68.4
PMC BMC-LongCLIP+ 512 16K 80.8 91.2 94.4 79.7 90.6 93.8

Acknowledgments

This work is supported by an NVIDIA Academic Grant.


Citation

@article{sun2025no,
title={No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models},
author={Sun, Min Woo and others},
journal={arXiv preprint arXiv:2510.03978},
year={2025}

@inproceedings{lozano2025biomedica,
  title={Biomedica: An open biomedical image-caption archive, dataset, and vision-language models derived from scientific literature},
  author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Chen, Liangyu and Nirschl, Jeffrey J and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and Rau, Anita and Katzer, Austin Wolfgang and others},
  booktitle={2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  pages={19724--19735},
  year={2025},
  organization={IEEE}
}

@article{lozano2025large,
  title={A large-scale vision-language dataset derived from open scientific literature to advance biomedical generalist ai},
  author={Lozano, Alejandro and Sun, Min Woo and Burgess, James and Nirschl, Jeffrey J and Polzak, Christopher and Zhang, Yuhui and Chen, Liangyu and Gu, Jeffrey and Lopez, Ivan and Aklilu, Josiah and others},
  journal={arXiv preprint arXiv:2503.22727},
  year={2025}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train BIOMEDICA/BMC-LongCLIP

Space using BIOMEDICA/BMC-LongCLIP 1

Collection including BIOMEDICA/BMC-LongCLIP

Papers for BIOMEDICA/BMC-LongCLIP