MemOCR-7B / README.md
yrshi's picture
Upload folder using huggingface_hub
e5acde6 verified
|
Raw History Blame Contribute Delete
5.02 kB
---
license: apache-2.0
language:
- en
base_model: Qwen/Qwen2.5-VL-7B-Instruct
tags:
- vision
- multimodal
- visual-question-answering
- memory-agent
- long-context
pipeline_tag: visual-question-answering
---
# MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning
<div align="center">
<img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_logo.png" alt="MemOCR Logo" width="720" style="margin-bottom: 18px;" />
<p align="center">
<a href="https://arxiv.org/abs/2601.21468"><img src="https://img.shields.io/badge/arXiv-2601.21468-b31b1b.svg" alt="arXiv"></a>
<a href="https://huggingface.co/papers/2601.21468"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging_Face-Paper-yellow" alt="HuggingFace Paper"></a>
<a href="https://meituan.github.io/MemOCR/"><img src="https://img.shields.io/badge/Project-Page-green" alt="Project Page"></a>
<a href="https://github.com/meituan/MemOCR"><img src="https://img.shields.io/badge/GitHub-Code-black?logo=github" alt="GitHub"></a>
<a href="./LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License"></a>
</p>
</div>
## 🌟 Model Description
MemOCR is a visual memory agent that dynamically adapts information density during memory drafting and reading, and optimizes visual layouts to highlight key information.
This checkpoint is fine-tuned from `Qwen2.5-VL-7B-Instruct` with budget-aware training objectives.
### Key Capabilities
- **Adaptive Information Density**: Dynamically adjusts memory content richness based on task requirements
- **Budget-Aware Memory**: Optimizes memory usage with explicit token budget constraints
- **Dual-Domain Architecture**: Separate memory drafting (text domain) and reading (vision domain) processes
- **Multi-Hop Reasoning**: Superior performance on complex question answering tasks
## πŸ—οΈ Architecture
<div align="center">
<img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_method_architecture.png" width="95%" alt="MemOCR Architecture">
</div>
MemOCR consists of two main components:
1. **Memory Drafting in Text Domain**: An LLM agent iteratively refines rich-text memory content based on question-answering feedback
2. **Memory Reading in Vision Domain**: A vision-language model processes rendered visual memory with optimized information density
The framework employs budget-aware training objectives to balance memory informativeness and token efficiency.
<div align="center">
<img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_method_augmentation.png" width="50%" alt="Data Augmentation Strategy">
</div>
## πŸ“Š Performance
### Main Results
<div align="center">
<img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_main_results.png" width="90%" alt="Main Results">
</div>
MemOCR achieves state-of-the-art performance across multiple multi-hop QA benchmarks:
- **HotpotQA**: Superior accuracy with efficient memory budgets
- **2WikiMultihopQA**: Strong multi-hop reasoning capabilities
- **NaturalQuestions & TriviaQA**: Excellent knowledge retrieval performance
### Ablation Studies
<div align="center">
<img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_ablation.png" width="75%" alt="Ablation Studies">
</div>
### Analysis: Information Density & Budget
<div align="center">
<img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_analysis_budget.png" width="45%" style="display:inline-block">
<img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_analysis_density.png" width="45%" style="display:inline-block">
</div>
## πŸš€ Usage
This model is designed to work with the MemOCR framework. Please refer to the [official repository](https://github.com/meituan/MemOCR) for detailed usage instructions.
## πŸ“š Citation
If you find MemOCR useful in your research, please consider citing:
```bibtex
@article{shi2026memocr,
title={MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning},
author={Yaorui Shi and Shugui Liu and Yu Yang and Wenyu Mao and Yuxin Chen and Qi GU and Hui Su and Xunliang Cai and Xiang Wang and An Zhang},
journal={arXiv preprint arXiv:2601.21468},
year={2026},
}
```
## πŸ™ Acknowledgements
This model is built upon:
- **[Qwen2.5-VL](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct)** as the vision-language model backbone
- **[veRL](https://github.com/volcengine/verl)** as the reinforcement learning training framework
- **[MemAgent](https://github.com/BytedTsinghua-SIA/MemAgent)** for the recurrent module and training dataset
## πŸ“„ License
This model is licensed under the Apache License 2.0. See the LICENSE file for details.
### Dataset License
Training and evaluation datasets are derived from:
- **HotpotQA, 2WikiMultihopQA, Natural Questions, TriviaQA**: Wikipedia-derived content licensed under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/)