|
Download README.md from meituan/MemOCR-7B: direct link, hf CLI and curl.
- Browser
- Download file 5.02 kB
-
https://huggingface.co/meituan/MemOCR-7B/resolve/main/README.md
- Command line
-
hf download hf://meituan/MemOCR-7B/README.md
-
curl -L -o README.md https://huggingface.co/meituan/MemOCR-7B/resolve/main/README.md
5.02 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| base_model: Qwen/Qwen2.5-VL-7B-Instruct | |
| tags: | |
| - vision | |
| - multimodal | |
| - visual-question-answering | |
| - memory-agent | |
| - long-context | |
| pipeline_tag: visual-question-answering | |
| # MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning | |
| <div align="center"> | |
| <img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_logo.png" alt="MemOCR Logo" width="720" style="margin-bottom: 18px;" /> | |
| <p align="center"> | |
| <a href="https://arxiv.org/abs/2601.21468"><img src="https://img.shields.io/badge/arXiv-2601.21468-b31b1b.svg" alt="arXiv"></a> | |
| <a href="https://huggingface.co/papers/2601.21468"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging_Face-Paper-yellow" alt="HuggingFace Paper"></a> | |
| <a href="https://meituan.github.io/MemOCR/"><img src="https://img.shields.io/badge/Project-Page-green" alt="Project Page"></a> | |
| <a href="https://github.com/meituan/MemOCR"><img src="https://img.shields.io/badge/GitHub-Code-black?logo=github" alt="GitHub"></a> | |
| <a href="./LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License"></a> | |
| </p> | |
| </div> | |
| ## π Model Description | |
| MemOCR is a visual memory agent that dynamically adapts information density during memory drafting and reading, and optimizes visual layouts to highlight key information. | |
| This checkpoint is fine-tuned from `Qwen2.5-VL-7B-Instruct` with budget-aware training objectives. | |
| ### Key Capabilities | |
| - **Adaptive Information Density**: Dynamically adjusts memory content richness based on task requirements | |
| - **Budget-Aware Memory**: Optimizes memory usage with explicit token budget constraints | |
| - **Dual-Domain Architecture**: Separate memory drafting (text domain) and reading (vision domain) processes | |
| - **Multi-Hop Reasoning**: Superior performance on complex question answering tasks | |
| ## ποΈ Architecture | |
| <div align="center"> | |
| <img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_method_architecture.png" width="95%" alt="MemOCR Architecture"> | |
| </div> | |
| MemOCR consists of two main components: | |
| 1. **Memory Drafting in Text Domain**: An LLM agent iteratively refines rich-text memory content based on question-answering feedback | |
| 2. **Memory Reading in Vision Domain**: A vision-language model processes rendered visual memory with optimized information density | |
| The framework employs budget-aware training objectives to balance memory informativeness and token efficiency. | |
| <div align="center"> | |
| <img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_method_augmentation.png" width="50%" alt="Data Augmentation Strategy"> | |
| </div> | |
| ## π Performance | |
| ### Main Results | |
| <div align="center"> | |
| <img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_main_results.png" width="90%" alt="Main Results"> | |
| </div> | |
| MemOCR achieves state-of-the-art performance across multiple multi-hop QA benchmarks: | |
| - **HotpotQA**: Superior accuracy with efficient memory budgets | |
| - **2WikiMultihopQA**: Strong multi-hop reasoning capabilities | |
| - **NaturalQuestions & TriviaQA**: Excellent knowledge retrieval performance | |
| ### Ablation Studies | |
| <div align="center"> | |
| <img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_ablation.png" width="75%" alt="Ablation Studies"> | |
| </div> | |
| ### Analysis: Information Density & Budget | |
| <div align="center"> | |
| <img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_analysis_budget.png" width="45%" style="display:inline-block"> | |
| <img src="https://raw.githubusercontent.com/meituan/MemOCR/main/docs/figs/memocr_analysis_density.png" width="45%" style="display:inline-block"> | |
| </div> | |
| ## π Usage | |
| This model is designed to work with the MemOCR framework. Please refer to the [official repository](https://github.com/meituan/MemOCR) for detailed usage instructions. | |
| ## π Citation | |
| If you find MemOCR useful in your research, please consider citing: | |
| ```bibtex | |
| @article{shi2026memocr, | |
| title={MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon Reasoning}, | |
| author={Yaorui Shi and Shugui Liu and Yu Yang and Wenyu Mao and Yuxin Chen and Qi GU and Hui Su and Xunliang Cai and Xiang Wang and An Zhang}, | |
| journal={arXiv preprint arXiv:2601.21468}, | |
| year={2026}, | |
| } | |
| ``` | |
| ## π Acknowledgements | |
| This model is built upon: | |
| - **[Qwen2.5-VL](https://huggingface.co/Qwen/Qwen2.5-VL-7B-Instruct)** as the vision-language model backbone | |
| - **[veRL](https://github.com/volcengine/verl)** as the reinforcement learning training framework | |
| - **[MemAgent](https://github.com/BytedTsinghua-SIA/MemAgent)** for the recurrent module and training dataset | |
| ## π License | |
| This model is licensed under the Apache License 2.0. See the LICENSE file for details. | |
| ### Dataset License | |
| Training and evaluation datasets are derived from: | |
| - **HotpotQA, 2WikiMultihopQA, Natural Questions, TriviaQA**: Wikipedia-derived content licensed under [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/) | |