Image-Text-to-Text
GGUF
English
Chinese
ocr
document-parsing
deepseek
llama.cpp
42model
conversational
Instructions to use 42ailab/DeepSeek-OCR-2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 42ailab/DeepSeek-OCR-2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use 42ailab/DeepSeek-OCR-2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "42ailab/DeepSeek-OCR-2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "42ailab/DeepSeek-OCR-2-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- Ollama
How to use 42ailab/DeepSeek-OCR-2-GGUF with Ollama:
ollama run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use 42ailab/DeepSeek-OCR-2-GGUF with Docker Model Runner:
docker model run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- Lemonade
How to use 42ailab/DeepSeek-OCR-2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.DeepSeek-OCR-2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
File size: 7,174 Bytes
5f3862a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 | ---
license: apache-2.0
base_model: deepseek-ai/DeepSeek-OCR-2
pipeline_tag: image-text-to-text
language:
- multilingual
tags:
- ocr
- document-parsing
- deepseek
- gguf
- llama.cpp
- 42model
frameworks:
- gguf
---
<!-- 这是 ModelScope 仓 `42ailab/DeepSeek-OCR-2-GGUF` 的 README。改这里 = 改仓首页。 -->
# DeepSeek-OCR-2 · 全本地文档识别(轻量档)
[](https://www.apache.org/licenses/LICENSE-2.0)
[](https://42model.com)
[](https://42model.com)
给一张文档图片或扫描件,本模型把整页读成**可编辑的文字与表格**——全部在你自己的电脑上完成,**不上云、免费、私密**;两个文件合计约 2.5 GB,普通笔记本也扛得住。
> **模型由[深度求索(DeepSeek)](https://modelscope.cn/models/deepseek-ai/DeepSeek-OCR-2)研发、以 Apache-2.0 开源**(DeepSeek-OCR 2: Visual Causal Flow)。**GGUF 量化档由社区贡献者 [sabafallah](https://huggingface.co/sabafallah/DeepSeek-OCR-2-GGUF) 制作**(Apache-2.0 允许再分发)。**本仓既不是新模型、也不是我们量化的**——我们做的是把这份社区量化档**逐字节镜像到 ModelScope**(国内可直连下载),并把它接进活水模型引擎、让它在桌面上一键可用。
## 一、它解决什么问题
把文档(合同、发票、论文、书籍、扫描件、截图)变成**可编辑、可检索、可喂给 AI** 的文字,是知识工作里最高频的一步:
- **转可编辑文档**:把图片 / 扫描件里的内容变成文字与表格,直接改、直接搜;
- **建检索与问答**:把纸质与图片资料数字化,喂给 AI 做总结、问答;
- **随手取字**:截图里的文字、拍下来的资料,不必手打重录。
过去要拿到像样的文档识别,往往得把文件**上传到别人的服务器**——既有隐私与合规顾虑,也可能产生费用。本模型让这一步**完全在本地完成**。而它的另一个好处是**小**:合计约 2.5 GB、8 GB 内存的电脑即可运行,是我们精选库里的**轻量 OCR 档**。
## 二、它是怎么做到的
**识别能力来自深度求索的 DeepSeek-OCR 2**,其公开的设计要点:
- **看整页、不切碎**:先把整页图片压成**少量「视觉词」**(每页约 256–1120 个),再一次性读出正文与版面——跨栏、跨行的结构不容易错乱。这条「用视觉token压缩长文档」的思路来自它的前作 DeepSeek-OCR(Contexts Optical Compression);
- **分辨率自适应**:按页面复杂度在 768×768 与 1024×1024 之间动态取图,简单页少花算力、复杂页看得更细;
- **同时给文字和版面**:既能只取纯文字,也能连版面位置一起输出。
**我们做的**(不含模型训练,也不含量化):
- 把社区制作的 GGUF 量化档**逐字节镜像**到 ModelScope,每个文件的 sha256 与来源完全一致(国内下载不必翻墙);
- **接进引擎**:识别结果里的版面标注由引擎自动剥掉,你拿到的是干净正文;在 Apple 显卡(Metal)上,引擎会**自动把视觉部分放到 CPU** 跑——这一代 DeepSeek-OCR 架构在 Metal 上有已知精度问题,规避后无需你操心;
- 沿用与上游**一致的 Apache-2.0 许可**。
## 三、效果如何
模型本体的完整评测请以**上游论文**为准([arXiv:2601.20552](https://arxiv.org/abs/2601.20552))——我们不转述未经核实的分数。这里只列**我们自己能背书的两项核验**:
| 我们核验的项目 | 结果 |
|---|---|
| 镜像忠实性(我方 ModelScope 档 vs 社区来源档,逐字节) | 两个文件 sha256 与大小**完全一致**,未做任何再压缩 |
| 引擎可跑性(我们捆绑的推理引擎版本) | DeepSeek-OCR / OCR-2 为**原生支持架构**,无需自建或切换分支版本 |
我们**未**重跑公开文档解析基准,因此不给出「量化前后掉几分」这类数字。
## 四、局限与下一步
- **定位是轻量档**:约 2.5 GB、省内存,胜在小和快装;对版面极复杂的文档,精选库里更大的 OCR 档会更稳。
- **装饰性竖排文字**(如竖排日文标题)等罕见版面仍可能读串——这是当前一代文档 OCR 的共性难点。
- **Apple 显卡(Metal)**:视觉部分在 Metal 上有已知精度问题,引擎已**自动规避**(视觉走 CPU、文字部分仍用 GPU 加速),普通用户无需关心。
- 下一步:持续跟进上游模型更新;评估自产量化档以进一步压缩体积。
## 五、怎么下载与使用
本模型专为 [活水模型(42model)](https://42model.com) 打包,推荐通过它获取:
**桌面版**
1. 在「模型库」→「OCR」里找到 **DeepSeek-OCR-v2**,下载;
2. 点「开始使用」把它设为当前 OCR 模型。
之后在「OCR」能力里把图片 / 扫描件转成文字即可。
## 文件与许可
| 文件 | 作用 | 大小 |
|---|---|---|
| `deepseek-ocr-2-Q4_K_M.gguf` | 文本解码器(Q4_K_M 量化) | ~1.95 GB |
| `mmproj-deepseek-ocr-2-q8_0.gguf` | 视觉编码器(Q8_0,接近无损) | ~512 MB |
各文件 sha256 见 ModelScope 文件页,可自行校验。
**许可**:模型本体为 DeepSeek-OCR 2,© DeepSeek,**Apache-2.0**(官方源:[GitHub](https://github.com/deepseek-ai/DeepSeek-OCR-2) · [ModelScope](https://modelscope.cn/models/deepseek-ai/DeepSeek-OCR-2) · [论文](https://arxiv.org/abs/2601.20552))。GGUF 量化档由社区贡献者 [sabafallah](https://huggingface.co/sabafallah/DeepSeek-OCR-2-GGUF) 制作;本仓为该量化档的镜像,同样遵循 [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0),使用即表示同意上游许可条款。
## 引用
请引用上游深度求索(DeepSeek-OCR 2 及其前作):
```bibtex
@article{wei2026deepseek,
title={DeepSeek-OCR 2: Visual Causal Flow},
author={Wei, Haoran and Sun, Yaofeng and Li, Yukun},
journal={arXiv preprint arXiv:2601.20552},
year={2026}
}
@article{wei2025deepseek,
title={DeepSeek-OCR: Contexts Optical Compression},
author={Wei, Haoran and Sun, Yaofeng and Li, Yukun},
journal={arXiv preprint arXiv:2510.18234},
year={2025}
}
```
本仓只做镜像与本地适配,不主张任何模型或量化成果的署名。
联系我们:**contact@42ailab.com**
## 关于我们
**[活水 AI 实验室(42ailab)](https://42ailab.com)** — 探索智能边界的 AI 创新实验室,以认知科学为基石,推动 AI 与人类智能的深度融合,真正理解并增强智能 —— 碳基的,也是硅基的。
**[活水模型(42model)](https://42model.com)** — 由活水 AI 实验室出品的高性能本地 AI 大模型推理引擎,让翻译、转写、识别、对话、编程等 AI 能力在你的本机免费私密运行;并可借云端算力微调专属模型,回传本机运行。
|