Image-Text-to-Text
GGUF
English
Chinese
ocr
document-parsing
deepseek
llama.cpp
42model
conversational
Instructions to use 42ailab/DeepSeek-OCR-2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use 42ailab/DeepSeek-OCR-2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use 42ailab/DeepSeek-OCR-2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "42ailab/DeepSeek-OCR-2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "42ailab/DeepSeek-OCR-2-GGUF", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- Ollama
How to use 42ailab/DeepSeek-OCR-2-GGUF with Ollama:
ollama run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use 42ailab/DeepSeek-OCR-2-GGUF with Docker Model Runner:
docker model run hf.co/42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
- Lemonade
How to use 42ailab/DeepSeek-OCR-2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull 42ailab/DeepSeek-OCR-2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.DeepSeek-OCR-2-GGUF-Q4_K_M
List all available models
lemonade list
- Atomic Chat
|
Download README_zh.md from 42ailab/DeepSeek-OCR-2-GGUF: direct link, hf CLI and curl.
- Browser
- Download file 7.17 kB
-
https://huggingface.co/42ailab/DeepSeek-OCR-2-GGUF/resolve/main/README_zh.md
- Command line
-
hf download hf://42ailab/DeepSeek-OCR-2-GGUF/README_zh.md
-
curl -L -o README_zh.md https://huggingface.co/42ailab/DeepSeek-OCR-2-GGUF/resolve/main/README_zh.md
7.17 kB
| license: apache-2.0 | |
| base_model: deepseek-ai/DeepSeek-OCR-2 | |
| pipeline_tag: image-text-to-text | |
| language: | |
| - multilingual | |
| tags: | |
| - ocr | |
| - document-parsing | |
| - deepseek | |
| - gguf | |
| - llama.cpp | |
| - 42model | |
| frameworks: | |
| - gguf | |
| <!-- 这是 ModelScope 仓 `42ailab/DeepSeek-OCR-2-GGUF` 的 README。改这里 = 改仓首页。 --> | |
| # DeepSeek-OCR-2 · 全本地文档识别(轻量档) | |
| [](https://www.apache.org/licenses/LICENSE-2.0) | |
| [](https://42model.com) | |
| [](https://42model.com) | |
| 给一张文档图片或扫描件,本模型把整页读成**可编辑的文字与表格**——全部在你自己的电脑上完成,**不上云、免费、私密**;两个文件合计约 2.5 GB,普通笔记本也扛得住。 | |
| > **模型由[深度求索(DeepSeek)](https://modelscope.cn/models/deepseek-ai/DeepSeek-OCR-2)研发、以 Apache-2.0 开源**(DeepSeek-OCR 2: Visual Causal Flow)。**GGUF 量化档由社区贡献者 [sabafallah](https://huggingface.co/sabafallah/DeepSeek-OCR-2-GGUF) 制作**(Apache-2.0 允许再分发)。**本仓既不是新模型、也不是我们量化的**——我们做的是把这份社区量化档**逐字节镜像到 ModelScope**(国内可直连下载),并把它接进活水模型引擎、让它在桌面上一键可用。 | |
| ## 一、它解决什么问题 | |
| 把文档(合同、发票、论文、书籍、扫描件、截图)变成**可编辑、可检索、可喂给 AI** 的文字,是知识工作里最高频的一步: | |
| - **转可编辑文档**:把图片 / 扫描件里的内容变成文字与表格,直接改、直接搜; | |
| - **建检索与问答**:把纸质与图片资料数字化,喂给 AI 做总结、问答; | |
| - **随手取字**:截图里的文字、拍下来的资料,不必手打重录。 | |
| 过去要拿到像样的文档识别,往往得把文件**上传到别人的服务器**——既有隐私与合规顾虑,也可能产生费用。本模型让这一步**完全在本地完成**。而它的另一个好处是**小**:合计约 2.5 GB、8 GB 内存的电脑即可运行,是我们精选库里的**轻量 OCR 档**。 | |
| ## 二、它是怎么做到的 | |
| **识别能力来自深度求索的 DeepSeek-OCR 2**,其公开的设计要点: | |
| - **看整页、不切碎**:先把整页图片压成**少量「视觉词」**(每页约 256–1120 个),再一次性读出正文与版面——跨栏、跨行的结构不容易错乱。这条「用视觉token压缩长文档」的思路来自它的前作 DeepSeek-OCR(Contexts Optical Compression); | |
| - **分辨率自适应**:按页面复杂度在 768×768 与 1024×1024 之间动态取图,简单页少花算力、复杂页看得更细; | |
| - **同时给文字和版面**:既能只取纯文字,也能连版面位置一起输出。 | |
| **我们做的**(不含模型训练,也不含量化): | |
| - 把社区制作的 GGUF 量化档**逐字节镜像**到 ModelScope,每个文件的 sha256 与来源完全一致(国内下载不必翻墙); | |
| - **接进引擎**:识别结果里的版面标注由引擎自动剥掉,你拿到的是干净正文;在 Apple 显卡(Metal)上,引擎会**自动把视觉部分放到 CPU** 跑——这一代 DeepSeek-OCR 架构在 Metal 上有已知精度问题,规避后无需你操心; | |
| - 沿用与上游**一致的 Apache-2.0 许可**。 | |
| ## 三、效果如何 | |
| 模型本体的完整评测请以**上游论文**为准([arXiv:2601.20552](https://arxiv.org/abs/2601.20552))——我们不转述未经核实的分数。这里只列**我们自己能背书的两项核验**: | |
| | 我们核验的项目 | 结果 | | |
| |---|---| | |
| | 镜像忠实性(我方 ModelScope 档 vs 社区来源档,逐字节) | 两个文件 sha256 与大小**完全一致**,未做任何再压缩 | | |
| | 引擎可跑性(我们捆绑的推理引擎版本) | DeepSeek-OCR / OCR-2 为**原生支持架构**,无需自建或切换分支版本 | | |
| 我们**未**重跑公开文档解析基准,因此不给出「量化前后掉几分」这类数字。 | |
| ## 四、局限与下一步 | |
| - **定位是轻量档**:约 2.5 GB、省内存,胜在小和快装;对版面极复杂的文档,精选库里更大的 OCR 档会更稳。 | |
| - **装饰性竖排文字**(如竖排日文标题)等罕见版面仍可能读串——这是当前一代文档 OCR 的共性难点。 | |
| - **Apple 显卡(Metal)**:视觉部分在 Metal 上有已知精度问题,引擎已**自动规避**(视觉走 CPU、文字部分仍用 GPU 加速),普通用户无需关心。 | |
| - 下一步:持续跟进上游模型更新;评估自产量化档以进一步压缩体积。 | |
| ## 五、怎么下载与使用 | |
| 本模型专为 [活水模型(42model)](https://42model.com) 打包,推荐通过它获取: | |
| **桌面版** | |
| 1. 在「模型库」→「OCR」里找到 **DeepSeek-OCR-v2**,下载; | |
| 2. 点「开始使用」把它设为当前 OCR 模型。 | |
| 之后在「OCR」能力里把图片 / 扫描件转成文字即可。 | |
| ## 文件与许可 | |
| | 文件 | 作用 | 大小 | | |
| |---|---|---| | |
| | `deepseek-ocr-2-Q4_K_M.gguf` | 文本解码器(Q4_K_M 量化) | ~1.95 GB | | |
| | `mmproj-deepseek-ocr-2-q8_0.gguf` | 视觉编码器(Q8_0,接近无损) | ~512 MB | | |
| 各文件 sha256 见 ModelScope 文件页,可自行校验。 | |
| **许可**:模型本体为 DeepSeek-OCR 2,© DeepSeek,**Apache-2.0**(官方源:[GitHub](https://github.com/deepseek-ai/DeepSeek-OCR-2) · [ModelScope](https://modelscope.cn/models/deepseek-ai/DeepSeek-OCR-2) · [论文](https://arxiv.org/abs/2601.20552))。GGUF 量化档由社区贡献者 [sabafallah](https://huggingface.co/sabafallah/DeepSeek-OCR-2-GGUF) 制作;本仓为该量化档的镜像,同样遵循 [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0),使用即表示同意上游许可条款。 | |
| ## 引用 | |
| 请引用上游深度求索(DeepSeek-OCR 2 及其前作): | |
| ```bibtex | |
| @article{wei2026deepseek, | |
| title={DeepSeek-OCR 2: Visual Causal Flow}, | |
| author={Wei, Haoran and Sun, Yaofeng and Li, Yukun}, | |
| journal={arXiv preprint arXiv:2601.20552}, | |
| year={2026} | |
| } | |
| @article{wei2025deepseek, | |
| title={DeepSeek-OCR: Contexts Optical Compression}, | |
| author={Wei, Haoran and Sun, Yaofeng and Li, Yukun}, | |
| journal={arXiv preprint arXiv:2510.18234}, | |
| year={2025} | |
| } | |
| ``` | |
| 本仓只做镜像与本地适配,不主张任何模型或量化成果的署名。 | |
| 联系我们:**contact@42ailab.com** | |
| ## 关于我们 | |
| **[活水 AI 实验室(42ailab)](https://42ailab.com)** — 探索智能边界的 AI 创新实验室,以认知科学为基石,推动 AI 与人类智能的深度融合,真正理解并增强智能 —— 碳基的,也是硅基的。 | |
| **[活水模型(42model)](https://42model.com)** — 由活水 AI 实验室出品的高性能本地 AI 大模型推理引擎,让翻译、转写、识别、对话、编程等 AI 能力在你的本机免费私密运行;并可借云端算力微调专属模型,回传本机运行。 | |