模型

Unlimited-OCR

baidu/Unlimited-OCR

查看上游原文 ↗
上游访问:公开 本站服务:可咨询 许可证:mit 上游版本:2a06ebf2d6f6

中文简介

Unlimited-OCR 是百度发布的视觉文字识别与长文档解析模型,重点面向一次性、长流程的页面内容解析。上游提供论文、代码和模型资料,标签显示其支持多语言 OCR;对表格、公式、版面和长文档的具体效果应结合原论文、测试样例和实际业务文档复核。

UPSTREAM README

上游模型卡 / 数据集卡

在 Hugging Face 查看原文 ↗

Unlimited-OCR 是百度发布的视觉文字识别与长文档解析模型,重点面向一次性、长流程的页面内容解析。上游提供论文、代码和模型资料,标签显示其支持多语言 OCR;对表格、公式、版面和长文档的具体效果应结合原论文、测试样例和实际业务文档复核。

已有简体中文译文 · 本站中文整理 · 2026-07-23 14:50

Unlimited OCR Works Welcome the Era of One-shot Long-horizon Parsing. Release [2026/07/21] 🤝 Thanks to the ms-swift community for their support, our model now supports training with ms-swift. [2026/07/03] 🤝 Thanks to the Baidu Cloud team for their support. Our model is now available on Baidu Cloud. [2026/06/28] 🤝 Thanks to the vLLM community and Tianyu Guo for their support, our model now supports vLLM inference. [2026/06/24] 🤝 Thanks to AK for creating a demo for us. It is now available at Hugging Face Spaces. [2026/06/23] 📄 Our paper is now available on arXiv. [2026/06/23] 🤝 Thanks to the ModelScope community for their support. Our model is now available at ModelScope. [2026/06/22] 🚀 We present Unlimited-OCR, aiming to push Deepseek-OCR one step further. Inference Transformers Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.3 + CUDA12.9: vLLM Please refer to the official vLLM recipe for deployment details: Recipe:** https://recipes.vllm.ai/baidu/Unlimited-OCR Docker Images Use the following Docker images depending on your GPU platform: Default (CUDA 13.0):** For Hopper GPUs (CUDA 12.9)** SGLang Set up the environment (uv-managed virtualenv). Install the local SGLang wheel first, then pin `kernels==0.9.0` and install PyMuPDF for PDF-to-image conversion: Start the SGLang server: Send streaming requests to the OpenAI-compatible API: Visualization Acknowledgement We would like to thank Deepseek-OCR, Deepseek-OCR-2, PaddleOCR for their valuable models and ideas. Citation ```bibtex @misc{yin2026unlimitedocrworks, title={Unlimited OCR Works}, author={Youyang Yin and Huanhuan Liu and YY and Qunyi Xie and Chaorun Liu and Shiqi Yang and Shaohua Wang and Zhanlong Liu and Hao Zou and Jinyue Chen and Shu Wei and Jingjing Wu and Mingxin Huang and Zhen Wu and Guibin Wang and Tengyu Du and Lei Jia}, year={2026}, eprint={2606.23050}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2606.23050}, }

公开页仅展示原文摘录;完整模型卡或数据集卡请前往上游仓库查看。

上游文件元数据

  • .gitattributes302 B
  • assets/baidu.png10.85 KB
  • assets/long-horizon-ocr.gif78.37 MB
  • assets/Unlimited-OCR.png103.79 KB
  • config.json2.81 KB
  • configuration_deepseek_v2.py10.47 KB
  • conversation.py9.04 KB
  • deepencoder.py37.12 KB
  • LICENSE1.04 KB
  • model-00001-of-000001.safetensors6.21 GB
  • model.safetensors.index.json251.57 KB
  • modeling_deepseekv2.py88.05 KB
第三方资源声明

本页面为橙子AI科技的中文整理与服务说明,不代表资源作者或平台官方页面。实际许可、访问和使用条件以上游原文为准。