Baidu's Unlimited-OCR model, released on July 3, 2026, advances optical character recognition by offering one-shot long-horizon parsing for single images and multi-page documents like PDFs. Hosted on Hugging Face, it demonstrated strong performance on the ParseBench benchmark, achieving 86.81 for text content accuracy.
This new model, detailed in an arXiv paper published on June 23, 2026, pushes the boundaries of existing OCR technology. Traditional OCR often struggles with complex layouts and multi-page documents, requiring significant pre-processing.
Unlimited-OCR tackles these challenges by processing entire documents more cohesively. The release signifies Baidu's continued investment in core AI capabilities, alongside its broader AI efforts like the integration into Apple Intelligence for Chinese users.
Unlimited-OCR: Baidu's New Vision-Language Model
Baidu's Unlimited-OCR is a 3-billion-parameter vision-language model designed for optical character recognition, specializing in processing lengthy and complex documents, including multi-page PDFs. Released on July 3, 2026, it supports advanced document parsing beyond traditional OCR tools.The model leverages deep learning to understand document structure and content across large inputs. This capability allows it to handle diverse formats, from scanned images to multi-page financial reports. The primary goal is to provide a comprehensive solution for extracting information from unstructured visual data.
It builds on previous work like Deepseek-OCR and PaddleOCR, integrating lessons learned from these projects. Its development highlights the increasing sophistication of AI models in handling real-world data challenges. This trend is observed across various sectors as highlighted in articles discussing the rise of AI agents and their complex applications, such as Agents, Deepfakes, and the Messy Reality of the AI Boom.
How Does Unlimited-OCR Boost Document Parsing?
Unlimited-OCR significantly improves document parsing by enabling one-shot long-horizon analysis. It processes extended documents without breaking them into smaller chunks. This capability is crucial for maintaining context and accuracy across complex layouts and multi-page files.
Traditional OCR systems often process documents page by page or in smaller sections. This approach can lose critical context when information spans across boundaries. Unlimited-OCR's long-horizon parsing addresses this limitation. It ensures a more holistic understanding of the document, even for intricate layouts.
For instance, it can parse entire legal contracts or technical manuals in a single operation. The model's design focuses on reducing the need for manual intervention and post-processing. This significantly improves efficiency for tasks like data entry and content digitization.
Deployment Flexibility and Performance Benchmarks
The model offers diverse deployment options, including Hugging Face Transformers, vLLM, and SGLang. This makes it accessible across various AI infrastructure setups. Performance on the ParseBench benchmark highlights its effectiveness in text content and layout extraction, with a text content score of 86.81%.Developers can integrate Unlimited-OCR using standard Transformers libraries for NVIDIA GPUs. Tested requirements include Python 3.12.3 and CUDA 12.9. For high-throughput inference, vLLM support was added on June 28, 2026. This allows for optimized serving of the model in production environments.
Additionally, SGLang provides a robust framework for launching the model server. The model requires specific Python packages like PyMuPDF for efficient PDF processing. Performance metrics on the ParseBench benchmark offer insights into the model's capabilities in specific areas:
Metric | Value (%) |
|---|---|
Mean | 46.17 |
Text Content | 86.81 |
Layout | 71.52 |
Table | 70.21 |
Chart | 1.34 |
Text Formatting | 0.97 |
These results, updated as of early July 2026, indicate strong performance in extracting text and understanding document layout and tables. The 3-billion parameter model, with BF16 tensor type, has seen significant adoption, recording over 2.1 million downloads last month alone. This wide adoption demonstrates its utility in practical applications.








