DeepSeek-OCR is a model that investigates the role of vision encoders from a large language model (LLM) perspective, exploring the boundaries of visual-text compression. The model achieves near-lossless OCR compression at approximately 10× ratios and retains 60% accuracy at 20× compression. It supports various modes, including native resolution and dynamic resolution, and can be used for tasks such as extracting text from images and converting documents to markdown. The model is available on GitHub and has been tested with various benchmarks.