Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.
$ npx -y skills add paddlepaddle/paddleocr --agent claude-code
Repo: paddlepaddle/paddleocr
What's inside

English | 简体中文 | 繁體中文 | 日本語 | 한국어 | Français | Русский | Español | العربية
PaddleOCR converts PDF documents and images into structured, LLM-ready data (JSON/Markdown) with industry-leading accuracy. With 70k+ Stars and trusted by top-tier projects like Dify, RAGFlow, and Cherry Studio, PaddleOCR is the bedrock for building intelligent RAG and Agentic applications.
Transforming messy visuals into structured data for the LLM era.
The global gold standard for high-speed, multilingual text spotting.
PP-OCRv6 highlights:
PaddleOCR-VL-1.6 highlights:
PaddleOCR-VL series, PP-StructureV3, and PP-DocTranslation now support exporting parsed results to DOCX for convenient viewing and editing in Microsoft Word.PaddleOCR.js, the official browser inference SDK that supports running PP-OCRv5 directly in the browser.Showing a partial view of a very large repo.
FAQ
paddleocr is a Claude Code plugin with 2 hand-picked skills for data work, indexed on Flowy. Install it with the command on its page. It includes paddleocr-doc-parsing, paddleocr-text-recognition. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it