PaddleOCR Skills

English | 简体中文

skills.sh

Upstream refactored the skills in PR #18090 (2026-06-03) — it removed the bundled scripts/ and references/ and switched to the official paddleocr api CLI. This mirror intentionally keeps the script-based version (which still works and is offline-friendly via uv), and additionally documents the CLI as an alternative path. See Two Ways to Run the Skills below.

Discover

Included Skills

Skill Use case Entry script
paddleocr-text-recognition Extract text from images, scans, and PDF files ocr_caller.py
paddleocr-doc-parsing Parse complex documents into Markdown/structured output layout_caller.py

Two Ways to Run the Skills

Each skill can be invoked in two ways. The bundled scripts are the default path in this mirror; the paddleocr CLI is the upstream-canonical alternative introduced in #18090.

Scripts (default) paddleocr CLI (alternative)
Install Just uv — deps are resolved from PEP 723 inline metadata pip install "paddleocr>=3.7.0"
Required env vars Per-skill PADDLEOCR_OCR_API_URL / PADDLEOCR_DOC_PARSING_API_URL + PADDLEOCR_ACCESS_TOKEN PADDLEOCR_ACCESS_TOKEN only (the CLI resolves the endpoint internally)
Output format {ok, text, result, error} envelope, auto-saved to a temp file {jobId, pages:[...]} printed to stdout
Page selection (PDF) Pre-split with scripts/split_pdf.py Native --page_ranges "1-5,10"
Best for Skills runtimes, airgapped / offline-friendly setups, no extra install Environments that already ship paddleocr, or where you want the upstream flow

The two paths return different output shapes and read different environment variables — they are not output-compatible. Pick one per workflow.

Detailed usage and per-skill examples live in each SKILL.md:

Requirements

  • Python 3.9 or later
  • uv
  • Internet access
  • PaddleOCR official API credentials from paddleocr.com

The scripts use PEP 723 inline dependency metadata, so there are no separate requirements.txt files to install.

Configuration

Set the environment variables required by the skill you want to use:

Skill Required Optional
paddleocr-text-recognition PADDLEOCR_OCR_API_URL ending with /ocr, PADDLEOCR_ACCESS_TOKEN PADDLEOCR_OCR_TIMEOUT
paddleocr-doc-parsing PADDLEOCR_DOC_PARSING_API_URL ending with /layout-parsing, PADDLEOCR_ACCESS_TOKEN PADDLEOCR_DOC_PARSING_TIMEOUT

The PADDLEOCR_*_API_URL variables above are only required by the bundled scripts. If you use the paddleocr CLI instead, only PADDLEOCR_ACCESS_TOKEN is needed — see the "Alternative: paddleocr CLI" section in each SKILL.md.

Local Usage

Run commands from the corresponding skill directory.

cd skills/paddleocr-text-recognition
uv run scripts/ocr_caller.py --file-path "/path/to/image-or-document.pdf" --pretty
cd skills/paddleocr-doc-parsing
uv run scripts/layout_caller.py --file-path "/path/to/document.pdf" --pretty

Install into AI Apps

One-prompt installation (easiest)

Copy the entire prompt below into Codex, Claude Code, Cursor, OpenCode, OpenClaw, or another AI agent that can operate a terminal:

Install both Agent Skills from https://github.com/Aidenwu0209/PaddleOCR-Skills on this computer.
1. Detect the current supported agent and check Node.js/npx, Python 3.9+, and uv. If something is missing, explain it and use its official installer. Do not use sudo or change unrelated system settings without my permission.
2. Install all skills globally for the detected agent with: npx skills add Aidenwu0209/PaddleOCR-Skills --skill '*' -g -y
3. Run npx skills list -g --json and verify that both paddleocr-text-recognition and paddleocr-doc-parsing are installed.
4. Do not invent, expose, or log a PaddleOCR token. Stop at credential configuration, show me the official https://www.paddleocr.com link, and tell me exactly which endpoint/token values are still required.
5. Report the commands used, install paths, and verification result.

ClawHub also provides a small setup skill that performs the same guarded repository installation and verification flow:

openclaw skills install @aidenwu0209/paddleocr-skills-setup

Use the skills.sh CLI to choose skills and target agents interactively:

npx skills add Aidenwu0209/PaddleOCR-Skills

Or install both included skills globally for a specific agent:

Agent Command
Codex npx skills add Aidenwu0209/PaddleOCR-Skills --agent codex --skill '*' -g -y
Claude Code npx skills add Aidenwu0209/PaddleOCR-Skills --agent claude-code --skill '*' -g -y
GitHub Copilot npx skills add Aidenwu0209/PaddleOCR-Skills --agent github-copilot --skill '*' -g -y
OpenClaw npx skills add Aidenwu0209/PaddleOCR-Skills --agent openclaw --skill '*' -g -y

GitHub CLI 2.90.0+ also provides native Agent Skills installation:

gh skill install Aidenwu0209/PaddleOCR-Skills --all --agent github-copilot --scope user

For Claude Code plugin development or local testing, clone the repository and load its plugin.json:

git clone https://github.com/Aidenwu0209/PaddleOCR-Skills.git
claude --plugin-dir ./PaddleOCR-Skills

To install directly from a local checkout instead:

npx skills add ./skills/paddleocr-text-recognition -g -y
npx skills add ./skills/paddleocr-doc-parsing -g -y

Or install through OpenClaw:

clawhub install paddleocr-text-recognition
clawhub install paddleocr-doc-parsing

Documentation

License

Apache-2.0. See LICENSE.