OCR-MCP
Complete AI OCR webapp and MCP server. A web app with a streamlined dashboard (drag-and-drop or scanner, pick an engine, click one button) and a FastMCP 3.1 MCP server for agentic IDEsClaude, Cursor, Windsurfso agents can run OCR, preprocessing, and workflows as tools. Same 14 engines, WIA scanner (Windows), and pipelines; one repo.
Topics: ocr, mcp, fastmcp, document-processing, scanner, wia, pdf, computer-vision, model-context-protocol, llm
What it does
- Web app React (
web_sota/) + FastAPI (backend/app.py): upload or scan, pick engine, get text/PDF/JSON. Ports 10858 (Vite) and 10859 (API). In-app Help (/help) documents the web UI, the MCP server, and OCR backends. - MCP server FastMCP 3.1 stdio: tools for OCR, preprocessing, scanner, workflows. Sampling defaults to local Ollama (
http://127.0.0.1:11434/v1, modelllama3.2) no cloud API key. SetOCR_SAMPLING_USE_CLIENT_LLM=1to use the host IDEs LLM instead. Mistral OCR usesMISTRAL_API_KEYwhen you call that backend. See AI_FEATURES.md.
Features: 14 backends (Unlimited-OCR, PaddleOCR-VL-1.5, Nemotron VL 8B, DeepSeek-OCR-2, MinerU2.5-Pro, Mistral OCR) Auto backend selection Preprocessing (deskew, enhance, crop) Layout & table extraction Quality assessment WIA scanner Auto-Scan watcher (detect documents on flatbed, auto-OCR) Batch & pipelines Multi-format export
Docs
| Doc | Description |
|---|---|
| Install | Install, run MCP, Web UI (start.ps1, ports 10858/10859), PyYAML notes, client config |
| Backend deps | Web FastAPI backend: same venv as ocr-mcp, pyproject.toml, PyTorch, OCR_AUTO_INSTALL_DEPS |
| Technical | Architecture, tools, config, development, packaging |
| OCR models | Engines, capabilities, hardware (see also AI_MODELS.md) |
| Backend requirements | Per-model pip packages, system deps, env/config |
| MCP toolset matrix | Portmanteau tools, operation status, corpus v0 |
| AI features | Sampling, SEP-1577, agentic workflows, prompts |
| Webapp redesign | July 2026 redesign: dashboard-first workflow, removed legacy frontend, 5-page sidebar |
| Book scanning | Home book scanning guide: V-cradle, CZUR, auto-scan integration |
| Book pipeline | Spine-to-EPUB pipeline plan: cut, scan, OCR, chapter detect, EPUB, Calibre |
| In-app Help | Source for /help: webapp vs MCP vs backends (mirrors INSTALL / TECHNICAL) |
| SOTA Compliance | Verified SOTA v12.0 Architecture |
Also: JUSTFILE.md (just recipes) OCR-MCP_MASTER_PLAN.md (roadmap) tests/README.md (testing)
Quick Start
git clone https://github.com/sandraschi/ocr-mcp
cd ocr-mcp
just
This opens an interactive dashboard showing all available commands. Run just bootstrap to install dependencies, then just serve or just dev to start.
Manual Setup
If you don't have just installed:
🛡️ Industrial Quality Stack
This project adheres to SOTA 14.1 industrial standards for high-fidelity agentic orchestration:
- Python (Core): Ruff for linting and formatting. Zero-tolerance for
printstatements in core handlers (T201). - Webapp (UI): Biome for sub-millisecond linting. Strict
noConsoleLogenforcement. - Protocol Compliance: Hardened
stdout/stderrisolation to ensure crash-resistant JSON-RPC communication. - Automation: Justfile recipes for all fleet operations (
just lint,just fix,just dev). - Security: Automated audits via
banditandsafety.
License
MIT see LICENSE.
No comments yet
Be the first to share your take.