No description
  • Python 87.8%
  • HTML 3.7%
  • JavaScript 3.2%
  • CSS 2.6%
  • Dockerfile 1.4%
  • Other 1.3%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-24 02:45:16 +00:00
.claude code 2026-09-24 02:45:16 +00:00
scripts code 2026-09-24 02:45:16 +00:00
server code 2026-09-24 02:45:16 +00:00
tests code 2026-09-24 02:45:16 +00:00
web code 2026-09-24 02:45:16 +00:00
.dockerignore code 2026-09-24 02:45:16 +00:00
.env.example code 2026-09-24 02:45:16 +00:00
.gitignore code 2026-09-24 02:45:16 +00:00
docker-compose.yml code 2026-09-24 02:45:16 +00:00
pytest.ini code 2026-09-24 02:45:16 +00:00
README.md first commit 2026-09-24 02:43:18 +00:00
requirements-dev.txt code 2026-09-24 02:45:16 +00:00

Polish OCR Comparison

Side-by-side OCR comparison for Polish text — open-source engines + major cloud APIs.

Engines

Engine Type Polish support
Tesseract local pol traineddata
EasyOCR local (GPU) pl language pack
docTR local (GPU) Latin-script multilingual (Polish diacritics)
Surya local (GPU) 90+ languages, auto-detect
RysOCR local VLM (GPU) Polish LoRA on PaddleOCR-VL
Qwen2.5-VL local VLM (GPU) general multimodal LLM as OCR; QWEN_VL_MODEL (default 3B)
Google Cloud Vision API language_hints: [pl]
AWS Textract API Latin script / auto
Azure Document Intelligence API locale: pl-PL
Mistral OCR API native Polish support

Local engines work out of the box. Cloud APIs need credentials in .env (see .env.example); they'll show as offline until configured.

Local LLM OCR: rysocr (PaddleOCR-VL + Polish LoRA) and qwen_vl (Qwen2.5-VL instruct). Both are vision-language models prompted for transcription — not classic detect→recognize OCR stacks.

By default OCR_LOCAL_ONLY=1 — only local engines load on first run (faster startup, no API panels). Set OCR_LOCAL_ONLY=0 in .env when you're ready to compare cloud APIs.

Run

cd ~/ocr-webapp
cp .env.example .env   # optional — fill in API keys, set OCR_LOCAL_ONLY=0 for cloud
docker compose up --build

Open http://localhost:8082 (override with OCR_PORT in .env)

First boot downloads model weights (EasyOCR, docTR, Surya, RysOCR, Qwen2.5-VL — cached in Docker volumes).

Usage

  1. Drop, paste, or browse an image with Polish text
  2. Click Recognize
  3. Compare results + latency per engine

Google credentials in Docker

If using Google Vision, mount your service-account JSON:

# docker-compose.yml (add under ocr service)
volumes:
  - ./secrets/gcp-sa.json:/secrets/gcp.json:ro
environment:
  - GOOGLE_APPLICATION_CREDENTIALS=/secrets/gcp.json

Tests

cd ~/ocr-webapp
pip install -r requirements-dev.txt
pytest -v

Integration tests (real engines/APIs) are skipped unless:

RUN_OCR_INTEGRATION=1 pytest -v -m integration
  • Default host port 8082 (OCR_PORT in .env) — ASR app uses 8080
  • docTR has no dedicated pl checkpoint; it uses a Latin-script multilingual CRNN (works well for Polish text)
  • RysOCR downloads PaddleOCR-VL + LoRA adapter on first boot (~several GB); requires transformers==4.57.6
  • Max upload size: 20 MB