- Python 87.8%
- HTML 3.7%
- JavaScript 3.2%
- CSS 2.6%
- Dockerfile 1.4%
- Other 1.3%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .claude | ||
| scripts | ||
| server | ||
| tests | ||
| web | ||
| .dockerignore | ||
| .env.example | ||
| .gitignore | ||
| docker-compose.yml | ||
| pytest.ini | ||
| README.md | ||
| requirements-dev.txt | ||
Polish OCR Comparison
Side-by-side OCR comparison for Polish text — open-source engines + major cloud APIs.
Engines
| Engine | Type | Polish support |
|---|---|---|
| Tesseract | local | pol traineddata |
| EasyOCR | local (GPU) | pl language pack |
| docTR | local (GPU) | Latin-script multilingual (Polish diacritics) |
| Surya | local (GPU) | 90+ languages, auto-detect |
| RysOCR | local VLM (GPU) | Polish LoRA on PaddleOCR-VL |
| Qwen2.5-VL | local VLM (GPU) | general multimodal LLM as OCR; QWEN_VL_MODEL (default 3B) |
| Google Cloud Vision | API | language_hints: [pl] |
| AWS Textract | API | Latin script / auto |
| Azure Document Intelligence | API | locale: pl-PL |
| Mistral OCR | API | native Polish support |
Local engines work out of the box. Cloud APIs need credentials in .env (see .env.example); they'll show as offline until configured.
Local LLM OCR: rysocr (PaddleOCR-VL + Polish LoRA) and qwen_vl (Qwen2.5-VL instruct). Both are vision-language models prompted for transcription — not classic detect→recognize OCR stacks.
By default OCR_LOCAL_ONLY=1 — only local engines load on first run (faster startup, no API panels). Set OCR_LOCAL_ONLY=0 in .env when you're ready to compare cloud APIs.
Run
cd ~/ocr-webapp
cp .env.example .env # optional — fill in API keys, set OCR_LOCAL_ONLY=0 for cloud
docker compose up --build
Open http://localhost:8082 (override with OCR_PORT in .env)
First boot downloads model weights (EasyOCR, docTR, Surya, RysOCR, Qwen2.5-VL — cached in Docker volumes).
Usage
- Drop, paste, or browse an image with Polish text
- Click Recognize
- Compare results + latency per engine
Google credentials in Docker
If using Google Vision, mount your service-account JSON:
# docker-compose.yml (add under ocr service)
volumes:
- ./secrets/gcp-sa.json:/secrets/gcp.json:ro
environment:
- GOOGLE_APPLICATION_CREDENTIALS=/secrets/gcp.json
Tests
cd ~/ocr-webapp
pip install -r requirements-dev.txt
pytest -v
Integration tests (real engines/APIs) are skipped unless:
RUN_OCR_INTEGRATION=1 pytest -v -m integration
- Default host port 8082 (
OCR_PORTin.env) — ASR app uses 8080 - docTR has no dedicated
plcheckpoint; it uses a Latin-script multilingual CRNN (works well for Polish text) - RysOCR downloads PaddleOCR-VL + LoRA adapter on first boot (~several GB); requires
transformers==4.57.6 - Max upload size: 20 MB