No description
  • Python 65.2%
  • JavaScript 16.9%
  • CSS 6.8%
  • Shell 6.5%
  • HTML 3.2%
  • Other 1.4%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-24 02:58:52 +00:00
.claude webapp 2026-09-24 02:58:52 +00:00
scripts webapp 2026-09-24 02:58:52 +00:00
server webapp 2026-09-24 02:58:52 +00:00
web webapp 2026-09-24 02:58:52 +00:00
.env.example webapp 2026-09-24 02:58:52 +00:00
.gitignore webapp 2026-09-24 02:58:52 +00:00
docker-compose.yml webapp 2026-09-24 02:58:52 +00:00
README.md webapp 2026-09-24 02:58:52 +00:00

Polish ASR — Parakeet TDT v3 (OpenVINO, CPU, Cloud Run)

Record audio → see a live-updating transcript as you speak → get the final accurate transcript when you press Stop. Uses nvidia/parakeet-tdt-0.6b-v3.

The model is compiled to ONNX then OpenVINO IR and served on CPU — no CUDA/GPU, no NeMo/torch at runtime. Concurrent users are handled via a pool of OpenVINO InferRequests per compiled model (see server/asr_engine.py).

There's no Polish-capable streaming-trained Parakeet checkpoint from NVIDIA (their streaming variants are English-only or multi-speaker/diarization models), and this model wasn't trained with limited attention context, so true low-latency token-by-token streaming isn't available without an accuracy hit. Instead, server/main.py periodically re-decodes the whole buffer recorded so far and pushes it as an interim preview (is_final: false) while the final decode on Stop is unchanged/full-accuracy (is_final: true).

Run locally

docker compose up --build

Open http://localhost:8080

Config

Env Default Meaning
LANGUAGE pl display only
MODELS_DIR /app/models where encoder/decoder_joint IR + tokenizer live
OV_DEVICE CPU OpenVINO device
OV_POOL_SIZE 1 in Cloud Run (4 in docker-compose) number of InferRequests per model — bounds real concurrency; extra requests queue. Kept low in Cloud Run because each connected client periodically re-decodes its whole growing buffer, so concurrent users multiply simultaneous heavy decodes fast — see scripts/deploy_cloud_run.sh

Re-exporting the model

server/models/ already contains the exported IR (encoder.xml/.bin, decoder_joint.xml/.bin, tokenizer.model). To regenerate them from scratch (e.g. after a model update):

python -m venv /tmp/export-venv && source /tmp/export-venv/bin/activate
pip install -r scripts/requirements-export.txt
python scripts/export_openvino.py

This is a separate, heavyweight toolchain (NeMo + torch) that is not part of the runtime container.

Deploying to Cloud Run

./scripts/deploy_cloud_run.sh

See that script for the project/region/service name defaults.