- Python 50.6%
- Shell 17.4%
- TypeScript 9.6%
- CSS 8.3%
- JavaScript 4.9%
- Other 9.2%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| .claude | ||
| api | ||
| deploy | ||
| ingestion | ||
| webapp | ||
| .dockerignore | ||
| .env | ||
| .gcloudignore | ||
| API.md | ||
| cloudbuild.yaml | ||
| docker-compose.yml | ||
| Makefile | ||
| README.md | ||
Word Search
Polish word/referent search: embed short referent phrases with a Polish retrieval model, store them in Postgres/pgvector, search by cosine similarity, and optionally re-rank the top candidates with Gemini to fix homonym confusion (e.g. "pilot" the TV remote vs. "pilot" the person who flies planes).
Three moving parts:
webapp/ Static TS/HTML frontend (search box + add-word form)
api/ FastAPI service — /api/status, /api/words, /api/search + serves webapp/
ingestion/ Embedding model, CLI (jina-pg), Postgres schema, batch ingest jobs
deploy/ Cloud Run deploy scripts + env files (not committed with real secrets — see below)
api/ and ingestion/ share the jina_pg package (embedding + DB code lives in
ingestion/src/jina_pg, the API imports it as a dependency).
How a search request actually works
GET /api/search?query=... in api/src/word_api/main.py:
- Embed the query —
jina_pg.embedder.OpenVINOEmbedder.embed_query()runs the query through PolDense-150M (OPI-PIB/PolDense-150M), prefixed with[query]:(queries and indexed passages use different prefixes by the model's own convention). ~25-35ms steady-state on CPU. - Vector search —
jina_pg.db.Database.search()runs a pgvector cosine-similarity query against Postgres, returningRAG_POOL_SIZE(default 50) candidates if grounding is on, or justlimitif it's off. ~2.5s — this opens a fresh connection per request (see Known limitations). - Grounding (optional, on by default) —
api/src/word_api/grounder.pysends the candidate list to Gemini (gemini-3.1-flash-litevia Vertex AI) and asks it to return only the IDs that genuinely match the query's intent, structured as JSON. This is what fixes the homonym problem that pure embedding similarity can't. ~1-3s. - Candidates get mapped back to full rows, truncated to
limit, and returned.
GROUNDING_ENABLED=false skips steps 3-4's Gemini call entirely (much faster, no
disambiguation).
Model: PolDense-150M, compiled to OpenVINO
The model is OPI-PIB/PolDense-150M (ModernBERT / ettin-encoder, 149M params,
768-dim, CLS-token pooling + L2 normalize). It beat MMLW-v2 on this corpus —
especially homonym disambiguation — at ~2× lower query latency and ~⅓ the IR size.
Earlier experiments (see ingestion/sql/00[3-6]_*.sql) were reverted; 006 switches
the column from vector(1024) → vector(768) and requires a re-ingest.
It's compiled once at Docker build time via
ingestion/scripts/export_openvino.py:
PyTorch → ONNX (legacy TorchScript exporter, dynamo=False — dynamo IRs break
batch>1) → OpenVINO IR (fp16). The runtime image ships no torch/CUDA — just
the OpenVINO runtime and a tokenizer.
Inference is served through jina_pg.embedder.OpenVINOEmbedder
(embedder.py): a pool of pre-created OpenVINO
InferRequest objects, borrowed/returned via a queue.Queue. A single InferRequest
isn't safe to call from two threads at once, but the pool means concurrent requests each
get their own — this is what makes it safe for multiple simultaneous users to hit one
Cloud Run instance (FastAPI's sync def endpoints run on a thread pool, so concurrent
HTTP requests are real concurrent threads). Pool size auto-detects via OpenVINO's
OPTIMAL_NUMBER_OF_INFER_REQUESTS, or set OPENVINO_NUM_INFER_REQUESTS to pin it.
To run jina-pg or the API outside Docker, export the model once:
cd ingestion
pip install -r requirements-export.txt --index-url https://download.pytorch.org/whl/cpu --extra-index-url https://pypi.org/simple/
python scripts/export_openvino.py --output-dir ../models/poldense-150m-openvino
Docker builds do this automatically in a model-export build stage — see api/Dockerfile.
Local development
cp deploy/db.env.example deploy/db.env # fill in real Postgres creds (or run the
# optional local one: docker compose --profile local-db up -d postgres)
make up # syncs env, builds, runs the API on :8001 (API_PORT)
UI: http://localhost:8001 · Swagger docs: http://localhost:8001/docs
See api/README.md and ingestion/README.md for service-specific detail, and API.md for the API contract (for teams integrating against this service rather than running it).
Deployment (Cloud Run)
cp deploy/cloudrun.env.example deploy/cloudrun.env # fill in real values
make deploy
deploy/deploy.sh does not build/push the image itself (that step is commented out —
build with docker build -f api/Dockerfile -t <artifact-registry-image> . and
docker push, or wire up gcloud builds submit --config=cloudbuild.yaml yourself, then
run make deploy to point Cloud Run at that image + sync env vars/sizing).
Current production config (deploy/cloudrun.env):
- Service account:
backend-service-account@snappy-helper-438617-d0.iam.gserviceaccount.com. The originally-intendedwordsearch-api@...service account doesn't exist and creating it requiresiam.serviceAccounts.create, which the deploying identity doesn't have — if you get that permission later, create a narrower-scoped SA (needs onlyroles/aiplatform.userfor Gemini) and swapSERVICE_ACCOUNTincloudrun.env. - Sizing: 4Gi memory / 2 CPU, min-instances=1 (always-on, no cold start for normal traffic — see cost notes below), max-instances=3.
- Gemini:
gemini-3.1-flash-litevia Vertex AI. This model is only available on Vertex AI'sglobalendpoint —GOOGLE_CLOUD_LOCATION=globalis required; a regional endpoint likeus-central1returns a 404 model-not-found.
Cost profile (Cloud Run + Vertex AI, us-central1/global, as of mid-2026)
Rough estimate, not measured from GCP's billing console — check Cloud Billing reports against actual traffic for ground truth, especially the Gemini token counts which are estimated from the prompt-building code, not GCP's tokenizer.
Fixed cost (min-instances=1 keeps one instance warm 24/7):
- ~4 GiB × ~2.59M seconds/month × $0.0000025/GiB-s ≈ $25-26/month, memory only — Cloud Run doesn't bill CPU while an instance is idle (default "CPU allocated during request processing" mode, not "CPU always allocated").
Marginal cost per 1000 grounded searches (steady-state, ~3.7s/request, 2 vCPU / 4Gi):
| Component | Rate | Per request | Per 1000 |
|---|---|---|---|
| Cloud Run CPU | $0.000024/vCPU-s | 2 × 3.7s | ~$0.18 |
| Cloud Run memory | $0.0000025/GiB-s | 4 × 3.7s | ~$0.04 |
| Cloud Run requests | $0.40/1M | — | ~$0.0004 |
| Gemini input (~1,250 tok) | $0.25/1M tok | ~$0.31 | |
| Gemini output (~75 tok) | $1.50/1M tok | ~$0.11 | |
| Total | ~$0.64 / 1000 searches |
GCP's free tier (180,000 vCPU-s, 360,000 GiB-s, 2M requests per month) covers roughly the
first 24,000 grounded searches/month of Cloud Run compute, so at low-to-moderate volume
the practical marginal cost is closer to just the Gemini portion ($0.40-0.45/1000),
stacked on top of the ~$25/month idle baseline. GROUNDING_ENABLED=false search requests
are essentially free at that volume (no Gemini call, DB + embed only).
Known limitations / things the next engineer should know
- DB connections aren't pooled —
jina_pg.db.Database.connect()opens a fresh Postgres connection (with SSL handshake) per request. This is the majority of the ~2.5s baseline latency on every request, grounded or not. A connection pool (e.g.psycopg_pool) would be the highest-leverage next optimization. - Postgres is a single external host, not managed by GCP (
178.105.237.149indeploy/db.env/deploy/cloudrun.env) — Cloud Run needs outbound access to it, and there's no HA/failover. - Grounding is a real per-request cost and latency hit (~1-3s with
gemini-3.1-flash-lite, was ~11-16s withgemini-2.5-flash). It's what makes homonym disambiguation work, but if traffic grows a lot, consider replacing it with a local cross-encoder reranker (compiled to OpenVINO the same way PolDense is) for sub-second, no-network-call reranking — flagged as a possible follow-up, not yet built. - No auth on the API —
--allow-unauthenticated, CORS is*. Fine for an internal tool; revisit before exposing this beyond the current product surface. - See API.md for the actual request/response contract.