No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-24 02:41:30 +00:00
eval Port icon pipeline from FLUX.2 Klein to Krea 2 2026-08-09 01:40:45 +00:00
lora_runs/configs Port icon pipeline from FLUX.2 Klein to Krea 2 2026-08-09 01:40:45 +00:00
scripts Wire exploratory sketch style into the enhancement pipeline 2026-08-18 02:07:04 +00:00
.gitignore Port icon pipeline from FLUX.2 Klein to Krea 2 2026-08-09 01:40:45 +00:00
README.md first commit 2026-09-24 02:41:30 +00:00
requirements.txt Port icon pipeline from FLUX.2 Klein to Krea 2 2026-08-09 01:40:45 +00:00

grafiki-krea-2 — Krea 2 icon generation

Port of grafiki-flux-v2 (Polish speech-therapy flashcard icons — black silhouette / outline / photorealistic styles, ~1196 words x 4 samples each) from Flux2KleinPipeline (FLUX.2 Klein 4B) to Krea2Pipeline (Krea 2).

Why this port exists

FLUX.2 Klein is a distilled model: classifier-free guidance is hard-disabled, so negative_prompt never worked, on inference or on training samples. The entire prompt architecture in scripts/style_prompts.py — every clause phrased as a positive assertion instead of a negative, the narrow "structural openings" vs. "dots, speckles, stippling" split, the double restatement of "white background" — is a workaround for that one missing feature. The source repo's comments document three separate measured failures this caused (an "interior detail" clause that raised the interior-speckle rate instead of lowering it; an absolute no-interior-white ban that killed a pupil, shelf divisions, and a bowl's steam because it couldn't distinguish structure from decoration; a shell-wide animal-eye grant that put eyes on glasses of water and hearts) — all traceable to having no real negative channel to put a ban in instead.

Krea 2 removes that constraint. krea/Krea-2-Raw is undistilled and restores true CFG, so negative_prompt is live. This port moves the ban half of each of those three workarounds out of the positive prompt shell and into NEGATIVES, while preserving every measured-history comment from the source file (they're the receipts for why the shell is shaped the way it is — read scripts/style_prompts.py before changing it further).

Raw vs. Turbo

Krea 2 ships as two materially different checkpoints (scripts/krea_config.py's VARIANTS):

variant model id steps guidance negative_prompt
raw (default) krea/Krea-2-Raw 52 3.5 live — real CFG
turbo krea/Krea-2-Turbo 8 0.0 silently inert

diffusers.Krea2Pipeline only runs its negative-branch when guidance_scale > 0 — the docstring says negative_prompt is "Ignored when guidance_scale <= 0". Turbo's guidance is exactly 0, so passing a negative prompt there compiles, runs, and produces images — it just has zero effect, silently. model_pipe.py warns once (not an error — Turbo's speed may be worth it for a quick preview) whenever a non-empty negative_prompt meets guidance_scale <= 0. raw is the default variant for exactly this reason: quality + working negatives is the whole point of this migration.

License: Krea 2 ships under the "Krea 2 Community License" — commercial use is fine; enterprises over 50 seats need a separate agreement with Krea. Check the license text on the model card before a commercial deployment.

Layout

grafiki-krea-2/
├── scripts/
│   ├── krea_config.py                    # ROOT, STYLES, VARIANTS (raw/turbo), LoRA paths
│   ├── model_pipe.py                     # load pipe, load/fuse LoRA, generate (+ negative_prompt);
│   │                                      #   load_pipe_resident/encode_prompts/generate_from_embeds
│   │                                      #   for --resident (see "High-throughput generation")
│   ├── style_prompts.py                  # TRAINING_PROMPTS, BASE_PROMPTS, PROMPTS, NEGATIVES
│   ├── sweep_cfg.py                      # GPU-only: calibrate guidance x negative before a real run
│   ├── run_baselines.py                  # 12-word EVAL_WORDS smoke test, base model
│   ├── generate_words.py                 # full CSV run, base model or trained LoRA
│   ├── export_word_prompts.py            # dump resolved prompts to TSV for human review
│   ├── train.sh                          # LoRA training entry point (ai-toolkit) — see below
│   ├── setup_datasets.sh                 # prepare -> caption -> augment -> validate -> clear cache
│   ├── prepare_datasets.py               # bake grafiki_v2 PNGs onto white (black = binarized)
│   ├── image_utils.py                    # rgb_on_white / silhouette_on_white
│   ├── write_captions.py                 # per-image .txt captions from TRAINING_PROMPTS
│   ├── augment_datasets.py               # horizontal-flip augmentation (+ caption copies)
│   ├── validate_datasets.py              # white border / ink coverage checks, fails the run
│   ├── clear_latent_cache.py             # drop ai-toolkit latent caches after caption edits
│   ├── sync_training_configs.py          # write lora_runs/configs/krea2_style_{style}.yaml
│   ├── lora_convert.py                   # ai-toolkit LoRA keys -> diffusers keys (see below)
│   ├── eval_checkpoints.py               # render EVAL_WORDS with every saved checkpoint
│   ├── audit_style.py                    # bulk PNG QA (speckle/blob/ground-slab/etc. flags)
│   ├── despeckle.py                      # post-process cleanup for speckled black-style icons
│   ├── gallery.py                        # local HTML viewer for generated_words/
│   ├── generate_prompt_enhancements_llm.py  # per-word depiction/scene review (Gemini/Groq)
│   ├── word_db.py                        # referent DB reader (hints + categories); --sync from GCS
│   ├── translate_from_db.py              # ReferentName_en from the DB's context hints
│   ├── audit_duplicate_depictions.py     # words asked to be drawn as the same picture
│   ├── build_review_set.py / render_review_set.py / review_app.py  # before/after review
│   ├── llm_backend.py                    # shared structured-output LLM plumbing
│   ├── attention_backend.py              # flash-attention backend selection
│   ├── compile_backend.py                # optional torch.compile
│   └── test_translated.v3.csv            # word list + Enhance_*/Scene columns (from source repo)
├── api/                                  # HTTP service (Cloud Run + GPU) — see api/README.md
│   ├── main.py                           # FastAPI app: /v1/icons, /v1/prompt, /healthz, /readyz
│   ├── pipeline.py                       # translate -> expand -> prompt -> render -> upload -> sign
│   ├── translate.py / expand.py          # the two LLM stages, one word per request
│   ├── catalog.py                        # reuse the reviewed CSV instead of re-asking the LLM
│   ├── generator.py                      # one resident pipeline, background warmup, GPU lock
│   ├── storage.py                        # GCS upload + V4 signed URLs (IAM signBlob on Cloud Run)
│   └── prefetch.py                       # download weights for a baked image / a models bucket
├── Dockerfile / Makefile                 # build + deploy the service (`make help`)
├── tests/                                # api/ test suite (`make test`; GPU + GCS stubbed)
├── black/ outline/ real/                 # training datasets, built by setup_datasets.sh (gitignored)
├── lora_runs/configs/                    # generated training configs; runs land in lora_runs/
├── eval/                                 # audit/baseline/sweep output (gitignored, dirs kept)
└── requirements.txt / requirements-api.txt

Prerequisites

  • NVIDIA GPU. Raw's 52 steps is not free — see the throughput note below.
  • HF access for gated models (HF_TOKEN), if Krea 2 requires it on your account.
  • diffusers from git main — Krea2Pipeline is not on a PyPI release yet.
pip install -r requirements.txt

Quickstart

1. Calibrate guidance before doing anything else. Krea 2 Raw's guidance_scale is a real, live knob for the first time in this pipeline's history (Klein never had one) — it has not been tuned against this word list or this prompt shell:

cd scripts
python3 sweep_cfg.py --device cuda:0
# eval/cfg_sweep/black/sweep.png — rows = guidance x negative on/off, columns = EVAL_WORDS

2. Run the 12-word eval baseline once you've picked a guidance value from the sweep:

python3 run_baselines.py --device cuda:0 --guidance 3.5
# eval/baselines/krea2_raw/{black,outline,real}/{word}.png + a grid

3. Full CSV run (1196 words x 4 samples x 3 styles — see the throughput warning below):

python3 generate_words.py --device cuda:0
# generated_words/krea2/baseline_{style}/{rowidx:04d}_{slug}_s{k}.png

4. Audit the results:

python3 audit_style.py

Word sense: the referent database

ReferentName_en used to be derived from the Polish word alone, which cannot resolve a word whose sense isn't recoverable from its spelling. The case that forced this: Polish ring is the boxing ring, the gloss came out ring, and every stage downstream faithfully drew a diamond engagement ring.

gs://lora-review/db.sqlite carries the missing fact — every referent has two patient-facing Polish hints and usually a category:

ring   [Obiekty sportowe]
  To kwadratowa arena otoczona linami, na której odbywają się walki bokserskie.
  Dwóch zawodników spotyka się tam, aby stoczyć sportowy pojedynek.
python3 word_db.py --sync          # gs://lora-review/db.sqlite -> data/db.sqlite
python3 word_db.py ring            # what the database knows about a word
python3 translate_from_db.py --dry-run    # re-derive every gloss from the hints

The hints feed both language stages: translate_from_db.py picks the sense for the gloss, and generate_prompt_enhancements_llm.py pastes the same context under each word when deciding the depiction (it degrades to gloss-only when data/db.sqlite is absent, so nothing breaks on a machine that has never synced it). ring now translates to "boxing ring" and renders as a square ring with corner posts and ropes in all four styles.

The database also covers the collision class audit_duplicate_depictions.py reports: 24 words share 12 English glosses (piec/piekarnik, helikopter/śmigłowiec, gniew/złość), and the hints are what let a translator tell those pairs apart rather than assigning one gloss to both.

HTTP API (Cloud Run)

api/ wraps the same pipeline as a service: one Polish word in, four signed GCS URLs out.

make bootstrap        # APIs, Artifact Registry, buckets, service account, IAM
make sync-weights     # this machine's HF cache -> the models bucket Cloud Run mounts
make build push deploy
make call WORD=zamek
POST /v1/icons  {"word": "zamek", "style": "black"}
  -> translation (Polish -> English gloss)      api/translate.py, or the reviewed CSV
  -> expansion   (gloss -> depiction hint)      api/expand.py, or the reviewed CSV
  -> prompt      style_prompts.base_prompt + NEGATIVES
  -> generation  4 images, one pipe() call per seed
  -> upload      gs://$GCS_BUCKET/icons/<style>/<date>/<word>-<batch>/s{k}.png
  -> V4 signed URLs in the response

It imports style_prompts.py / krea_config.py / model_pipe.py / llm_backend.py / generate_prompt_enhancements_llm.py rather than restating any of them, and it never loads a LoRA (base-model prompting only, same as generate_words.py's default).

Two things to know before deploying, both covered in api/README.md:

  • Krea 2 does not fit on Cloud Run's GPU unquantized. The only accelerator on offer is one L4 (24 GB) and the transformer alone is 26.3 GB in bf16 — enable_model_cpu_offload() can't rescue that, since it keeps the running component fully resident. The service defaults to QUANTIZE=nf4 (bitsandbytes), which puts the transformer at ~7 GB and the text encoder at ~2.5 GB. That is lossy and unaudited: every speckle/blob number in style_prompts.py was measured in bf16 on an A100, so re-run audit_style.py against a deployed batch before trusting the quantized output to match.
  • Signed URLs need roles/iam.serviceAccountTokenCreator on the service account itself. Cloud Run's runtime credentials carry no private key, so signing goes through the IAM signBlob API. make iam grants it; without it every request fails after the render.

Settings defaults to the raw variant (52 steps, live negatives); the Makefile deploys turbo (8 steps, negatives silently inert — see krea_config.py) because a synchronous HTTP request on an L4 is a different tradeoff from a batch job. Cloud Run's request timeout caps at 3600s.

High-throughput generation (--resident)

--cpu-offload (pipe.enable_model_cpu_offload()) trades VRAM for a PCIe transfer of the entire 26.3 GB transformer on every pipe() call. On --model turbo that transfer is most of the wall clock — Turbo only runs 8 denoising steps, so the compute is cheap and the transfer dominates. Measured on an A100-SXM4-40GB (39.5 GB usable): 52.5 s/image with --cpu-offload on Turbo. A full black-style run (1196 words x 4 samples = 4784 images) at that rate is 2.9 days.

--resident inverts the tradeoff: it pins the 26.3 GB transformer + 0.5 GB vae on the GPU permanently (25.9 GB, ~13.6 GB headroom) and only ever moves the 8.9 GB text_encoder — to GPU to encode a chunk of prompts (--embed-chunk, default 64), then back to CPU before any pipe() forward runs. That amortises the encoder's round-trip to near-zero per image instead of paying the transformer's much larger round-trip on every single one. Measured: 17.6–17.8 s/image on Turbo (~3x faster) — the same 4784-image black run drops from 2.9 days to ~23 hours.

python3 generate_words.py --model turbo --resident --device cuda:0
# same output layout as the default path:
# generated_words/krea2/baseline_{style}/{rowidx:04d}_{slug}_s{k}.png

Notes / constraints:

  • Batch size is forced to 1 image per pipe() call, on both paths. num_images_per_prompt

    = 2 OOMs at 1024px on this card — activations run ~10 GB/image on top of the 25.9 GB resident transformer+vae, and there's only ~13.6 GB of headroom. model_pipe.generate_from_embeds() loops over seeds one image at a time rather than batching.

  • --resident is Turbo-oriented. On --model raw, CFG runs two sequential forwards per step (not batched) and 52 steps x 2 forwards already dominates the wall clock, so pinning the transformer only saves ~11% there — real, but nowhere near Turbo's ~3x. raw isn't blocked, it's just not the point of the flag.
  • Negatives are encoded once per style (not per word) and skipped entirely on a variant/guidance combo where CFG doesn't run (e.g. Turbo's guidance_scale=0.0) — encoding a negative there would cost a real text_encoder round-trip for zero effect.
  • --resident and --use-lora/--loras cannot be combined yet — fuse_lora() against a resident pipe has never been exercised (there is still no trained Krea 2 LoRA to test it against); generate_words.py raises a clear SystemExit rather than risk a silently-wrong fuse.
  • --resident and --cpu-offload are mutually exclusive (pick one VRAM strategy) — generate_words.py raises a clear SystemExit if both are passed.
  • Implementation: scripts/model_pipe.py's load_pipe_resident() / encode_prompts() / generate_from_embeds(). load_pipe_resident() also has to pin pipe.device / pipe._execution_device to the GPU via a one-off subclass — DiffusionPipeline.device otherwise resolves to the device of the alphabetically-first nn.Module component (text_encoder for Krea2Pipeline), which is CPU here, and pipe() would try to build latents on CPU while the transformer runs on GPU.

LoRA training

Ported from grafiki-flux-v2's training setup (train.sh / sync_training_configs.py / setup_datasets.sh and friends) onto Krea 2 Raw. Same trainer (ostris ai-toolkit), same tiny hand-built datasets, same per-style hyperparameters; what changed is documented under "Differences from the FLUX.2 Klein setup" below.

Prerequisites

  • An ai-toolkit checkout with Krea 2 support. krea_config.MODEL_ARCH = "krea2" resolves to extensions_built_in/diffusion_models/krea2/ upstream, which landed well after the Klein work — the checkout grafiki-flux-v2 used almost certainly predates it. train.sh refuses to start if the directory is missing and tells you to update:

    git -C ~/grafiki/zimage_lora_pipeline/ai-toolkit pull
    

    Point elsewhere with TOOLKIT=/path/to/ai-toolkit ./train.sh if you'd rather not move the checkout grafiki-flux-v2 shares.

  • ~/grafiki_v2/{black,outline,real} — the source PNGs (krea_config.V2_ROOT, override with V2_ROOT=). Nothing trains straight out of that directory; it gets baked into this repo first.

  • python-dotenv in whatever interpreter runs ai-toolkit (train.sh checks, and takes PYTHON=/path/to/venv/bin/python if the toolkit has its own environment — the default is plain python3, which is what the Klein LoRAs were trained under).

  • Training downloads krea/Krea-2-Raw's raw.safetensors plus Qwen/Qwen3-VL-4B-Instruct and the Qwen/Qwen-Image VAE separately — ai-toolkit loads the three components itself and does not reuse whatever Krea2Pipeline.from_pretrained() already cached for inference.

Run it

cd scripts
./train.sh                 # black, then outline, then real
./train.sh black           # one style

train.sh runs setup_datasets.sh and sync_training_configs.py first, every time, so the datasets, captions and configs on disk always match the code that generated them. Checkpoints land in lora_runs/krea2_style_{style}/krea2_style_{style}_{step:09d}.safetensors (every 50 steps, latest 12 kept) with training sample images beside them in samples/.

What setup_datasets.sh does, in order — each step is also runnable on its own:

script what it does
prepare_datasets.py bakes ~/grafiki_v2/{style}/*.png into {style}/ — black is binarized to a hard silhouette on white, outline/real are just composited onto white. Excludes kolo_black (inverted source).
write_captions.py writes one .txt per image from style_prompts.TRAINING_PROMPTS, so every caption carries its kreablksty / kreaoutsty / kreaphotsty trigger
augment_datasets.py adds a horizontal flip of every image, sharing the base caption → 34 / 40 / 50 images
validate_datasets.py white-border, ink-coverage and all-black checks; exits non-zero and stops the run on a failure
clear_latent_cache.py drops ai-toolkit's cached latents, which are keyed on filename and would otherwise survive a caption edit

After training

python3 eval_checkpoints.py --style black --device cuda:0
# eval/ckpt/krea2/black/grids/{word}.png — one row per word, one column per saved checkpoint

Pick the best step per style off those grids and write it into krea_config.LORA_INFER_STEP (the values there now are placeholders inherited from the Klein repo). Then:

python3 generate_words.py --use-lora black --device cuda:0

--use-lora auto-discovers the checkpoint nearest LORA_INFER_STEP; --loras black=/path.safetensors overrides it explicitly, and --lora-scale (default krea_config.DEFAULT_LORA_SCALE = 0.45) sets the fuse strength. Note that --use-lora still can't be combined with --resident — see the constraint list under "High-throughput generation".

Differences from the FLUX.2 Klein setup

Three of these are traps rather than preferences; they're the reason the two repos' training configs are not interchangeable.

  • Sample guidance is offset by +1. ai-toolkit's Krea2Model.generate_single_image() normalizes the value itself (guidance = max(0.0, guidance_scale - 1.0)) before calling its own pipeline, while diffusers.Krea2Pipeline takes the 0-normalized number directly — which is what krea_config.VARIANTS["raw"].guidance = 3.5 is. So sync_training_configs.py writes guidance_scale: 4.5 into the YAML. Copying the 3.5 straight across would silently render every training sample at an effective 2.5 and make the LoRA look weaker than it is.
  • lora_convert.py exists, and inference depends on it. Training and inference run two different implementations of the same model: ai-toolkit's SingleStreamDiT and diffusers' Krea2Transformer2DModel. ai-toolkit saves LoRA keys against its own module names (diffusion_model.blocks.0.attn.wq.…), and unlike the FLUX path, Krea2LoraLoaderMixin ships no converter — an unconverted checkpoint resolves nothing and silently generates base-model output. model_pipe.load_lora() therefore rewrites the keys in memory (transformer.transformer_blocks.0.attn.to_q.…) before handing them to peft. The file on disk stays in ai-toolkit form, which is what ComfyUI and ai-toolkit's own tooling read. python3 lora_convert.py ckpt.safetensors -o converted.safetensors exports a diffusers-native copy if you need one; the mapping is verified module-by-module against both implementations.
  • Training samples finally use a negative prompt. Klein's configs carried neg: "" because CFG was hard-disabled there. Krea 2 Raw restores it, so samples render through style_prompts.negative_for(style) — the same negatives generate_words.py uses.
  • Samples are rendered every 100 steps, checkpoints still every 50. One Krea 2 Raw sample is ~5x the cost of a Klein one (52 steps and a second CFG forward per step, vs. 20 steps and no CFG), and four of them every 50 steps would eat a real fraction of a 300-step run.
  • The model is quantized during training (quantize + quantize_te, qfloat8). Krea 2's denoiser is ~26 GB in bf16 — with 8-bit Adam state, activations and the Qwen3-VL text encoder there is no room on a 40 GB card otherwise. low_vram (block swapping) stays off; it trades a lot of speed for VRAM that quantization already freed.
  • Rank/alpha are 32, not 8 (krea_config.LORA_RANK/LORA_ALPHA) — the source repo underfit at rank 8 on this dataset size. Keep the two equal: ai-toolkit writes no alpha entries, so peft falls back to alpha == rank, and a mismatch would make the fused strength differ from what training saw.

Step counts, LR (2e-5), caption dropout, per-style resolutions, adamw8bit, EMA at 0.99 and the flowmatch/weighted-timestep settings are all carried over unchanged — see the comments on krea_config.TRAIN_STEPS, which are a starting point to re-cut against eval_checkpoints.py, not tuned Krea 2 values.

What is NOT yet done

  • No Krea 2 LoRA has been trained yet — the pipeline to train one now exists (see "LoRA training" below), but nothing has been run through it. lora_runs/ is empty, so generate_words.py still defaults to base-model + prompts, and krea_config.LORA_INFER_STEP is still placeholder values carried over from the FLUX.2 Klein repo. Re-cut those from eval_checkpoints.py grids after a real run.
  • krea_config.LORA_RANK = LORA_ALPHA = 32 is untested at that value. The source repo used rank 8 on the same ~34-image dataset and it underfit; 32 is a deliberate increase, not a carried-over default, but nobody has seen what it does on Krea 2 yet.
  • The rebalanced prompt shell (style_prompts.py) is unmeasured on Krea 2. Every speckle rate, blob rate, and failure-mode number in that file's comments was measured on FLUX.2 Klein. Moving ban clauses from the positive shell into NEGATIVES is a reasoned port, not a validated one — run audit_style.py against a real Krea 2 Raw batch and compare before trusting the current balance, especially for the black style's speckle rate.
  • Throughput: Raw's 52 steps/image is ~2.6x Klein's 20 steps/image (Turbo's 8 steps is ~0.4x, but negatives don't work there). A full 1196-word x 4-sample x 3-style run is a large job even before that multiplier — calibrate on sweep_cfg.py and run_baselines.py's 12-word set first, and consider --cpu-offload if VRAM is tight rather than reaching for Turbo by default, since Turbo silently drops your negative prompts. If VRAM is not tight (a free 40 GB-class card), prefer --resident over --cpu-offload on Turbo runs — see "High-throughput generation" above: 52.5 -> 17.6 s/image measured, turning the 4784-image black run from 2.9 days into ~23 hours.

Notes

  • style_prompts.NEGATIVES now matters. It was dead weight under Klein — kept only as documentation of "what wrong looks like" per style. On Krea 2 Raw it is genuinely subtracted from the prediction at every denoising step. Read the migration comment directly above NEGATIVES in style_prompts.py for exactly which clauses moved out of the positive shell and why, and negative_for(style) to fetch it.
  • model_pipe.load_pipe(device, *, variant="raw", cpu_offload=False) resolves the model id, step count, and guidance default from krea_config.VARIANTS[variant] rather than a single flat module constant — Krea 2 ships two checkpoints with genuinely different behavior, unlike Klein's one.
  • model_pipe.load_pipe_resident(device, *, variant) is the --resident counterpart — see "High-throughput generation" above. Pair it with encode_prompts() (text_encoder to GPU, encode a chunk, back to CPU) and generate_from_embeds() (prompt_embeds/prompt_embeds_mask instead of a prompt string), not generate()/generate_batch().
  • generate_words.py / run_baselines.py / sweep_cfg.py all take --model {raw,turbo} and --negative/--no-negative (default: on); passing negatives with --model turbo prints a warning (from model_pipe.py, once per process) since they're inert there.
  • Trigger tokens are kreablksty / kreaoutsty / kreaphotsty (renamed from the source repo's flxblksty / flxoutsty / flxphotsty so a future Krea 2 LoRA's vocabulary can't collide with, or be confused for, the old FLUX.2 Klein LoRA's).
  • ATTENTION_BACKEND env var selects the flash-attention variant (auto-detects Hopper for FA3); TORCH_COMPILE=1 opts into torch.compile on the transformer — both ported unchanged from the source repo, neither is Klein-specific.
  • generate_prompt_enhancements_llm.py / llm_backend.py / test_translated.v3.csv are carried over verbatim — the Polish-language word list and its per-word depiction/scene review don't depend on which diffusion model renders them.