- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| eval | ||
| lora_runs/configs | ||
| scripts | ||
| .gitignore | ||
| README.md | ||
| requirements.txt | ||
grafiki-krea-2 — Krea 2 icon generation
Port of grafiki-flux-v2 (Polish speech-therapy flashcard icons — black
silhouette / outline / photorealistic styles, ~1196 words x 4 samples each) from
Flux2KleinPipeline (FLUX.2 Klein 4B) to Krea2Pipeline (Krea 2).
Why this port exists
FLUX.2 Klein is a distilled model: classifier-free guidance is hard-disabled, so
negative_prompt never worked, on inference or on training samples. The entire prompt
architecture in scripts/style_prompts.py — every clause phrased as a positive assertion
instead of a negative, the narrow "structural openings" vs. "dots, speckles, stippling" split,
the double restatement of "white background" — is a workaround for that one missing feature.
The source repo's comments document three separate measured failures this caused (an
"interior detail" clause that raised the interior-speckle rate instead of lowering it; an
absolute no-interior-white ban that killed a pupil, shelf divisions, and a bowl's steam because
it couldn't distinguish structure from decoration; a shell-wide animal-eye grant that put eyes
on glasses of water and hearts) — all traceable to having no real negative channel to put a ban
in instead.
Krea 2 removes that constraint. krea/Krea-2-Raw is undistilled and restores true CFG, so
negative_prompt is live. This port moves the ban half of each of those three workarounds out
of the positive prompt shell and into NEGATIVES, while preserving every measured-history
comment from the source file (they're the receipts for why the shell is shaped the way it
is — read scripts/style_prompts.py before changing it further).
Raw vs. Turbo
Krea 2 ships as two materially different checkpoints (scripts/krea_config.py's VARIANTS):
| variant | model id | steps | guidance | negative_prompt |
|---|---|---|---|---|
raw (default) |
krea/Krea-2-Raw |
52 | 3.5 | live — real CFG |
turbo |
krea/Krea-2-Turbo |
8 | 0.0 | silently inert |
diffusers.Krea2Pipeline only runs its negative-branch when guidance_scale > 0 — the
docstring says negative_prompt is "Ignored when guidance_scale <= 0". Turbo's guidance is
exactly 0, so passing a negative prompt there compiles, runs, and produces images — it just has
zero effect, silently. model_pipe.py warns once (not an error — Turbo's speed may be worth
it for a quick preview) whenever a non-empty negative_prompt meets guidance_scale <= 0.
raw is the default variant for exactly this reason: quality + working negatives is the whole
point of this migration.
License: Krea 2 ships under the "Krea 2 Community License" — commercial use is fine; enterprises over 50 seats need a separate agreement with Krea. Check the license text on the model card before a commercial deployment.
Layout
grafiki-krea-2/
├── scripts/
│ ├── krea_config.py # ROOT, STYLES, VARIANTS (raw/turbo), LoRA paths
│ ├── model_pipe.py # load pipe, load/fuse LoRA, generate (+ negative_prompt);
│ │ # load_pipe_resident/encode_prompts/generate_from_embeds
│ │ # for --resident (see "High-throughput generation")
│ ├── style_prompts.py # TRAINING_PROMPTS, BASE_PROMPTS, PROMPTS, NEGATIVES
│ ├── sweep_cfg.py # GPU-only: calibrate guidance x negative before a real run
│ ├── run_baselines.py # 12-word EVAL_WORDS smoke test, base model
│ ├── generate_words.py # full CSV run, base model or trained LoRA
│ ├── export_word_prompts.py # dump resolved prompts to TSV for human review
│ ├── train.sh # LoRA training entry point (ai-toolkit) — see below
│ ├── setup_datasets.sh # prepare -> caption -> augment -> validate -> clear cache
│ ├── prepare_datasets.py # bake grafiki_v2 PNGs onto white (black = binarized)
│ ├── image_utils.py # rgb_on_white / silhouette_on_white
│ ├── write_captions.py # per-image .txt captions from TRAINING_PROMPTS
│ ├── augment_datasets.py # horizontal-flip augmentation (+ caption copies)
│ ├── validate_datasets.py # white border / ink coverage checks, fails the run
│ ├── clear_latent_cache.py # drop ai-toolkit latent caches after caption edits
│ ├── sync_training_configs.py # write lora_runs/configs/krea2_style_{style}.yaml
│ ├── lora_convert.py # ai-toolkit LoRA keys -> diffusers keys (see below)
│ ├── eval_checkpoints.py # render EVAL_WORDS with every saved checkpoint
│ ├── audit_style.py # bulk PNG QA (speckle/blob/ground-slab/etc. flags)
│ ├── despeckle.py # post-process cleanup for speckled black-style icons
│ ├── gallery.py # local HTML viewer for generated_words/
│ ├── generate_prompt_enhancements_llm.py # per-word depiction/scene review (Gemini/Groq)
│ ├── word_db.py # referent DB reader (hints + categories); --sync from GCS
│ ├── translate_from_db.py # ReferentName_en from the DB's context hints
│ ├── audit_duplicate_depictions.py # words asked to be drawn as the same picture
│ ├── build_review_set.py / render_review_set.py / review_app.py # before/after review
│ ├── llm_backend.py # shared structured-output LLM plumbing
│ ├── attention_backend.py # flash-attention backend selection
│ ├── compile_backend.py # optional torch.compile
│ └── test_translated.v3.csv # word list + Enhance_*/Scene columns (from source repo)
├── api/ # HTTP service (Cloud Run + GPU) — see api/README.md
│ ├── main.py # FastAPI app: /v1/icons, /v1/prompt, /healthz, /readyz
│ ├── pipeline.py # translate -> expand -> prompt -> render -> upload -> sign
│ ├── translate.py / expand.py # the two LLM stages, one word per request
│ ├── catalog.py # reuse the reviewed CSV instead of re-asking the LLM
│ ├── generator.py # one resident pipeline, background warmup, GPU lock
│ ├── storage.py # GCS upload + V4 signed URLs (IAM signBlob on Cloud Run)
│ └── prefetch.py # download weights for a baked image / a models bucket
├── Dockerfile / Makefile # build + deploy the service (`make help`)
├── tests/ # api/ test suite (`make test`; GPU + GCS stubbed)
├── black/ outline/ real/ # training datasets, built by setup_datasets.sh (gitignored)
├── lora_runs/configs/ # generated training configs; runs land in lora_runs/
├── eval/ # audit/baseline/sweep output (gitignored, dirs kept)
└── requirements.txt / requirements-api.txt
Prerequisites
- NVIDIA GPU. Raw's 52 steps is not free — see the throughput note below.
- HF access for gated models (
HF_TOKEN), if Krea 2 requires it on your account. diffusersfrom git main —Krea2Pipelineis not on a PyPI release yet.
pip install -r requirements.txt
Quickstart
1. Calibrate guidance before doing anything else. Krea 2 Raw's guidance_scale is a real,
live knob for the first time in this pipeline's history (Klein never had one) — it has not been
tuned against this word list or this prompt shell:
cd scripts
python3 sweep_cfg.py --device cuda:0
# eval/cfg_sweep/black/sweep.png — rows = guidance x negative on/off, columns = EVAL_WORDS
2. Run the 12-word eval baseline once you've picked a guidance value from the sweep:
python3 run_baselines.py --device cuda:0 --guidance 3.5
# eval/baselines/krea2_raw/{black,outline,real}/{word}.png + a grid
3. Full CSV run (1196 words x 4 samples x 3 styles — see the throughput warning below):
python3 generate_words.py --device cuda:0
# generated_words/krea2/baseline_{style}/{rowidx:04d}_{slug}_s{k}.png
4. Audit the results:
python3 audit_style.py
Word sense: the referent database
ReferentName_en used to be derived from the Polish word alone, which cannot resolve a word
whose sense isn't recoverable from its spelling. The case that forced this: Polish ring is the
boxing ring, the gloss came out ring, and every stage downstream faithfully drew a diamond
engagement ring.
gs://lora-review/db.sqlite carries the missing fact — every referent has two patient-facing
Polish hints and usually a category:
ring [Obiekty sportowe]
To kwadratowa arena otoczona linami, na której odbywają się walki bokserskie.
Dwóch zawodników spotyka się tam, aby stoczyć sportowy pojedynek.
python3 word_db.py --sync # gs://lora-review/db.sqlite -> data/db.sqlite
python3 word_db.py ring # what the database knows about a word
python3 translate_from_db.py --dry-run # re-derive every gloss from the hints
The hints feed both language stages: translate_from_db.py picks the sense for the gloss, and
generate_prompt_enhancements_llm.py pastes the same context under each word when deciding the
depiction (it degrades to gloss-only when data/db.sqlite is absent, so nothing breaks on a
machine that has never synced it). ring now translates to "boxing ring" and renders as a square
ring with corner posts and ropes in all four styles.
The database also covers the collision class audit_duplicate_depictions.py reports: 24 words
share 12 English glosses (piec/piekarnik, helikopter/śmigłowiec, gniew/złość), and the
hints are what let a translator tell those pairs apart rather than assigning one gloss to both.
HTTP API (Cloud Run)
api/ wraps the same pipeline as a service: one Polish word in, four signed GCS URLs out.
make bootstrap # APIs, Artifact Registry, buckets, service account, IAM
make sync-weights # this machine's HF cache -> the models bucket Cloud Run mounts
make build push deploy
make call WORD=zamek
POST /v1/icons {"word": "zamek", "style": "black"}
-> translation (Polish -> English gloss) api/translate.py, or the reviewed CSV
-> expansion (gloss -> depiction hint) api/expand.py, or the reviewed CSV
-> prompt style_prompts.base_prompt + NEGATIVES
-> generation 4 images, one pipe() call per seed
-> upload gs://$GCS_BUCKET/icons/<style>/<date>/<word>-<batch>/s{k}.png
-> V4 signed URLs in the response
It imports style_prompts.py / krea_config.py / model_pipe.py / llm_backend.py /
generate_prompt_enhancements_llm.py rather than restating any of them, and it never loads a
LoRA (base-model prompting only, same as generate_words.py's default).
Two things to know before deploying, both covered in api/README.md:
- Krea 2 does not fit on Cloud Run's GPU unquantized. The only accelerator on offer is one
L4 (24 GB) and the transformer alone is 26.3 GB in bf16 —
enable_model_cpu_offload()can't rescue that, since it keeps the running component fully resident. The service defaults toQUANTIZE=nf4(bitsandbytes), which puts the transformer at ~7 GB and the text encoder at ~2.5 GB. That is lossy and unaudited: every speckle/blob number instyle_prompts.pywas measured in bf16 on an A100, so re-runaudit_style.pyagainst a deployed batch before trusting the quantized output to match. - Signed URLs need
roles/iam.serviceAccountTokenCreatoron the service account itself. Cloud Run's runtime credentials carry no private key, so signing goes through the IAMsignBlobAPI.make iamgrants it; without it every request fails after the render.
Settings defaults to the raw variant (52 steps, live negatives); the Makefile deploys turbo
(8 steps, negatives silently inert — see krea_config.py) because a synchronous HTTP request on
an L4 is a different tradeoff from a batch job. Cloud Run's request timeout caps at 3600s.
High-throughput generation (--resident)
--cpu-offload (pipe.enable_model_cpu_offload()) trades VRAM for a PCIe transfer of the
entire 26.3 GB transformer on every pipe() call. On --model turbo that transfer is most
of the wall clock — Turbo only runs 8 denoising steps, so the compute is cheap and the transfer
dominates. Measured on an A100-SXM4-40GB (39.5 GB usable): 52.5 s/image with --cpu-offload
on Turbo. A full black-style run (1196 words x 4 samples = 4784 images) at that rate is 2.9
days.
--resident inverts the tradeoff: it pins the 26.3 GB transformer + 0.5 GB vae on the GPU
permanently (25.9 GB, ~13.6 GB headroom) and only ever moves the 8.9 GB text_encoder — to
GPU to encode a chunk of prompts (--embed-chunk, default 64), then back to CPU before any
pipe() forward runs. That amortises the encoder's round-trip to near-zero per image instead of
paying the transformer's much larger round-trip on every single one. Measured: 17.6–17.8
s/image on Turbo (~3x faster) — the same 4784-image black run drops from 2.9 days to ~23
hours.
python3 generate_words.py --model turbo --resident --device cuda:0
# same output layout as the default path:
# generated_words/krea2/baseline_{style}/{rowidx:04d}_{slug}_s{k}.png
Notes / constraints:
- Batch size is forced to 1 image per
pipe()call, on both paths.num_images_per_prompt= 2 OOMs at 1024px on this card — activations run ~10 GB/image on top of the 25.9 GB resident transformer+vae, and there's only ~13.6 GB of headroom.
model_pipe.generate_from_embeds()loops over seeds one image at a time rather than batching. --residentis Turbo-oriented. On--model raw, CFG runs two sequential forwards per step (not batched) and 52 steps x 2 forwards already dominates the wall clock, so pinning the transformer only saves ~11% there — real, but nowhere near Turbo's ~3x.rawisn't blocked, it's just not the point of the flag.- Negatives are encoded once per style (not per word) and skipped entirely on a
variant/guidance combo where CFG doesn't run (e.g. Turbo's
guidance_scale=0.0) — encoding a negative there would cost a realtext_encoderround-trip for zero effect. --residentand--use-lora/--lorascannot be combined yet —fuse_lora()against a resident pipe has never been exercised (there is still no trained Krea 2 LoRA to test it against);generate_words.pyraises a clearSystemExitrather than risk a silently-wrong fuse.--residentand--cpu-offloadare mutually exclusive (pick one VRAM strategy) —generate_words.pyraises a clearSystemExitif both are passed.- Implementation:
scripts/model_pipe.py'sload_pipe_resident()/encode_prompts()/generate_from_embeds().load_pipe_resident()also has to pinpipe.device/pipe._execution_deviceto the GPU via a one-off subclass —DiffusionPipeline.deviceotherwise resolves to the device of the alphabetically-firstnn.Modulecomponent (text_encoderforKrea2Pipeline), which is CPU here, andpipe()would try to build latents on CPU while the transformer runs on GPU.
LoRA training
Ported from grafiki-flux-v2's training setup (train.sh / sync_training_configs.py /
setup_datasets.sh and friends) onto Krea 2 Raw. Same trainer (ostris ai-toolkit), same
tiny hand-built datasets, same per-style hyperparameters; what changed is documented under
"Differences from the FLUX.2 Klein setup" below.
Prerequisites
-
An ai-toolkit checkout with Krea 2 support.
krea_config.MODEL_ARCH = "krea2"resolves toextensions_built_in/diffusion_models/krea2/upstream, which landed well after the Klein work — the checkoutgrafiki-flux-v2used almost certainly predates it.train.shrefuses to start if the directory is missing and tells you to update:git -C ~/grafiki/zimage_lora_pipeline/ai-toolkit pullPoint elsewhere with
TOOLKIT=/path/to/ai-toolkit ./train.shif you'd rather not move the checkoutgrafiki-flux-v2shares. -
~/grafiki_v2/{black,outline,real}— the source PNGs (krea_config.V2_ROOT, override withV2_ROOT=). Nothing trains straight out of that directory; it gets baked into this repo first. -
python-dotenvin whatever interpreter runs ai-toolkit (train.shchecks, and takesPYTHON=/path/to/venv/bin/pythonif the toolkit has its own environment — the default is plainpython3, which is what the Klein LoRAs were trained under). -
Training downloads
krea/Krea-2-Raw'sraw.safetensorsplusQwen/Qwen3-VL-4B-Instructand theQwen/Qwen-ImageVAE separately — ai-toolkit loads the three components itself and does not reuse whateverKrea2Pipeline.from_pretrained()already cached for inference.
Run it
cd scripts
./train.sh # black, then outline, then real
./train.sh black # one style
train.sh runs setup_datasets.sh and sync_training_configs.py first, every time, so the
datasets, captions and configs on disk always match the code that generated them. Checkpoints
land in lora_runs/krea2_style_{style}/krea2_style_{style}_{step:09d}.safetensors (every 50
steps, latest 12 kept) with training sample images beside them in samples/.
What setup_datasets.sh does, in order — each step is also runnable on its own:
| script | what it does |
|---|---|
prepare_datasets.py |
bakes ~/grafiki_v2/{style}/*.png into {style}/ — black is binarized to a hard silhouette on white, outline/real are just composited onto white. Excludes kolo_black (inverted source). |
write_captions.py |
writes one .txt per image from style_prompts.TRAINING_PROMPTS, so every caption carries its kreablksty / kreaoutsty / kreaphotsty trigger |
augment_datasets.py |
adds a horizontal flip of every image, sharing the base caption → 34 / 40 / 50 images |
validate_datasets.py |
white-border, ink-coverage and all-black checks; exits non-zero and stops the run on a failure |
clear_latent_cache.py |
drops ai-toolkit's cached latents, which are keyed on filename and would otherwise survive a caption edit |
After training
python3 eval_checkpoints.py --style black --device cuda:0
# eval/ckpt/krea2/black/grids/{word}.png — one row per word, one column per saved checkpoint
Pick the best step per style off those grids and write it into krea_config.LORA_INFER_STEP
(the values there now are placeholders inherited from the Klein repo). Then:
python3 generate_words.py --use-lora black --device cuda:0
--use-lora auto-discovers the checkpoint nearest LORA_INFER_STEP; --loras black=/path.safetensors
overrides it explicitly, and --lora-scale (default krea_config.DEFAULT_LORA_SCALE = 0.45)
sets the fuse strength. Note that --use-lora still can't be combined with --resident — see
the constraint list under "High-throughput generation".
Differences from the FLUX.2 Klein setup
Three of these are traps rather than preferences; they're the reason the two repos' training configs are not interchangeable.
- Sample guidance is offset by +1. ai-toolkit's
Krea2Model.generate_single_image()normalizes the value itself (guidance = max(0.0, guidance_scale - 1.0)) before calling its own pipeline, whilediffusers.Krea2Pipelinetakes the 0-normalized number directly — which is whatkrea_config.VARIANTS["raw"].guidance = 3.5is. Sosync_training_configs.pywritesguidance_scale: 4.5into the YAML. Copying the 3.5 straight across would silently render every training sample at an effective 2.5 and make the LoRA look weaker than it is. lora_convert.pyexists, and inference depends on it. Training and inference run two different implementations of the same model: ai-toolkit'sSingleStreamDiTand diffusers'Krea2Transformer2DModel. ai-toolkit saves LoRA keys against its own module names (diffusion_model.blocks.0.attn.wq.…), and unlike the FLUX path,Krea2LoraLoaderMixinships no converter — an unconverted checkpoint resolves nothing and silently generates base-model output.model_pipe.load_lora()therefore rewrites the keys in memory (transformer.transformer_blocks.0.attn.to_q.…) before handing them to peft. The file on disk stays in ai-toolkit form, which is what ComfyUI and ai-toolkit's own tooling read.python3 lora_convert.py ckpt.safetensors -o converted.safetensorsexports a diffusers-native copy if you need one; the mapping is verified module-by-module against both implementations.- Training samples finally use a negative prompt. Klein's configs carried
neg: ""because CFG was hard-disabled there. Krea 2 Raw restores it, so samples render throughstyle_prompts.negative_for(style)— the same negativesgenerate_words.pyuses. - Samples are rendered every 100 steps, checkpoints still every 50. One Krea 2 Raw sample is ~5x the cost of a Klein one (52 steps and a second CFG forward per step, vs. 20 steps and no CFG), and four of them every 50 steps would eat a real fraction of a 300-step run.
- The model is quantized during training (
quantize+quantize_te, qfloat8). Krea 2's denoiser is ~26 GB in bf16 — with 8-bit Adam state, activations and the Qwen3-VL text encoder there is no room on a 40 GB card otherwise.low_vram(block swapping) stays off; it trades a lot of speed for VRAM that quantization already freed. - Rank/alpha are 32, not 8 (
krea_config.LORA_RANK/LORA_ALPHA) — the source repo underfit at rank 8 on this dataset size. Keep the two equal: ai-toolkit writes noalphaentries, so peft falls back to alpha == rank, and a mismatch would make the fused strength differ from what training saw.
Step counts, LR (2e-5), caption dropout, per-style resolutions, adamw8bit, EMA at 0.99 and
the flowmatch/weighted-timestep settings are all carried over unchanged — see the comments on
krea_config.TRAIN_STEPS, which are a starting point to re-cut against eval_checkpoints.py,
not tuned Krea 2 values.
What is NOT yet done
- No Krea 2 LoRA has been trained yet — the pipeline to train one now exists (see "LoRA
training" below), but nothing has been run through it.
lora_runs/is empty, sogenerate_words.pystill defaults to base-model + prompts, andkrea_config.LORA_INFER_STEPis still placeholder values carried over from the FLUX.2 Klein repo. Re-cut those fromeval_checkpoints.pygrids after a real run. krea_config.LORA_RANK = LORA_ALPHA = 32is untested at that value. The source repo used rank 8 on the same ~34-image dataset and it underfit; 32 is a deliberate increase, not a carried-over default, but nobody has seen what it does on Krea 2 yet.- The rebalanced prompt shell (
style_prompts.py) is unmeasured on Krea 2. Every speckle rate, blob rate, and failure-mode number in that file's comments was measured on FLUX.2 Klein. Moving ban clauses from the positive shell intoNEGATIVESis a reasoned port, not a validated one — runaudit_style.pyagainst a real Krea 2 Raw batch and compare before trusting the current balance, especially for the black style's speckle rate. - Throughput: Raw's 52 steps/image is ~2.6x Klein's 20 steps/image (Turbo's 8 steps is
~0.4x, but negatives don't work there). A full 1196-word x 4-sample x 3-style run is a large
job even before that multiplier — calibrate on
sweep_cfg.pyandrun_baselines.py's 12-word set first, and consider--cpu-offloadif VRAM is tight rather than reaching for Turbo by default, since Turbo silently drops your negative prompts. If VRAM is not tight (a free 40 GB-class card), prefer--residentover--cpu-offloadon Turbo runs — see "High-throughput generation" above: 52.5 -> 17.6 s/image measured, turning the 4784-image black run from 2.9 days into ~23 hours.
Notes
style_prompts.NEGATIVESnow matters. It was dead weight under Klein — kept only as documentation of "what wrong looks like" per style. On Krea 2 Raw it is genuinely subtracted from the prediction at every denoising step. Read the migration comment directly aboveNEGATIVESinstyle_prompts.pyfor exactly which clauses moved out of the positive shell and why, andnegative_for(style)to fetch it.model_pipe.load_pipe(device, *, variant="raw", cpu_offload=False)resolves the model id, step count, and guidance default fromkrea_config.VARIANTS[variant]rather than a single flat module constant — Krea 2 ships two checkpoints with genuinely different behavior, unlike Klein's one.model_pipe.load_pipe_resident(device, *, variant)is the--residentcounterpart — see "High-throughput generation" above. Pair it withencode_prompts()(text_encoder to GPU, encode a chunk, back to CPU) andgenerate_from_embeds()(prompt_embeds/prompt_embeds_maskinstead of apromptstring), notgenerate()/generate_batch().generate_words.py/run_baselines.py/sweep_cfg.pyall take--model {raw,turbo}and--negative/--no-negative(default: on); passing negatives with--model turboprints a warning (frommodel_pipe.py, once per process) since they're inert there.- Trigger tokens are
kreablksty/kreaoutsty/kreaphotsty(renamed from the source repo'sflxblksty/flxoutsty/flxphotstyso a future Krea 2 LoRA's vocabulary can't collide with, or be confused for, the old FLUX.2 Klein LoRA's). ATTENTION_BACKENDenv var selects the flash-attention variant (auto-detects Hopper for FA3);TORCH_COMPILE=1opts intotorch.compileon the transformer — both ported unchanged from the source repo, neither is Klein-specific.generate_prompt_enhancements_llm.py/llm_backend.py/test_translated.v3.csvare carried over verbatim — the Polish-language word list and its per-word depiction/scene review don't depend on which diffusion model renders them.