- Python 90.3%
- Shell 9.7%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| eval | ||
| results_svg | ||
| scripts | ||
| .gitignore | ||
| api_models.json | ||
| API_MODELS.md | ||
| benchmark.log | ||
| benchmark_new_models.log | ||
| benchmark_retry.log | ||
| EDGE_CASES.md | ||
| manifest.json | ||
| manifest_svg.json | ||
| MODELS.md | ||
| PIPELINE.md | ||
| README.md | ||
| SOTA.md | ||
| svg_models.json | ||
| SVG_MODELS.md | ||
grafiki-mixed — multi-model AAC icon benchmark
Compare commercial-friendly (and reference NC) image models against our production baseline
FLUX.2 Klein Base 4B (~/grafiki-flux-v2), using the same v1 prompt shell + Polish prompt-enhancement pipeline.
| Doc | Contents |
|---|---|
| SOTA.md | Local models that beat Klein Base 4B, tiered |
| API_MODELS.md | Icon-native + API SOTA (Recraft SVG, Iconly, GPT Image, fal.ai, …) |
| SVG_MODELS.md | SVG registry + benchmark_svg_models.py (Recraft, OmniSVG, StarVector, …) |
| MODELS.md | Full local benchmark registry + licenses |
| EDGE_CASES.md | Polish AAC failure modes (zima, pilot, …) |
| PIPELINE.md | Prompt enhancement flow from grafiki-flux-v2 |
Quick start
cd ~/grafiki-mixed/scripts
# Polish edge cases — the stress test (enhanced prompts)
python3 benchmark_models.py --eval-set edge --prompt-mode enhanced \
--models flux2_klein_base_4b qwen_image_2512 ovis_image_7b
# General pictogram sanity (plain base prompts)
python3 benchmark_models.py --eval-set general --prompt-mode base \
--models flux2_klein_base_4b z_image_turbo krea2_turbo
# Prefetch gated/large models first
HF_HUB_ENABLE_HF_TRANSFER=0 python3 benchmark_models.py --download-only --models sd35_large flux2_klein_9b
SVG models (native .svg output — Recraft API, OmniSVG, …): see SVG_MODELS.md.
export RECRAFT_API_KEY=...
python3 benchmark_svg_models.py --models recraft_v3_vector --styles black outline
Prompts import from ~/grafiki-flux-v2/scripts/style_prompts.py (BASE_PROMPTS, base_prompt, with_hint).
Edge-case ground truth: eval/polish_edge_cases.json.
Results → results/{model}/{style}/[edge/]{slug}.png
Grids → grids/
Baseline vs challengers (short)
You have: FLUX.2 Klein Base 4B — best speed/VRAM/commercial combo for ~1200-word batches.
Commercial models likely better on quality (see SOTA.md): Qwen-Image-2512, Ovis-Image-7B, Krea-2-Raw, SD 3.5 Large, HiDream-I1, Ideogram 4 (HF NC weights).
Quality ceiling (NC, eval only): FLUX.2-dev, FLUX.1-dev, FLUX.2-klein-9B.
Not expected to beat Klein on icons: SDXL (ecosystem legacy), SD 3.5 Medium (VRAM niche).
Manifest
manifest.json tracks last benchmark run per model. Re-run benchmark_models.py to refresh after adding models from scripts/benchmark_models.py MODELS list.