No description
  • Python 90.3%
  • Shell 9.7%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-24 02:46:30 +00:00
eval benchmarking code 2026-09-24 02:46:30 +00:00
results_svg benchmarking code 2026-09-24 02:46:30 +00:00
scripts benchmarking code 2026-09-24 02:46:30 +00:00
.gitignore benchmarking code 2026-09-24 02:46:30 +00:00
api_models.json benchmarking code 2026-09-24 02:46:30 +00:00
API_MODELS.md benchmarking code 2026-09-24 02:46:30 +00:00
benchmark.log benchmarking code 2026-09-24 02:46:30 +00:00
benchmark_new_models.log benchmarking code 2026-09-24 02:46:30 +00:00
benchmark_retry.log benchmarking code 2026-09-24 02:46:30 +00:00
EDGE_CASES.md benchmarking code 2026-09-24 02:46:30 +00:00
manifest.json benchmarking code 2026-09-24 02:46:30 +00:00
manifest_svg.json benchmarking code 2026-09-24 02:46:30 +00:00
MODELS.md benchmarking code 2026-09-24 02:46:30 +00:00
PIPELINE.md benchmarking code 2026-09-24 02:46:30 +00:00
README.md first commit 2026-09-24 02:43:54 +00:00
SOTA.md benchmarking code 2026-09-24 02:46:30 +00:00
svg_models.json benchmarking code 2026-09-24 02:46:30 +00:00
SVG_MODELS.md benchmarking code 2026-09-24 02:46:30 +00:00

grafiki-mixed — multi-model AAC icon benchmark

Compare commercial-friendly (and reference NC) image models against our production baseline
FLUX.2 Klein Base 4B (~/grafiki-flux-v2), using the same v1 prompt shell + Polish prompt-enhancement pipeline.

Doc Contents
SOTA.md Local models that beat Klein Base 4B, tiered
API_MODELS.md Icon-native + API SOTA (Recraft SVG, Iconly, GPT Image, fal.ai, …)
SVG_MODELS.md SVG registry + benchmark_svg_models.py (Recraft, OmniSVG, StarVector, …)
MODELS.md Full local benchmark registry + licenses
EDGE_CASES.md Polish AAC failure modes (zima, pilot, …)
PIPELINE.md Prompt enhancement flow from grafiki-flux-v2

Quick start

cd ~/grafiki-mixed/scripts

# Polish edge cases — the stress test (enhanced prompts)
python3 benchmark_models.py --eval-set edge --prompt-mode enhanced \
  --models flux2_klein_base_4b qwen_image_2512 ovis_image_7b

# General pictogram sanity (plain base prompts)
python3 benchmark_models.py --eval-set general --prompt-mode base \
  --models flux2_klein_base_4b z_image_turbo krea2_turbo

# Prefetch gated/large models first
HF_HUB_ENABLE_HF_TRANSFER=0 python3 benchmark_models.py --download-only --models sd35_large flux2_klein_9b

SVG models (native .svg output — Recraft API, OmniSVG, …): see SVG_MODELS.md.

export RECRAFT_API_KEY=...
python3 benchmark_svg_models.py --models recraft_v3_vector --styles black outline

Prompts import from ~/grafiki-flux-v2/scripts/style_prompts.py (BASE_PROMPTS, base_prompt, with_hint).

Edge-case ground truth: eval/polish_edge_cases.json.

Results → results/{model}/{style}/[edge/]{slug}.png
Grids → grids/

Baseline vs challengers (short)

You have: FLUX.2 Klein Base 4B — best speed/VRAM/commercial combo for ~1200-word batches.

Commercial models likely better on quality (see SOTA.md): Qwen-Image-2512, Ovis-Image-7B, Krea-2-Raw, SD 3.5 Large, HiDream-I1, Ideogram 4 (HF NC weights).

Quality ceiling (NC, eval only): FLUX.2-dev, FLUX.1-dev, FLUX.2-klein-9B.

Not expected to beat Klein on icons: SDXL (ecosystem legacy), SD 3.5 Medium (VRAM niche).

Manifest

manifest.json tracks last benchmark run per model. Re-run benchmark_models.py to refresh after adding models from scripts/benchmark_models.py MODELS list.