Nurymanau's picture
Upload 73 files
91a3ee4 verified
|
Raw History Blame Contribute Delete
3.31 kB

Quality pilot — 2026-09-24

Three public prompts and criteria were frozen in quality-cases.json before execution. Both backends used the same Unsloth transformer/encoder and BF16 VAE, 512×512, 16 Euler steps, guidance 1, seed 42 and the explicit MLX sigma schedule in quality-sigmas.json. The system prompt/prefix removal was checked against sd.cpp source. The reference used commit 88411ef1e0688ff2df1010aeeb5d92b2d8cea2be, Metal, mmap, stage placement, a 5 GiB budget and 16×16-latent VAE tiles with 25% overlap.

Initial random noise differs between implementations. sd.cpp has its exact text prefix cache enabled; MLX does not. Background applications and thermal state were not controlled. These are observed complete-run times, not a causal speedup benchmark, pixel-parity test or general quality ranking.

Case Backend Wall time (s) Peak process GiB Peak swap growth MiB
portrait mlx 200.55 4.717 52.19
portrait sdcpp 331.30 4.994 0.00
lettering mlx 193.70 4.719 0.00
lettering sdcpp 337.36 4.994 0.00
composition mlx 199.32 4.724 0.00
composition sdcpp 319.46 4.994 0.00

Visual assessment

Non-blind manual review, one image per case/backend. All six images met the listed criteria in this small pilot; this is not a population-level accuracy score.

Case Criteria MLX image Original GGUF image
Portrait Coherent adult face; blue coat/white mug; two plausible hands; no severe visible deformation. Some fingers are occluded. MLX GGUF
Lettering Exact HELLO MAC and 2026; cream background/red circle; no extra words. MLX GGUF
Composition Red cube left of blue sphere; plant between; two adults repairing a recognizable telescope; light/window right. MLX GGUF

Differences in pose, typography, texture and framing are not attributed to the backend because initial noise differs. Exact criterion decisions and image hashes are in verification/quality-assessment.json; timings are in verification/quality-summaries.json.

Failed attempts retained

The first original-GGUF comparison completed all 16 denoising steps, but its 32×32-latent VAE tiles exceeded the 5 GiB manager budget while the transformer was still resident. Reducing tiles to 16×16 with 25% overlap fixed decoder admission. The successful comparison used this explicit setting for every source case.

The first MLX portrait attempt stopped at +492.38 MiB system-swap growth after 7.65 s. One retry with identical runtime settings completed at +52.19 MiB; lettering and composition then completed without growth. This does not establish universal zero-swap behavior on a busy 16 GiB machine. The +256 MiB guard stayed in place, and OpenCode/Qoder were not closed. Baseline swap was already nonzero.

The examples were generated before the selected-row embedding optimization. All three complete encoder outputs after that change match the original BF16 bit patterns and masks exactly; see OPTIMIZATION.md. No model weight bytes changed.