Instructions to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("krea/Krea-2-Turbo", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("lvladikov/Krea2-Turbo-Distill-2step-LoRA") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Upload 2 files
Browse files- DETAILED-README.md +1142 -0
- README.md +0 -0
DETAILED-README.md
ADDED
|
@@ -0,0 +1,1142 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: krea/Krea-2-Turbo
|
| 3 |
+
base_model_relation: adapter
|
| 4 |
+
license: other
|
| 5 |
+
license_name: krea-2-community-license
|
| 6 |
+
license_link: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-2step-LoRA/blob/main/LICENSE.pdf
|
| 7 |
+
library_name: diffusers
|
| 8 |
+
tags:
|
| 9 |
+
- lora
|
| 10 |
+
- text-to-image
|
| 11 |
+
- distillation
|
| 12 |
+
- step-distillation
|
| 13 |
+
- distribution-matching
|
| 14 |
+
- krea-2
|
| 15 |
+
pipeline_tag: text-to-image
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
# Krea 2 Turbo β 2-Step Distillation LoRA
|
| 19 |
+
|
| 20 |
+
> π§ͺ **A fast-preview adapter, from a project still in training.** When the subject is close and fills a good part of the
|
| 21 |
+
> frame β a portrait, a single figure, an object seen up close β two steps already hold up well, and you can rely on
|
| 22 |
+
> this adapter for those images. Small subjects are where it still falls short, people and objects alike: faces in a
|
| 23 |
+
> crowd, figures in a wide scene, the machines at the back of a gym β anything that takes up little of the frame can
|
| 24 |
+
> come out ghosted or smeared. For those, and whenever quality matters more than speed, use the
|
| 25 |
+
> [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Every figure on this page measures the
|
| 26 |
+
> adapter honestly against 4-step and 8-step renders. Training continues one recipe change at a time, and a later
|
| 27 |
+
> checkpoint replaces this one only when the sweeps and I visually agree it is better. Known issues β see
|
| 28 |
+
> [Known issues](#known-issues).
|
| 29 |
+
>
|
| 30 |
+
> π **The saved steps can also go into resolution.** A small subject is simply one that covers few pixels, so a larger
|
| 31 |
+
> render makes the same subject bigger β and at a quarter of the teacher's steps, renders up to 2048Γ2048, Krea's
|
| 32 |
+
> published maximum recommended resolution and beyond the largest size this adapter was trained at (1440Γ1440), come
|
| 33 |
+
> within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past
|
| 34 |
+
> 2048Γ2048, stock Krea 2 itself begins to duplicate subjects β a property of the base model, with or without this
|
| 35 |
+
> adapter.
|
| 36 |
+
>
|
| 37 |
+
> π **Also compatible with Krea 2 Raw.** With some prompts it works very well on Krea 2 Raw too, at 7 steps+ β see
|
| 38 |
+
> [Using it on Raw](#using-it-on-raw) and the [dedicated experiment](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md).
|
| 39 |
+
|
| 40 |
+
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
|
| 41 |
+
down to 2** β Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β aiming at
|
| 42 |
+
the best quality two steps can give. Two steps give up more than four: this adapter is for **fast previews and drafts**
|
| 43 |
+
at half the 4-step adapter's cost and a quarter of the teacher's, and the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the
|
| 44 |
+
recommendation for quality renders.
|
| 45 |
+
|
| 46 |
+
- π― **The aim** β the best two-step quality this base can give, at every one of the same 12 resolutions, measured
|
| 47 |
+
against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
|
| 48 |
+
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas.
|
| 49 |
+
- β±οΈ **4.2Γ faster denoising** β the model runs twice instead of eight times, and denoising is the part this adapter
|
| 50 |
+
changes: **81.4 s β 19.5 s** measured at 1024Γ1024 on the same prompts, the adapter's own cost per call within
|
| 51 |
+
measurement noise. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
|
| 52 |
+
[Performance](#performance).
|
| 53 |
+
- π **Distribution matching, not imitation** β the training objective that got the renders improving again after the
|
| 54 |
+
4-step project's recipe had stopped helping at two steps (see [Method](#method)).
|
| 55 |
+
- π£οΈ **Prompt-conditioned throughout** β both scores in the distribution match, the teacher's and the fake adapter's, are
|
| 56 |
+
evaluated on each prompt's own conditioning, so the student is matched to what the teacher makes _for that prompt_,
|
| 57 |
+
not to a prompt-free look. There is no separate adherence term: instead a vision-language judge checks every
|
| 58 |
+
checkpoint β each render scored alone against the prompt's objects, counts, attributes and relations, with the
|
| 59 |
+
teacher scored the same way β and a term would only be added if that meter showed adherence slipping. The one part of
|
| 60 |
+
training that looks at images without their prompt is the artefact critic (see [Method](#method)), and it only judges
|
| 61 |
+
whether fine structure looks like the teacher's.
|
| 62 |
+
- π **12 trained resolutions** β multi-aspect from 512Γ512 up to 1440Γ1440, each with its sweep.
|
| 63 |
+
- π **Drop-in, no exceptions** β a plain LoRA sampled by stock Euler at sigmas `[1.0, 0.7595]` in diffusers, ComfyUI
|
| 64 |
+
or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this
|
| 65 |
+
project.
|
| 66 |
+
- 𧬠**Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β rank 64 on the same 228 modules; a second adapter exists during training
|
| 67 |
+
only and never ships.
|
| 68 |
+
- π² **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
|
| 69 |
+
re-run.
|
| 70 |
+
- π’ **17,464 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β all of them
|
| 71 |
+
drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
|
| 72 |
+
training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
|
| 73 |
+
4-step project, read again at the two sigmas this schedule uses.
|
| 74 |
+
- π
**7 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β and the project continues.
|
| 75 |
+
- π **26 recipe adjustments** across two methods so far β seven of trajectory distillation before the switch, nineteen of distribution matching since β each kept only when the renders did not get worse.
|
| 76 |
+
- π₯οΈ **One RTX 3090**, and a recipe shaped by its 24 GB.
|
| 77 |
+
|
| 78 |
+
[](assets/poster.jpg)
|
| 79 |
+
|
| 80 |
+
_All of the above were created with this LoRA at 2 steps: the 15 test prompts, Krea 2 Turbo + the
|
| 81 |
+
LoRA, seed 4242, each at one of its trained resolutions. Click for full size. The side-by-side
|
| 82 |
+
comparisons with the 8-step teacher are in [Examples](#examples)._
|
| 83 |
+
|
| 84 |
+
## Files
|
| 85 |
+
|
| 86 |
+
| file | what it is |
|
| 87 |
+
| -------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 88 |
+
| `krea2_turbo_2step_rank_64_lora.safetensors` | the LoRA in diffusers key format β see [Inference with diffusers](#inference-with-diffusers); also for MLX or anything that reads safetensors |
|
| 89 |
+
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | the same weights under ComfyUI's key names β see [ComfyUI](#comfyui) |
|
| 90 |
+
| `krea2_turbo_2step_lora_t2i.json` | a ready ComfyUI workflow, stock nodes only |
|
| 91 |
+
| `krea2_raw_7step_lora_experiment_t2i.json` | the Krea 2 Raw 7-step experiment's ComfyUI workflow ([Using it on Raw](#using-it-on-raw)) |
|
| 92 |
+
| [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **the quick place to check which checkpoint the two weight files are based on.** The pair above keeps its names and is updated in place as better checkpoints ship; this file always says what they are today. Every published checkpoint also sits in [`_archive/checkpoints/`](_archive/checkpoints) under its number |
|
| 93 |
+
| `LICENSE.pdf` | the Krea 2 Community License Agreement, which covers this adapter β see [License](#license) |
|
| 94 |
+
| `NOTICE.txt` | the attribution notice the license requires of a derivative |
|
| 95 |
+
|
| 96 |
+
The two weight files are one adapter β only the key names differ. Both carry the training details
|
| 97 |
+
in their safetensors metadata: base model, lineage, method, the checkpoint and the inference settings. Their file names
|
| 98 |
+
never change; when a better checkpoint ships they are replaced in place, and
|
| 99 |
+
[`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) is the quick
|
| 100 |
+
place to check which checkpoint the current files are based on.
|
| 101 |
+
|
| 102 |
+
## Where it stands
|
| 103 |
+
|
| 104 |
+
| | |
|
| 105 |
+
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 106 |
+
| lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β 2-step trajectory distillation β distribution matching β a spectral match against the teacher's own images β an artefact critic and detail terms β three critics taking turns |
|
| 107 |
+
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
|
| 108 |
+
| what it gives | usable two-step renders at every trained resolution: fine detail at or just above the teacher's β from 1 megapixel up, closer to the teacher than the 4-step adapter β with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 1 point missing out of 240). A judge asked which render follows the prompt better still prefers the 8-step teacher on 11 of 45, against 6 for the 4-step adapter, mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
|
| 109 |
+
|
| 110 |
+
### Known issues
|
| 111 |
+
|
| 112 |
+
The usual costs of two steps, in this order of how often they show. **Small subjects are the weak spot, people and objects alike**: a portrait-sized face or an object seen up close holds up, while small or distant subjects β faces in a crowd, a figure in a wide scene, the machines at the back of a room β can come out ghosted, smeared or misshapen, since at that size a whole subject is only a few of the blocks the model works in. Fine structure can come out soft or a few pixels out of register β feathers, hair strands, signage, the surface of a distant object β most at 1280Γ1280 and above, and a faint doubled contour can show on limbs. On stylised prompts, **how the prompt says the image should look is followed less faithfully than what should be in it**: crisp anime linework, energetic brush strokes, the fingerprints in clay or a matte-painting finish come out closer to a generic rendering than the teacher's. On busy action or crowd scenes the composition can repeat itself β an extra hand or held object, a figure duplicated in a crowd β where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. Skin reads slightly smoother and less saturated than the teacher's, and colour overall runs a little under the teacher's at the largest sizes; freckles tend to gather into clusters rather than separate dots. A fine grain remains on the most textured subjects at the largest sizes, lighter than in the previous checkpoint. Every one of these is being worked on; none is hidden in the sweeps or the examples.
|
| 113 |
+
|
| 114 |
+
## How I got here
|
| 115 |
+
|
| 116 |
+
The [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) closed its page with a promise: a 2-step LoRA as the next project, and a guess at the lever it would
|
| 117 |
+
need β matching the teacher's _distribution_ rather than its trajectory. That guess turned out to be the whole story.
|
| 118 |
+
|
| 119 |
+
The project began where the [4-step one](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) ended, from its final weights, and ran the same recipe at two steps:
|
| 120 |
+
**progressive distillation** on the recorded teacher trajectories, each student call covering four teacher steps, with
|
| 121 |
+
the LADD-style critic as the finisher. Well into that run, every number had stopped moving and the pictures had a
|
| 122 |
+
signature the numbers could not see: doubled contours on faces and limbs, soft fine texture, crowds averaged into
|
| 123 |
+
translucent overlaps. Several variations followed β the critic re-weighted, judged per token, a heavier hand on the
|
| 124 |
+
final call, the student's own first-step output fed into its second β and each traded one of those faults for another
|
| 125 |
+
without moving past them. A capacity probe ruled out adapter rank; a learning-rate shock ruled out the optimiser.
|
| 126 |
+
|
| 127 |
+
The reason is structural, and worth stating plainly because it decides the whole design. A regression loss asks the
|
| 128 |
+
student to land on the teacher's _specific_ image for each prompt. When a two-step jump is wide enough that several
|
| 129 |
+
images are plausible, the answer that minimises the squared error is their average β and the average of two sharp
|
| 130 |
+
images is a blurred one with doubled edges. Every earlier recipe rewarded that average. Tuning its weights could not
|
| 131 |
+
change what it rewarded.
|
| 132 |
+
|
| 133 |
+
**Distribution matching** asks a different question: not "does your image match this one" but "would the teacher
|
| 134 |
+
plausibly have produced your image". The first run of that objective, on top of the trajectory-distilled weights, produced
|
| 135 |
+
in a fraction of the old recipe's training what all of it never had β and it did so while every latent
|
| 136 |
+
distance to the teacher _rose_, which is exactly what a mode-seeking objective predicts and what a mean-seeking metric
|
| 137 |
+
punishes. The distances are reported on this page; they are not optimised for, and they are not what decides a
|
| 138 |
+
checkpoint. Pictures are, at fixed seeds, at every resolution, with faces viewed at 1:1.
|
| 139 |
+
|
| 140 |
+
The recipe adjustments so far, each made on the measurement of the one before:
|
| 141 |
+
|
| 142 |
+
1. progressive distillation at two steps from the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s weights, with the LADD critic β the baseline
|
| 143 |
+
2. the critic made to judge the first call's endpoint against the teacher's mid-states, which removed the gross
|
| 144 |
+
ghosting and left the doubled contours
|
| 145 |
+
3. the critic's weight, its per-token form and the student's own first-step output as the second call's input β
|
| 146 |
+
tried one at a time; the objective flip that mattered was not among them
|
| 147 |
+
4. **distribution matching** (the DMD2 family) as the primary objective, trajectory regression demoted to an anchor at
|
| 148 |
+
half weight, the critic switched off to read the new term alone
|
| 149 |
+
5. the running average of the weights restarted at the objective switch, after it was caught averaging two lineages
|
| 150 |
+
into composites
|
| 151 |
+
6. a spectral match of the student's image against the teacher's own, on the whole latent and on decoded pixel
|
| 152 |
+
windows β the term that finally reached the grain at large resolutions, after a critic, per-resolution weights and
|
| 153 |
+
a filtered push had each been tried against it and retired
|
| 154 |
+
7. the decoded-window spectral term given a much lighter hand β capped at a quarter of its earlier strength, which kept
|
| 155 |
+
the detail and took some grain out of flat areas
|
| 156 |
+
8. the fake-score adapter updated four times per student step instead of twice, so it keeps up with what the student
|
| 157 |
+
currently makes; faint straight-line artefacts that had begun to appear on flat illustrated areas went away with it
|
| 158 |
+
9. **an artefact critic** β a small head reading the frozen base model's own mid-network features, trained to tell the
|
| 159 |
+
teacher's finished images from the student's, with its push restricted to structure finer than 32 pixels and held
|
| 160 |
+
well below the distribution term β aimed at what the distribution term leaves behind: melted small faces and dense
|
| 161 |
+
detail smeared into blotches
|
| 162 |
+
10. four detail terms together: the anchor counting the fine-detail part of its error twice; a one-sided floor that
|
| 163 |
+
stops the finest detail dropping below the level real photographs carry; a second-call target made by the teacher
|
| 164 |
+
finishing the image from the student's own first-call output, so the target shares the student's layout; and a
|
| 165 |
+
smoothness limit on the fake-score adapter, so the distribution push keeps pointing at detail rather than away from it
|
| 166 |
+
11. **three critics taking turns** β the artefact critic joined by one whose real examples are half real photographs and one
|
| 167 |
+
weighted toward faces, one of them pushing on each step while the others keep training in between, because the three did
|
| 168 |
+
not fit in memory side by side
|
| 169 |
+
|
| 170 |
+
Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
|
| 171 |
+
a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
|
| 172 |
+
not place exactly where the teacher had it.
|
| 173 |
+
|
| 174 |
+
## chk00017464 vs chk00013663
|
| 175 |
+
|
| 176 |
+
`chk00013663` (published 12 Sep 2026) was the first public checkpoint: distribution matching with the spectral match, and nothing
|
| 177 |
+
yet aimed at the faults the distribution term leaves behind. `chk00017464` (14 Sep 2026) is 3,801 training samples later, and every
|
| 178 |
+
one of those samples went to those faults β the grain and grid pattern at large sizes, small faces, dense detail β through five recipe
|
| 179 |
+
changes, each kept only after its own look at the renders:
|
| 180 |
+
|
| 181 |
+
1. the decoded-window spectral term brought down to a quarter of its strength
|
| 182 |
+
2. the fake-score adapter updated four times per student step instead of twice
|
| 183 |
+
3. the artefact critic, reading the frozen base model's own features
|
| 184 |
+
4. four detail terms: the anchor counting fine-detail error twice, the one-sided photo floor, the teacher's finish of the student's
|
| 185 |
+
first call as the second call's target, and a smoothness limit on the fake adapter
|
| 186 |
+
5. three critics taking turns: the artefact critic, a photo critic and a face critic
|
| 187 |
+
|
| 188 |
+
Two ideas were tried and taken back out along the way: the distribution term confined to the second call's noise range, and a detail
|
| 189 |
+
pyramid compared pixel by pixel against the teacher.
|
| 190 |
+
|
| 191 |
+
**Detail at large sizes β the headline.** Every one of the 12 sweep resolutions Γ 15 prompts measured against the 8-step teacher, as in
|
| 192 |
+
[Measured against the teacher](#measured-against-the-teacher). Above 1 megapixel the excess fine energy two steps used to put into
|
| 193 |
+
images is roughly halved, and those sizes are now closer to the teacher than the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) on
|
| 194 |
+
fine texture and on both grid bands:
|
| 195 |
+
|
| 196 |
+
| 1.00 = the teacher | fine texture | 16-px band | 8-px band |
|
| 197 |
+
| --- | --- | --- | --- |
|
| 198 |
+
| 1280Γ1280 | 1.20 β **1.09** | 1.16 β **1.08** | 1.25 β **1.13** |
|
| 199 |
+
| 1440Γ1280 | 1.21 β **1.08** | 1.19 β **1.09** | 1.28 β **1.14** |
|
| 200 |
+
| 1440Γ1440 | 1.34 β **1.17** | 1.25 β **1.08** | 1.42 β **1.21** |
|
| 201 |
+
|
| 202 |
+
Fine-texture energy and both grid bands come closer to the teacher at 10 of the 12 resolutions, the blur-invariant ghosting index at
|
| 203 |
+
10 of 12, the micro-ghost index at all 12, and the grain in flat areas β skies, walls, out-of-focus backgrounds β at 11 of 12
|
| 204 |
+
(1440Γ1440: 1.84Γ the teacher's β 1.57Γ, median of the 15 prompts). It shows where it should: hair renders as more individual strands, and the token-grid
|
| 205 |
+
texture that read as grain on birds, food and foliage at 1440Γ1440 is lighter.
|
| 206 |
+
|
| 207 |
+
**Faces and skin.** Freckles on the test portrait gather into lighter, more dot-like clusters than before β better, not yet the
|
| 208 |
+
teacher's separate dots β and eyes look about the same: clean irises, lashes softer than the teacher's.
|
| 209 |
+
|
| 210 |
+
**What did not improve.** The judge that asks which of two renders follows the prompt better prefers the 8-step teacher on 11 of 45
|
| 211 |
+
against 6 before, mostly on how a stylised prompt says the picture should look: crisp linework, energetic brush strokes, the texture
|
| 212 |
+
of clay. The blind rubric that checks each render alone for the prompt's objects, counts, attributes and relations moves from 0 to 1
|
| 213 |
+
point missing out of 240, and on the same 15 fresh prompts the previous checkpoint was measured on, the two checkpoints come out level
|
| 214 |
+
(1 win, 13 ties, 1 loss). Colour runs slightly lower (saturation 0.93Γ the teacher's across the sweep, from 0.94Γ; 0.86Γ at
|
| 215 |
+
1440Γ1440), and at 512Γ512 the two checkpoints are level. Both are what the next recipe changes target.
|
| 216 |
+
|
| 217 |
+
| axis | `chk00013663` | `chk00017464` |
|
| 218 |
+
| --- | --- | --- |
|
| 219 |
+
| fine texture vs the teacher, 1280Β² / 1440Β² | 1.20 / 1.34 | **1.09 / 1.17** |
|
| 220 |
+
| 16-px grid band, 1280Β² / 1440Β² | 1.16 / 1.25 | **1.08 / 1.08** |
|
| 221 |
+
| grain in flat areas, sweep median | 1.36Γ | **1.26Γ** |
|
| 222 |
+
| distance to the teacher, sweep mean | 0.413 | **0.406** |
|
| 223 |
+
| judge prefers the teacher (of 45) | **6** | 11 |
|
| 224 |
+
| blind adherence rubric, points missing of 240 | **0** | 1 |
|
| 225 |
+
| saturation vs the teacher, sweep mean | **0.94Γ** | 0.93Γ |
|
| 226 |
+
| training samples in the 2-step stages | 13,663 | 17,464 |
|
| 227 |
+
|
| 228 |
+
## Measured against the teacher
|
| 229 |
+
|
| 230 |
+
Every number here compares a render of this LoRA at 2 steps with the 8-step reference render of the **same prompt at the
|
| 231 |
+
same seed**, across the 15 test prompts and the 12 trained resolutions. Two other columns are measured the same way, so
|
| 232 |
+
the figures have a floor and a ceiling around them: **stock Krea 2 Turbo at 2 steps**, which is what the base model does
|
| 233 |
+
without the adapter, and the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) at its own 4 steps, which is the better tool and the thing worth
|
| 234 |
+
being compared against.
|
| 235 |
+
|
| 236 |
+
**Detail, band by band.** Fine-texture energy and the two grid bands, as a ratio to the teacher's own (1.00 = the
|
| 237 |
+
teacher):
|
| 238 |
+
|
| 239 |
+
| bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
|
| 240 |
+
| --- | --- | --- | --- | --- | --- |
|
| 241 |
+
| 512Γ512 | 1.02 | 1.02 | 1.05 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
|
| 242 |
+
| 768Γ1024 | 1.02 | 0.98 | 1.07 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
|
| 243 |
+
| 1024Γ1024 | 1.06 | 1.07 | 1.12 | 1.11 Β· 1.17 Β· 1.16 | 0.39 |
|
| 244 |
+
| 1280Γ1280 | 1.09 | 1.08 | 1.13 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
|
| 245 |
+
| 1440Γ1440 | 1.17 | 1.08 | 1.21 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
|
| 246 |
+
|
| 247 |
+
Two steps without the adapter carry 0.57Γ the teacher's fine detail at 512Γ512 and **less than half** (0.39β0.42Γ) at the four larger sizes. With it, the detail
|
| 248 |
+
sits at or just above the teacher's everywhere β and from 1 megapixel up it is closer to the teacher than the 4-step
|
| 249 |
+
adapter, which carries more excess fine energy there.
|
| 250 |
+
|
| 251 |
+
**Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
|
| 252 |
+
in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
|
| 253 |
+
|
| 254 |
+
| bucket | wins | ties | losses |
|
| 255 |
+
| --- | --- | --- | --- |
|
| 256 |
+
| 512Γ512 | 1 | 10 | 4 |
|
| 257 |
+
| 1280Γ1280 | 1 | 10 | 4 |
|
| 258 |
+
| 1440Γ1440 | 0 | 12 | 3 |
|
| 259 |
+
|
| 260 |
+
Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
|
| 261 |
+
stylised prompts and on how a prompt says the picture should look β crisp linework, energetic brush strokes, the texture of
|
| 262 |
+
clay β more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
|
| 263 |
+
attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
|
| 264 |
+
training prompt bank and never rendered before, the judge returned 0 wins, 10 ties, 5 losses; on another 15 fresh prompts, 1 win,
|
| 265 |
+
13 ties, 1 loss.
|
| 266 |
+
|
| 267 |
+
**Checked for the damage this kind of training can do.** Saturation sits at 0.99Γ the teacher's at 768Γ1024 and 0.89Γ
|
| 268 |
+
at the larger sizes; edge detail 0.93β0.96Γ; skin texture inside detected faces 1.07Γ at 768Γ1024 and 0.85Γ at
|
| 269 |
+
1440Γ1440, with skin saturation 0.92Γ and 0.77β0.85Γ at the larger sizes. The honest reading of those last two: **skin
|
| 270 |
+
is the softest and least saturated part of this adapter's output at large sizes, and colour overall runs a little under
|
| 271 |
+
the teacher's there.** Fine detail in flat regions β skies, walls, out-of-focus backgrounds β runs 1.53β1.85Γ the
|
| 272 |
+
teacher's, which is where two steps put grain that eight steps do not.
|
| 273 |
+
|
| 274 |
+
**Distance to the teacher**, as a plain pixel measure, is 0.39β0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
|
| 275 |
+
0.30β0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
|
| 276 |
+
not the teacher's render of it β see [Known issues](#known-issues).
|
| 277 |
+
|
| 278 |
+
## Usage
|
| 279 |
+
|
| 280 |
+
| setting | value |
|
| 281 |
+
| -------------- | ----------------------------------------------- |
|
| 282 |
+
| base model | Krea 2 Turbo |
|
| 283 |
+
| LoRA scale | 1.0 |
|
| 284 |
+
| steps | **2** |
|
| 285 |
+
| guidance / CFG | **0.0** (Turbo is CFG-free; do not enable it) |
|
| 286 |
+
| timestep shift | **mu = 1.15**, fixed (Turbo's deployment shift) |
|
| 287 |
+
|
| 288 |
+
The 2 sampling sigmas are Turbo's own deployment grid: `[1.0, 0.7595]` β the first and the middle
|
| 289 |
+
of the 4-step grid, so the model is evaluated at two points it already knows.
|
| 290 |
+
|
| 291 |
+
## Using it on Raw
|
| 292 |
+
|
| 293 |
+
**This LoRA is trained on Krea 2 Turbo, against Turbo as its own teacher, and for Turbo.** Every layer it targets also
|
| 294 |
+
exists in Krea 2 Raw, so it will load there without complaint β but that is a side effect of the shared architecture, not
|
| 295 |
+
a supported mode.
|
| 296 |
+
|
| 297 |
+
It still does something useful there. On Turbo the adapter runs a quarter of the teacher's steps β 2 of its 8 β so on
|
| 298 |
+
Raw the same quarter of its usual 28 steps, **7**, is the natural place to start, and for most prompts it is a good one.
|
| 299 |
+
For a quick preview, some prompts hold together at even 4β5 steps β not as good as 7 steps or more, but not broken
|
| 300 |
+
either. Keep the guidance **light**: `guidance_scale=1.0` in diffusers, which is **cfg 2.0 in ComfyUI**. Heavier
|
| 301 |
+
guidance, such as 4.5, crushes most images into near-black frames at 7 steps, and with no guidance at all the pictures
|
| 302 |
+
come out flat.
|
| 303 |
+
|
| 304 |
+
[](assets/resolution_sweeps/raw-LoRA-7steps-experiment/1024x768/portrait.jpg)
|
| 305 |
+
[](assets/resolution_sweeps/raw-LoRA-7steps-experiment/1024x768/pizza.jpg)
|
| 306 |
+
|
| 307 |
+
_Krea 2 Raw + this LoRA, 7 steps, light guidance, empty negative prompt, seed 4242, 1024Γ768. Click for full size._
|
| 308 |
+
|
| 309 |
+
Results are **mixed and subject-dependent**. Some prompts come through as finished pictures; others do not β an
|
| 310 |
+
underexposed night street, a cityscape with less detail than Turbo gives at 2 steps β and for those a few more steps may
|
| 311 |
+
help. The full experiment, with all 15 test prompts at 1024Γ768, what worked and what did not, and a ComfyUI workflow,
|
| 312 |
+
is in its [dedicated README](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md), in
|
| 313 |
+
[`assets/resolution_sweeps/raw-LoRA-7steps-experiment/`](assets/resolution_sweeps/raw-LoRA-7steps-experiment). The same
|
| 314 |
+
prompts on stock Raw at the same 7 steps and light guidance, without the LoRA, are in
|
| 315 |
+
[`_raw-base-NO-LoRA-7step-cfg1/`](assets/resolution_sweeps/raw-LoRA-7steps-experiment/_raw-base-NO-LoRA-7step-cfg1).
|
| 316 |
+
|
| 317 |
+
[](assets/resolution_sweeps/raw-LoRA-7steps-experiment/workflow_preview_raw.jpg)
|
| 318 |
+
|
| 319 |
+
_The experiment's ComfyUI workflow,
|
| 320 |
+
[`krea2_raw_7step_lora_experiment_t2i.json`](krea2_raw_7step_lora_experiment_t2i.json): Krea 2 Raw + this LoRA, 7 steps,
|
| 321 |
+
cfg 2.0. Click for full size._
|
| 322 |
+
|
| 323 |
+
## Inference with diffusers
|
| 324 |
+
|
| 325 |
+
Krea 2 Turbo has a native diffusers pipeline, `Krea2Pipeline`, in diffusers from source β the same
|
| 326 |
+
setup as Krea's own model card. The LoRA loads through that pipeline's standard LoRA loader, with
|
| 327 |
+
one line of preparation: diffusers expects a `transformer.` prefix on every key, and it does not read
|
| 328 |
+
the per-module `alpha` entries this file carries (alpha equals the rank, so dropping them changes
|
| 329 |
+
nothing β the scale stays 1.0). Why the file is laid out the way it is, and what each runtime
|
| 330 |
+
expects, is in [The keys, runtime by runtime](#the-keys-runtime-by-runtime).
|
| 331 |
+
|
| 332 |
+
```
|
| 333 |
+
pip install git+https://github.com/huggingface/diffusers.git
|
| 334 |
+
```
|
| 335 |
+
|
| 336 |
+
```python
|
| 337 |
+
import torch
|
| 338 |
+
from diffusers import Krea2Pipeline
|
| 339 |
+
from huggingface_hub import hf_hub_download
|
| 340 |
+
from safetensors.torch import load_file
|
| 341 |
+
|
| 342 |
+
pipe = Krea2Pipeline.from_pretrained("krea/Krea-2-Turbo", torch_dtype=torch.bfloat16).to("cuda") # "mps" on Apple Silicon
|
| 343 |
+
|
| 344 |
+
lora = hf_hub_download("lvladikov/Krea2-Turbo-Distill-2step-LoRA", "krea2_turbo_2step_rank_64_lora.safetensors")
|
| 345 |
+
state = load_file(lora)
|
| 346 |
+
state = {f"transformer.{k}": v for k, v in state.items() if not k.endswith(".alpha")}
|
| 347 |
+
pipe.load_lora_weights(state, adapter_name="2step")
|
| 348 |
+
|
| 349 |
+
image = pipe("a fox in the snow", num_inference_steps=2, guidance_scale=0.0).images[0]
|
| 350 |
+
image.save("krea2_2step.png")
|
| 351 |
+
```
|
| 352 |
+
|
| 353 |
+
- **`num_inference_steps=2` is the whole configuration.** The pipeline applies Turbo's fixed timestep
|
| 354 |
+
shift (mu = 1.15) on its own and evaluates the model at Ο = 1.0 and 0.7595 β exactly the two points
|
| 355 |
+
in [Usage](#usage) that the LoRA was trained on. Keep `guidance_scale=0.0`.
|
| 356 |
+
- **Strength:** `pipe.set_adapters(["2step"], adapter_weights=[0.75])`. Stock Turbo is one call away
|
| 357 |
+
for a side-by-side: `pipe.unload_lora_weights()` and `num_inference_steps=8`.
|
| 358 |
+
- Use the diffusers file, not the `_comfyui` one: diffusers' Krea 2 key converter reads Krea's
|
| 359 |
+
reference-trainer naming, not ComfyUI's `lora_down`/`lora_up`.
|
| 360 |
+
|
| 361 |
+
## ComfyUI
|
| 362 |
+
|
| 363 |
+
A pre-converted file (`..._comfyui.safetensors`) and a ready workflow sit in the repo root. **No custom nodes** β stock
|
| 364 |
+
ComfyUI only.
|
| 365 |
+
|
| 366 |
+
[](assets/workflow_preview.jpg)
|
| 367 |
+
|
| 368 |
+
| file | put it in |
|
| 369 |
+
| ----------------------------------------------------------------------------------------------------------------------- | ---------------------------------- |
|
| 370 |
+
| [`krea2_turbo_2step_rank_64_lora_comfyui.safetensors`](krea2_turbo_2step_rank_64_lora_comfyui.safetensors) | `ComfyUI/models/loras/` |
|
| 371 |
+
| `krea2_turbo_bf16.safetensors` β [Comfy-Org/Krea-2](https://huggingface.co/Comfy-Org/Krea-2/tree/main/diffusion_models) | `ComfyUI/models/diffusion_models/` |
|
| 372 |
+
| `qwen3vl_4b_bf16.safetensors` β [same repo](https://huggingface.co/Comfy-Org/Krea-2/tree/main/text_encoders) | `ComfyUI/models/text_encoders/` |
|
| 373 |
+
| `qwen_image_vae.safetensors` β [same repo](https://huggingface.co/Comfy-Org/Krea-2/tree/main/vae) | `ComfyUI/models/vae/` |
|
| 374 |
+
|
| 375 |
+
Then load [`krea2_turbo_2step_lora_t2i.json`](krea2_turbo_2step_lora_t2i.json).
|
| 376 |
+
|
| 377 |
+
**The workflow is full bf16, with no quantisation anywhere.** bf16 needs no backend-specific
|
| 378 |
+
kernel, so it runs unchanged on CUDA, Apple Silicon and CPU β one workflow, no platform caveats,
|
| 379 |
+
nothing that depends on which device a component happens to land on.
|
| 380 |
+
|
| 381 |
+
Smaller builds work too; both loaders accept any variant, just set the matching filename:
|
| 382 |
+
|
| 383 |
+
| diffusion model | size | NVIDIA | Apple Silicon |
|
| 384 |
+
| ----------------------------------------- | ------- | ------ | ------------- |
|
| 385 |
+
| **`krea2_turbo_bf16`** (workflow default) | 26.3 GB | β
| β
|
|
| 386 |
+
| `krea2_turbo_int8_convrot` | 13.5 GB | β
| β
|
|
| 387 |
+
| `krea2_turbo_fp8_scaled` | 13.1 GB | β
| β |
|
| 388 |
+
|
| 389 |
+
The text encoder ships as bf16 (8.9 GB) or fp8 (5.2 GB) only β **there is no int8 text encoder**, so
|
| 390 |
+
a fully matched int8 pair is not possible.
|
| 391 |
+
|
| 392 |
+
**The LoRA is independent of the base build.** It is applied on top of the diffusion model by
|
| 393 |
+
ComfyUI's own loader, which handles any dequantisation, so a quantised or otherwise optimised build
|
| 394 |
+
of Krea 2 Turbo behaves just as bf16 does. Please use whichever variant suits your hardware β set it
|
| 395 |
+
in the **Load Diffusion Model** node and leave the rest of the workflow untouched. The workflow
|
| 396 |
+
ships bf16 simply because it is the one build guaranteed to run everywhere.
|
| 397 |
+
|
| 398 |
+
> π **`fp8_scaled` does not work on Apple Silicon.** MPS has no `Float8_e4m3fn` support, so the run
|
| 399 |
+
> dies at the sampler with _"Trying to convert Float8_e4m3fn to the MPS backend but it does not have
|
| 400 |
+
> support for that dtype"_. That failure is the weight dtype, not the workflow or the LoRA β the
|
| 401 |
+
> graph executes fine right up to the sampler. The fp8 _text encoder_ does run on MPS, but only
|
| 402 |
+
> because ComfyUI places it on CPU; the workflow does not rely on that.
|
| 403 |
+
|
| 404 |
+
> π **The `_comfyui` file carries the same weights as the diffusers file** β only the key names differ.
|
| 405 |
+
|
| 406 |
+
**Why a separate file.** ComfyUI addresses the transformer by its own layer names, so the adapter
|
| 407 |
+
needs a key remap: `transformer_blocks.0.attn.to_q.lora_A` becomes
|
| 408 |
+
`diffusion_model.blocks.0.attn.wq.lora_down`. **The tensors are bit-identical** β nothing is
|
| 409 |
+
requantised or rescaled, only renamed. The mapping was verified against Comfy-Org's own Krea 2 LoRA:
|
| 410 |
+
all 456 tensors land on keys that file also uses, with matching shapes.
|
| 411 |
+
|
| 412 |
+
`alpha` keys are omitted, as in Comfy's own file. ComfyUI defaults alpha to the rank when absent,
|
| 413 |
+
giving scale = alpha/rank = 1.0 β exactly what alpha 64 at rank 64 encodes.
|
| 414 |
+
|
| 415 |
+
### Settings
|
| 416 |
+
|
| 417 |
+
| | |
|
| 418 |
+
| ------------------- | ------------------ |
|
| 419 |
+
| steps | **2** |
|
| 420 |
+
| **cfg** | **1.0** |
|
| 421 |
+
| sampler / scheduler | `euler` / `simple` |
|
| 422 |
+
| LoRA strength | 1.0 |
|
| 423 |
+
|
| 424 |
+
> βοΈ **`cfg 1.0`, not `0.0`.** ComfyUI expresses "no classifier-free guidance" as cfg **1.0**, whereas
|
| 425 |
+
> diffusers expresses the same thing as guidance **0.0**. They mean the same: one forward pass per
|
| 426 |
+
> step, no negative branch. Setting 0.0 in ComfyUI is not the same thing and will not give you
|
| 427 |
+
> Turbo's intended behaviour. That is also why the workflow's negative input is a
|
| 428 |
+
> `ConditioningZeroOut` β at cfg 1.0 it is never evaluated, so there is nothing to write in it.
|
| 429 |
+
|
| 430 |
+
To compare against stock Turbo, set steps back to 8 and bypass the LoRA node with `Ctrl+B`.
|
| 431 |
+
|
| 432 |
+
## Performance
|
| 433 |
+
|
| 434 |
+
Measured on this machine (Apple Silicon, MLX, bf16) at 1024Γ1024, two prompts per configuration, each run on its own with
|
| 435 |
+
the model already loaded, so the numbers are the render itself and not a model load:
|
| 436 |
+
|
| 437 |
+
| configuration | denoising | per model call | GPU peak |
|
| 438 |
+
| --- | --- | --- | --- |
|
| 439 |
+
| Krea 2 Turbo β 8 steps (the reference) | **81.4 s** | 10.2 s | 25.2 GiB |
|
| 440 |
+
| Krea 2 Turbo β 2 steps, no LoRA | 20.4 s | 10.2 s | 25.2 GiB |
|
| 441 |
+
| **Krea 2 Turbo β 2 steps + this LoRA** | **19.5 s** | 9.8 s | 25.2 GiB |
|
| 442 |
+
|
| 443 |
+
**Denoising is 4.2Γ faster than the 8-step reference** β two model calls instead of eight. The adapter's own cost per call did
|
| 444 |
+
not show up in this measurement: the runs with it came in marginally faster than those without, which is measurement noise, not a
|
| 445 |
+
speed-up. A rank-64 low-rank product is small beside the transformer it is added to, and it adds no measurable memory.
|
| 446 |
+
|
| 447 |
+
Denoising is the part the step count changes. What a complete render costs on top of it β encoding the prompt, decoding
|
| 448 |
+
the latent, writing the file β is the same whether you run two steps or eight, and it depends on your pipeline, so the
|
| 449 |
+
end-to-end figure on your machine will sit below 4.2Γ and rise toward it as the render gets larger.
|
| 450 |
+
|
| 451 |
+
**By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
|
| 452 |
+
size; sweep renders run one at a time, not the controlled measurement above):
|
| 453 |
+
|
| 454 |
+
| resolution | denoising (2 calls) | resolution | denoising (2 calls) |
|
| 455 |
+
| --- | --- | --- | --- |
|
| 456 |
+
| 512Γ512 | 6.4 s | 1024Γ1024 | 20.4 s |
|
| 457 |
+
| 512Γ768 / 768Γ512 | 8.9 s / 9.0 s | 1280Γ960 / 960Γ1280 | 23.2 s / 24.0 s |
|
| 458 |
+
| 768Γ768 | 11.9 s | 1280Γ1280 | 32.2 s |
|
| 459 |
+
| 768Γ1024 / 1024Γ768 | 14.9 s / 15.1 s | 1440Γ1280 / 1440Γ1440 | 34.8 s / 39.8 s |
|
| 460 |
+
|
| 461 |
+
## LoRA strength
|
| 462 |
+
|
| 463 |
+
Use **1.0**. That is the value the adapter was trained at, and where the 2-step output sits closest to the 8-step
|
| 464 |
+
reference on every measurement in this page.
|
| 465 |
+
|
| 466 |
+
[](assets/portrait_strength_sweep.jpg)
|
| 467 |
+
|
| 468 |
+
The three panels individually: [0.5](assets/portrait_strength_0.5.jpg) Β· [1.0](assets/portrait_strength_1.0.jpg) Β· [1.5](assets/portrait_strength_1.5.jpg) β 1024Γ1024, seed 4242, 2 steps.
|
| 469 |
+
|
| 470 |
+
What the dial scales is this adapter's whole job at two steps: turning a pair of coarse calls into a finished image. So
|
| 471 |
+
it behaves differently from a 4-step adapter's strength control, where the base render is already coherent and the LoRA
|
| 472 |
+
only adds the missing texture.
|
| 473 |
+
|
| 474 |
+
| strength | what happens |
|
| 475 |
+
| --- | --- |
|
| 476 |
+
| **below 1.0** | the correction is only partly applied β softer skin and hair, less fine structure, closer to what two steps look like without the adapter. There is less reason to reach for it here than at four steps, where the base render stands on its own |
|
| 477 |
+
| **1.0** | the trained point, and the recommendation |
|
| 478 |
+
| **1.0β1.5** | extrapolation past training: texture grows denser than the subject warrants and fine structure begins to read as wiry rather than sharp. Usable if you want that look, on a prompt-by-prompt basis |
|
| 479 |
+
| **above 1.5** | not recommended, and not measured here |
|
| 480 |
+
|
| 481 |
+
If a render is not giving you what you want, **reach for steps before strength**: this adapter is trained for two, and the
|
| 482 |
+
[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) is the better tool whenever quality
|
| 483 |
+
matters more than speed.
|
| 484 |
+
|
| 485 |
+
## File format and compatibility
|
| 486 |
+
|
| 487 |
+
A plain `.safetensors` file β **not tied to any framework or backend**. It is weights plus a naming
|
| 488 |
+
convention, so it loads under PyTorch (CUDA, MPS or CPU), MLX on Apple Silicon, or anything else
|
| 489 |
+
that can read safetensors and do a matrix multiply.
|
| 490 |
+
|
| 491 |
+
| | |
|
| 492 |
+
| --------------- | --------------------------------------------- |
|
| 493 |
+
| container | `safetensors` |
|
| 494 |
+
| adapter weights | bf16 (`lora_A`, `lora_B`) |
|
| 495 |
+
| `alpha` | fp32 scalar per module, **64.0** |
|
| 496 |
+
| rank | 64 β effective scale `alpha / rank` = **1.0** |
|
| 497 |
+
|
| 498 |
+
Keys are **diffusers module paths** with PEFT-style suffixes:
|
| 499 |
+
|
| 500 |
+
```
|
| 501 |
+
transformer_blocks.0.attn.to_gate.lora_A.weight (64, 6144)
|
| 502 |
+
transformer_blocks.0.attn.to_gate.lora_B.weight (6144, 64)
|
| 503 |
+
transformer_blocks.0.attn.to_gate.alpha scalar
|
| 504 |
+
time_embed.linear_2.lora_A.weight ...
|
| 505 |
+
```
|
| 506 |
+
|
| 507 |
+
applied the standard way:
|
| 508 |
+
|
| 509 |
+
```
|
| 510 |
+
W' = W + (alpha / rank) Β· (B @ A)
|
| 511 |
+
```
|
| 512 |
+
|
| 513 |
+
The one thing to watch when porting is **naming, not framework**. Runtimes that use their own layer
|
| 514 |
+
names β ComfyUI, for instance, calls these `diffusion_model.blocks.N.attn.gate` with
|
| 515 |
+
`lora_down`/`lora_up` β need a key remap first. The tensors themselves need no conversion.
|
| 516 |
+
|
| 517 |
+
### The keys, runtime by runtime
|
| 518 |
+
|
| 519 |
+
The module paths in the file are the ones diffusers' `Krea2Transformer2DModel` uses for its layers β
|
| 520 |
+
`transformer_blocks.N.attn.to_q`, `ff.up`, `time_embed.linear_2`, and so on β so they name the right
|
| 521 |
+
tensors in any runtime that follows the diffusers architecture. What differs between runtimes is the
|
| 522 |
+
wrapping around those paths:
|
| 523 |
+
|
| 524 |
+
| runtime | what it expects | what to do |
|
| 525 |
+
| ---------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- |
|
| 526 |
+
| **diffusers** (`pipe.load_lora_weights`) | every key prefixed with the pipeline component it belongs to β `transformer.` here β because one pipeline LoRA file may carry adapters for several components; the scale comes from the adapter config, so the loader drops `alpha` | add the prefix and drop `.alpha`, as in the snippet above |
|
| 527 |
+
| **ComfyUI** | its own layer names, `diffusion_model.blocks.N.attn.wq` with `lora_down`/`lora_up`, no `alpha` keys (absent alpha defaults to the rank) | use the `_comfyui` file |
|
| 528 |
+
| **MLX and custom loaders** | nothing in particular | read the keys as they are and apply `W + (alpha/rank)Β·(B @ A)` |
|
| 529 |
+
|
| 530 |
+
In every case the tensors are the same 456 bf16 matrices; only the names around them change.
|
| 531 |
+
|
| 532 |
+
## Method
|
| 533 |
+
|
| 534 |
+
**Distribution matching with a trajectory anchor**, Krea 2 Turbo as its own teacher, on the recorded 8-step
|
| 535 |
+
trajectories.
|
| 536 |
+
|
| 537 |
+
The student makes two calls, at Ο = 1.0 and Ο = 0.7595 β the first and fifth points of the teacher's 8-step grid at
|
| 538 |
+
mu = 1.15 β and stock Euler carries it between them. That grid is what makes the objective a drop-in: Euler's first step
|
| 539 |
+
from pure noise lands _exactly_ on the flow-matching interpolant at Ο = 0.7595 with the same noise and the student's
|
| 540 |
+
own clean-image prediction as the data point. So the student's first-call output is a legitimate image prediction that
|
| 541 |
+
can be judged as an image, and the second call is fed from it during training the way it will be at inference.
|
| 542 |
+
|
| 543 |
+
**The distribution term.** For an image the student produces, two denoisers estimate how it should be cleaned up from a
|
| 544 |
+
freshly noised copy: the frozen teacher, and a second small adapter on the same frozen base β the _fake score_ β that
|
| 545 |
+
is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
|
| 546 |
+
the image more like the teacher's work and less like the student's habits, and the student is pushed that way
|
| 547 |
+
(the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
|
| 548 |
+
rank 32, starts as an exact copy of the teacher, updates four times per student step β often enough to keep up with a
|
| 549 |
+
student that is still changing β and is discarded at the end.
|
| 550 |
+
|
| 551 |
+
**The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
|
| 552 |
+
student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
|
| 553 |
+
it is what the earlier trajectory distillation had already satisfied β which is why the first distribution-matching run
|
| 554 |
+
moved so far so fast.
|
| 555 |
+
|
| 556 |
+
**Per resolution.** The distribution term's push grows with resolution: the fake adapter sees few large-bucket samples
|
| 557 |
+
and under-fits fine structure there, and a per-pixel normaliser lands harder as pixel counts grow. A full 12-bucket
|
| 558 |
+
sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
|
| 559 |
+
1.4β1.8Γ the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
|
| 560 |
+
not hold, and the spectral match replaced it. The same sweep of the published checkpoint, its running-average weights, fixed
|
| 561 |
+
seed, 15 prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
|
| 562 |
+
first distribution-matching checkpoint on the same prompts):
|
| 563 |
+
|
| 564 |
+
| bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
|
| 565 |
+
| --------- | --------------------------- | --------------- | -------------- | ----------------------- |
|
| 566 |
+
| 512x512 | 1.02 (1.16) | 1.02 (1.23) | 1.05 (1.23) | 0.41 (0.44) |
|
| 567 |
+
| 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.06 (1.33) | 0.39 (0.41) |
|
| 568 |
+
| 768x512 | 1.02 (1.22) | 1.08 (1.37) | 1.04 (1.26) | 0.44 (0.47) |
|
| 569 |
+
| 768x768 | 1.09 (1.43) | 1.05 (1.46) | 1.14 (1.50) | 0.40 (0.43) |
|
| 570 |
+
| 768x1024 | 1.02 (1.36) | 0.98 (1.38) | 1.07 (1.42) | 0.39 (0.43) |
|
| 571 |
+
| 1024x768 | 1.01 (1.35) | 0.98 (1.40) | 1.04 (1.38) | 0.43 (0.46) |
|
| 572 |
+
| 1024x1024 | 1.06 (1.47) | 1.07 (1.55) | 1.12 (1.56) | 0.41 (0.44) |
|
| 573 |
+
| 1280x960 | 1.03 (1.50) | 1.04 (1.59) | 1.07 (1.56) | 0.40 (0.42) |
|
| 574 |
+
| 960x1280 | 1.10 (1.60) | 1.05 (1.63) | 1.19 (1.70) | 0.40 (0.43) |
|
| 575 |
+
| 1280x1280 | 1.09 (1.65) | 1.08 (1.68) | 1.13 (1.69) | 0.41 (0.44) |
|
| 576 |
+
| 1440x1280 | 1.08 (1.59) | 1.08 (1.66) | 1.14 (1.66) | 0.40 (0.42) |
|
| 577 |
+
| 1440x1440 | 1.17 (1.81) | 1.08 (1.79) | 1.21 (1.89) | 0.39 (0.41) |
|
| 578 |
+
|
| 579 |
+
### The spectral match
|
| 580 |
+
|
| 581 |
+
Distribution matching delivers its push one latent token at a time, and a token is a 16Γ16-pixel block. At large
|
| 582 |
+
resolutions the push that tells the student to add detail also, unintentionally, adds detail in the shape of that
|
| 583 |
+
block: a faint checkerboard at the 16-pixel and 8-pixel periods that reads as grain on birds, food and foliage. It is
|
| 584 |
+
measurable in the images' Fourier spectra β the student carried about twice the teacher's energy at exactly those two
|
| 585 |
+
periods β and several ways of reshaping or re-weighting the push (a critic, per-region weights, a filtered push) were
|
| 586 |
+
tried and retired: none of them could reach a pattern that lives inside the token.
|
| 587 |
+
|
| 588 |
+
What can is a term that judges the **picture**, not the push. For every second-call sample, the student's estimate of
|
| 589 |
+
the finished image and the teacher's own image for the same prompt and seed are compared through their radial power
|
| 590 |
+
spectra β how much energy each has at every spatial scale β as an L1 on the log spectrum, in two places: on the whole
|
| 591 |
+
latent every sample, and on one random 256-pixel window decoded through the VAE so colour and decoder behaviour are
|
| 592 |
+
judged too. The comparison is two-sided, so too much fine energy and too little are both penalised; a blurred image
|
| 593 |
+
does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
|
| 594 |
+
refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
|
| 595 |
+
without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too. The whole-latent comparison trains at that
|
| 596 |
+
strength; the decoded window was later brought down to a quarter of it, which kept the detail and removed some grain.
|
| 597 |
+
|
| 598 |
+
### The artefact critic
|
| 599 |
+
|
| 600 |
+
Distribution matching improves what the fake adapter can see, and the fake adapter learns from the student's own
|
| 601 |
+
images β so where the student smears something, the fake learns the smear and the push stops correcting it. The two
|
| 602 |
+
places that shows most are small faces, which come out melted, and dense content such as the goods on a market stall or
|
| 603 |
+
the shelves of a shop seen through its window, which comes out as coloured blotches.
|
| 604 |
+
|
| 605 |
+
A critic breaks that loop by looking at finished images instead. It is a small head on the frozen base model's
|
| 606 |
+
mid-network features β every adapter switched off, the forward pass stopped halfway, so no adapter can learn to fool
|
| 607 |
+
it and nothing it learns leaks into the fake adapter. Its real examples are the teacher's own finished images for
|
| 608 |
+
other prompts at the same resolution; its fake examples are the student's final images. Both are lightly re-noised
|
| 609 |
+
first, in the low-noise range where fine structure lives, and the critic reads them with an empty prompt, so it judges
|
| 610 |
+
only whether the structure looks like the teacher's. Its push on the student is filtered to periods finer than 32
|
| 611 |
+
pixels β below that it can rebuild a face or a shelf, above it it could move layout, colour or pose, which it must not
|
| 612 |
+
β and capped at a quarter of the distribution term's strength per sample. Lazy gradient regularisation and spectral
|
| 613 |
+
normalisation keep the head from overshooting. On the largest steps, where memory is tightest, the critic and the
|
| 614 |
+
teacher's finishing pass below take turns instead of sharing a step.
|
| 615 |
+
|
| 616 |
+
### Critics in turn
|
| 617 |
+
|
| 618 |
+
One critic holds one idea of what is wrong. The artefact critic was joined by two more on the same frozen mid-network
|
| 619 |
+
features and under the same rules β lightly re-noised inputs, an empty prompt, a push that is filtered and capped β each
|
| 620 |
+
aimed at a different fault:
|
| 621 |
+
|
| 622 |
+
- **A photo critic.** Half of its real examples are real photographs and half the teacher's finished images, so it learns
|
| 623 |
+
what fine texture looks like in a photograph as well as in the teacher's rendering of one. Its push is filtered to
|
| 624 |
+
periods finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher
|
| 625 |
+
does not.
|
| 626 |
+
- **A face critic.** Its real examples are the teacher's finished images of prompts with faces, the face regions counted at
|
| 627 |
+
full weight and the rest at half, so its push concentrates on what small and mid-sized faces lose first.
|
| 628 |
+
|
| 629 |
+
All three on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
|
| 630 |
+
others train so none goes stale before its turn comes back. The two new heads started from the artefact critic's weights and
|
| 631 |
+
trained on their own before they were allowed to push.
|
| 632 |
+
|
| 633 |
+
### Detail terms
|
| 634 |
+
|
| 635 |
+
Four smaller terms sit on top, each capped relative to the distribution term so none of them can take over:
|
| 636 |
+
|
| 637 |
+
- **A detail-weighted anchor.** The trajectory regression counts the fine-detail part of its error β everything finer
|
| 638 |
+
than 32 pixels β twice, so the anchor stops tolerating softness it used to average away.
|
| 639 |
+
- **A one-sided photo floor.** On the decoded window, the student's energy at periods of 3β10 pixels may not fall below
|
| 640 |
+
the teacher's plus the margin real photographs carry over it at those scales. That margin is measured once from a
|
| 641 |
+
pool of real photographs and clamped, and the term only ever pushes upward to that floor, never past it β so it
|
| 642 |
+
lifts detail that is missing without adding grain that is not.
|
| 643 |
+
- **The teacher's finish as a target.** Every second step, the teacher itself runs its remaining steps starting from
|
| 644 |
+
the student's own first-call output. The result is a finished image that shares the student's layout, and the second
|
| 645 |
+
call is pulled gently toward it β a target that lines up with what the student actually drew, where the recorded
|
| 646 |
+
trajectory might have drawn something else.
|
| 647 |
+
- **A smoothness limit on the fake adapter.** The fake adapter's fine-detail energy is kept below the teacher's at the
|
| 648 |
+
point where the distribution term is measured, so the difference between the two keeps pointing toward detail.
|
| 649 |
+
|
| 650 |
+
## What the LoRA touches
|
| 651 |
+
|
| 652 |
+
Rank **64**, alpha = rank, bf16, the same **228 modules** as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA): the 8 attention and feed-forward linears
|
| 653 |
+
of all 28 transformer blocks, plus the four global linears β `time_embed.linear_1`, `time_embed.linear_2`,
|
| 654 |
+
`time_mod_proj`, `final_layer.linear` β that a step-count change needs most. Nothing about the base model changes.
|
| 655 |
+
|
| 656 |
+
## Training data
|
| 657 |
+
|
| 658 |
+
The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β Krea 2 Turbo's own 8-step run at mu = 1.15 and
|
| 659 |
+
guidance 0.0, every latent and velocity stored β serve unchanged: a 2-step chord is two of the 4-step chords end to
|
| 660 |
+
end. 203 held-out prompts measure the studentβteacher gap on unseen prompts and never receive a gradient. The spectral
|
| 661 |
+
match and the artefact critic read the teacher's finals for the training prompts. The 43,044 real-photo crops of the
|
| 662 |
+
[4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) enter in two places: as one precomputed statistic β how much fine-detail energy they carry at 3β10
|
| 663 |
+
pixels relative to the teacher, clamped β which sets the photo floor, and as half of the photo critic's real examples. Only that
|
| 664 |
+
critic's head sees them; the student and the fake adapter never do, and receive only its filtered, capped push.
|
| 665 |
+
|
| 666 |
+
## Resolutions
|
| 667 |
+
|
| 668 |
+
The same 12 buckets as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA), interleaved in proportion to their remaining samples:
|
| 669 |
+
|
| 670 |
+
| | | |
|
| 671 |
+
| --------- | --------- | --------- |
|
| 672 |
+
| 512Γ512 | 512Γ768 | 768Γ512 |
|
| 673 |
+
| 768Γ768 | 768Γ1024 | 1024Γ768 |
|
| 674 |
+
| 1024Γ1024 | 960Γ1280 | 1280Γ960 |
|
| 675 |
+
| 1280Γ1280 | 1440Γ1280 | 1440Γ1440 |
|
| 676 |
+
|
| 677 |
+
## Resolution sweeps
|
| 678 |
+
|
| 679 |
+
The [Examples](#examples) below are all 768Γ1024. A single resolution is not enough to judge an adapter of this kind:
|
| 680 |
+
the shard pool it trains on is never evenly spread across buckets, and adapters carry recency bias, so one can be
|
| 681 |
+
strong at the size it saw most while quietly softer at the ones it barely saw. Only rendering every bucket shows that,
|
| 682 |
+
which is why every cut of this run is rendered across the buckets and why the per-resolution table under
|
| 683 |
+
[Method](#method) above exists.
|
| 684 |
+
|
| 685 |
+
[`assets/resolution_sweeps/`](assets/resolution_sweeps) holds the evidence β same 15 prompts, same seed, one folder per
|
| 686 |
+
resolution, one image per prompt, so any image can be compared 1:1 with its twin in the next tree:
|
| 687 |
+
|
| 688 |
+
- [`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step) β the **official Krea 2 Turbo 8-step reference renders**:
|
| 689 |
+
the stock model, no LoRA, at its native settings (8 steps, guidance 0.0). The teacher this LoRA is distilled from and
|
| 690 |
+
the quality bar it is measured against.
|
| 691 |
+
- [`_turbo-base-NO-LoRA-2step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-2step) β **stock Turbo at 2 steps, no
|
| 692 |
+
LoRA**: the do-nothing baseline, what two steps look like before the adapter. Every claim on this page is checked
|
| 693 |
+
against both trees.
|
| 694 |
+
- [`2step-LoRA/`](assets/resolution_sweeps/2step-LoRA) β **this LoRA at 2 steps**, the same tree: the renders of the
|
| 695 |
+
checkpoint published here, replaced whenever a better one ships. Which checkpoint that is today is in
|
| 696 |
+
[`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md), and every
|
| 697 |
+
published checkpoint's tree is kept under its own number in [`_archive/resolution_sweeps/`](_archive/resolution_sweeps).
|
| 698 |
+
- [`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step) and
|
| 699 |
+
[`1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme) β the out-of-spec single-step material of the
|
| 700 |
+
[bonus section](#bonus-the-1-step-extreme-test) further down.
|
| 701 |
+
|
| 702 |
+
```
|
| 703 |
+
assets/resolution_sweeps/
|
| 704 |
+
βββ _teacher-8step/ the official 8-step stock-Turbo reference renders (no LoRA)
|
| 705 |
+
β βββ 512x512/ one folder per resolution
|
| 706 |
+
β β βββ portrait.jpg
|
| 707 |
+
β β βββ kingfisher.jpg
|
| 708 |
+
β β βββ β¦ 13 more, one per test prompt
|
| 709 |
+
β β βββ snowleopard.jpg
|
| 710 |
+
β βββ 768x1024/
|
| 711 |
+
β βββ 1024x1024/
|
| 712 |
+
β βββ 1280x1280/
|
| 713 |
+
β βββ 1440x1280/ β¦and the remaining buckets
|
| 714 |
+
β βββ 1440x1440/
|
| 715 |
+
βββ _turbo-base-NO-LoRA-2step/ the same tree, stock Turbo at 2 steps β the floor
|
| 716 |
+
βββ 2step-LoRA/ the same tree, rendered with this LoRA at 2 steps β the published checkpoint
|
| 717 |
+
βββ _turbo-base-NO-LoRA-1step/ stock Turbo at ONE step (bonus section)
|
| 718 |
+
βββ 1step-LoRA-extreme/ this LoRA at ONE step, with side-by-side strips (bonus section)
|
| 719 |
+
```
|
| 720 |
+
|
| 721 |
+
Two ways to read them, both useful:
|
| 722 |
+
|
| 723 |
+
- π **down the sweep** β does the adapter hold together across every resolution, or is it strong at one size and soft
|
| 724 |
+
at others?
|
| 725 |
+
- π― **against the teacher and the floor** β open the same `<WxH>/<prompt>.jpg` under
|
| 726 |
+
[`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step) to see how close two steps with the LoRA get to the
|
| 727 |
+
full 8-step render, and under [`_turbo-base-NO-LoRA-2step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-2step) to
|
| 728 |
+
see what those two steps looked like without it.
|
| 729 |
+
|
| 730 |
+
The resolutions are exactly the training buckets listed under [Resolutions](#resolutions).
|
| 731 |
+
|
| 732 |
+
## Hardware
|
| 733 |
+
|
| 734 |
+
One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned
|
| 735 |
+
host memory above 0.3 megapixels; the student, the fake adapter, the spectral and detail terms, the critic and the teacher's finishing pass each
|
| 736 |
+
build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
|
| 737 |
+
does not fit fails loudly. A full step with every term live and every critic pushing reserves about 22.4 GB at 1440Γ1440,
|
| 738 |
+
of 24. The price of the objective is throughput: **about 107 training samples an hour** measured over the recipe's complete
|
| 739 |
+
run, against the [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β more than four times the cost per sample, and so far a small
|
| 740 |
+
fraction of the samples.
|
| 741 |
+
|
| 742 |
+
**Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
|
| 743 |
+
The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
|
| 744 |
+
disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
|
| 745 |
+
of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
|
| 746 |
+
alongside to imitate the student β and that second adapter takes four optimiser steps of its own per student step.
|
| 747 |
+
The spectral and detail terms decode part of the image out of the latent to compare its texture with the teacher's,
|
| 748 |
+
each critic reads half the network twice more, and every second step the teacher finishes the image from the student's
|
| 749 |
+
first call.
|
| 750 |
+
|
| 751 |
+
Counted in whole model runs per training sample, the difference is roughly **two there against about a dozen here**. None of
|
| 752 |
+
that difference is the teacher generating anything: its renders were recorded once for the 4-step project and are
|
| 753 |
+
read from disk by both. The extra work is the objective itself, and it bought the only thing that mattered. Run at
|
| 754 |
+
two steps, the 4-step project's recipe reached a point where more training changed nothing: the measurements sat
|
| 755 |
+
flat and every new checkpoint had the same faults as the one before β doubled contours on faces and limbs, soft
|
| 756 |
+
fine texture, crowds blurred into one another. Distribution matching is the change that made each new checkpoint
|
| 757 |
+
visibly better than the last again.
|
| 758 |
+
|
| 759 |
+
## How it is judged
|
| 760 |
+
|
| 761 |
+
At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
|
| 762 |
+
15 fixed prompts across four resolutions (512Γ512, 768Γ1024, 1280Γ1280, 1440Γ1440); milestone checkpoints get the same
|
| 763 |
+
render at all 12 buckets, which is where the per-resolution table above comes from. Every image is measured against the
|
| 764 |
+
teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
|
| 765 |
+
and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
|
| 766 |
+
seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
|
| 767 |
+
scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
|
| 768 |
+
identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
|
| 769 |
+
checkpoint. Latent distances β the held-out chord gap and the two-step rollout
|
| 770 |
+
error β are recorded but never used to keep or stop a run: this project's clearest lesson is that they reward blur, and
|
| 771 |
+
a run that improved them while its pictures collapsed was stopped by its pictures. I have the last word at every gate,
|
| 772 |
+
looking at the renders.
|
| 773 |
+
|
| 774 |
+
## Examples
|
| 775 |
+
|
| 776 |
+
Every sheet below is three renders at the **same seed**: the base model as shipped, the base model at two steps
|
| 777 |
+
_without_ the LoRA, and the same two steps _with_ it. Each panel is captioned with its own steps, CFG and NFE. Click
|
| 778 |
+
any image for the full-size version.
|
| 779 |
+
|
| 780 |
+
**NFE** = _number of function evaluations_: how many times the model itself is run, and the honest unit of cost β
|
| 781 |
+
steps are not, because a step with CFG runs the model twice (once conditional, once unconditional). Turbo is CFG-free,
|
| 782 |
+
so here NFE equals steps: the 8-step reference costs 8, and this LoRA's 2 steps cost 2. Wall-clock tracks NFE.
|
| 783 |
+
|
| 784 |
+
**How to read these sheets.** Compare the **second and third panels** β they run the same step count and differ only
|
| 785 |
+
by the adapter, so that pair isolates what the LoRA does. The first panel is the quality bar, not a pixel-level
|
| 786 |
+
target: changing the step count moves the sampling trajectory by itself, so the full-step render often differs in
|
| 787 |
+
pose and framing from both reduced-step panels regardless of whether the LoRA is loaded. Same seed throughout; the
|
| 788 |
+
seed fixes the starting noise, not the destination. Because two steps give up more than four, each prompt also gets a
|
| 789 |
+
**two-panel sheet against the teacher alone** β the LoRA's render beside the 8-step one, nothing else in the frame β
|
| 790 |
+
which is the comparison this project is judged on.
|
| 791 |
+
|
| 792 |
+
### Krea 2 Turbo β 8 steps β 2 steps
|
| 793 |
+
|
| 794 |
+
One section per prompt: the three-panel comparison first, then the two-panel comparison against the teacher, then the
|
| 795 |
+
individual renders β click any image for full size.
|
| 796 |
+
|
| 797 |
+
### Portrait of a young woman with freckles and windswept auburn hair, soft window light, shallow depth of field, photograph, sharp detail
|
| 798 |
+
|
| 799 |
+
**3-way comparison** β one image, all three renders side by side
|
| 800 |
+
|
| 801 |
+
[](assets/portrait_compare_turbo.jpg)
|
| 802 |
+
|
| 803 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 804 |
+
|
| 805 |
+
[](assets/portrait_compare_teacher.jpg)
|
| 806 |
+
|
| 807 |
+
**Individual frames** β click any panel to open that render full size
|
| 808 |
+
|
| 809 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 810 |
+
| ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
|
| 811 |
+
| [](assets/portrait_turbo_8step.jpg) | [](assets/portrait_turbo_2step.jpg) | [](assets/portrait_turbo_2step_lora.jpg) |
|
| 812 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 813 |
+
|
| 814 |
+
---
|
| 815 |
+
|
| 816 |
+
### A kingfisher bird bursting out of water with spread wings, water droplets frozen mid-air, iridescent blue and orange feathers, high-speed photography
|
| 817 |
+
|
| 818 |
+
**3-way comparison** β one image, all three renders side by side
|
| 819 |
+
|
| 820 |
+
[](assets/kingfisher_compare_turbo.jpg)
|
| 821 |
+
|
| 822 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 823 |
+
|
| 824 |
+
[](assets/kingfisher_compare_teacher.jpg)
|
| 825 |
+
|
| 826 |
+
**Individual frames** β click any panel to open that render full size
|
| 827 |
+
|
| 828 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 829 |
+
| ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
|
| 830 |
+
| [](assets/kingfisher_turbo_8step.jpg) | [](assets/kingfisher_turbo_2step.jpg) | [](assets/kingfisher_turbo_2step_lora.jpg) |
|
| 831 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 832 |
+
|
| 833 |
+
---
|
| 834 |
+
|
| 835 |
+
### Rainy night city street with glowing neon shop signs and readable text, wet asphalt reflections, pedestrians with umbrellas, cinematic
|
| 836 |
+
|
| 837 |
+
**3-way comparison** β one image, all three renders side by side
|
| 838 |
+
|
| 839 |
+
[](assets/neonstreet_compare_turbo.jpg)
|
| 840 |
+
|
| 841 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 842 |
+
|
| 843 |
+
[](assets/neonstreet_compare_teacher.jpg)
|
| 844 |
+
|
| 845 |
+
**Individual frames** β click any panel to open that render full size
|
| 846 |
+
|
| 847 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 848 |
+
| ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
|
| 849 |
+
| [](assets/neonstreet_turbo_8step.jpg) | [](assets/neonstreet_turbo_2step.jpg) | [](assets/neonstreet_turbo_2step_lora.jpg) |
|
| 850 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 851 |
+
|
| 852 |
+
---
|
| 853 |
+
|
| 854 |
+
### Busy outdoor street market crowded with many people browsing colorful fruit and vegetable stalls, awnings, midday sun, wide shot, photorealistic
|
| 855 |
+
|
| 856 |
+
**3-way comparison** β one image, all three renders side by side
|
| 857 |
+
|
| 858 |
+
[](assets/market_compare_turbo.jpg)
|
| 859 |
+
|
| 860 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 861 |
+
|
| 862 |
+
[](assets/market_compare_teacher.jpg)
|
| 863 |
+
|
| 864 |
+
**Individual frames** β click any panel to open that render full size
|
| 865 |
+
|
| 866 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 867 |
+
| --------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
|
| 868 |
+
| [](assets/market_turbo_8step.jpg) | [](assets/market_turbo_2step.jpg) | [](assets/market_turbo_2step_lora.jpg) |
|
| 869 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 870 |
+
|
| 871 |
+
---
|
| 872 |
+
|
| 873 |
+
### Overhead shot of a rustic wood-fired pizza with bubbling melted cheese, basil leaves, charred crust, on a dark wooden table, food photography
|
| 874 |
+
|
| 875 |
+
**3-way comparison** β one image, all three renders side by side
|
| 876 |
+
|
| 877 |
+
[](assets/pizza_compare_turbo.jpg)
|
| 878 |
+
|
| 879 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 880 |
+
|
| 881 |
+
[](assets/pizza_compare_teacher.jpg)
|
| 882 |
+
|
| 883 |
+
**Individual frames** β click any panel to open that render full size
|
| 884 |
+
|
| 885 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 886 |
+
| ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
|
| 887 |
+
| [](assets/pizza_turbo_8step.jpg) | [](assets/pizza_turbo_2step.jpg) | [](assets/pizza_turbo_2step_lora.jpg) |
|
| 888 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 889 |
+
|
| 890 |
+
---
|
| 891 |
+
|
| 892 |
+
### A young swordsman leaping through falling cherry blossoms, dynamic action pose, anime key visual, crisp linework, vivid colors
|
| 893 |
+
|
| 894 |
+
**3-way comparison** β one image, all three renders side by side
|
| 895 |
+
|
| 896 |
+
[](assets/swordsman_compare_turbo.jpg)
|
| 897 |
+
|
| 898 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 899 |
+
|
| 900 |
+
[](assets/swordsman_compare_teacher.jpg)
|
| 901 |
+
|
| 902 |
+
**Individual frames** β click any panel to open that render full size
|
| 903 |
+
|
| 904 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 905 |
+
| --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
|
| 906 |
+
| [](assets/swordsman_turbo_8step.jpg) | [](assets/swordsman_turbo_2step.jpg) | [](assets/swordsman_turbo_2step_lora.jpg) |
|
| 907 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 908 |
+
|
| 909 |
+
---
|
| 910 |
+
|
| 911 |
+
### A giant mecha standing in a rain-soaked city plaza, anime style, panel lining, glowing cockpit, dramatic low angle
|
| 912 |
+
|
| 913 |
+
**3-way comparison** β one image, all three renders side by side
|
| 914 |
+
|
| 915 |
+
[](assets/mecha_compare_turbo.jpg)
|
| 916 |
+
|
| 917 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 918 |
+
|
| 919 |
+
[](assets/mecha_compare_teacher.jpg)
|
| 920 |
+
|
| 921 |
+
**Individual frames** β click any panel to open that render full size
|
| 922 |
+
|
| 923 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 924 |
+
| ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
|
| 925 |
+
| [](assets/mecha_turbo_8step.jpg) | [](assets/mecha_turbo_2step.jpg) | [](assets/mecha_turbo_2step_lora.jpg) |
|
| 926 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 927 |
+
|
| 928 |
+
---
|
| 929 |
+
|
| 930 |
+
### A fox in a red scarf reading a book under a mushroom, children's storybook illustration, watercolour texture, soft edges
|
| 931 |
+
|
| 932 |
+
**3-way comparison** β one image, all three renders side by side
|
| 933 |
+
|
| 934 |
+
[](assets/storybookfox_compare_turbo.jpg)
|
| 935 |
+
|
| 936 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 937 |
+
|
| 938 |
+
[](assets/storybookfox_compare_teacher.jpg)
|
| 939 |
+
|
| 940 |
+
**Individual frames** β click any panel to open that render full size
|
| 941 |
+
|
| 942 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 943 |
+
| --------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- |
|
| 944 |
+
| [](assets/storybookfox_turbo_8step.jpg) | [](assets/storybookfox_turbo_2step.jpg) | [](assets/storybookfox_turbo_2step_lora.jpg) |
|
| 945 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 946 |
+
|
| 947 |
+
---
|
| 948 |
+
|
| 949 |
+
### A curious young inventor girl with oversized goggles, 3D animated film style, subsurface skin, soft studio lighting, shallow depth of field
|
| 950 |
+
|
| 951 |
+
**3-way comparison** β one image, all three renders side by side
|
| 952 |
+
|
| 953 |
+
[](assets/inventor_compare_turbo.jpg)
|
| 954 |
+
|
| 955 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 956 |
+
|
| 957 |
+
[](assets/inventor_compare_teacher.jpg)
|
| 958 |
+
|
| 959 |
+
**Individual frames** β click any panel to open that render full size
|
| 960 |
+
|
| 961 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 962 |
+
| ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
|
| 963 |
+
| [](assets/inventor_turbo_8step.jpg) | [](assets/inventor_turbo_2step.jpg) | [](assets/inventor_turbo_2step_lora.jpg) |
|
| 964 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 965 |
+
|
| 966 |
+
---
|
| 967 |
+
|
| 968 |
+
### A claymation chef holding a tiny cake, visible fingerprints in the clay, miniature set, tilt-shift
|
| 969 |
+
|
| 970 |
+
**3-way comparison** β one image, all three renders side by side
|
| 971 |
+
|
| 972 |
+
[](assets/claychef_compare_turbo.jpg)
|
| 973 |
+
|
| 974 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 975 |
+
|
| 976 |
+
[](assets/claychef_compare_teacher.jpg)
|
| 977 |
+
|
| 978 |
+
**Individual frames** β click any panel to open that render full size
|
| 979 |
+
|
| 980 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 981 |
+
| ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
|
| 982 |
+
| [](assets/claychef_turbo_8step.jpg) | [](assets/claychef_turbo_2step.jpg) | [](assets/claychef_turbo_2step_lora.jpg) |
|
| 983 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 984 |
+
|
| 985 |
+
---
|
| 986 |
+
|
| 987 |
+
### A gleaming white colony ship in orbit above a turquoise ocean planet, smooth curved hull, glowing cyan engine rings, brilliant sunlight, clean sci-fi concept art, bold simple shapes, vivid colors
|
| 988 |
+
|
| 989 |
+
**3-way comparison** β one image, all three renders side by side
|
| 990 |
+
|
| 991 |
+
[](assets/colonyship_compare_turbo.jpg)
|
| 992 |
+
|
| 993 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 994 |
+
|
| 995 |
+
[](assets/colonyship_compare_teacher.jpg)
|
| 996 |
+
|
| 997 |
+
**Individual frames** β click any panel to open that render full size
|
| 998 |
+
|
| 999 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1000 |
+
| ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
|
| 1001 |
+
| [](assets/colonyship_turbo_8step.jpg) | [](assets/colonyship_turbo_2step.jpg) | [](assets/colonyship_turbo_2step_lora.jpg) |
|
| 1002 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1003 |
+
|
| 1004 |
+
---
|
| 1005 |
+
|
| 1006 |
+
### A sleek winged drone gliding between glowing futuristic skyscrapers at night, bright lit avenue far below, deep blue sky above, digital matte painting, bold clean forms, vivid colors
|
| 1007 |
+
|
| 1008 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1009 |
+
|
| 1010 |
+
[](assets/megacity_compare_turbo.jpg)
|
| 1011 |
+
|
| 1012 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1013 |
+
|
| 1014 |
+
[](assets/megacity_compare_teacher.jpg)
|
| 1015 |
+
|
| 1016 |
+
**Individual frames** β click any panel to open that render full size
|
| 1017 |
+
|
| 1018 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1019 |
+
| ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
|
| 1020 |
+
| [](assets/megacity_turbo_8step.jpg) | [](assets/megacity_turbo_2step.jpg) | [](assets/megacity_turbo_2step_lora.jpg) |
|
| 1021 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1022 |
+
|
| 1023 |
+
---
|
| 1024 |
+
|
| 1025 |
+
### A storm sorceress channelling lightning, video-game splash art, bold rim lighting, energetic brush strokes, high contrast
|
| 1026 |
+
|
| 1027 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1028 |
+
|
| 1029 |
+
[](assets/sorceress_compare_turbo.jpg)
|
| 1030 |
+
|
| 1031 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1032 |
+
|
| 1033 |
+
[](assets/sorceress_compare_teacher.jpg)
|
| 1034 |
+
|
| 1035 |
+
**Individual frames** β click any panel to open that render full size
|
| 1036 |
+
|
| 1037 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1038 |
+
| --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
|
| 1039 |
+
| [](assets/sorceress_turbo_8step.jpg) | [](assets/sorceress_turbo_2step.jpg) | [](assets/sorceress_turbo_2step_lora.jpg) |
|
| 1040 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1041 |
+
|
| 1042 |
+
---
|
| 1043 |
+
|
| 1044 |
+
### A formula 1 futuristic looking racing car beefed up with a lot of technology mid-corner on a wet track, motion blur background, photorealistic motorsport photography
|
| 1045 |
+
|
| 1046 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1047 |
+
|
| 1048 |
+
[](assets/racecar_compare_turbo.jpg)
|
| 1049 |
+
|
| 1050 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1051 |
+
|
| 1052 |
+
[](assets/racecar_compare_teacher.jpg)
|
| 1053 |
+
|
| 1054 |
+
**Individual frames** β click any panel to open that render full size
|
| 1055 |
+
|
| 1056 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1057 |
+
| ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
| 1058 |
+
| [](assets/racecar_turbo_8step.jpg) | [](assets/racecar_turbo_2step.jpg) | [](assets/racecar_turbo_2step_lora.jpg) |
|
| 1059 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1060 |
+
|
| 1061 |
+
---
|
| 1062 |
+
|
| 1063 |
+
### A snow leopard walking along a rocky ridge in falling snow, telephoto wildlife photograph, natural light
|
| 1064 |
+
|
| 1065 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1066 |
+
|
| 1067 |
+
[](assets/snowleopard_compare_turbo.jpg)
|
| 1068 |
+
|
| 1069 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1070 |
+
|
| 1071 |
+
[](assets/snowleopard_compare_teacher.jpg)
|
| 1072 |
+
|
| 1073 |
+
**Individual frames** β click any panel to open that render full size
|
| 1074 |
+
|
| 1075 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1076 |
+
| ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
|
| 1077 |
+
| [](assets/snowleopard_turbo_8step.jpg) | [](assets/snowleopard_turbo_2step.jpg) | [](assets/snowleopard_turbo_2step_lora.jpg) |
|
| 1078 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1079 |
+
|
| 1080 |
+
---
|
| 1081 |
+
|
| 1082 |
+
## Notes and limitations
|
| 1083 |
+
|
| 1084 |
+
- π― **Krea 2 Turbo only**, at **2 steps**, guidance **0.0** (cfg 1.0 in ComfyUI), **mu = 1.15** β the two training
|
| 1085 |
+
sigmas are anchored to that grid.
|
| 1086 |
+
- π§ͺ **Not a finished adapter.** Usable at 2 steps for previews and drafts; not the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s quality, which remains the recommendation for quality renders. Training continues, and a later checkpoint replaces this file only when the sweeps and I visually agree it is better.
|
| 1087 |
+
|
| 1088 |
+
## Bonus: the 1-step extreme test
|
| 1089 |
+
|
| 1090 |
+
> **This is an extreme, out-of-spec experiment β not recommended for any use.** This LoRA is trained for 2 steps; at
|
| 1091 |
+
> 1 step it is half its trained step count and an eighth of the teacher's.
|
| 1092 |
+
|
| 1093 |
+
Stock Krea 2 Turbo and Turbo + this LoRA, each run at just **one step** β a single call β on the same 15 prompts, the
|
| 1094 |
+
same seed, at all 12 resolutions of the sweep, the LoRA at its normal strength (1.0). Stock Turbo returns a smear at one
|
| 1095 |
+
step: a colour field with a ghost of the subject in it. With the LoRA the same single call returns a coherent picture β
|
| 1096 |
+
the subject, the composition, the lighting and the colours are all there. What is missing is the fine detail the second
|
| 1097 |
+
step adds: skin is soft, hair and fur come out streaked rather than in strands, and the finest structure (feathers,
|
| 1098 |
+
falling snow, small text) is largely absent. That makes one step a rough **preview of composition and colour** at an
|
| 1099 |
+
eighth of the teacher's cost, and nothing more.
|
| 1100 |
+
|
| 1101 |
+
The full set is in [`assets/resolution_sweeps/1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme), one
|
| 1102 |
+
folder per resolution: the LoRA's single-step render as `<prompt>.jpg` and the side-by-side strip as
|
| 1103 |
+
`<prompt>_comparison.jpg` (native on the left, LoRA on the right); stock Turbo's single-step renders are in
|
| 1104 |
+
[`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step). A few samples at 768Γ1024:
|
| 1105 |
+
|
| 1106 |
+
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/portrait_comparison.jpg)
|
| 1107 |
+
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/inventor_comparison.jpg)
|
| 1108 |
+
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/mecha_comparison.jpg)
|
| 1109 |
+
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/pizza_comparison.jpg)
|
| 1110 |
+
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/racecar_comparison.jpg)
|
| 1111 |
+
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/sorceress_comparison.jpg)
|
| 1112 |
+
|
| 1113 |
+
_Still out-of-spec and still edge case preview-only β this is one step, not the 2-step regime the rest of this page measures._
|
| 1114 |
+
|
| 1115 |
+
> π **A 1-step adapter is a possible follow-on project.** The picture above is why: that a single call already holds
|
| 1116 |
+
> together with an adapter trained for two suggests a dedicated one-step adapter is worth attempting once this one
|
| 1117 |
+
> ships β as a booster on top of this LoRA rather than a replacement, trained by distribution matching alone (at one
|
| 1118 |
+
> step there is no trajectory left to regress), and judged on seed variety as much as on detail, since one-step students
|
| 1119 |
+
> are the ones that collapse to a favourite. Expectations set accordingly: a usable preview at an eighth of the
|
| 1120 |
+
> teacher's cost, not the quality bar.
|
| 1121 |
+
|
| 1122 |
+
## What's next
|
| 1123 |
+
|
| 1124 |
+
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any
|
| 1125 |
+
resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). The next step
|
| 1126 |
+
is already running, and it goes after what this checkpoint still gets wrong: a critic that judges the first call against the
|
| 1127 |
+
teacher's own intermediate state from the same noise, so a first call that blends two layouts is caught where the blend
|
| 1128 |
+
happens; a critic weighted toward wherever the teacher put fine detail; the photo critic and the photo floor limited to
|
| 1129 |
+
photographic prompts, so illustration, anime and 3D renders are no longer pulled toward photographic grain; detail held to the
|
| 1130 |
+
teacher region by region, with a ceiling as well as a floor; a focus on the eyes, nose and lips of faces so they sharpen while
|
| 1131 |
+
skin stays the teacher's; a colour floor, so colour at large sizes stops falling below the teacher's; and a more even mix of
|
| 1132 |
+
resolutions. After it comes prompt adherence β a critic that learns whether an image belongs to its own prompt, and stylised
|
| 1133 |
+
prompts drawn more often β held to the rule that none of it may cost the sharpness this checkpoint gained. A better checkpoint
|
| 1134 |
+
replaces this one when the sweeps and I visually agree, the same discipline as the
|
| 1135 |
+
[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
|
| 1136 |
+
one is the fast preview.
|
| 1137 |
+
|
| 1138 |
+
## License
|
| 1139 |
+
|
| 1140 |
+
The adapter is a derivative of [Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo) and is covered by the
|
| 1141 |
+
**Krea 2 Community License Agreement** (`LICENSE.pdf` in this repository), as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) is. The attribution
|
| 1142 |
+
notice the license requires of a derivative ships with the files.
|
README.md
CHANGED
|
The diff for this file is too large to render.
See raw diff
|
|
|