Instructions to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("krea/Krea-2-Turbo", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("lvladikov/Krea2-Turbo-Distill-2step-LoRA") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
chk17464 release
Browse files
README.md
CHANGED
|
@@ -17,11 +17,22 @@ pipeline_tag: text-to-image
|
|
| 17 |
|
| 18 |
# Krea 2 Turbo β 2-Step Distillation LoRA
|
| 19 |
|
| 20 |
-
> π§ͺ **A fast-preview adapter, from a project still in training.**
|
| 21 |
-
>
|
| 22 |
-
>
|
| 23 |
-
>
|
| 24 |
-
>
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 25 |
|
| 26 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
|
| 27 |
down to 2** β Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β aiming at
|
|
@@ -32,9 +43,9 @@ recommendation for quality renders.
|
|
| 32 |
- π― **The aim** β the best two-step quality this base can give, at every one of the same 12 resolutions, measured
|
| 33 |
against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
|
| 34 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas.
|
| 35 |
-
- β±οΈ **
|
| 36 |
-
changes: **
|
| 37 |
-
|
| 38 |
[Performance](#performance).
|
| 39 |
- π **Distribution matching, not imitation** β the training objective that got the renders improving again after the
|
| 40 |
4-step project's recipe had stopped helping at two steps (see [Method](#method)).
|
|
@@ -42,21 +53,23 @@ recommendation for quality renders.
|
|
| 42 |
evaluated on each prompt's own conditioning, so the student is matched to what the teacher makes _for that prompt_,
|
| 43 |
not to a prompt-free look. There is no separate adherence term: instead a vision-language judge checks every
|
| 44 |
checkpoint β each render scored alone against the prompt's objects, counts, attributes and relations, with the
|
| 45 |
-
teacher scored the same way β and a term would only be added if that meter showed adherence slipping.
|
|
|
|
|
|
|
| 46 |
- π **12 trained resolutions** β multi-aspect from 512Γ512 up to 1440Γ1440, each with its sweep.
|
| 47 |
-
- π **Drop-in, no exceptions** β a plain LoRA sampled by stock Euler at sigmas `[1.0, 0.
|
| 48 |
or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this
|
| 49 |
project.
|
| 50 |
- 𧬠**Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β rank 64 on the same 228 modules; a second adapter exists during training
|
| 51 |
only and never ships.
|
| 52 |
- π² **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
|
| 53 |
re-run.
|
| 54 |
-
- π’ **
|
| 55 |
drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
|
| 56 |
training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
|
| 57 |
4-step project, read again at the two sigmas this schedule uses.
|
| 58 |
-
- π
**
|
| 59 |
-
- π **
|
| 60 |
- π₯οΈ **One RTX 3090**, and a recipe shaped by its 24 GB.
|
| 61 |
|
| 62 |
[](assets/poster.jpg)
|
|
@@ -86,13 +99,13 @@ place to check which checkpoint the current files are based on.
|
|
| 86 |
|
| 87 |
| | |
|
| 88 |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 89 |
-
| lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β 2-step trajectory distillation β distribution matching β a spectral match against the teacher's own images
|
| 90 |
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
|
| 91 |
-
| what it gives | usable two-step renders at every trained resolution: fine detail
|
| 92 |
|
| 93 |
### Known issues
|
| 94 |
|
| 95 |
-
The usual costs of two steps, in this order of how often they show:
|
| 96 |
|
| 97 |
## How I got here
|
| 98 |
|
|
@@ -134,6 +147,79 @@ The recipe adjustments so far, each made on the measurement of the one before:
|
|
| 134 |
6. a spectral match of the student's image against the teacher's own, on the whole latent and on decoded pixel
|
| 135 |
windows β the term that finally reached the grain at large resolutions, after a critic, per-resolution weights and
|
| 136 |
a filtered push had each been tried against it and retired
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 137 |
|
| 138 |
## Measured against the teacher
|
| 139 |
|
|
@@ -148,36 +234,41 @@ teacher):
|
|
| 148 |
|
| 149 |
| bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
|
| 150 |
| --- | --- | --- | --- | --- | --- |
|
| 151 |
-
| 512Γ512 | 1.
|
| 152 |
-
| 768Γ1024 | 1.
|
| 153 |
-
| 1024Γ1024 | 1.
|
| 154 |
-
| 1280Γ1280 | 1.
|
| 155 |
-
| 1440Γ1440 | 1.
|
| 156 |
|
| 157 |
Two steps without the adapter carry **less than half** the teacher's fine detail at every size. With it, the detail
|
| 158 |
-
|
|
|
|
| 159 |
|
| 160 |
**Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
|
| 161 |
in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
|
| 162 |
|
| 163 |
| bucket | wins | ties | losses |
|
| 164 |
| --- | --- | --- | --- |
|
| 165 |
-
| 512Γ512 |
|
| 166 |
-
| 1280Γ1280 | 1 |
|
| 167 |
-
| 1440Γ1440 |
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
| 177 |
-
|
| 178 |
-
|
| 179 |
-
|
| 180 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 181 |
not the teacher's render of it β see [Known issues](#known-issues).
|
| 182 |
|
| 183 |
## Usage
|
|
@@ -190,7 +281,7 @@ not the teacher's render of it β see [Known issues](#known-issues).
|
|
| 190 |
| guidance / CFG | **0.0** (Turbo is CFG-free; do not enable it) |
|
| 191 |
| timestep shift | **mu = 1.15**, fixed (Turbo's deployment shift) |
|
| 192 |
|
| 193 |
-
The 2 sampling sigmas are Turbo's own deployment grid: `[1.0, 0.
|
| 194 |
of the 4-step grid, so the model is evaluated at two points it already knows.
|
| 195 |
|
| 196 |
## Inference with diffusers
|
|
@@ -309,18 +400,27 @@ the model already loaded, so the numbers are the render itself and not a model l
|
|
| 309 |
|
| 310 |
| configuration | denoising | per model call | GPU peak |
|
| 311 |
| --- | --- | --- | --- |
|
| 312 |
-
| Krea 2 Turbo β 8 steps (the reference) | **
|
| 313 |
-
| Krea 2 Turbo β 2 steps, no LoRA | 20.
|
| 314 |
-
| **Krea 2 Turbo β 2 steps + this LoRA** | **
|
| 315 |
|
| 316 |
-
**Denoising is
|
| 317 |
-
|
| 318 |
-
small beside the transformer it is added to
|
| 319 |
-
schedule would pay it.
|
| 320 |
|
| 321 |
Denoising is the part the step count changes. What a complete render costs on top of it β encoding the prompt, decoding
|
| 322 |
the latent, writing the file β is the same whether you run two steps or eight, and it depends on your pipeline, so the
|
| 323 |
-
end-to-end figure on your machine will sit below
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 324 |
|
| 325 |
## LoRA strength
|
| 326 |
|
|
@@ -398,9 +498,9 @@ In every case the tensors are the same 456 bf16 matrices; only the names around
|
|
| 398 |
**Distribution matching with a trajectory anchor**, Krea 2 Turbo as its own teacher, on the recorded 8-step
|
| 399 |
trajectories.
|
| 400 |
|
| 401 |
-
The student makes two calls, at Ο = 1.0 and Ο = 0.
|
| 402 |
mu = 1.15 β and stock Euler carries it between them. That grid is what makes the objective a drop-in: Euler's first step
|
| 403 |
-
from pure noise lands _exactly_ on the flow-matching interpolant at Ο = 0.
|
| 404 |
own clean-image prediction as the data point. So the student's first-call output is a legitimate image prediction that
|
| 405 |
can be judged as an image, and the second call is fed from it during training the way it will be at inference.
|
| 406 |
|
|
@@ -409,7 +509,8 @@ freshly noised copy: the frozen teacher, and a second small adapter on the same
|
|
| 409 |
is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
|
| 410 |
the image more like the teacher's work and less like the student's habits, and the student is pushed that way
|
| 411 |
(the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
|
| 412 |
-
rank 32, starts as an exact copy of the teacher, updates
|
|
|
|
| 413 |
|
| 414 |
**The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
|
| 415 |
student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
|
|
@@ -420,24 +521,24 @@ moved so far so fast.
|
|
| 420 |
and under-fits fine structure there, and a per-pixel normaliser lands harder as pixel counts grow. A full 12-bucket
|
| 421 |
sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
|
| 422 |
1.4β1.8Γ the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
|
| 423 |
-
not hold, and the spectral match replaced it. The same sweep of the
|
| 424 |
-
prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
|
| 425 |
first distribution-matching checkpoint on the same prompts):
|
| 426 |
|
| 427 |
| bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
|
| 428 |
| --------- | --------------------------- | --------------- | -------------- | ----------------------- |
|
| 429 |
-
| 512x512 | 1.
|
| 430 |
-
| 512x768 | 1.01 (1.23) | 1.
|
| 431 |
-
| 768x512 | 1.
|
| 432 |
-
| 768x768 | 1.09 (1.43) | 1.
|
| 433 |
-
| 768x1024 | 1.
|
| 434 |
-
| 1024x768 | 1.
|
| 435 |
-
| 1024x1024 | 1.
|
| 436 |
-
| 1280x960 | 1.
|
| 437 |
-
| 960x1280 | 1.
|
| 438 |
-
| 1280x1280 | 1.
|
| 439 |
-
| 1440x1280 | 1.
|
| 440 |
-
| 1440x1440 | 1.
|
| 441 |
|
| 442 |
### The spectral match
|
| 443 |
|
|
@@ -455,8 +556,60 @@ latent every sample, and on one random 256-pixel window decoded through the VAE
|
|
| 455 |
judged too. The comparison is two-sided, so too much fine energy and too little are both penalised; a blurred image
|
| 456 |
does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
|
| 457 |
refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
|
| 458 |
-
without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too
|
| 459 |
-
the
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 460 |
|
| 461 |
## What the LoRA touches
|
| 462 |
|
|
@@ -469,8 +622,10 @@ of all 28 transformer blocks, plus the four global linears β `time_embed.linea
|
|
| 469 |
The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β Krea 2 Turbo's own 8-step run at mu = 1.15 and
|
| 470 |
guidance 0.0, every latent and velocity stored β serve unchanged: a 2-step chord is two of the 4-step chords end to
|
| 471 |
end. 203 held-out prompts measure the studentβteacher gap on unseen prompts and never receive a gradient. The spectral
|
| 472 |
-
match
|
| 473 |
-
|
|
|
|
|
|
|
| 474 |
|
| 475 |
## Resolutions
|
| 476 |
|
|
@@ -541,22 +696,23 @@ The resolutions are exactly the training buckets listed under [Resolutions](#res
|
|
| 541 |
## Hardware
|
| 542 |
|
| 543 |
One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned
|
| 544 |
-
host memory above 0.3 megapixels; the student, the fake adapter
|
| 545 |
-
in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
|
| 546 |
-
does not fit fails loudly. A full step with every term live reserves about
|
| 547 |
-
the objective is throughput: **
|
| 548 |
-
[4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β
|
| 549 |
-
|
| 550 |
|
| 551 |
**Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
|
| 552 |
The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
|
| 553 |
disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
|
| 554 |
of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
|
| 555 |
-
alongside to imitate the student β and that second adapter takes
|
| 556 |
-
|
| 557 |
-
|
|
|
|
| 558 |
|
| 559 |
-
Counted in whole model runs per training sample, the difference is roughly **two there against
|
| 560 |
that difference is the teacher generating anything: its renders were recorded once for the 4-step project and are
|
| 561 |
read from disk by both. The extra work is the objective itself, and it bought the only thing that mattered. Run at
|
| 562 |
two steps, the 4-step project's recipe reached a point where more training changed nothing: the measurements sat
|
|
@@ -570,7 +726,8 @@ At regular intervals, both the live weights and their running average are pulled
|
|
| 570 |
15 fixed prompts across four resolutions (512Γ512, 768Γ1024, 1280Γ1280, 1440Γ1440); milestone checkpoints get the same
|
| 571 |
render at all 12 buckets, which is where the per-resolution table above comes from. Every image is measured against the
|
| 572 |
teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
|
| 573 |
-
and flat-region grain, saturation, faces cut out at 1:1,
|
|
|
|
| 574 |
scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
|
| 575 |
identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
|
| 576 |
checkpoint. Latent distances β the held-out chord gap and the two-step rollout
|
|
@@ -928,12 +1085,19 @@ _Still out-of-spec and still edge case preview-only β this is one step, not th
|
|
| 928 |
|
| 929 |
## What's next
|
| 930 |
|
| 931 |
-
Training continues from
|
| 932 |
-
|
| 933 |
-
|
| 934 |
-
|
| 935 |
-
|
| 936 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 937 |
|
| 938 |
## License
|
| 939 |
|
|
|
|
| 17 |
|
| 18 |
# Krea 2 Turbo β 2-Step Distillation LoRA
|
| 19 |
|
| 20 |
+
> π§ͺ **A fast-preview adapter, from a project still in training.** When the subject is close and fills a good part of the
|
| 21 |
+
> frame β a portrait, a single figure, an object seen up close β two steps already hold up well, and you can rely on
|
| 22 |
+
> this adapter for those images. Small subjects are where it still falls short, people and objects alike: faces in a
|
| 23 |
+
> crowd, figures in a wide scene, the machines at the back of a gym β anything that takes up little of the frame can
|
| 24 |
+
> come out ghosted or smeared. For those, and whenever quality matters more than speed, use the
|
| 25 |
+
> [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Every figure on this page measures the
|
| 26 |
+
> adapter honestly against 4-step and 8-step renders. Training continues one recipe change at a time, and a later
|
| 27 |
+
> checkpoint replaces this one only when the sweeps and I visually agree it is better. Known issues β see
|
| 28 |
+
> [Known issues](#known-issues).
|
| 29 |
+
>
|
| 30 |
+
> π **The saved steps can also go into resolution.** A small subject is simply one that covers few pixels, so a larger
|
| 31 |
+
> render makes the same subject bigger β and at a quarter of the teacher's steps, renders up to 2048Γ2048, Krea's
|
| 32 |
+
> published maximum recommended resolution and beyond the largest size this adapter was trained at (1440Γ1440), come
|
| 33 |
+
> within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past
|
| 34 |
+
> 2048Γ2048, stock Krea 2 itself begins to duplicate subjects β a property of the base model, with or without this
|
| 35 |
+
> adapter.
|
| 36 |
|
| 37 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
|
| 38 |
down to 2** β Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β aiming at
|
|
|
|
| 43 |
- π― **The aim** β the best two-step quality this base can give, at every one of the same 12 resolutions, measured
|
| 44 |
against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
|
| 45 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas.
|
| 46 |
+
- β±οΈ **4.2Γ faster denoising** β the model runs twice instead of eight times, and denoising is the part this adapter
|
| 47 |
+
changes: **81.4 s β 19.5 s** measured at 1024Γ1024 on the same prompts, the adapter's own cost per call within
|
| 48 |
+
measurement noise. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
|
| 49 |
[Performance](#performance).
|
| 50 |
- π **Distribution matching, not imitation** β the training objective that got the renders improving again after the
|
| 51 |
4-step project's recipe had stopped helping at two steps (see [Method](#method)).
|
|
|
|
| 53 |
evaluated on each prompt's own conditioning, so the student is matched to what the teacher makes _for that prompt_,
|
| 54 |
not to a prompt-free look. There is no separate adherence term: instead a vision-language judge checks every
|
| 55 |
checkpoint β each render scored alone against the prompt's objects, counts, attributes and relations, with the
|
| 56 |
+
teacher scored the same way β and a term would only be added if that meter showed adherence slipping. The one part of
|
| 57 |
+
training that looks at images without their prompt is the artefact critic (see [Method](#method)), and it only judges
|
| 58 |
+
whether fine structure looks like the teacher's.
|
| 59 |
- π **12 trained resolutions** β multi-aspect from 512Γ512 up to 1440Γ1440, each with its sweep.
|
| 60 |
+
- π **Drop-in, no exceptions** β a plain LoRA sampled by stock Euler at sigmas `[1.0, 0.7595]` in diffusers, ComfyUI
|
| 61 |
or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this
|
| 62 |
project.
|
| 63 |
- 𧬠**Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β rank 64 on the same 228 modules; a second adapter exists during training
|
| 64 |
only and never ships.
|
| 65 |
- π² **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
|
| 66 |
re-run.
|
| 67 |
+
- π’ **17,464 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β all of them
|
| 68 |
drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
|
| 69 |
training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
|
| 70 |
4-step project, read again at the two sigmas this schedule uses.
|
| 71 |
+
- π
**7 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β and the project continues.
|
| 72 |
+
- π **26 recipe adjustments** across two methods so far β seven of trajectory distillation before the switch, nineteen of distribution matching since β each kept only when the renders did not get worse.
|
| 73 |
- π₯οΈ **One RTX 3090**, and a recipe shaped by its 24 GB.
|
| 74 |
|
| 75 |
[](assets/poster.jpg)
|
|
|
|
| 99 |
|
| 100 |
| | |
|
| 101 |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 102 |
+
| lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β 2-step trajectory distillation β distribution matching β a spectral match against the teacher's own images β an artefact critic and detail terms β three critics taking turns |
|
| 103 |
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
|
| 104 |
+
| what it gives | usable two-step renders at every trained resolution: fine detail at or just above the teacher's β from 1 megapixel up, closer to the teacher than the 4-step adapter β with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 1 point missing out of 240). A judge asked which render follows the prompt better still prefers the 8-step teacher on 11 of 45, against 6 for the 4-step adapter, mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
|
| 105 |
|
| 106 |
### Known issues
|
| 107 |
|
| 108 |
+
The usual costs of two steps, in this order of how often they show. **Small subjects are the weak spot, people and objects alike**: a portrait-sized face or an object seen up close holds up, while small or distant subjects β faces in a crowd, a figure in a wide scene, the machines at the back of a room β can come out ghosted, smeared or misshapen, since at that size a whole subject is only a few of the blocks the model works in. Fine structure can come out soft or a few pixels out of register β feathers, hair strands, signage, the surface of a distant object β most at 1280Γ1280 and above, and a faint doubled contour can show on limbs. On stylised prompts, **how the prompt says the image should look is followed less faithfully than what should be in it**: crisp anime linework, energetic brush strokes, the fingerprints in clay or a matte-painting finish come out closer to a generic rendering than the teacher's. On busy action or crowd scenes the composition can repeat itself β an extra hand or held object, a figure duplicated in a crowd β where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. Skin reads slightly smoother and less saturated than the teacher's, and colour overall runs a little under the teacher's at the largest sizes; freckles tend to gather into clusters rather than separate dots. A fine grain remains on the most textured subjects at the largest sizes, lighter than in the previous checkpoint. Every one of these is being worked on; none is hidden in the sweeps or the examples.
|
| 109 |
|
| 110 |
## How I got here
|
| 111 |
|
|
|
|
| 147 |
6. a spectral match of the student's image against the teacher's own, on the whole latent and on decoded pixel
|
| 148 |
windows β the term that finally reached the grain at large resolutions, after a critic, per-resolution weights and
|
| 149 |
a filtered push had each been tried against it and retired
|
| 150 |
+
7. the decoded-window spectral term given a much lighter hand β capped at a quarter of its earlier strength, which kept
|
| 151 |
+
the detail and took some grain out of flat areas
|
| 152 |
+
8. the fake-score adapter updated four times per student step instead of twice, so it keeps up with what the student
|
| 153 |
+
currently makes; faint straight-line artefacts that had begun to appear on flat illustrated areas went away with it
|
| 154 |
+
9. **an artefact critic** β a small head reading the frozen base model's own mid-network features, trained to tell the
|
| 155 |
+
teacher's finished images from the student's, with its push restricted to structure finer than 32 pixels and held
|
| 156 |
+
well below the distribution term β aimed at what the distribution term leaves behind: melted small faces and dense
|
| 157 |
+
detail smeared into blotches
|
| 158 |
+
10. four detail terms together: the anchor counting the fine-detail part of its error twice; a one-sided floor that
|
| 159 |
+
stops the finest detail dropping below the level real photographs carry; a second-call target made by the teacher
|
| 160 |
+
finishing the image from the student's own first-call output, so the target shares the student's layout; and a
|
| 161 |
+
smoothness limit on the fake-score adapter, so the distribution push keeps pointing at detail rather than away from it
|
| 162 |
+
11. **three critics taking turns** β the artefact critic joined by one whose real examples are half real photographs and one
|
| 163 |
+
weighted toward faces, one of them pushing on each step while the others keep training in between, because the three did
|
| 164 |
+
not fit in memory side by side
|
| 165 |
+
|
| 166 |
+
Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
|
| 167 |
+
a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
|
| 168 |
+
not place exactly where the teacher had it.
|
| 169 |
+
|
| 170 |
+
## chk00017464 vs chk00013663
|
| 171 |
+
|
| 172 |
+
`chk00013663` (published 12 Sep 2026) was the first public checkpoint: distribution matching with the spectral match, and nothing
|
| 173 |
+
yet aimed at the faults the distribution term leaves behind. `chk00017464` (14 Sep 2026) is 3,801 training samples later, and every
|
| 174 |
+
one of those samples went to those faults β the grain and grid pattern at large sizes, small faces, dense detail β through five recipe
|
| 175 |
+
changes, each kept only after its own look at the renders:
|
| 176 |
+
|
| 177 |
+
1. the decoded-window spectral term brought down to a quarter of its strength
|
| 178 |
+
2. the fake-score adapter updated four times per student step instead of twice
|
| 179 |
+
3. the artefact critic, reading the frozen base model's own features
|
| 180 |
+
4. four detail terms: the anchor counting fine-detail error twice, the one-sided photo floor, the teacher's finish of the student's
|
| 181 |
+
first call as the second call's target, and a smoothness limit on the fake adapter
|
| 182 |
+
5. three critics taking turns: the artefact critic, a photo critic and a face critic
|
| 183 |
+
|
| 184 |
+
Two ideas were tried and taken back out along the way: the distribution term confined to the second call's noise range, and a detail
|
| 185 |
+
pyramid compared pixel by pixel against the teacher.
|
| 186 |
+
|
| 187 |
+
**Detail at large sizes β the headline.** Every one of the 12 sweep resolutions Γ 15 prompts measured against the 8-step teacher, as in
|
| 188 |
+
[Measured against the teacher](#measured-against-the-teacher). Above 1 megapixel the excess fine energy two steps used to put into
|
| 189 |
+
images is roughly halved, and those sizes are now closer to the teacher than the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) on
|
| 190 |
+
fine texture and on both grid bands:
|
| 191 |
+
|
| 192 |
+
| 1.00 = the teacher | fine texture | 16-px band | 8-px band |
|
| 193 |
+
| --- | --- | --- | --- |
|
| 194 |
+
| 1280Γ1280 | 1.20 β **1.09** | 1.16 β **1.08** | 1.25 β **1.13** |
|
| 195 |
+
| 1440Γ1280 | 1.21 β **1.08** | 1.19 β **1.09** | 1.28 β **1.14** |
|
| 196 |
+
| 1440Γ1440 | 1.34 β **1.17** | 1.25 β **1.08** | 1.42 β **1.21** |
|
| 197 |
+
|
| 198 |
+
Fine-texture energy and both grid bands come closer to the teacher at 10 of the 12 resolutions, the blur-invariant ghosting index at
|
| 199 |
+
10 of 12, the micro-ghost index at all 12, and the grain in flat areas β skies, walls, out-of-focus backgrounds β at 11 of 12
|
| 200 |
+
(1440Γ1440: 1.84Γ the teacher's β 1.57Γ, median of the 15 prompts). It shows where it should: hair renders as more individual strands, and the token-grid
|
| 201 |
+
texture that read as grain on birds, food and foliage at 1440Γ1440 is lighter.
|
| 202 |
+
|
| 203 |
+
**Faces and skin.** Freckles on the test portrait gather into lighter, more dot-like clusters than before β better, not yet the
|
| 204 |
+
teacher's separate dots β and eyes look about the same: clean irises, lashes softer than the teacher's.
|
| 205 |
+
|
| 206 |
+
**What did not improve.** The judge that asks which of two renders follows the prompt better prefers the 8-step teacher on 11 of 45
|
| 207 |
+
against 6 before, mostly on how a stylised prompt says the picture should look: crisp linework, energetic brush strokes, the texture
|
| 208 |
+
of clay. The blind rubric that checks each render alone for the prompt's objects, counts, attributes and relations moves from 0 to 1
|
| 209 |
+
point missing out of 240, and on the same 15 fresh prompts the previous checkpoint was measured on, the two checkpoints come out level
|
| 210 |
+
(1 win, 13 ties, 1 loss). Colour runs slightly lower (saturation 0.93Γ the teacher's across the sweep, from 0.94Γ; 0.86Γ at
|
| 211 |
+
1440Γ1440), and at 512Γ512 the two checkpoints are level. Both are what the next recipe changes target.
|
| 212 |
+
|
| 213 |
+
| axis | `chk00013663` | `chk00017464` |
|
| 214 |
+
| --- | --- | --- |
|
| 215 |
+
| fine texture vs the teacher, 1280Β² / 1440Β² | 1.20 / 1.34 | **1.09 / 1.17** |
|
| 216 |
+
| 16-px grid band, 1280Β² / 1440Β² | 1.16 / 1.25 | **1.08 / 1.08** |
|
| 217 |
+
| grain in flat areas, sweep median | 1.36Γ | **1.26Γ** |
|
| 218 |
+
| distance to the teacher, sweep mean | 0.413 | **0.406** |
|
| 219 |
+
| judge prefers the teacher (of 45) | **6** | 11 |
|
| 220 |
+
| blind adherence rubric, points missing of 240 | **0** | 1 |
|
| 221 |
+
| saturation vs the teacher, sweep mean | **0.94Γ** | 0.93Γ |
|
| 222 |
+
| training samples in the 2-step stages | 13,663 | 17,464 |
|
| 223 |
|
| 224 |
## Measured against the teacher
|
| 225 |
|
|
|
|
| 234 |
|
| 235 |
| bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
|
| 236 |
| --- | --- | --- | --- | --- | --- |
|
| 237 |
+
| 512Γ512 | 1.02 | 1.02 | 1.05 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
|
| 238 |
+
| 768Γ1024 | 1.02 | 0.98 | 1.07 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
|
| 239 |
+
| 1024Γ1024 | 1.06 | 1.07 | 1.12 | 1.11 Β· 1.17 Β· 1.16 | 0.39 |
|
| 240 |
+
| 1280Γ1280 | 1.09 | 1.08 | 1.13 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
|
| 241 |
+
| 1440Γ1440 | 1.17 | 1.08 | 1.21 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
|
| 242 |
|
| 243 |
Two steps without the adapter carry **less than half** the teacher's fine detail at every size. With it, the detail
|
| 244 |
+
sits at or just above the teacher's everywhere β and from 1 megapixel up it is closer to the teacher than the 4-step
|
| 245 |
+
adapter, which carries more excess fine energy there.
|
| 246 |
|
| 247 |
**Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
|
| 248 |
in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
|
| 249 |
|
| 250 |
| bucket | wins | ties | losses |
|
| 251 |
| --- | --- | --- | --- |
|
| 252 |
+
| 512Γ512 | 1 | 10 | 4 |
|
| 253 |
+
| 1280Γ1280 | 1 | 10 | 4 |
|
| 254 |
+
| 1440Γ1440 | 0 | 12 | 3 |
|
| 255 |
+
|
| 256 |
+
Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
|
| 257 |
+
stylised prompts and on how a prompt says the picture should look β crisp linework, energetic brush strokes, the texture of
|
| 258 |
+
clay β more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
|
| 259 |
+
attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
|
| 260 |
+
training prompt bank and never rendered before, the judge returned 0 wins, 10 ties, 5 losses; on another 15 fresh prompts, 1 win,
|
| 261 |
+
13 ties, 1 loss.
|
| 262 |
+
|
| 263 |
+
**Checked for the damage this kind of training can do.** Saturation sits at 0.99Γ the teacher's at 768Γ1024 and 0.89Γ
|
| 264 |
+
at the larger sizes; edge detail 0.93β0.96Γ; skin texture inside detected faces 1.07Γ at 768Γ1024 and 0.85Γ at
|
| 265 |
+
1440Γ1440, with skin saturation 0.92Γ and 0.77β0.85Γ at the larger sizes. The honest reading of those last two: **skin
|
| 266 |
+
is the softest and least saturated part of this adapter's output at large sizes, and colour overall runs a little under
|
| 267 |
+
the teacher's there.** Fine detail in flat regions β skies, walls, out-of-focus backgrounds β runs 1.53β1.85Γ the
|
| 268 |
+
teacher's, which is where two steps put grain that eight steps do not.
|
| 269 |
+
|
| 270 |
+
**Distance to the teacher**, as a plain pixel measure, is 0.39β0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
|
| 271 |
+
0.30β0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
|
| 272 |
not the teacher's render of it β see [Known issues](#known-issues).
|
| 273 |
|
| 274 |
## Usage
|
|
|
|
| 281 |
| guidance / CFG | **0.0** (Turbo is CFG-free; do not enable it) |
|
| 282 |
| timestep shift | **mu = 1.15**, fixed (Turbo's deployment shift) |
|
| 283 |
|
| 284 |
+
The 2 sampling sigmas are Turbo's own deployment grid: `[1.0, 0.7595]` β the first and the middle
|
| 285 |
of the 4-step grid, so the model is evaluated at two points it already knows.
|
| 286 |
|
| 287 |
## Inference with diffusers
|
|
|
|
| 400 |
|
| 401 |
| configuration | denoising | per model call | GPU peak |
|
| 402 |
| --- | --- | --- | --- |
|
| 403 |
+
| Krea 2 Turbo β 8 steps (the reference) | **81.4 s** | 10.2 s | 25.2 GiB |
|
| 404 |
+
| Krea 2 Turbo β 2 steps, no LoRA | 20.4 s | 10.2 s | 25.2 GiB |
|
| 405 |
+
| **Krea 2 Turbo β 2 steps + this LoRA** | **19.5 s** | 9.8 s | 25.2 GiB |
|
| 406 |
|
| 407 |
+
**Denoising is 4.2Γ faster than the 8-step reference** β two model calls instead of eight. The adapter's own cost per call did
|
| 408 |
+
not show up in this measurement: the runs with it came in marginally faster than those without, which is measurement noise, not a
|
| 409 |
+
speed-up. A rank-64 low-rank product is small beside the transformer it is added to, and it adds no measurable memory.
|
|
|
|
| 410 |
|
| 411 |
Denoising is the part the step count changes. What a complete render costs on top of it β encoding the prompt, decoding
|
| 412 |
the latent, writing the file β is the same whether you run two steps or eight, and it depends on your pipeline, so the
|
| 413 |
+
end-to-end figure on your machine will sit below 4.2Γ and rise toward it as the render gets larger.
|
| 414 |
+
|
| 415 |
+
**By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
|
| 416 |
+
size; sweep renders run one at a time, not the controlled measurement above):
|
| 417 |
+
|
| 418 |
+
| resolution | denoising (2 calls) | resolution | denoising (2 calls) |
|
| 419 |
+
| --- | --- | --- | --- |
|
| 420 |
+
| 512Γ512 | 6.4 s | 1024Γ1024 | 20.4 s |
|
| 421 |
+
| 512Γ768 / 768Γ512 | 8.9 s / 9.0 s | 1280Γ960 / 960Γ1280 | 23.2 s / 24.0 s |
|
| 422 |
+
| 768Γ768 | 11.9 s | 1280Γ1280 | 32.2 s |
|
| 423 |
+
| 768Γ1024 / 1024Γ768 | 14.9 s / 15.1 s | 1440Γ1280 / 1440Γ1440 | 34.8 s / 39.8 s |
|
| 424 |
|
| 425 |
## LoRA strength
|
| 426 |
|
|
|
|
| 498 |
**Distribution matching with a trajectory anchor**, Krea 2 Turbo as its own teacher, on the recorded 8-step
|
| 499 |
trajectories.
|
| 500 |
|
| 501 |
+
The student makes two calls, at Ο = 1.0 and Ο = 0.7595 β the first and fifth points of the teacher's 8-step grid at
|
| 502 |
mu = 1.15 β and stock Euler carries it between them. That grid is what makes the objective a drop-in: Euler's first step
|
| 503 |
+
from pure noise lands _exactly_ on the flow-matching interpolant at Ο = 0.7595 with the same noise and the student's
|
| 504 |
own clean-image prediction as the data point. So the student's first-call output is a legitimate image prediction that
|
| 505 |
can be judged as an image, and the second call is fed from it during training the way it will be at inference.
|
| 506 |
|
|
|
|
| 509 |
is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
|
| 510 |
the image more like the teacher's work and less like the student's habits, and the student is pushed that way
|
| 511 |
(the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
|
| 512 |
+
rank 32, starts as an exact copy of the teacher, updates four times per student step β often enough to keep up with a
|
| 513 |
+
student that is still changing β and is discarded at the end.
|
| 514 |
|
| 515 |
**The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
|
| 516 |
student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
|
|
|
|
| 521 |
and under-fits fine structure there, and a per-pixel normaliser lands harder as pixel counts grow. A full 12-bucket
|
| 522 |
sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
|
| 523 |
1.4β1.8Γ the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
|
| 524 |
+
not hold, and the spectral match replaced it. The same sweep of the published checkpoint, its running-average weights, fixed
|
| 525 |
+
seed, 15 prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
|
| 526 |
first distribution-matching checkpoint on the same prompts):
|
| 527 |
|
| 528 |
| bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
|
| 529 |
| --------- | --------------------------- | --------------- | -------------- | ----------------------- |
|
| 530 |
+
| 512x512 | 1.02 (1.16) | 1.02 (1.23) | 1.05 (1.23) | 0.41 (0.44) |
|
| 531 |
+
| 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.06 (1.33) | 0.39 (0.41) |
|
| 532 |
+
| 768x512 | 1.02 (1.22) | 1.08 (1.37) | 1.04 (1.26) | 0.44 (0.47) |
|
| 533 |
+
| 768x768 | 1.09 (1.43) | 1.05 (1.46) | 1.14 (1.50) | 0.40 (0.43) |
|
| 534 |
+
| 768x1024 | 1.02 (1.36) | 0.98 (1.38) | 1.07 (1.42) | 0.39 (0.43) |
|
| 535 |
+
| 1024x768 | 1.01 (1.35) | 0.98 (1.40) | 1.04 (1.38) | 0.43 (0.46) |
|
| 536 |
+
| 1024x1024 | 1.06 (1.47) | 1.07 (1.55) | 1.12 (1.56) | 0.41 (0.44) |
|
| 537 |
+
| 1280x960 | 1.03 (1.50) | 1.04 (1.59) | 1.07 (1.56) | 0.40 (0.42) |
|
| 538 |
+
| 960x1280 | 1.10 (1.60) | 1.05 (1.63) | 1.19 (1.70) | 0.40 (0.43) |
|
| 539 |
+
| 1280x1280 | 1.09 (1.65) | 1.08 (1.68) | 1.13 (1.69) | 0.41 (0.44) |
|
| 540 |
+
| 1440x1280 | 1.08 (1.59) | 1.08 (1.66) | 1.14 (1.66) | 0.40 (0.42) |
|
| 541 |
+
| 1440x1440 | 1.17 (1.81) | 1.08 (1.79) | 1.21 (1.89) | 0.39 (0.41) |
|
| 542 |
|
| 543 |
### The spectral match
|
| 544 |
|
|
|
|
| 556 |
judged too. The comparison is two-sided, so too much fine energy and too little are both penalised; a blurred image
|
| 557 |
does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
|
| 558 |
refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
|
| 559 |
+
without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too. The whole-latent comparison trains at that
|
| 560 |
+
strength; the decoded window was later brought down to a quarter of it, which kept the detail and removed some grain.
|
| 561 |
+
|
| 562 |
+
### The artefact critic
|
| 563 |
+
|
| 564 |
+
Distribution matching improves what the fake adapter can see, and the fake adapter learns from the student's own
|
| 565 |
+
images β so where the student smears something, the fake learns the smear and the push stops correcting it. The two
|
| 566 |
+
places that shows most are small faces, which come out melted, and dense content such as the goods on a market stall or
|
| 567 |
+
the shelves of a shop seen through its window, which comes out as coloured blotches.
|
| 568 |
+
|
| 569 |
+
A critic breaks that loop by looking at finished images instead. It is a small head on the frozen base model's
|
| 570 |
+
mid-network features β every adapter switched off, the forward pass stopped halfway, so no adapter can learn to fool
|
| 571 |
+
it and nothing it learns leaks into the fake adapter. Its real examples are the teacher's own finished images for
|
| 572 |
+
other prompts at the same resolution; its fake examples are the student's final images. Both are lightly re-noised
|
| 573 |
+
first, in the low-noise range where fine structure lives, and the critic reads them with an empty prompt, so it judges
|
| 574 |
+
only whether the structure looks like the teacher's. Its push on the student is filtered to periods finer than 32
|
| 575 |
+
pixels β below that it can rebuild a face or a shelf, above it it could move layout, colour or pose, which it must not
|
| 576 |
+
β and capped at a quarter of the distribution term's strength per sample. Lazy gradient regularisation and spectral
|
| 577 |
+
normalisation keep the head from overshooting. On the largest steps, where memory is tightest, the critic and the
|
| 578 |
+
teacher's finishing pass below take turns instead of sharing a step.
|
| 579 |
+
|
| 580 |
+
### Critics in turn
|
| 581 |
+
|
| 582 |
+
One critic holds one idea of what is wrong. The artefact critic was joined by two more on the same frozen mid-network
|
| 583 |
+
features and under the same rules β lightly re-noised inputs, an empty prompt, a push that is filtered and capped β each
|
| 584 |
+
aimed at a different fault:
|
| 585 |
+
|
| 586 |
+
- **A photo critic.** Half of its real examples are real photographs and half the teacher's finished images, so it learns
|
| 587 |
+
what fine texture looks like in a photograph as well as in the teacher's rendering of one. Its push is filtered to
|
| 588 |
+
periods finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher
|
| 589 |
+
does not.
|
| 590 |
+
- **A face critic.** Its real examples are the teacher's finished images of prompts with faces, the face regions counted at
|
| 591 |
+
full weight and the rest at half, so its push concentrates on what small and mid-sized faces lose first.
|
| 592 |
+
|
| 593 |
+
All three on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
|
| 594 |
+
others train so none goes stale before its turn comes back. The two new heads started from the artefact critic's weights and
|
| 595 |
+
trained on their own before they were allowed to push.
|
| 596 |
+
|
| 597 |
+
### Detail terms
|
| 598 |
+
|
| 599 |
+
Four smaller terms sit on top, each capped relative to the distribution term so none of them can take over:
|
| 600 |
+
|
| 601 |
+
- **A detail-weighted anchor.** The trajectory regression counts the fine-detail part of its error β everything finer
|
| 602 |
+
than 32 pixels β twice, so the anchor stops tolerating softness it used to average away.
|
| 603 |
+
- **A one-sided photo floor.** On the decoded window, the student's energy at periods of 3β10 pixels may not fall below
|
| 604 |
+
the teacher's plus the margin real photographs carry over it at those scales. That margin is measured once from a
|
| 605 |
+
pool of real photographs and clamped, and the term only ever pushes upward to that floor, never past it β so it
|
| 606 |
+
lifts detail that is missing without adding grain that is not.
|
| 607 |
+
- **The teacher's finish as a target.** Every second step, the teacher itself runs its remaining steps starting from
|
| 608 |
+
the student's own first-call output. The result is a finished image that shares the student's layout, and the second
|
| 609 |
+
call is pulled gently toward it β a target that lines up with what the student actually drew, where the recorded
|
| 610 |
+
trajectory might have drawn something else.
|
| 611 |
+
- **A smoothness limit on the fake adapter.** The fake adapter's fine-detail energy is kept below the teacher's at the
|
| 612 |
+
point where the distribution term is measured, so the difference between the two keeps pointing toward detail.
|
| 613 |
|
| 614 |
## What the LoRA touches
|
| 615 |
|
|
|
|
| 622 |
The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β Krea 2 Turbo's own 8-step run at mu = 1.15 and
|
| 623 |
guidance 0.0, every latent and velocity stored β serve unchanged: a 2-step chord is two of the 4-step chords end to
|
| 624 |
end. 203 held-out prompts measure the studentβteacher gap on unseen prompts and never receive a gradient. The spectral
|
| 625 |
+
match and the artefact critic read the teacher's finals for the training prompts. The 43,044 real-photo crops of the
|
| 626 |
+
[4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) enter in two places: as one precomputed statistic β how much fine-detail energy they carry at 3β10
|
| 627 |
+
pixels relative to the teacher, clamped β which sets the photo floor, and as half of the photo critic's real examples. Only that
|
| 628 |
+
critic's head sees them; the student and the fake adapter never do, and receive only its filtered, capped push.
|
| 629 |
|
| 630 |
## Resolutions
|
| 631 |
|
|
|
|
| 696 |
## Hardware
|
| 697 |
|
| 698 |
One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned
|
| 699 |
+
host memory above 0.3 megapixels; the student, the fake adapter, the spectral and detail terms, the critic and the teacher's finishing pass each
|
| 700 |
+
build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
|
| 701 |
+
does not fit fails loudly. A full step with every term live and every critic pushing reserves about 22.4 GB at 1440Γ1440,
|
| 702 |
+
of 24. The price of the objective is throughput: **about 107 training samples an hour** measured over the recipe's complete
|
| 703 |
+
run, against the [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β more than four times the cost per sample, and so far a small
|
| 704 |
+
fraction of the samples.
|
| 705 |
|
| 706 |
**Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
|
| 707 |
The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
|
| 708 |
disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
|
| 709 |
of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
|
| 710 |
+
alongside to imitate the student β and that second adapter takes four optimiser steps of its own per student step.
|
| 711 |
+
The spectral and detail terms decode part of the image out of the latent to compare its texture with the teacher's,
|
| 712 |
+
each critic reads half the network twice more, and every second step the teacher finishes the image from the student's
|
| 713 |
+
first call.
|
| 714 |
|
| 715 |
+
Counted in whole model runs per training sample, the difference is roughly **two there against about a dozen here**. None of
|
| 716 |
that difference is the teacher generating anything: its renders were recorded once for the 4-step project and are
|
| 717 |
read from disk by both. The extra work is the objective itself, and it bought the only thing that mattered. Run at
|
| 718 |
two steps, the 4-step project's recipe reached a point where more training changed nothing: the measurements sat
|
|
|
|
| 726 |
15 fixed prompts across four resolutions (512Γ512, 768Γ1024, 1280Γ1280, 1440Γ1440); milestone checkpoints get the same
|
| 727 |
render at all 12 buckets, which is where the per-resolution table above comes from. Every image is measured against the
|
| 728 |
teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
|
| 729 |
+
and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
|
| 730 |
+
seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
|
| 731 |
scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
|
| 732 |
identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
|
| 733 |
checkpoint. Latent distances β the held-out chord gap and the two-step rollout
|
|
|
|
| 1085 |
|
| 1086 |
## What's next
|
| 1087 |
|
| 1088 |
+
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any
|
| 1089 |
+
resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). The next step
|
| 1090 |
+
is already running, and it goes after what this checkpoint still gets wrong: a critic that judges the first call against the
|
| 1091 |
+
teacher's own intermediate state from the same noise, so a first call that blends two layouts is caught where the blend
|
| 1092 |
+
happens; a critic weighted toward wherever the teacher put fine detail; the photo critic and the photo floor limited to
|
| 1093 |
+
photographic prompts, so illustration, anime and 3D renders are no longer pulled toward photographic grain; detail held to the
|
| 1094 |
+
teacher region by region, with a ceiling as well as a floor; a focus on the eyes, nose and lips of faces so they sharpen while
|
| 1095 |
+
skin stays the teacher's; a colour floor, so colour at large sizes stops falling below the teacher's; and a more even mix of
|
| 1096 |
+
resolutions. After it comes prompt adherence β a critic that learns whether an image belongs to its own prompt, and stylised
|
| 1097 |
+
prompts drawn more often β held to the rule that none of it may cost the sharpness this checkpoint gained. A better checkpoint
|
| 1098 |
+
replaces this one when the sweeps and I visually agree, the same discipline as the
|
| 1099 |
+
[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
|
| 1100 |
+
one is the fast preview.
|
| 1101 |
|
| 1102 |
## License
|
| 1103 |
|
krea2_turbo_2step_rank_64_lora.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:cc7a7d65ca04070b2aa1dbf60b3cd0cf210fde2a93d4d7d4fc18bdadf23060d3
|
| 3 |
+
size 438161144
|
krea2_turbo_2step_rank_64_lora_checkpoint_info.md
CHANGED
|
@@ -1,13 +1,13 @@
|
|
| 1 |
# Which checkpoint is this?
|
| 2 |
|
| 3 |
`krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
|
| 4 |
-
**
|
| 5 |
`_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
|
| 6 |
keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
|
| 7 |
|
| 8 |
| file | SHA-256 | size |
|
| 9 |
| --- | --- | --- |
|
| 10 |
-
| `krea2_turbo_2step_rank_64_lora.safetensors` | `
|
| 11 |
-
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `
|
| 12 |
|
| 13 |
-
Updated:
|
|
|
|
| 1 |
# Which checkpoint is this?
|
| 2 |
|
| 3 |
`krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
|
| 4 |
+
**chk00017464** β the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00017464.safetensors` (and its
|
| 5 |
`_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
|
| 6 |
keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
|
| 7 |
|
| 8 |
| file | SHA-256 | size |
|
| 9 |
| --- | --- | --- |
|
| 10 |
+
| `krea2_turbo_2step_rank_64_lora.safetensors` | `cc7a7d65ca04070b2aa1dbf60b3cd0cf210fde2a93d4d7d4fc18bdadf23060d3` | 418M |
|
| 11 |
+
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `5c02cac5de27dcb40ea98caee5fd5668ac1a928a105a1fb29c02eb9b0b1e5862` | 418M |
|
| 12 |
|
| 13 |
+
Updated: 14 Sep 2026
|
krea2_turbo_2step_rank_64_lora_comfyui.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:5c02cac5de27dcb40ea98caee5fd5668ac1a928a105a1fb29c02eb9b0b1e5862
|
| 3 |
+
size 438141680
|