lvladikov commited on
Commit
5c3d67c
Β·
verified Β·
1 Parent(s): a13553d

chk17464 release

Browse files
README.md CHANGED
@@ -17,11 +17,22 @@ pipeline_tag: text-to-image
17
 
18
  # Krea 2 Turbo β€” 2-Step Distillation LoRA
19
 
20
- > πŸ§ͺ **A fast-preview adapter, from a project still in training.** The published checkpoint files give usable two-step
21
- > renders and are measured honestly below 4-step or 8-step renders; they are not the
22
- > [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s quality, and that adapter remains the recommendation for quality renders. Training
23
- > continues one recipe change at a time, and a later checkpoint replaces this one only when the sweeps and I visually
24
- > agree it is better. Known issues β€” see [Known issues](#known-issues).
 
 
 
 
 
 
 
 
 
 
 
25
 
26
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
27
  down to 2** β€” Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at
@@ -32,9 +43,9 @@ recommendation for quality renders.
32
  - 🎯 **The aim** β€” the best two-step quality this base can give, at every one of the same 12 resolutions, measured
33
  against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
34
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas.
35
- - ⏱️ **3.8Γ— faster denoising** β€” the model runs twice instead of eight times, and denoising is the part this adapter
36
- changes: **79.9 s β†’ 20.9 s** measured at 1024Γ—1024 on the same prompts, the adapter itself costing about 3.5% per
37
- call. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
38
  [Performance](#performance).
39
  - πŸ“Š **Distribution matching, not imitation** β€” the training objective that got the renders improving again after the
40
  4-step project's recipe had stopped helping at two steps (see [Method](#method)).
@@ -42,21 +53,23 @@ recommendation for quality renders.
42
  evaluated on each prompt's own conditioning, so the student is matched to what the teacher makes _for that prompt_,
43
  not to a prompt-free look. There is no separate adherence term: instead a vision-language judge checks every
44
  checkpoint β€” each render scored alone against the prompt's objects, counts, attributes and relations, with the
45
- teacher scored the same way β€” and a term would only be added if that meter showed adherence slipping.
 
 
46
  - πŸ“ **12 trained resolutions** β€” multi-aspect from 512Γ—512 up to 1440Γ—1440, each with its sweep.
47
- - πŸ”Œ **Drop-in, no exceptions** β€” a plain LoRA sampled by stock Euler at sigmas `[1.0, 0.5128]` in diffusers, ComfyUI
48
  or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this
49
  project.
50
  - 🧬 **Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β€” rank 64 on the same 228 modules; a second adapter exists during training
51
  only and never ships.
52
  - 🎲 **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
53
  re-run.
54
- - πŸ”’ **13,663 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β€” all of them
55
  drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
56
  training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
57
  4-step project, read again at the two sigmas this schedule uses.
58
- - πŸ“… **5 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β€” and the project continues.
59
- - πŸ” **15 recipe adjustments** across two methods so far β€” seven of trajectory distillation before the switch, eight of distribution matching since.
60
  - πŸ–₯️ **One RTX 3090**, and a recipe shaped by its 24 GB.
61
 
62
  [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
@@ -86,13 +99,13 @@ place to check which checkpoint the current files are based on.
86
 
87
  | | |
88
  | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
89
- | lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β†’ 2-step trajectory distillation β†’ distribution matching β†’ a spectral match against the teacher's own images on top |
90
  | this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
91
- | what it gives | usable two-step renders at every trained resolution: fine detail within a few percent of what the 4-step adapter carries, and a prompt-adherence judge that calls it a loss against the 8-step teacher on 6 of 45 renders β€” the same count the 4-step adapter scores. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
92
 
93
  ### Known issues
94
 
95
- The usual costs of two steps, in this order of how often they show: fine structure comes out soft or a few pixels out of register β€” feathers, skin texture, hair strands, signage, the surface of a distant object β€” most at 1280Γ—1280 and above; a faint doubled contour on faces and limbs. **Faces depend on how much of the frame they occupy**: a portrait-sized face holds up, while small or distant faces β€” a crowd, a figure in a wide scene β€” lose their features first and can come out misshapen, since at that size a whole face is only a few of the blocks the model works in. On busy action or crowd scenes the composition can also repeat itself β€” an extra hand or held object, a figure duplicated in a crowd β€” where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. At the largest sizes a fine grain remains on the most textured subjects and skin reads slightly smoother and less saturated than the teacher's. Every one of these is being worked on; none is hidden in the sweeps or the examples.
96
 
97
  ## How I got here
98
 
@@ -134,6 +147,79 @@ The recipe adjustments so far, each made on the measurement of the one before:
134
  6. a spectral match of the student's image against the teacher's own, on the whole latent and on decoded pixel
135
  windows β€” the term that finally reached the grain at large resolutions, after a critic, per-resolution weights and
136
  a filtered push had each been tried against it and retired
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
137
 
138
  ## Measured against the teacher
139
 
@@ -148,36 +234,41 @@ teacher):
148
 
149
  | bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
150
  | --- | --- | --- | --- | --- | --- |
151
- | 512Γ—512 | 1.00 | 1.01 | 1.05 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
152
- | 768Γ—1024 | 1.04 | 1.01 | 1.10 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
153
- | 1024Γ—1024 | 1.11 | 1.12 | 1.15 | 1.11 Β· 1.17 Β· 1.19 | 0.39 |
154
- | 1280Γ—1280 | 1.20 | 1.16 | 1.25 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
155
- | 1440Γ—1440 | 1.34 | 1.25 | 1.42 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
156
 
157
  Two steps without the adapter carry **less than half** the teacher's fine detail at every size. With it, the detail
158
- lands within a few percent of what the 4-step adapter carries β€” at 1440Γ—1440 slightly above it.
 
159
 
160
  **Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
161
  in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
162
 
163
  | bucket | wins | ties | losses |
164
  | --- | --- | --- | --- |
165
- | 512Γ—512 | 3 | 10 | 2 |
166
- | 1280Γ—1280 | 1 | 12 | 2 |
167
- | 1440Γ—1440 | 2 | 11 | 2 |
168
-
169
- Six losses out of 45 β€” the same count the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores against the same teacher. On 15 prompts drawn
170
- fresh from the training prompt bank and never rendered before, the same judge returned 1 win, 13 ties, 1 loss.
171
-
172
- **Checked for the damage this kind of training can do.** Saturation sits at 1.01Γ— the teacher's at 768Γ—1024 and 0.91Γ—
173
- at the larger sizes; edge detail 0.92–0.96Γ—; skin texture inside detected faces 0.96Γ— at 768Γ—1024 and 0.85Γ— at
174
- 1440Γ—1440, with skin saturation 0.93Γ— and 0.81Γ—. The honest reading of those last two: **skin is the softest and least
175
- saturated part of this adapter's output, and it gets softer as the render gets larger.** Fine detail in flat regions β€”
176
- skies, walls, out-of-focus backgrounds β€” runs 1.65–1.88Γ— the teacher's, which is where two steps put grain that eight
177
- steps do not.
178
-
179
- **Distance to the teacher**, as a plain pixel measure, is 0.39–0.42 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
180
- 0.32–0.34. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
 
 
 
 
181
  not the teacher's render of it β€” see [Known issues](#known-issues).
182
 
183
  ## Usage
@@ -190,7 +281,7 @@ not the teacher's render of it β€” see [Known issues](#known-issues).
190
  | guidance / CFG | **0.0** (Turbo is CFG-free; do not enable it) |
191
  | timestep shift | **mu = 1.15**, fixed (Turbo's deployment shift) |
192
 
193
- The 2 sampling sigmas are Turbo's own deployment grid: `[1.0, 0.51284]` β€” the first and the middle
194
  of the 4-step grid, so the model is evaluated at two points it already knows.
195
 
196
  ## Inference with diffusers
@@ -309,18 +400,27 @@ the model already loaded, so the numbers are the render itself and not a model l
309
 
310
  | configuration | denoising | per model call | GPU peak |
311
  | --- | --- | --- | --- |
312
- | Krea 2 Turbo β€” 8 steps (the reference) | **79.9 s** | 9.99 s | 25.2 GiB |
313
- | Krea 2 Turbo β€” 2 steps, no LoRA | 20.2 s | 10.12 s | 25.2 GiB |
314
- | **Krea 2 Turbo β€” 2 steps + this LoRA** | **20.9 s** | 10.47 s | 25.2 GiB |
315
 
316
- **Denoising is 3.8Γ— faster than the 8-step reference** β€” two model calls instead of eight. The adapter itself costs about
317
- **3.5% more per call** (10.47 s against 10.12 s) and no measurable extra memory, because a rank-64 low-rank product is
318
- small beside the transformer it is added to; that overhead is paid twice here instead of the four or eight times a longer
319
- schedule would pay it.
320
 
321
  Denoising is the part the step count changes. What a complete render costs on top of it β€” encoding the prompt, decoding
322
  the latent, writing the file β€” is the same whether you run two steps or eight, and it depends on your pipeline, so the
323
- end-to-end figure on your machine will sit below 3.8Γ— and rise toward it as the render gets larger.
 
 
 
 
 
 
 
 
 
 
324
 
325
  ## LoRA strength
326
 
@@ -398,9 +498,9 @@ In every case the tensors are the same 456 bf16 matrices; only the names around
398
  **Distribution matching with a trajectory anchor**, Krea 2 Turbo as its own teacher, on the recorded 8-step
399
  trajectories.
400
 
401
- The student makes two calls, at Οƒ = 1.0 and Οƒ = 0.5128 β€” the first and fifth points of the teacher's 8-step grid at
402
  mu = 1.15 β€” and stock Euler carries it between them. That grid is what makes the objective a drop-in: Euler's first step
403
- from pure noise lands _exactly_ on the flow-matching interpolant at Οƒ = 0.5128 with the same noise and the student's
404
  own clean-image prediction as the data point. So the student's first-call output is a legitimate image prediction that
405
  can be judged as an image, and the second call is fed from it during training the way it will be at inference.
406
 
@@ -409,7 +509,8 @@ freshly noised copy: the frozen teacher, and a second small adapter on the same
409
  is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
410
  the image more like the teacher's work and less like the student's habits, and the student is pushed that way
411
  (the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
412
- rank 32, starts as an exact copy of the teacher, updates twice per student step, and is discarded at the end.
 
413
 
414
  **The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
415
  student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
@@ -420,24 +521,24 @@ moved so far so fast.
420
  and under-fits fine structure there, and a per-pixel normaliser lands harder as pixel counts grow. A full 12-bucket
421
  sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
422
  1.4–1.8Γ— the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
423
- not hold, and the spectral match replaced it. The same sweep of the current checkpoint, live weights, fixed seed, 15
424
- prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
425
  first distribution-matching checkpoint on the same prompts):
426
 
427
  | bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
428
  | --------- | --------------------------- | --------------- | -------------- | ----------------------- |
429
- | 512x512 | 1.00 (1.16) | 1.01 (1.23) | 1.05 (1.23) | 0.42 (0.44) |
430
- | 512x768 | 1.01 (1.23) | 1.08 (1.37) | 1.06 (1.33) | 0.39 (0.41) |
431
- | 768x512 | 1.04 (1.22) | 1.08 (1.37) | 1.09 (1.26) | 0.45 (0.47) |
432
- | 768x768 | 1.09 (1.43) | 1.06 (1.46) | 1.15 (1.50) | 0.40 (0.43) |
433
- | 768x1024 | 1.04 (1.36) | 1.01 (1.38) | 1.10 (1.42) | 0.39 (0.43) |
434
- | 1024x768 | 1.04 (1.35) | 1.03 (1.40) | 1.07 (1.38) | 0.43 (0.46) |
435
- | 1024x1024 | 1.11 (1.47) | 1.12 (1.55) | 1.18 (1.56) | 0.42 (0.44) |
436
- | 1280x960 | 1.09 (1.50) | 1.07 (1.59) | 1.16 (1.56) | 0.42 (0.42) |
437
- | 960x1280 | 1.20 (1.60) | 1.16 (1.63) | 1.30 (1.70) | 0.41 (0.43) |
438
- | 1280x1280 | 1.20 (1.65) | 1.16 (1.68) | 1.25 (1.69) | 0.42 (0.44) |
439
- | 1440x1280 | 1.20 (1.59) | 1.19 (1.66) | 1.28 (1.66) | 0.40 (0.42) |
440
- | 1440x1440 | 1.34 (1.81) | 1.25 (1.79) | 1.42 (1.89) | 0.40 (0.41) |
441
 
442
  ### The spectral match
443
 
@@ -455,8 +556,60 @@ latent every sample, and on one random 256-pixel window decoded through the VAE
455
  judged too. The comparison is two-sided, so too much fine energy and too little are both penalised; a blurred image
456
  does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
457
  refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
458
- without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too, and that is
459
- the setting the current run trains with.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
460
 
461
  ## What the LoRA touches
462
 
@@ -469,8 +622,10 @@ of all 28 transformer blocks, plus the four global linears β€” `time_embed.linea
469
  The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β€” Krea 2 Turbo's own 8-step run at mu = 1.15 and
470
  guidance 0.0, every latent and velocity stored β€” serve unchanged: a 2-step chord is two of the 4-step chords end to
471
  end. 203 held-out prompts measure the student–teacher gap on unseen prompts and never receive a gradient. The spectral
472
- match reads the teacher's finals for the training prompts; the 43,044 real-photo crops of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) are
473
- not used in the current run.
 
 
474
 
475
  ## Resolutions
476
 
@@ -541,22 +696,23 @@ The resolutions are exactly the training buckets listed under [Resolutions](#res
541
  ## Hardware
542
 
543
  One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned
544
- host memory above 0.3 megapixels; the student, the fake adapter and the spectral term each build and free their own graph
545
- in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
546
- does not fit fails loudly. A full step with every term live reserves about 21.4 GB at 1440Γ—1440, of 24. The price of
547
- the objective is throughput: **188 training samples an hour** measured over a complete 10-hour run, against the
548
- [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β€” two and a half times the
549
- cost per sample, and so far a small fraction of the samples.
550
 
551
  **Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
552
  The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
553
  disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
554
  of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
555
- alongside to imitate the student β€” and that second adapter takes two optimiser steps of its own per student step.
556
- A third term then decodes part of the image out of the latent to compare its texture with the teacher's, which costs
557
- another pass through the decoder.
 
558
 
559
- Counted in whole model runs per training sample, the difference is roughly **two there against seven here**. None of
560
  that difference is the teacher generating anything: its renders were recorded once for the 4-step project and are
561
  read from disk by both. The extra work is the objective itself, and it bought the only thing that mattered. Run at
562
  two steps, the 4-step project's recipe reached a point where more training changed nothing: the measurements sat
@@ -570,7 +726,8 @@ At regular intervals, both the live weights and their running average are pulled
570
  15 fixed prompts across four resolutions (512Γ—512, 768Γ—1024, 1280Γ—1280, 1440Γ—1440); milestone checkpoints get the same
571
  render at all 12 buckets, which is where the per-resolution table above comes from. Every image is measured against the
572
  teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
573
- and flat-region grain, saturation, faces cut out at 1:1, a graded judge, a pairwise preference against the teacher, and a blind rubric that
 
574
  scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
575
  identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
576
  checkpoint. Latent distances β€” the held-out chord gap and the two-step rollout
@@ -928,12 +1085,19 @@ _Still out-of-spec and still edge case preview-only β€” this is one step, not th
928
 
929
  ## What's next
930
 
931
- Training continues from the released checkpoint, one recipe change at a time, each kept only if the pictures do not
932
- degrade at any resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). The
933
- levers on the table are the ones the remaining defects point at: fine structure that comes out soft or a few pixels
934
- out of register at the largest sizes, small faces in crowds, signage. A better checkpoint replaces this one when the
935
- sweeps and I visually agree, the same discipline as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the
936
- [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this one is the fast preview.
 
 
 
 
 
 
 
937
 
938
  ## License
939
 
 
17
 
18
  # Krea 2 Turbo β€” 2-Step Distillation LoRA
19
 
20
+ > πŸ§ͺ **A fast-preview adapter, from a project still in training.** When the subject is close and fills a good part of the
21
+ > frame β€” a portrait, a single figure, an object seen up close β€” two steps already hold up well, and you can rely on
22
+ > this adapter for those images. Small subjects are where it still falls short, people and objects alike: faces in a
23
+ > crowd, figures in a wide scene, the machines at the back of a gym β€” anything that takes up little of the frame can
24
+ > come out ghosted or smeared. For those, and whenever quality matters more than speed, use the
25
+ > [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Every figure on this page measures the
26
+ > adapter honestly against 4-step and 8-step renders. Training continues one recipe change at a time, and a later
27
+ > checkpoint replaces this one only when the sweeps and I visually agree it is better. Known issues β€” see
28
+ > [Known issues](#known-issues).
29
+ >
30
+ > πŸ“ **The saved steps can also go into resolution.** A small subject is simply one that covers few pixels, so a larger
31
+ > render makes the same subject bigger β€” and at a quarter of the teacher's steps, renders up to 2048Γ—2048, Krea's
32
+ > published maximum recommended resolution and beyond the largest size this adapter was trained at (1440Γ—1440), come
33
+ > within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past
34
+ > 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects β€” a property of the base model, with or without this
35
+ > adapter.
36
 
37
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
38
  down to 2** β€” Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at
 
43
  - 🎯 **The aim** β€” the best two-step quality this base can give, at every one of the same 12 resolutions, measured
44
  against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
45
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas.
46
+ - ⏱️ **4.2Γ— faster denoising** β€” the model runs twice instead of eight times, and denoising is the part this adapter
47
+ changes: **81.4 s β†’ 19.5 s** measured at 1024Γ—1024 on the same prompts, the adapter's own cost per call within
48
+ measurement noise. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
49
  [Performance](#performance).
50
  - πŸ“Š **Distribution matching, not imitation** β€” the training objective that got the renders improving again after the
51
  4-step project's recipe had stopped helping at two steps (see [Method](#method)).
 
53
  evaluated on each prompt's own conditioning, so the student is matched to what the teacher makes _for that prompt_,
54
  not to a prompt-free look. There is no separate adherence term: instead a vision-language judge checks every
55
  checkpoint β€” each render scored alone against the prompt's objects, counts, attributes and relations, with the
56
+ teacher scored the same way β€” and a term would only be added if that meter showed adherence slipping. The one part of
57
+ training that looks at images without their prompt is the artefact critic (see [Method](#method)), and it only judges
58
+ whether fine structure looks like the teacher's.
59
  - πŸ“ **12 trained resolutions** β€” multi-aspect from 512Γ—512 up to 1440Γ—1440, each with its sweep.
60
+ - πŸ”Œ **Drop-in, no exceptions** β€” a plain LoRA sampled by stock Euler at sigmas `[1.0, 0.7595]` in diffusers, ComfyUI
61
  or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this
62
  project.
63
  - 🧬 **Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β€” rank 64 on the same 228 modules; a second adapter exists during training
64
  only and never ships.
65
  - 🎲 **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
66
  re-run.
67
+ - πŸ”’ **17,464 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β€” all of them
68
  drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
69
  training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
70
  4-step project, read again at the two sigmas this schedule uses.
71
+ - πŸ“… **7 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β€” and the project continues.
72
+ - πŸ” **26 recipe adjustments** across two methods so far β€” seven of trajectory distillation before the switch, nineteen of distribution matching since β€” each kept only when the renders did not get worse.
73
  - πŸ–₯️ **One RTX 3090**, and a recipe shaped by its 24 GB.
74
 
75
  [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
 
99
 
100
  | | |
101
  | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
102
+ | lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β†’ 2-step trajectory distillation β†’ distribution matching β†’ a spectral match against the teacher's own images β†’ an artefact critic and detail terms β†’ three critics taking turns |
103
  | this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
104
+ | what it gives | usable two-step renders at every trained resolution: fine detail at or just above the teacher's β€” from 1 megapixel up, closer to the teacher than the 4-step adapter β€” with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 1 point missing out of 240). A judge asked which render follows the prompt better still prefers the 8-step teacher on 11 of 45, against 6 for the 4-step adapter, mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
105
 
106
  ### Known issues
107
 
108
+ The usual costs of two steps, in this order of how often they show. **Small subjects are the weak spot, people and objects alike**: a portrait-sized face or an object seen up close holds up, while small or distant subjects β€” faces in a crowd, a figure in a wide scene, the machines at the back of a room β€” can come out ghosted, smeared or misshapen, since at that size a whole subject is only a few of the blocks the model works in. Fine structure can come out soft or a few pixels out of register β€” feathers, hair strands, signage, the surface of a distant object β€” most at 1280Γ—1280 and above, and a faint doubled contour can show on limbs. On stylised prompts, **how the prompt says the image should look is followed less faithfully than what should be in it**: crisp anime linework, energetic brush strokes, the fingerprints in clay or a matte-painting finish come out closer to a generic rendering than the teacher's. On busy action or crowd scenes the composition can repeat itself β€” an extra hand or held object, a figure duplicated in a crowd β€” where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. Skin reads slightly smoother and less saturated than the teacher's, and colour overall runs a little under the teacher's at the largest sizes; freckles tend to gather into clusters rather than separate dots. A fine grain remains on the most textured subjects at the largest sizes, lighter than in the previous checkpoint. Every one of these is being worked on; none is hidden in the sweeps or the examples.
109
 
110
  ## How I got here
111
 
 
147
  6. a spectral match of the student's image against the teacher's own, on the whole latent and on decoded pixel
148
  windows β€” the term that finally reached the grain at large resolutions, after a critic, per-resolution weights and
149
  a filtered push had each been tried against it and retired
150
+ 7. the decoded-window spectral term given a much lighter hand β€” capped at a quarter of its earlier strength, which kept
151
+ the detail and took some grain out of flat areas
152
+ 8. the fake-score adapter updated four times per student step instead of twice, so it keeps up with what the student
153
+ currently makes; faint straight-line artefacts that had begun to appear on flat illustrated areas went away with it
154
+ 9. **an artefact critic** β€” a small head reading the frozen base model's own mid-network features, trained to tell the
155
+ teacher's finished images from the student's, with its push restricted to structure finer than 32 pixels and held
156
+ well below the distribution term β€” aimed at what the distribution term leaves behind: melted small faces and dense
157
+ detail smeared into blotches
158
+ 10. four detail terms together: the anchor counting the fine-detail part of its error twice; a one-sided floor that
159
+ stops the finest detail dropping below the level real photographs carry; a second-call target made by the teacher
160
+ finishing the image from the student's own first-call output, so the target shares the student's layout; and a
161
+ smoothness limit on the fake-score adapter, so the distribution push keeps pointing at detail rather than away from it
162
+ 11. **three critics taking turns** β€” the artefact critic joined by one whose real examples are half real photographs and one
163
+ weighted toward faces, one of them pushing on each step while the others keep training in between, because the three did
164
+ not fit in memory side by side
165
+
166
+ Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
167
+ a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
168
+ not place exactly where the teacher had it.
169
+
170
+ ## chk00017464 vs chk00013663
171
+
172
+ `chk00013663` (published 12 Sep 2026) was the first public checkpoint: distribution matching with the spectral match, and nothing
173
+ yet aimed at the faults the distribution term leaves behind. `chk00017464` (14 Sep 2026) is 3,801 training samples later, and every
174
+ one of those samples went to those faults β€” the grain and grid pattern at large sizes, small faces, dense detail β€” through five recipe
175
+ changes, each kept only after its own look at the renders:
176
+
177
+ 1. the decoded-window spectral term brought down to a quarter of its strength
178
+ 2. the fake-score adapter updated four times per student step instead of twice
179
+ 3. the artefact critic, reading the frozen base model's own features
180
+ 4. four detail terms: the anchor counting fine-detail error twice, the one-sided photo floor, the teacher's finish of the student's
181
+ first call as the second call's target, and a smoothness limit on the fake adapter
182
+ 5. three critics taking turns: the artefact critic, a photo critic and a face critic
183
+
184
+ Two ideas were tried and taken back out along the way: the distribution term confined to the second call's noise range, and a detail
185
+ pyramid compared pixel by pixel against the teacher.
186
+
187
+ **Detail at large sizes β€” the headline.** Every one of the 12 sweep resolutions Γ— 15 prompts measured against the 8-step teacher, as in
188
+ [Measured against the teacher](#measured-against-the-teacher). Above 1 megapixel the excess fine energy two steps used to put into
189
+ images is roughly halved, and those sizes are now closer to the teacher than the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) on
190
+ fine texture and on both grid bands:
191
+
192
+ | 1.00 = the teacher | fine texture | 16-px band | 8-px band |
193
+ | --- | --- | --- | --- |
194
+ | 1280Γ—1280 | 1.20 β†’ **1.09** | 1.16 β†’ **1.08** | 1.25 β†’ **1.13** |
195
+ | 1440Γ—1280 | 1.21 β†’ **1.08** | 1.19 β†’ **1.09** | 1.28 β†’ **1.14** |
196
+ | 1440Γ—1440 | 1.34 β†’ **1.17** | 1.25 β†’ **1.08** | 1.42 β†’ **1.21** |
197
+
198
+ Fine-texture energy and both grid bands come closer to the teacher at 10 of the 12 resolutions, the blur-invariant ghosting index at
199
+ 10 of 12, the micro-ghost index at all 12, and the grain in flat areas β€” skies, walls, out-of-focus backgrounds β€” at 11 of 12
200
+ (1440Γ—1440: 1.84Γ— the teacher's β†’ 1.57Γ—, median of the 15 prompts). It shows where it should: hair renders as more individual strands, and the token-grid
201
+ texture that read as grain on birds, food and foliage at 1440Γ—1440 is lighter.
202
+
203
+ **Faces and skin.** Freckles on the test portrait gather into lighter, more dot-like clusters than before β€” better, not yet the
204
+ teacher's separate dots β€” and eyes look about the same: clean irises, lashes softer than the teacher's.
205
+
206
+ **What did not improve.** The judge that asks which of two renders follows the prompt better prefers the 8-step teacher on 11 of 45
207
+ against 6 before, mostly on how a stylised prompt says the picture should look: crisp linework, energetic brush strokes, the texture
208
+ of clay. The blind rubric that checks each render alone for the prompt's objects, counts, attributes and relations moves from 0 to 1
209
+ point missing out of 240, and on the same 15 fresh prompts the previous checkpoint was measured on, the two checkpoints come out level
210
+ (1 win, 13 ties, 1 loss). Colour runs slightly lower (saturation 0.93Γ— the teacher's across the sweep, from 0.94Γ—; 0.86Γ— at
211
+ 1440Γ—1440), and at 512Γ—512 the two checkpoints are level. Both are what the next recipe changes target.
212
+
213
+ | axis | `chk00013663` | `chk00017464` |
214
+ | --- | --- | --- |
215
+ | fine texture vs the teacher, 1280Β² / 1440Β² | 1.20 / 1.34 | **1.09 / 1.17** |
216
+ | 16-px grid band, 1280Β² / 1440Β² | 1.16 / 1.25 | **1.08 / 1.08** |
217
+ | grain in flat areas, sweep median | 1.36Γ— | **1.26Γ—** |
218
+ | distance to the teacher, sweep mean | 0.413 | **0.406** |
219
+ | judge prefers the teacher (of 45) | **6** | 11 |
220
+ | blind adherence rubric, points missing of 240 | **0** | 1 |
221
+ | saturation vs the teacher, sweep mean | **0.94Γ—** | 0.93Γ— |
222
+ | training samples in the 2-step stages | 13,663 | 17,464 |
223
 
224
  ## Measured against the teacher
225
 
 
234
 
235
  | bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
236
  | --- | --- | --- | --- | --- | --- |
237
+ | 512Γ—512 | 1.02 | 1.02 | 1.05 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
238
+ | 768Γ—1024 | 1.02 | 0.98 | 1.07 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
239
+ | 1024Γ—1024 | 1.06 | 1.07 | 1.12 | 1.11 Β· 1.17 Β· 1.16 | 0.39 |
240
+ | 1280Γ—1280 | 1.09 | 1.08 | 1.13 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
241
+ | 1440Γ—1440 | 1.17 | 1.08 | 1.21 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
242
 
243
  Two steps without the adapter carry **less than half** the teacher's fine detail at every size. With it, the detail
244
+ sits at or just above the teacher's everywhere β€” and from 1 megapixel up it is closer to the teacher than the 4-step
245
+ adapter, which carries more excess fine energy there.
246
 
247
  **Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
248
  in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
249
 
250
  | bucket | wins | ties | losses |
251
  | --- | --- | --- | --- |
252
+ | 512Γ—512 | 1 | 10 | 4 |
253
+ | 1280Γ—1280 | 1 | 10 | 4 |
254
+ | 1440Γ—1440 | 0 | 12 | 3 |
255
+
256
+ Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
257
+ stylised prompts and on how a prompt says the picture should look β€” crisp linework, energetic brush strokes, the texture of
258
+ clay β€” more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
259
+ attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
260
+ training prompt bank and never rendered before, the judge returned 0 wins, 10 ties, 5 losses; on another 15 fresh prompts, 1 win,
261
+ 13 ties, 1 loss.
262
+
263
+ **Checked for the damage this kind of training can do.** Saturation sits at 0.99Γ— the teacher's at 768Γ—1024 and 0.89Γ—
264
+ at the larger sizes; edge detail 0.93–0.96Γ—; skin texture inside detected faces 1.07Γ— at 768Γ—1024 and 0.85Γ— at
265
+ 1440Γ—1440, with skin saturation 0.92Γ— and 0.77–0.85Γ— at the larger sizes. The honest reading of those last two: **skin
266
+ is the softest and least saturated part of this adapter's output at large sizes, and colour overall runs a little under
267
+ the teacher's there.** Fine detail in flat regions β€” skies, walls, out-of-focus backgrounds β€” runs 1.53–1.85Γ— the
268
+ teacher's, which is where two steps put grain that eight steps do not.
269
+
270
+ **Distance to the teacher**, as a plain pixel measure, is 0.39–0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
271
+ 0.30–0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
272
  not the teacher's render of it β€” see [Known issues](#known-issues).
273
 
274
  ## Usage
 
281
  | guidance / CFG | **0.0** (Turbo is CFG-free; do not enable it) |
282
  | timestep shift | **mu = 1.15**, fixed (Turbo's deployment shift) |
283
 
284
+ The 2 sampling sigmas are Turbo's own deployment grid: `[1.0, 0.7595]` β€” the first and the middle
285
  of the 4-step grid, so the model is evaluated at two points it already knows.
286
 
287
  ## Inference with diffusers
 
400
 
401
  | configuration | denoising | per model call | GPU peak |
402
  | --- | --- | --- | --- |
403
+ | Krea 2 Turbo β€” 8 steps (the reference) | **81.4 s** | 10.2 s | 25.2 GiB |
404
+ | Krea 2 Turbo β€” 2 steps, no LoRA | 20.4 s | 10.2 s | 25.2 GiB |
405
+ | **Krea 2 Turbo β€” 2 steps + this LoRA** | **19.5 s** | 9.8 s | 25.2 GiB |
406
 
407
+ **Denoising is 4.2Γ— faster than the 8-step reference** β€” two model calls instead of eight. The adapter's own cost per call did
408
+ not show up in this measurement: the runs with it came in marginally faster than those without, which is measurement noise, not a
409
+ speed-up. A rank-64 low-rank product is small beside the transformer it is added to, and it adds no measurable memory.
 
410
 
411
  Denoising is the part the step count changes. What a complete render costs on top of it β€” encoding the prompt, decoding
412
  the latent, writing the file β€” is the same whether you run two steps or eight, and it depends on your pipeline, so the
413
+ end-to-end figure on your machine will sit below 4.2Γ— and rise toward it as the render gets larger.
414
+
415
+ **By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
416
+ size; sweep renders run one at a time, not the controlled measurement above):
417
+
418
+ | resolution | denoising (2 calls) | resolution | denoising (2 calls) |
419
+ | --- | --- | --- | --- |
420
+ | 512Γ—512 | 6.4 s | 1024Γ—1024 | 20.4 s |
421
+ | 512Γ—768 / 768Γ—512 | 8.9 s / 9.0 s | 1280Γ—960 / 960Γ—1280 | 23.2 s / 24.0 s |
422
+ | 768Γ—768 | 11.9 s | 1280Γ—1280 | 32.2 s |
423
+ | 768Γ—1024 / 1024Γ—768 | 14.9 s / 15.1 s | 1440Γ—1280 / 1440Γ—1440 | 34.8 s / 39.8 s |
424
 
425
  ## LoRA strength
426
 
 
498
  **Distribution matching with a trajectory anchor**, Krea 2 Turbo as its own teacher, on the recorded 8-step
499
  trajectories.
500
 
501
+ The student makes two calls, at Οƒ = 1.0 and Οƒ = 0.7595 β€” the first and fifth points of the teacher's 8-step grid at
502
  mu = 1.15 β€” and stock Euler carries it between them. That grid is what makes the objective a drop-in: Euler's first step
503
+ from pure noise lands _exactly_ on the flow-matching interpolant at Οƒ = 0.7595 with the same noise and the student's
504
  own clean-image prediction as the data point. So the student's first-call output is a legitimate image prediction that
505
  can be judged as an image, and the second call is fed from it during training the way it will be at inference.
506
 
 
509
  is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
510
  the image more like the teacher's work and less like the student's habits, and the student is pushed that way
511
  (the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
512
+ rank 32, starts as an exact copy of the teacher, updates four times per student step β€” often enough to keep up with a
513
+ student that is still changing β€” and is discarded at the end.
514
 
515
  **The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
516
  student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
 
521
  and under-fits fine structure there, and a per-pixel normaliser lands harder as pixel counts grow. A full 12-bucket
522
  sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
523
  1.4–1.8Γ— the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
524
+ not hold, and the spectral match replaced it. The same sweep of the published checkpoint, its running-average weights, fixed
525
+ seed, 15 prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
526
  first distribution-matching checkpoint on the same prompts):
527
 
528
  | bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
529
  | --------- | --------------------------- | --------------- | -------------- | ----------------------- |
530
+ | 512x512 | 1.02 (1.16) | 1.02 (1.23) | 1.05 (1.23) | 0.41 (0.44) |
531
+ | 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.06 (1.33) | 0.39 (0.41) |
532
+ | 768x512 | 1.02 (1.22) | 1.08 (1.37) | 1.04 (1.26) | 0.44 (0.47) |
533
+ | 768x768 | 1.09 (1.43) | 1.05 (1.46) | 1.14 (1.50) | 0.40 (0.43) |
534
+ | 768x1024 | 1.02 (1.36) | 0.98 (1.38) | 1.07 (1.42) | 0.39 (0.43) |
535
+ | 1024x768 | 1.01 (1.35) | 0.98 (1.40) | 1.04 (1.38) | 0.43 (0.46) |
536
+ | 1024x1024 | 1.06 (1.47) | 1.07 (1.55) | 1.12 (1.56) | 0.41 (0.44) |
537
+ | 1280x960 | 1.03 (1.50) | 1.04 (1.59) | 1.07 (1.56) | 0.40 (0.42) |
538
+ | 960x1280 | 1.10 (1.60) | 1.05 (1.63) | 1.19 (1.70) | 0.40 (0.43) |
539
+ | 1280x1280 | 1.09 (1.65) | 1.08 (1.68) | 1.13 (1.69) | 0.41 (0.44) |
540
+ | 1440x1280 | 1.08 (1.59) | 1.08 (1.66) | 1.14 (1.66) | 0.40 (0.42) |
541
+ | 1440x1440 | 1.17 (1.81) | 1.08 (1.79) | 1.21 (1.89) | 0.39 (0.41) |
542
 
543
  ### The spectral match
544
 
 
556
  judged too. The comparison is two-sided, so too much fine energy and too little are both penalised; a blurred image
557
  does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
558
  refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
559
+ without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too. The whole-latent comparison trains at that
560
+ strength; the decoded window was later brought down to a quarter of it, which kept the detail and removed some grain.
561
+
562
+ ### The artefact critic
563
+
564
+ Distribution matching improves what the fake adapter can see, and the fake adapter learns from the student's own
565
+ images β€” so where the student smears something, the fake learns the smear and the push stops correcting it. The two
566
+ places that shows most are small faces, which come out melted, and dense content such as the goods on a market stall or
567
+ the shelves of a shop seen through its window, which comes out as coloured blotches.
568
+
569
+ A critic breaks that loop by looking at finished images instead. It is a small head on the frozen base model's
570
+ mid-network features β€” every adapter switched off, the forward pass stopped halfway, so no adapter can learn to fool
571
+ it and nothing it learns leaks into the fake adapter. Its real examples are the teacher's own finished images for
572
+ other prompts at the same resolution; its fake examples are the student's final images. Both are lightly re-noised
573
+ first, in the low-noise range where fine structure lives, and the critic reads them with an empty prompt, so it judges
574
+ only whether the structure looks like the teacher's. Its push on the student is filtered to periods finer than 32
575
+ pixels β€” below that it can rebuild a face or a shelf, above it it could move layout, colour or pose, which it must not
576
+ β€” and capped at a quarter of the distribution term's strength per sample. Lazy gradient regularisation and spectral
577
+ normalisation keep the head from overshooting. On the largest steps, where memory is tightest, the critic and the
578
+ teacher's finishing pass below take turns instead of sharing a step.
579
+
580
+ ### Critics in turn
581
+
582
+ One critic holds one idea of what is wrong. The artefact critic was joined by two more on the same frozen mid-network
583
+ features and under the same rules β€” lightly re-noised inputs, an empty prompt, a push that is filtered and capped β€” each
584
+ aimed at a different fault:
585
+
586
+ - **A photo critic.** Half of its real examples are real photographs and half the teacher's finished images, so it learns
587
+ what fine texture looks like in a photograph as well as in the teacher's rendering of one. Its push is filtered to
588
+ periods finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher
589
+ does not.
590
+ - **A face critic.** Its real examples are the teacher's finished images of prompts with faces, the face regions counted at
591
+ full weight and the rest at half, so its push concentrates on what small and mid-sized faces lose first.
592
+
593
+ All three on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
594
+ others train so none goes stale before its turn comes back. The two new heads started from the artefact critic's weights and
595
+ trained on their own before they were allowed to push.
596
+
597
+ ### Detail terms
598
+
599
+ Four smaller terms sit on top, each capped relative to the distribution term so none of them can take over:
600
+
601
+ - **A detail-weighted anchor.** The trajectory regression counts the fine-detail part of its error β€” everything finer
602
+ than 32 pixels β€” twice, so the anchor stops tolerating softness it used to average away.
603
+ - **A one-sided photo floor.** On the decoded window, the student's energy at periods of 3–10 pixels may not fall below
604
+ the teacher's plus the margin real photographs carry over it at those scales. That margin is measured once from a
605
+ pool of real photographs and clamped, and the term only ever pushes upward to that floor, never past it β€” so it
606
+ lifts detail that is missing without adding grain that is not.
607
+ - **The teacher's finish as a target.** Every second step, the teacher itself runs its remaining steps starting from
608
+ the student's own first-call output. The result is a finished image that shares the student's layout, and the second
609
+ call is pulled gently toward it β€” a target that lines up with what the student actually drew, where the recorded
610
+ trajectory might have drawn something else.
611
+ - **A smoothness limit on the fake adapter.** The fake adapter's fine-detail energy is kept below the teacher's at the
612
+ point where the distribution term is measured, so the difference between the two keeps pointing toward detail.
613
 
614
  ## What the LoRA touches
615
 
 
622
  The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β€” Krea 2 Turbo's own 8-step run at mu = 1.15 and
623
  guidance 0.0, every latent and velocity stored β€” serve unchanged: a 2-step chord is two of the 4-step chords end to
624
  end. 203 held-out prompts measure the student–teacher gap on unseen prompts and never receive a gradient. The spectral
625
+ match and the artefact critic read the teacher's finals for the training prompts. The 43,044 real-photo crops of the
626
+ [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) enter in two places: as one precomputed statistic β€” how much fine-detail energy they carry at 3–10
627
+ pixels relative to the teacher, clamped β€” which sets the photo floor, and as half of the photo critic's real examples. Only that
628
+ critic's head sees them; the student and the fake adapter never do, and receive only its filtered, capped push.
629
 
630
  ## Resolutions
631
 
 
696
  ## Hardware
697
 
698
  One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned
699
+ host memory above 0.3 megapixels; the student, the fake adapter, the spectral and detail terms, the critic and the teacher's finishing pass each
700
+ build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
701
+ does not fit fails loudly. A full step with every term live and every critic pushing reserves about 22.4 GB at 1440Γ—1440,
702
+ of 24. The price of the objective is throughput: **about 107 training samples an hour** measured over the recipe's complete
703
+ run, against the [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β€” more than four times the cost per sample, and so far a small
704
+ fraction of the samples.
705
 
706
  **Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
707
  The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
708
  disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
709
  of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
710
+ alongside to imitate the student β€” and that second adapter takes four optimiser steps of its own per student step.
711
+ The spectral and detail terms decode part of the image out of the latent to compare its texture with the teacher's,
712
+ each critic reads half the network twice more, and every second step the teacher finishes the image from the student's
713
+ first call.
714
 
715
+ Counted in whole model runs per training sample, the difference is roughly **two there against about a dozen here**. None of
716
  that difference is the teacher generating anything: its renders were recorded once for the 4-step project and are
717
  read from disk by both. The extra work is the objective itself, and it bought the only thing that mattered. Run at
718
  two steps, the 4-step project's recipe reached a point where more training changed nothing: the measurements sat
 
726
  15 fixed prompts across four resolutions (512Γ—512, 768Γ—1024, 1280Γ—1280, 1440Γ—1440); milestone checkpoints get the same
727
  render at all 12 buckets, which is where the per-resolution table above comes from. Every image is measured against the
728
  teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
729
+ and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
730
+ seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
731
  scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
732
  identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
733
  checkpoint. Latent distances β€” the held-out chord gap and the two-step rollout
 
1085
 
1086
  ## What's next
1087
 
1088
+ Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any
1089
+ resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). The next step
1090
+ is already running, and it goes after what this checkpoint still gets wrong: a critic that judges the first call against the
1091
+ teacher's own intermediate state from the same noise, so a first call that blends two layouts is caught where the blend
1092
+ happens; a critic weighted toward wherever the teacher put fine detail; the photo critic and the photo floor limited to
1093
+ photographic prompts, so illustration, anime and 3D renders are no longer pulled toward photographic grain; detail held to the
1094
+ teacher region by region, with a ceiling as well as a floor; a focus on the eyes, nose and lips of faces so they sharpen while
1095
+ skin stays the teacher's; a colour floor, so colour at large sizes stops falling below the teacher's; and a more even mix of
1096
+ resolutions. After it comes prompt adherence β€” a critic that learns whether an image belongs to its own prompt, and stylised
1097
+ prompts drawn more often β€” held to the rule that none of it may cost the sharpness this checkpoint gained. A better checkpoint
1098
+ replaces this one when the sweeps and I visually agree, the same discipline as the
1099
+ [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
1100
+ one is the fast preview.
1101
 
1102
  ## License
1103
 
krea2_turbo_2step_rank_64_lora.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:59306ece2d3d55186c1498752a04816f19d9acb195d219e330e1f4c59228b614
3
- size 438160888
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cc7a7d65ca04070b2aa1dbf60b3cd0cf210fde2a93d4d7d4fc18bdadf23060d3
3
+ size 438161144
krea2_turbo_2step_rank_64_lora_checkpoint_info.md CHANGED
@@ -1,13 +1,13 @@
1
  # Which checkpoint is this?
2
 
3
  `krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
4
- **chk00013663** β€” the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00013663.safetensors` (and its
5
  `_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
6
  keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
7
 
8
  | file | SHA-256 | size |
9
  | --- | --- | --- |
10
- | `krea2_turbo_2step_rank_64_lora.safetensors` | `59306ece2d3d55186c1498752a04816f19d9acb195d219e330e1f4c59228b614` | 418M |
11
- | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `0887ac53a78a015c22c6fc6dfa2f76e9fca7e8c19ac074d65ad12146cbacb999` | 418M |
12
 
13
- Updated: 12 Sep 2026
 
1
  # Which checkpoint is this?
2
 
3
  `krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
4
+ **chk00017464** β€” the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00017464.safetensors` (and its
5
  `_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
6
  keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
7
 
8
  | file | SHA-256 | size |
9
  | --- | --- | --- |
10
+ | `krea2_turbo_2step_rank_64_lora.safetensors` | `cc7a7d65ca04070b2aa1dbf60b3cd0cf210fde2a93d4d7d4fc18bdadf23060d3` | 418M |
11
+ | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `5c02cac5de27dcb40ea98caee5fd5668ac1a928a105a1fb29c02eb9b0b1e5862` | 418M |
12
 
13
+ Updated: 14 Sep 2026
krea2_turbo_2step_rank_64_lora_comfyui.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:0887ac53a78a015c22c6fc6dfa2f76e9fca7e8c19ac074d65ad12146cbacb999
3
- size 438141424
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5c02cac5de27dcb40ea98caee5fd5668ac1a928a105a1fb29c02eb9b0b1e5862
3
+ size 438141680