lvladikov commited on
Commit
63bf174
Β·
verified Β·
1 Parent(s): e007d17

chk41320 release

Browse files
DETAILED-README.md CHANGED
@@ -43,8 +43,8 @@ recommendation for quality renders.
43
  - 🎯 **The aim** β€” the best two-step quality this base can give, at every one of the same 12 resolutions, measured
44
  against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
45
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas.
46
- - ⏱️ **4.2Γ— faster denoising** β€” the model runs twice instead of eight times, and denoising is the part this adapter
47
- changes: **81.4 s β†’ 19.5 s** measured at 1024Γ—1024 on the same prompts, the adapter's own cost per call within
48
  measurement noise. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
49
  [Performance](#performance).
50
  - πŸ“Š **Distribution matching, not imitation** β€” the training objective that got the renders improving again after the
@@ -64,17 +64,17 @@ recommendation for quality renders.
64
  only and never ships.
65
  - 🎲 **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
66
  re-run.
67
- - πŸ”’ **31,600 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β€” all of them
68
  drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
69
  training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
70
  4-step project, read again at the two sigmas this schedule uses.
71
- - πŸ“… **18 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β€” and the project continues.
72
  - πŸ” **More than forty recipe adjustments** across two methods so far β€” seven of trajectory distillation before the switch, the rest of distribution matching since β€” each kept only when the renders did not get worse.
73
  - πŸ–₯️ **One RTX 3090**, and a recipe shaped by its 24 GB.
74
 
75
- [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
76
 
77
- _All of the above were created with this LoRA at 2 steps: the 15 test prompts, Krea 2 Turbo + the
78
  LoRA, seed 4242, each at one of its trained resolutions. Click for full size. The side-by-side
79
  comparisons with the 8-step teacher are in [Examples](#examples)._
80
 
@@ -99,9 +99,9 @@ place to check which checkpoint the current files are based on.
99
 
100
  | | |
101
  | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
102
- | lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β†’ 2-step trajectory distillation β†’ distribution matching β†’ a spectral match against the teacher's own images β†’ an artefact critic and detail terms β†’ three critics taking turns β†’ nine critics, detail and colour held to the teacher region by region, a guarded running average |
103
  | this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
104
- | what it gives | usable two-step renders at every trained resolution: fine detail and colour at the teacher's level β€” from 1 megapixel up, closer to the teacher than the 4-step adapter β€” with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 1 point missing out of 240). A judge asked which render follows the prompt better still prefers the 8-step teacher on 11 of 45, against 6 for the 4-step adapter, mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
105
 
106
  ### Known issues
107
 
@@ -183,11 +183,96 @@ The recipe adjustments so far, each made on the measurement of the one before:
183
  17. **a guarded running average** β€” the published weights are a running average of training, and averaging two layouts of
184
  the same prompt had produced doubled subjects: a short-lived excursion of the training weights is now kept out of the
185
  average and a lasting change taken in whole, and the average was restarted once, after a layout change it had blended
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
186
 
187
  Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
188
  a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
189
  not place exactly where the teacher had it.
190
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
191
  ## chk00031600 vs chk00017464
192
 
193
  `chk00017464` (published 14 Sep 2026) was the first checkpoint with critics and detail terms. `chk00031600` (25 Sep 2026) is
@@ -315,51 +400,51 @@ point missing out of 240, and on the same 15 fresh prompts the previous checkpoi
315
  ## Measured against the teacher
316
 
317
  Every number here compares a render of this LoRA at 2 steps with the 8-step reference render of the **same prompt at the
318
- same seed**, across the 15 test prompts and the 12 trained resolutions. Two other columns are measured the same way, so
319
  the figures have a floor and a ceiling around them: **stock Krea 2 Turbo at 2 steps**, which is what the base model does
320
  without the adapter, and the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) at its own 4 steps, which is the better tool and the thing worth
321
- being compared against.
322
 
323
  **Detail, band by band.** Fine-texture energy and the two grid bands, as a ratio to the teacher's own (1.00 = the
324
  teacher):
325
 
326
- | bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
327
  | --- | --- | --- | --- | --- | --- |
328
- | 512Γ—512 | 1.00 | 1.03 | 1.04 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
329
- | 768Γ—1024 | 0.95 | 0.93 | 0.98 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
330
- | 1024Γ—1024 | 1.00 | 1.03 | 1.05 | 1.11 Β· 1.17 Β· 1.16 | 0.39 |
331
- | 1280Γ—1280 | 0.95 | 0.95 | 0.96 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
332
- | 1440Γ—1440 | 1.08 | 1.05 | 1.13 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
333
 
334
- Two steps without the adapter carry 0.57Γ— the teacher's fine detail at 512Γ—512 and **less than half** (0.39–0.42Γ—) at the four larger sizes. With it, the detail
335
- sits at the teacher's level everywhere β€” between 0.93Γ— and 1.13Γ— of it β€” and from 1 megapixel up it is closer to the teacher
336
- than the 4-step adapter, which carries more excess fine energy there.
337
 
338
  **Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
339
  in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
340
 
341
  | bucket | wins | ties | losses |
342
  | --- | --- | --- | --- |
343
- | 512Γ—512 | 0 | 12 | 3 |
344
- | 1280Γ—1280 | 0 | 11 | 4 |
345
- | 1440Γ—1440 | 0 | 11 | 4 |
346
-
347
- Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
348
- stylised prompts and on how a prompt says the picture should look β€” crisp linework, energetic brush strokes, the texture of
349
- clay β€” more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
350
- attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
351
- training prompt bank for this checkpoint and never rendered before, the judge returned 1 win, 9 ties, 5 losses, and on five fresh
352
- black-and-white prompts 0 wins, 4 ties, 1 loss β€” with no colour cast in any of the five.
353
-
354
- **Checked for the damage this kind of training can do.** Saturation sits at 1.04Γ— the teacher's at 768Γ—1024 and 0.94–0.97Γ—
355
- at the larger sizes; edge detail 0.90–0.91Γ—; skin texture inside detected faces 1.04Γ— at 768Γ—1024 and 0.87Γ— / 0.81Γ— at
356
- 1280Γ—1280 / 1440Γ—1440, with skin saturation 0.97Γ— and 0.80–0.94Γ— at the larger sizes. The honest reading: **colour is now at
357
- the teacher's level at every size, and skin remains the softest part of this adapter's output at large sizes.** Fine detail
358
- in flat regions β€” skies, walls, out-of-focus backgrounds β€” runs 1.33–1.69Γ— the teacher's on this measure, which is where two
359
- steps put grain that eight steps do not.
360
-
361
- **Distance to the teacher**, as a plain pixel measure, is 0.39–0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
362
- 0.30–0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
363
  not the teacher's render of it β€” see [Known issues](#known-issues).
364
 
365
  ## Usage
@@ -497,28 +582,28 @@ the model already loaded, so the numbers are the render itself and not a model l
497
 
498
  | configuration | denoising | per model call | GPU peak |
499
  | --- | --- | --- | --- |
500
- | Krea 2 Turbo β€” 8 steps (the reference) | **81.4 s** | 10.2 s | 25.2 GiB |
501
- | Krea 2 Turbo β€” 2 steps, no LoRA | 20.4 s | 10.2 s | 25.2 GiB |
502
- | **Krea 2 Turbo β€” 2 steps + this LoRA** | **19.5 s** | 9.8 s | 25.2 GiB |
503
 
504
- **Denoising is 4.2Γ— faster than the 8-step reference** β€” two model calls instead of eight. The adapter's own cost per call did
505
  not show up in this measurement: the runs with it came in marginally faster than those without, which is measurement noise, not a
506
  speed-up. A rank-64 low-rank product is small beside the transformer it is added to, and it adds no measurable memory.
507
 
508
  Denoising is the part the step count changes. What a complete render costs on top of it β€” encoding the prompt, decoding
509
  the latent, writing the file β€” is the same whether you run two steps or eight, and it depends on your pipeline, so the
510
- end-to-end figure on your machine will sit below 4.2Γ— and rise toward it as the render gets larger.
511
 
512
- **By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
513
  size; sweep renders run one at a time, not the controlled measurement above, so a second or two either way between one sweep and
514
  the next is run-to-run variation β€” the adapter's shape, and so its cost, is the same at every checkpoint):
515
 
516
  | resolution | denoising (2 calls) | resolution | denoising (2 calls) |
517
  | --- | --- | --- | --- |
518
- | 512Γ—512 | 7.0 s | 1024Γ—1024 | 19.6 s |
519
- | 512Γ—768 / 768Γ—512 | 11.5 s / 10.9 s | 1280Γ—960 / 960Γ—1280 | 22.9 s / 24.2 s |
520
- | 768Γ—768 | 12.6 s | 1280Γ—1280 | 31.5 s |
521
- | 768Γ—1024 / 1024Γ—768 | 15.1 s / 15.1 s | 1440Γ—1280 / 1440Γ—1440 | 34.4 s / 39.6 s |
522
 
523
  ## LoRA strength
524
 
@@ -621,23 +706,23 @@ and under-fits fine structure there, and a per-pixel normaliser lands harder as
621
  sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
622
  1.4–1.8Γ— the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
623
  not hold, and the spectral match replaced it. The same sweep of the published checkpoint, its running-average weights, fixed
624
- seed, 15 prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
625
- first distribution-matching checkpoint on the same prompts):
626
 
627
  | bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
628
  | --------- | --------------------------- | --------------- | -------------- | ----------------------- |
629
- | 512x512 | 1.00 (1.16) | 1.03 (1.23) | 1.04 (1.23) | 0.42 (0.44) |
630
- | 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.04 (1.33) | 0.40 (0.41) |
631
- | 768x512 | 0.98 (1.22) | 1.03 (1.37) | 0.99 (1.26) | 0.44 (0.47) |
632
- | 768x768 | 1.03 (1.43) | 1.05 (1.46) | 1.05 (1.50) | 0.41 (0.43) |
633
- | 768x1024 | 0.95 (1.36) | 0.93 (1.38) | 0.98 (1.42) | 0.40 (0.43) |
634
- | 1024x768 | 1.01 (1.35) | 1.00 (1.40) | 1.02 (1.38) | 0.42 (0.46) |
635
- | 1024x1024 | 1.00 (1.47) | 1.03 (1.55) | 1.05 (1.56) | 0.42 (0.44) |
636
- | 1280x960 | 0.95 (1.50) | 0.98 (1.59) | 1.00 (1.56) | 0.42 (0.42) |
637
- | 960x1280 | 1.02 (1.60) | 0.99 (1.63) | 1.07 (1.70) | 0.42 (0.43) |
638
- | 1280x1280 | 0.95 (1.65) | 0.95 (1.68) | 0.96 (1.69) | 0.42 (0.44) |
639
- | 1440x1280 | 0.95 (1.59) | 0.97 (1.66) | 0.99 (1.66) | 0.40 (0.42) |
640
- | 1440x1440 | 1.08 (1.81) | 1.05 (1.79) | 1.13 (1.89) | 0.39 (0.41) |
641
 
642
  ### The spectral match
643
 
@@ -680,7 +765,7 @@ teacher's finishing pass below take turns instead of sharing a step.
680
 
681
  ### Critics in turn
682
 
683
- One critic holds one idea of what is wrong. The artefact critic is joined by eight more on the same frozen mid-network
684
  features and under the same rules β€” lightly re-noised inputs, a push that is filtered and capped β€” each aimed at a
685
  different fault:
686
 
@@ -688,9 +773,9 @@ different fault:
688
  what fine texture looks like in a photograph as well as in the teacher's rendering of one. It judges photographic prompts
689
  only, so illustration, anime and 3D renders are not pulled toward photographic grain, and its push is filtered to periods
690
  finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher does not.
691
- - **Two face critics, on photographs.** One reads faces of every size, the face regions counted at full weight and the rest at
692
- half, so its push concentrates on what small and mid-sized faces lose first; the other reads only large faces, 192 pixels and
693
- up, and pushes on every step.
694
  - **A structure critic.** It judges the first call β€” the layout, before any detail β€” against the teacher's own intermediate
695
  state from the same noise, at the high noise levels where layout is decided and at periods of 32 pixels and coarser only,
696
  so a first call that blends two layouts is caught where the blend happens.
@@ -701,7 +786,7 @@ different fault:
701
  - **A rollout critic.** Its real examples are the teacher's own finish from the student's second-call starting point, so the
702
  second call is judged against what the teacher would have made from the same start.
703
 
704
- All nine on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
705
  others train so none goes stale before its turn comes back. Each new head started from the artefact critic's weights and
706
  trained on its own before it was allowed to push.
707
 
@@ -716,11 +801,8 @@ Smaller terms sit on top, each capped relative to the distribution term so none
716
  pixels may not fall below the teacher's plus the margin real photographs carry over it at those scales, judged tile by tile
717
  on the window's textured tiles. That margin is measured once from a pool of real photographs and clamped, and the term
718
  only ever pushes upward to that floor, never past it β€” so it lifts detail that is missing without adding grain that is not.
719
- - **Floors at the teacher's own level.** Stylised prompts, and the flat tiles of every image, get a floor at the teacher's
720
- own 3–10-pixel energy instead, so a clay surface or a painted sky cannot fade below the teacher's and nothing is added
721
- above it.
722
- - **A ceiling.** Tile by tile, detail at 3–16 pixels may not climb past 1.3Γ— the teacher's β€” the counterpart of the floors,
723
- and what keeps grain from building up at large sizes.
724
  - **A direction-aware term.** The spectral comparison is also made orientation by orientation, so the student's fine detail
725
  runs in the same directions as the teacher's.
726
  - **Windows on features.** On photographs, half of the decoded windows are centred on an eye, the nose or the lips of a face,
@@ -745,8 +827,21 @@ monochrome prompt is never pushed toward colour.
745
  The published weights are a running average of training (decay 0.999), which smooths out the noise of single steps.
746
  Averaging has one failure: while the training weights move between two layouts of the same prompt, their average draws
747
  both β€” a doubled subject. A guard watches a fixed set of layout probes every ten steps; a move away from the trend that
748
- comes back within a few rounds is kept out of the average, and a lasting move is taken in whole β€” the guard never resets
749
- the average. The average itself was restarted once, by hand, after a layout change it had blended.
 
 
 
 
 
 
 
 
 
 
 
 
 
750
 
751
  ## What the LoRA touches
752
 
@@ -786,8 +881,9 @@ strong at the size it saw most while quietly softer at the ones it barely saw. O
786
  which is why every cut of this run is rendered across the buckets and why the per-resolution table under
787
  [Method](#method) above exists.
788
 
789
- [`assets/resolution_sweeps/`](assets/resolution_sweeps) holds the evidence β€” same 15 prompts, same seed, one folder per
790
- resolution, one image per prompt, so any image can be compared 1:1 with its twin in the next tree:
 
791
 
792
  - [`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step) β€” the **official Krea 2 Turbo 8-step reference renders**:
793
  the stock model, no LoRA, at its native settings (8 steps, guidance 0.0). The teacher this LoRA is distilled from and
@@ -806,8 +902,8 @@ assets/resolution_sweeps/
806
  β”‚ β”œβ”€β”€ 512x512/ one folder per resolution
807
  β”‚ β”‚ β”œβ”€β”€ portrait.jpg
808
  β”‚ β”‚ β”œβ”€β”€ kingfisher.jpg
809
- β”‚ β”‚ β”œβ”€β”€ … 13 more, one per test prompt
810
- β”‚ β”‚ └── snowleopard.jpg
811
  β”‚ β”œβ”€β”€ 768x1024/
812
  β”‚ β”œβ”€β”€ 1024x1024/
813
  β”‚ β”œβ”€β”€ 1280x1280/
@@ -834,15 +930,15 @@ One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's che
834
  host memory above 0.3 megapixels; the student, the fake adapter, the spectral and detail terms, the critic and the teacher's finishing pass each
835
  build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
836
  does not fit fails loudly. A full step with every term live and every critic pushing reserves about 22.4 GB at 1440Γ—1440,
837
- of 24. The price of the objective is throughput: **about 107 training samples an hour** measured over the recipe's complete
838
- run, against the [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β€” more than four times the cost per sample, and so far a small
839
- fraction of the samples.
840
 
841
  **Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
842
  The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
843
  disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
844
  of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
845
- alongside to imitate the student β€” and that second adapter takes four optimiser steps of its own per student step.
846
  The spectral and detail terms decode part of the image out of the latent to compare its texture with the teacher's,
847
  each critic reads half the network twice more, and every second step the teacher finishes the image from the student's
848
  first call.
@@ -858,12 +954,14 @@ visibly better than the last again.
858
  ## How it is judged
859
 
860
  At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
861
- 15 fixed prompts across four resolutions (512Γ—512, 1280Γ—1280, 1440Γ—1440, 1440Γ—1280), after a layout check at both 1440
862
  sizes that has to pass first; milestone checkpoints get the same render at all 12 buckets, which is where the
863
  per-resolution table above comes from. Every image is measured against the
864
  teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
865
  and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
866
- seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
 
 
867
  scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
868
  identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
869
  checkpoint. Latent distances β€” the held-out chord gap and the two-step rollout
@@ -1179,6 +1277,139 @@ individual renders β€” click any image for full size.
1179
 
1180
  ---
1181
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1182
  ## Notes and limitations
1183
 
1184
  - 🎯 **Krea 2 Turbo only**, at **2 steps**, guidance **0.0** (cfg 1.0 in ComfyUI), **mu = 1.15** β€” the two training
@@ -1191,8 +1422,9 @@ Training continues from this checkpoint, one recipe change at a time, each kept
1191
  resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Next, aimed
1192
  at what this checkpoint still gets wrong: small subjects in wide scenes; the finest edges and the texture of skin at the
1193
  largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two
1194
- plausible poses of the same subject can meet; and the tactile surface of stylised materials such as clay β€” each held to the
1195
- rule that none of it may cost the colour and the clean large sizes this checkpoint gained. A better checkpoint
 
1196
  replaces this one when the sweeps and I visually agree, the same discipline as the
1197
  [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
1198
  one is the fast preview.
 
43
  - 🎯 **The aim** β€” the best two-step quality this base can give, at every one of the same 12 resolutions, measured
44
  against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
45
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas.
46
+ - ⏱️ **4Γ— faster denoising** β€” the model runs twice instead of eight times, and denoising is the part this adapter
47
+ changes: **76.4 s β†’ 19.3 s** measured at 1024Γ—1024 on the same prompts, the adapter's own cost per call within
48
  measurement noise. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
49
  [Performance](#performance).
50
  - πŸ“Š **Distribution matching, not imitation** β€” the training objective that got the renders improving again after the
 
64
  only and never ships.
65
  - 🎲 **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
66
  re-run.
67
+ - πŸ”’ **41,320 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β€” all of them
68
  drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
69
  training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
70
  4-step project, read again at the two sigmas this schedule uses.
71
+ - πŸ“… **25 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β€” and the project continues.
72
  - πŸ” **More than forty recipe adjustments** across two methods so far β€” seven of trajectory distillation before the switch, the rest of distribution matching since β€” each kept only when the renders did not get worse.
73
  - πŸ–₯️ **One RTX 3090**, and a recipe shaped by its 24 GB.
74
 
75
+ [![The 22 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
76
 
77
+ _All of the above were created with this LoRA at 2 steps: the 22 test prompts, Krea 2 Turbo + the
78
  LoRA, seed 4242, each at one of its trained resolutions. Click for full size. The side-by-side
79
  comparisons with the 8-step teacher are in [Examples](#examples)._
80
 
 
99
 
100
  | | |
101
  | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
102
+ | lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β†’ 2-step trajectory distillation β†’ distribution matching β†’ a spectral match against the teacher's own images β†’ an artefact critic and detail terms β†’ three critics taking turns β†’ nine critics, detail and colour held to the teacher region by region, a guarded running average β†’ the teacher's own layout held at the large sizes, eight critics, every update to the average checked for doubled subjects |
103
  | this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
104
+ | what it gives | usable two-step renders at every trained resolution: fine detail and colour at the teacher's level β€” from 1 megapixel up, closer to the teacher than the 4-step adapter β€” with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 2 points missing out of 352, the teacher's own renders 1). A judge asked which render follows the prompt better still prefers the 8-step teacher on 12 of 66 (the 4-step adapter: 6 of 45, on the original 15 prompts), mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
105
 
106
  ### Known issues
107
 
 
183
  17. **a guarded running average** β€” the published weights are a running average of training, and averaging two layouts of
184
  the same prompt had produced doubled subjects: a short-lived excursion of the training weights is now kept out of the
185
  average and a lasting change taken in whole, and the average was restarted once, after a layout change it had blended
186
+ 18. **three of the newest additions taken back out** β€” the floors at the teacher's level for stylised prompts and for flat areas,
187
+ and a face critic for photographs, all of which had trained the last stretch of the previous checkpoint: a render of the
188
+ training weights drew one animal as two joined bodies, and the recipe went back to the one that had trained the weeks before,
189
+ with eight critics
190
+ 19. **a stricter guard on the running average** β€” a change in the training weights is taken into the average only once it has held
191
+ for six rounds instead of three; shorter swings are left out
192
+ 20. **the second call held to the teacher's layout** β€” the anchor to the teacher's recorded trajectory counts the coarse layout of
193
+ the second call, everything 64 pixels and up, twice
194
+ 21. **every update to the running average checked for doubled subjects first** β€” before a stretch of training goes into the
195
+ average, the average it would make renders a fixed probe at the large sizes, and a stretch that would put a second head on the
196
+ subject is held out
197
+ 22. **the distribution term halved at the large sizes** β€” at every size with a 1280- or 1440-pixel side, while the anchor to the
198
+ teacher's trajectory keeps its full weight there, so at those sizes the teacher's own layout carries twice the share it did;
199
+ the doubled subjects had formed at exactly the sizes trained most
200
+ 23. the running average rebuilt once more, from the training weights since the halving, after the guard had held it still through
201
+ a long change of pose
202
 
203
  Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
204
  a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
205
  not place exactly where the teacher had it.
206
 
207
+ ## chk00041320 vs chk00031600
208
+
209
+ `chk00031600` (published 25 Sep 2026) was the first checkpoint with a colour band, nine critics and a guarded running average.
210
+ `chk00041320` (2 Oct 2026) is 9,720 training samples later, and those samples went to what matters first in a picture β€” subjects
211
+ drawn twice, or two poses blended into one, at the large sizes β€” and to faces in busy scenes, through the changes numbered 18 to 23
212
+ under [How I got here](#how-i-got-here), each kept only after its own look at the renders:
213
+
214
+ 1. three of the newest additions taken back out, eight critics
215
+ 2. a stricter guard on the running average
216
+ 3. the second call held to the teacher's layout
217
+ 4. every update to the running average checked for doubled subjects
218
+ 5. the distribution term halved at the large sizes
219
+ 6. the running average rebuilt from the recent training
220
+
221
+ It is also the first checkpoint measured on the full 22-prompt sweep: the original 15, plus seven scenes of people, animals and
222
+ action added on 29 Sep 2026 because structure is what they test β€” five friends on a beach, a family at a table seen from above, three
223
+ kittens in a basket, two dogs in a tug of war, a show jumper, a pianist's hands, a flock of flamingos. Every figure below compares the
224
+ two checkpoints on those 22, each against the 8-step teacher.
225
+
226
+ **Structure β€” the headline.** A judge compares each of the 21 structure renders β€” the seven scenes at 1280Γ—1280, 1440Γ—1440 and
227
+ 1440Γ—1280 β€” with the teacher's for missing, extra or merged body parts and subjects, and every flag is checked by eye at 1:1. The
228
+ previous checkpoint has one real fault among them β€” a ghost saddle pad and boot behind the show jumper's neck at 1280Γ—1280 β€” and this
229
+ one has none, and the snow leopard the guard watches has come out as one animal at every large size in every checkpoint since the
230
+ average was rebuilt. Doubling at the level of fine structure falls too: the two ghosting indexes come closer to the teacher at 11 and
231
+ 10 of the 12 sizes, most at the large ones β€” at 1440Γ—1440 from 1.17Γ— the teacher's to 1.06Γ—.
232
+
233
+ **Faces in busy scenes.** The market scene at 1280Γ—1280 β€” vendors leaning over a stall β€” now comes out with whole faces where the
234
+ previous checkpoint drew a broken one, and its layout sits much closer to the teacher's (a correlation of the stall side with the
235
+ teacher's render: 0.71 β†’ 0.79). Inside the faces the sweep detects, skin texture comes closer to the teacher's at 1280Γ—1280 and
236
+ 1440Γ—1440 (0.89Γ— β†’ 0.90Γ—, 0.83Γ— β†’ 0.87Γ—), and skin colour at 1280Γ—1280 rises from 0.86Γ— of the teacher's to 0.95Γ—.
237
+
238
+ **Detail and colour.** At the two largest sizes the excess energy two steps put into the grid bands comes down toward the
239
+ teacher's level:
240
+
241
+ | 1.00 = the teacher | fine texture | 16-px band | 8-px band |
242
+ | --- | --- | --- | --- |
243
+ | 1440Γ—1280 | 1.01 β†’ 0.99 | 1.03 β†’ **1.00** | 1.10 β†’ **1.09** |
244
+ | 1440Γ—1440 | 1.07 β†’ **1.03** | 1.08 β†’ **1.05** | 1.13 β†’ **1.09** |
245
+
246
+ At 1280Γ—1280 and below the fine texture stays within a few percent of the teacher's, and the grain in flat areas β€” skies, walls,
247
+ out-of-focus backgrounds β€” comes closer to the teacher's at 8 of the 12 sizes: 1.03Γ— the teacher's across the sweep, from 1.07Γ—.
248
+ Edge detail rises at 768Γ—1024 and 1280Γ—1280 (0.91Γ— β†’ 0.94Γ—, 0.91Γ— β†’ 0.93Γ—), and colour stays at the teacher's level: 0.99Γ— across
249
+ the sweep, from 0.98Γ—.
250
+
251
+ **Prompt adherence.** The judge that asks which of two renders follows the prompt better prefers the teacher on 12 of 66 β€” the 22
252
+ prompts at 512Γ—512, 1280Γ—1280 and 1440Γ—1440 β€” against 15 for the previous checkpoint, and gives this one the only outright win. The
253
+ blind rubric, now 22 prompts at four sizes, finds the same 2 points missing out of 352 for both; the teacher's own renders lose 1. On
254
+ 15 prompts drawn fresh from the prompt bank for this checkpoint, plus 5 black-and-white ones: 0 wins, 14 ties, 1 loss, and 0 Β· 4 Β· 1
255
+ in black and white.
256
+
257
+ **Speed β€” unchanged.** The same adapter shape at the same cost: measured again at 1024Γ—1024, two steps with this LoRA took 19.3
258
+ and 19.2 s against 76.4 and 76.3 s for the teacher's eight β€” 4.0Γ—, two model calls instead of eight.
259
+
260
+ **What stays a limit of two steps.** Small subjects in wide scenes and the tactile surface of stylised materials such as clay stay
261
+ where two steps fall furthest short of eight. The figures are in [Measured against the teacher](#measured-against-the-teacher).
262
+
263
+ | axis | `chk00031600` | `chk00041320` |
264
+ | --- | --- | --- |
265
+ | structure renders with a fault the teacher's does not have (of 21, checked at 1:1) | 1 | **0** |
266
+ | ghosting, blur-invariant index, sweep mean | 1.24Γ— | **1.21Γ—** |
267
+ | market faces at 1280Β², layout closeness to the teacher | 0.71 | **0.79** |
268
+ | grain in flat areas, sweep mean | 1.07Γ— | **1.03Γ—** |
269
+ | fine texture vs the teacher, 1440Β² | 1.07 | **1.03** |
270
+ | judge prefers the teacher (of 66) | 15 | **12** |
271
+ | blind adherence rubric, points missing of 352 | 2 | 2 |
272
+ | saturation vs the teacher, sweep mean | 0.98Γ— | **0.99Γ—** |
273
+ | distance to the teacher, sweep mean | 0.41 | 0.41 |
274
+ | training samples in the 2-step stages | 31,600 | 41,320 |
275
+
276
  ## chk00031600 vs chk00017464
277
 
278
  `chk00017464` (published 14 Sep 2026) was the first checkpoint with critics and detail terms. `chk00031600` (25 Sep 2026) is
 
400
  ## Measured against the teacher
401
 
402
  Every number here compares a render of this LoRA at 2 steps with the 8-step reference render of the **same prompt at the
403
+ same seed**, across the 22 test prompts and the 12 trained resolutions. Two other columns are measured the same way, so
404
  the figures have a floor and a ceiling around them: **stock Krea 2 Turbo at 2 steps**, which is what the base model does
405
  without the adapter, and the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) at its own 4 steps, which is the better tool and the thing worth
406
+ being compared against β€” its renders exist for the original 15 prompts, so its figures are on those 15.
407
 
408
  **Detail, band by band.** Fine-texture energy and the two grid bands, as a ratio to the teacher's own (1.00 = the
409
  teacher):
410
 
411
+ | bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three, original 15) | stock 2-step, fine texture |
412
  | --- | --- | --- | --- | --- | --- |
413
+ | 512Γ—512 | 1.00 | 1.01 | 1.09 | 1.05 Β· 1.05 Β· 1.07 | 0.65 |
414
+ | 768Γ—1024 | 1.00 | 0.98 | 1.02 | 1.05 Β· 1.05 Β· 1.08 | 0.44 |
415
+ | 1024Γ—1024 | 1.06 | 1.10 | 1.13 | 1.11 Β· 1.17 Β· 1.16 | 0.42 |
416
+ | 1280Γ—1280 | 1.02 | 1.02 | 1.03 | 1.20 Β· 1.18 Β· 1.22 | 0.44 |
417
+ | 1440Γ—1440 | 1.03 | 1.05 | 1.09 | 1.24 Β· 1.23 Β· 1.32 | 0.52 |
418
 
419
+ Two steps without the adapter carry 0.65Γ— the teacher's fine detail at 512Γ—512 and **about half or less** (0.42–0.52Γ—) at the four
420
+ larger sizes. With it, the fine texture sits at the teacher's level everywhere β€” between 0.97Γ— and 1.09Γ— of it across all 12
421
+ sizes β€” and from 1 megapixel up it is closer to the teacher than the 4-step adapter, which carries more excess fine energy there.
422
 
423
  **Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
424
  in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
425
 
426
  | bucket | wins | ties | losses |
427
  | --- | --- | --- | --- |
428
+ | 512Γ—512 | 1 | 15 | 6 |
429
+ | 1280Γ—1280 | 0 | 18 | 4 |
430
+ | 1440Γ—1440 | 0 | 20 | 2 |
431
+
432
+ Twelve losses out of 66 (the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six of 45 against the same teacher on the original 15). They
433
+ gather on stylised prompts and on how a prompt says the picture should look β€” crisp linework, energetic brush strokes, the texture
434
+ of clay β€” more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
435
+ attributes and relations, with the teacher scored identically, finds 2 points missing out of 352 (the teacher's own renders lose 1).
436
+ On 15 prompts drawn fresh from the training prompt bank for this checkpoint and never rendered before, the judge returned 0 wins,
437
+ 14 ties and 1 loss, and on five fresh black-and-white prompts 0 wins, 4 ties, 1 loss β€” with no colour cast in any of the five.
438
+
439
+ **Checked for the damage this kind of training can do.** Saturation sits at 1.04Γ— the teacher's at 768Γ—1024 and 0.98Γ— at
440
+ the larger sizes; edge detail 0.92–0.94Γ—; skin texture inside detected faces 1.02Γ— at 768Γ—1024 and 0.90Γ— / 0.87Γ— at
441
+ 1280Γ—1280 / 1440Γ—1440, with skin saturation 0.98Γ— and 0.95Γ— at the larger sizes. The honest reading: **colour is at the
442
+ teacher's level at every size, and skin remains the softest part of this adapter's output at large sizes.** Fine detail in flat
443
+ regions β€” skies, walls, out-of-focus backgrounds β€” runs 1.32–1.57Γ— the teacher's on this measure, which is where two steps put
444
+ grain that eight steps do not.
445
+
446
+ **Distance to the teacher**, as a plain pixel measure, is 0.39–0.43 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
447
+ 0.30–0.37 (on the original 15). That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
448
  not the teacher's render of it β€” see [Known issues](#known-issues).
449
 
450
  ## Usage
 
582
 
583
  | configuration | denoising | per model call | GPU peak |
584
  | --- | --- | --- | --- |
585
+ | Krea 2 Turbo β€” 8 steps (the reference) | **76.4 s** | 9.5 s | 25.2 GiB |
586
+ | Krea 2 Turbo β€” 2 steps, no LoRA | 19.4 s | 9.7 s | 25.2 GiB |
587
+ | **Krea 2 Turbo β€” 2 steps + this LoRA** | **19.3 s** | 9.6 s | 25.2 GiB |
588
 
589
+ **Denoising is 4.0Γ— faster than the 8-step reference** β€” two model calls instead of eight. The adapter's own cost per call did
590
  not show up in this measurement: the runs with it came in marginally faster than those without, which is measurement noise, not a
591
  speed-up. A rank-64 low-rank product is small beside the transformer it is added to, and it adds no measurable memory.
592
 
593
  Denoising is the part the step count changes. What a complete render costs on top of it β€” encoding the prompt, decoding
594
  the latent, writing the file β€” is the same whether you run two steps or eight, and it depends on your pipeline, so the
595
+ end-to-end figure on your machine will sit below 4Γ— and rise toward it as the render gets larger.
596
 
597
+ **By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 22 test prompts per
598
  size; sweep renders run one at a time, not the controlled measurement above, so a second or two either way between one sweep and
599
  the next is run-to-run variation β€” the adapter's shape, and so its cost, is the same at every checkpoint):
600
 
601
  | resolution | denoising (2 calls) | resolution | denoising (2 calls) |
602
  | --- | --- | --- | --- |
603
+ | 512Γ—512 | 6.5 s | 1024Γ—1024 | 19.5 s |
604
+ | 512Γ—768 / 768Γ—512 | 10.7 s / 11.3 s | 1280Γ—960 / 960Γ—1280 | 22.8 s / 22.7 s |
605
+ | 768Γ—768 | 12.7 s | 1280Γ—1280 | 32.0 s |
606
+ | 768Γ—1024 / 1024Γ—768 | 15.1 s / 15.2 s | 1440Γ—1280 / 1440Γ—1440 | 33.3 s / 38.3 s |
607
 
608
  ## LoRA strength
609
 
 
706
  sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
707
  1.4–1.8Γ— the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
708
  not hold, and the spectral match replaced it. The same sweep of the published checkpoint, its running-average weights, fixed
709
+ seed, 22 prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
710
+ first distribution-matching checkpoint, on the original 15 prompts):
711
 
712
  | bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
713
  | --------- | --------------------------- | --------------- | -------------- | ----------------------- |
714
+ | 512x512 | 1.00 (1.16) | 1.01 (1.23) | 1.09 (1.23) | 0.43 (0.44) |
715
+ | 512x768 | 0.97 (1.23) | 0.99 (1.37) | 1.02 (1.33) | 0.41 (0.41) |
716
+ | 768x512 | 1.01 (1.22) | 1.07 (1.37) | 1.02 (1.26) | 0.42 (0.47) |
717
+ | 768x768 | 1.05 (1.43) | 1.07 (1.46) | 1.09 (1.50) | 0.40 (0.43) |
718
+ | 768x1024 | 1.00 (1.36) | 0.98 (1.38) | 1.02 (1.42) | 0.39 (0.43) |
719
+ | 1024x768 | 0.98 (1.35) | 0.98 (1.40) | 0.98 (1.38) | 0.40 (0.46) |
720
+ | 1024x1024 | 1.06 (1.47) | 1.10 (1.55) | 1.13 (1.56) | 0.42 (0.44) |
721
+ | 1280x960 | 0.97 (1.50) | 0.99 (1.59) | 0.98 (1.56) | 0.43 (0.42) |
722
+ | 960x1280 | 1.09 (1.60) | 1.02 (1.63) | 1.13 (1.70) | 0.42 (0.43) |
723
+ | 1280x1280 | 1.02 (1.65) | 1.02 (1.68) | 1.03 (1.69) | 0.43 (0.44) |
724
+ | 1440x1280 | 0.99 (1.59) | 1.00 (1.66) | 1.09 (1.66) | 0.40 (0.42) |
725
+ | 1440x1440 | 1.03 (1.81) | 1.05 (1.79) | 1.09 (1.89) | 0.40 (0.41) |
726
 
727
  ### The spectral match
728
 
 
765
 
766
  ### Critics in turn
767
 
768
+ One critic holds one idea of what is wrong. The artefact critic is joined by seven more on the same frozen mid-network
769
  features and under the same rules β€” lightly re-noised inputs, a push that is filtered and capped β€” each aimed at a
770
  different fault:
771
 
 
773
  what fine texture looks like in a photograph as well as in the teacher's rendering of one. It judges photographic prompts
774
  only, so illustration, anime and 3D renders are not pulled toward photographic grain, and its push is filtered to periods
775
  finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher does not.
776
+ - **A large-face critic, on photographs.** It reads only large faces, 192 pixels and up, and pushes on every step. A second
777
+ face critic, reading faces of every size, trained the previous checkpoint and was taken back out (item 18 under
778
+ [How I got here](#how-i-got-here)).
779
  - **A structure critic.** It judges the first call β€” the layout, before any detail β€” against the teacher's own intermediate
780
  state from the same noise, at the high noise levels where layout is decided and at periods of 32 pixels and coarser only,
781
  so a first call that blends two layouts is caught where the blend happens.
 
786
  - **A rollout critic.** Its real examples are the teacher's own finish from the student's second-call starting point, so the
787
  second call is judged against what the teacher would have made from the same start.
788
 
789
+ All eight on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
790
  others train so none goes stale before its turn comes back. Each new head started from the artefact critic's weights and
791
  trained on its own before it was allowed to push.
792
 
 
801
  pixels may not fall below the teacher's plus the margin real photographs carry over it at those scales, judged tile by tile
802
  on the window's textured tiles. That margin is measured once from a pool of real photographs and clamped, and the term
803
  only ever pushes upward to that floor, never past it β€” so it lifts detail that is missing without adding grain that is not.
804
+ - **A ceiling.** Tile by tile, detail at 3–16 pixels may not climb past 1.3Γ— the teacher's β€” the counterpart of the photo
805
+ floor, and what keeps grain from building up at large sizes.
 
 
 
806
  - **A direction-aware term.** The spectral comparison is also made orientation by orientation, so the student's fine detail
807
  runs in the same directions as the teacher's.
808
  - **Windows on features.** On photographs, half of the decoded windows are centred on an eye, the nose or the lips of a face,
 
827
  The published weights are a running average of training (decay 0.999), which smooths out the noise of single steps.
828
  Averaging has one failure: while the training weights move between two layouts of the same prompt, their average draws
829
  both β€” a doubled subject. A guard watches a fixed set of layout probes every ten steps; a move away from the trend that
830
+ comes back within six rounds is kept out of the average, and a lasting move is taken in whole β€” the guard never resets
831
+ the average. Before any stretch of training goes in, the average it would make renders a fixed probe at the large sizes, and a
832
+ stretch that would put a second head on the subject is held out, even when the training weights themselves drew a single one
833
+ throughout. The average itself
834
+ has been rebuilt twice, by hand: once after a layout change it had blended, and once from the training since the change below,
835
+ after the guard had held it still through a long change of pose.
836
+
837
+ ### Layout at the large sizes
838
+
839
+ The doubled subjects that two steps can draw β€” a second head, two bodies joined, two poses blended β€” formed at the sizes
840
+ trained most, 1280 pixels and up, and the distribution term is the one that pushes the student toward whatever the teacher would
841
+ plausibly draw, which at a two-step jump can be more than one layout at once. So at every size with a 1280- or 1440-pixel side the
842
+ distribution term pushes at half strength, while the anchor to the teacher's recorded trajectory keeps its full weight there: at
843
+ those sizes the teacher's own layout for the same noise carries twice the share it did. The anchor also counts the coarse layout of
844
+ the second call β€” everything 64 pixels and up β€” twice, since that is the call in which a second head appears.
845
 
846
  ## What the LoRA touches
847
 
 
881
  which is why every cut of this run is rendered across the buckets and why the per-resolution table under
882
  [Method](#method) above exists.
883
 
884
+ [`assets/resolution_sweeps/`](assets/resolution_sweeps) holds the evidence β€” same 22 prompts (the original 15, plus seven
885
+ scenes of people, animals and action added on 29 Sep 2026 to check structure), same seed, one folder per resolution, one
886
+ image per prompt, so any image can be compared 1:1 with its twin in the next tree:
887
 
888
  - [`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step) β€” the **official Krea 2 Turbo 8-step reference renders**:
889
  the stock model, no LoRA, at its native settings (8 steps, guidance 0.0). The teacher this LoRA is distilled from and
 
902
  β”‚ β”œβ”€β”€ 512x512/ one folder per resolution
903
  β”‚ β”‚ β”œβ”€β”€ portrait.jpg
904
  β”‚ β”‚ β”œβ”€β”€ kingfisher.jpg
905
+ β”‚ β”‚ β”œβ”€β”€ … 19 more, one per test prompt
906
+ β”‚ β”‚ └── flamingos.jpg
907
  β”‚ β”œβ”€β”€ 768x1024/
908
  β”‚ β”œβ”€β”€ 1024x1024/
909
  β”‚ β”œβ”€β”€ 1280x1280/
 
930
  host memory above 0.3 megapixels; the student, the fake adapter, the spectral and detail terms, the critic and the teacher's finishing pass each
931
  build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
932
  does not fit fails loudly. A full step with every term live and every critic pushing reserves about 22.4 GB at 1440Γ—1440,
933
+ of 24. The price of the objective is throughput: **about 58 training samples an hour** with every critic and detail term in
934
+ the recipe, against the [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β€” about eight times the cost per sample, and so far a
935
+ small fraction of the samples.
936
 
937
  **Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
938
  The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
939
  disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
940
  of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
941
+ alongside to imitate the student β€” and that second adapter takes three optimiser steps of its own per student step.
942
  The spectral and detail terms decode part of the image out of the latent to compare its texture with the teacher's,
943
  each critic reads half the network twice more, and every second step the teacher finishes the image from the student's
944
  first call.
 
954
  ## How it is judged
955
 
956
  At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
957
+ 22 fixed prompts across four resolutions (512Γ—512, 1280Γ—1280, 1440Γ—1440, 1440Γ—1280), after a layout check at three large
958
  sizes that has to pass first; milestone checkpoints get the same render at all 12 buckets, which is where the
959
  per-resolution table above comes from. Every image is measured against the
960
  teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
961
  and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
962
+ seen through their windows), straight-line artefacts, a structure judge that compares the seven scenes of people, animals
963
+ and action with the teacher's for missing, extra or merged parts (every flag checked by eye at 1:1), a graded judge, a pairwise
964
+ preference against the teacher, and a blind rubric that
965
  scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
966
  identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
967
  checkpoint. Latent distances β€” the held-out chord gap and the two-step rollout
 
1277
 
1278
  ---
1279
 
1280
+ ### A group photo of five friends standing side by side on a sunny beach, arms around each other's shoulders, all smiling at the camera; five clearly different people, each with a unique face, no twins or lookalikes: a tall bearded man in his forties, a young woman with curly red hair and freckles, an older East Asian man with grey hair and glasses, a Black woman with short natural hair, and a teenage boy with messy blond hair, photograph, sharp detail
1281
+
1282
+ **3-way comparison** β€” one image, all three renders side by side
1283
+
1284
+ [![group5 comparison](assets/thumbs/group5_compare_turbo.jpg)](assets/group5_compare_turbo.jpg)
1285
+
1286
+ **Against the teacher** β€” this LoRA at 2 steps beside the 8-step render
1287
+
1288
+ [![group5 vs teacher](assets/thumbs/group5_compare_teacher.jpg)](assets/group5_compare_teacher.jpg)
1289
+
1290
+ **Individual frames** β€” click any panel to open that render full size
1291
+
1292
+ | Turbo β€” 8 steps | Turbo β€” 2 steps, no LoRA | **Turbo β€” 2 steps + this LoRA** |
1293
+ | --------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
1294
+ | [![8 steps](assets/thumbs/group5_turbo_8step.jpg)](assets/group5_turbo_8step.jpg) | [![2 steps](assets/thumbs/group5_turbo_2step.jpg)](assets/group5_turbo_2step.jpg) | [![2 steps + LoRA](assets/thumbs/group5_turbo_2step_lora.jpg)](assets/group5_turbo_2step_lora.jpg) |
1295
+ | as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
1296
+
1297
+ ---
1298
+
1299
+ ### A family of four having dinner at a round wooden table, seen from directly above, plates, glasses and bowls of food, warm evening light, photograph
1300
+
1301
+ **3-way comparison** β€” one image, all three renders side by side
1302
+
1303
+ [![family4 comparison](assets/thumbs/family4_compare_turbo.jpg)](assets/family4_compare_turbo.jpg)
1304
+
1305
+ **Against the teacher** β€” this LoRA at 2 steps beside the 8-step render
1306
+
1307
+ [![family4 vs teacher](assets/thumbs/family4_compare_teacher.jpg)](assets/family4_compare_teacher.jpg)
1308
+
1309
+ **Individual frames** β€” click any panel to open that render full size
1310
+
1311
+ | Turbo β€” 8 steps | Turbo β€” 2 steps, no LoRA | **Turbo β€” 2 steps + this LoRA** |
1312
+ | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
1313
+ | [![8 steps](assets/thumbs/family4_turbo_8step.jpg)](assets/family4_turbo_8step.jpg) | [![2 steps](assets/thumbs/family4_turbo_2step.jpg)](assets/family4_turbo_2step.jpg) | [![2 steps + LoRA](assets/thumbs/family4_turbo_2step_lora.jpg)](assets/family4_turbo_2step_lora.jpg) |
1314
+ | as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
1315
+
1316
+ ---
1317
+
1318
+ ### Three kittens sitting side by side in a wicker basket with a tall arched handle over them, all looking at the camera, soft natural light, photograph, sharp detail
1319
+
1320
+ **3-way comparison** β€” one image, all three renders side by side
1321
+
1322
+ [![kittens3 comparison](assets/thumbs/kittens3_compare_turbo.jpg)](assets/kittens3_compare_turbo.jpg)
1323
+
1324
+ **Against the teacher** β€” this LoRA at 2 steps beside the 8-step render
1325
+
1326
+ [![kittens3 vs teacher](assets/thumbs/kittens3_compare_teacher.jpg)](assets/kittens3_compare_teacher.jpg)
1327
+
1328
+ **Individual frames** β€” click any panel to open that render full size
1329
+
1330
+ | Turbo β€” 8 steps | Turbo β€” 2 steps, no LoRA | **Turbo β€” 2 steps + this LoRA** |
1331
+ | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
1332
+ | [![8 steps](assets/thumbs/kittens3_turbo_8step.jpg)](assets/kittens3_turbo_8step.jpg) | [![2 steps](assets/thumbs/kittens3_turbo_2step.jpg)](assets/kittens3_turbo_2step.jpg) | [![2 steps + LoRA](assets/thumbs/kittens3_turbo_2step_lora.jpg)](assets/kittens3_turbo_2step_lora.jpg) |
1333
+ | as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
1334
+
1335
+ ---
1336
+
1337
+ ### Two golden retrievers playing tug of war with a red rope on a green lawn, action shot, photograph, sharp detail
1338
+
1339
+ **3-way comparison** β€” one image, all three renders side by side
1340
+
1341
+ [![dogs2 comparison](assets/thumbs/dogs2_compare_turbo.jpg)](assets/dogs2_compare_turbo.jpg)
1342
+
1343
+ **Against the teacher** β€” this LoRA at 2 steps beside the 8-step render
1344
+
1345
+ [![dogs2 vs teacher](assets/thumbs/dogs2_compare_teacher.jpg)](assets/dogs2_compare_teacher.jpg)
1346
+
1347
+ **Individual frames** β€” click any panel to open that render full size
1348
+
1349
+ | Turbo β€” 8 steps | Turbo β€” 2 steps, no LoRA | **Turbo β€” 2 steps + this LoRA** |
1350
+ | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
1351
+ | [![8 steps](assets/thumbs/dogs2_turbo_8step.jpg)](assets/dogs2_turbo_8step.jpg) | [![2 steps](assets/thumbs/dogs2_turbo_2step.jpg)](assets/dogs2_turbo_2step.jpg) | [![2 steps + LoRA](assets/thumbs/dogs2_turbo_2step_lora.jpg)](assets/dogs2_turbo_2step_lora.jpg) |
1352
+ | as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
1353
+
1354
+ ---
1355
+
1356
+ ### A horse and rider jumping over a wooden fence at a show jumping event, side view, sports photography, sharp detail
1357
+
1358
+ **3-way comparison** β€” one image, all three renders side by side
1359
+
1360
+ [![horserider comparison](assets/thumbs/horserider_compare_turbo.jpg)](assets/horserider_compare_turbo.jpg)
1361
+
1362
+ **Against the teacher** β€” this LoRA at 2 steps beside the 8-step render
1363
+
1364
+ [![horserider vs teacher](assets/thumbs/horserider_compare_teacher.jpg)](assets/horserider_compare_teacher.jpg)
1365
+
1366
+ **Individual frames** β€” click any panel to open that render full size
1367
+
1368
+ | Turbo β€” 8 steps | Turbo β€” 2 steps, no LoRA | **Turbo β€” 2 steps + this LoRA** |
1369
+ | ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
1370
+ | [![8 steps](assets/thumbs/horserider_turbo_8step.jpg)](assets/horserider_turbo_8step.jpg) | [![2 steps](assets/thumbs/horserider_turbo_2step.jpg)](assets/horserider_turbo_2step.jpg) | [![2 steps + LoRA](assets/thumbs/horserider_turbo_2step_lora.jpg)](assets/horserider_turbo_2step_lora.jpg) |
1371
+ | as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
1372
+
1373
+ ---
1374
+
1375
+ ### Close-up of a pianist's two hands playing the keys of a grand piano, dramatic side light, photograph, sharp detail
1376
+
1377
+ **3-way comparison** β€” one image, all three renders side by side
1378
+
1379
+ [![pianist comparison](assets/thumbs/pianist_compare_turbo.jpg)](assets/pianist_compare_turbo.jpg)
1380
+
1381
+ **Against the teacher** β€” this LoRA at 2 steps beside the 8-step render
1382
+
1383
+ [![pianist vs teacher](assets/thumbs/pianist_compare_teacher.jpg)](assets/pianist_compare_teacher.jpg)
1384
+
1385
+ **Individual frames** β€” click any panel to open that render full size
1386
+
1387
+ | Turbo β€” 8 steps | Turbo β€” 2 steps, no LoRA | **Turbo β€” 2 steps + this LoRA** |
1388
+ | ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
1389
+ | [![8 steps](assets/thumbs/pianist_turbo_8step.jpg)](assets/pianist_turbo_8step.jpg) | [![2 steps](assets/thumbs/pianist_turbo_2step.jpg)](assets/pianist_turbo_2step.jpg) | [![2 steps + LoRA](assets/thumbs/pianist_turbo_2step_lora.jpg)](assets/pianist_turbo_2step_lora.jpg) |
1390
+ | as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
1391
+
1392
+ ---
1393
+
1394
+ ### A flock of pink flamingos standing in shallow turquoise water with their reflections, wildlife photography, sharp detail
1395
+
1396
+ **3-way comparison** β€” one image, all three renders side by side
1397
+
1398
+ [![flamingos comparison](assets/thumbs/flamingos_compare_turbo.jpg)](assets/flamingos_compare_turbo.jpg)
1399
+
1400
+ **Against the teacher** β€” this LoRA at 2 steps beside the 8-step render
1401
+
1402
+ [![flamingos vs teacher](assets/thumbs/flamingos_compare_teacher.jpg)](assets/flamingos_compare_teacher.jpg)
1403
+
1404
+ **Individual frames** β€” click any panel to open that render full size
1405
+
1406
+ | Turbo β€” 8 steps | Turbo β€” 2 steps, no LoRA | **Turbo β€” 2 steps + this LoRA** |
1407
+ | --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
1408
+ | [![8 steps](assets/thumbs/flamingos_turbo_8step.jpg)](assets/flamingos_turbo_8step.jpg) | [![2 steps](assets/thumbs/flamingos_turbo_2step.jpg)](assets/flamingos_turbo_2step.jpg) | [![2 steps + LoRA](assets/thumbs/flamingos_turbo_2step_lora.jpg)](assets/flamingos_turbo_2step_lora.jpg) |
1409
+ | as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
1410
+
1411
+ ---
1412
+
1413
  ## Notes and limitations
1414
 
1415
  - 🎯 **Krea 2 Turbo only**, at **2 steps**, guidance **0.0** (cfg 1.0 in ComfyUI), **mu = 1.15** β€” the two training
 
1422
  resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Next, aimed
1423
  at what this checkpoint still gets wrong: small subjects in wide scenes; the finest edges and the texture of skin at the
1424
  largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two
1425
+ plausible poses of the same subject can meet; the fine detail of close-up faces β€” eyes and skin texture; and the tactile surface of
1426
+ stylised materials such as clay β€” each held to the rule that none of it may cost the structure, the colour and the clean large
1427
+ sizes this checkpoint gained. A better checkpoint
1428
  replaces this one when the sweeps and I visually agree, the same discipline as the
1429
  [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
1430
  one is the fast preview.
README.md CHANGED
@@ -17,28 +17,28 @@ pipeline_tag: text-to-image
17
 
18
  # Krea 2 Turbo β€” 2-Step Distillation LoRA
19
 
20
- **A quarter of the steps Β· 4.2Γ— faster denoising Β· fine detail at 0.95–1.08Γ— the teacher's across all 12 trained resolutions Β· colour at the teacher's level Β· 1 point missing of 240 on a blind prompt-adherence rubric Β· teacher preferred on 11 of 45 judged renders Β· 31,600 training samples on the 4-step project's recorded trajectories Β· 18 days on one RTX 3090 Β· still in training.**
21
 
22
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β€” Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
23
 
24
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
25
- - ⏱️ **4.2Γ— faster denoising** β€” 81.4 s β†’ 19.5 s at 1024Γ—1024; the adapter's own cost per call is within measurement noise
26
- - 🎯 **Fine detail at the teacher's level** β€” **0.95–1.08Γ—** the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: **0.39–0.57Γ—**); from 1 megapixel up, closer to the teacher than the 4-step adapter
27
  - πŸ“Š **Distribution matching, not imitation** β€” matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
28
- - πŸ—£οΈ **Prompt-conditioned throughout** β€” teacher and fake scores both read each prompt's conditioning; a blind rubric finds **1 point missing of 240** (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on **11 of 45** (4-step adapter: 6), mostly on style
29
  - πŸ“ **12 trained resolutions** β€” multi-aspect from 512Γ—512 up to 1440Γ—1440
30
  - πŸ”Œ **Drop-in, no exceptions** β€” plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
31
  - 🧬 **Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β€” rank 64 on the same 228 modules
32
  - 🎲 **13,750 recorded teacher trajectories** from the 4-step project, reused β€” not one new teacher run
33
- - πŸ”’ **31,600 training samples** in the 2-step stages, on top of the 4-step LoRA's 78,000
34
- - πŸ“… **18 days** from the first 2-step launch to this checkpoint, on a single RTX 3090 β€” training continues
35
  - πŸ” **More than forty recipe adjustments** across two methods β€” each kept only when the renders did not get worse
36
 
37
  > πŸ§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β€” a portrait, a single figure, an object up close β€” hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
38
  >
39
  > πŸ“ **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ—2048 β€” Krea's published maximum recommended resolution, beyond this adapter's largest trained size β€” come within easy reach. Past 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
40
 
41
- [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
42
 
43
  ---
44
 
@@ -111,13 +111,13 @@ Trained on Turbo, for Turbo β€” it loads on **Krea 2 Raw** because the architect
111
 
112
  | | denoise | per model call | GPU peak |
113
  | ----------------------------- | ---------- | -------------- | -------- |
114
- | Turbo 8 steps (quality bar) | 81.4 s | 10.2 s | 25.2 GiB |
115
- | Turbo 2 steps, no LoRA | 20.4 s | 10.2 s | 25.2 GiB |
116
- | **Turbo 2 steps + this LoRA** | **19.5 s** | 9.8 s | 25.2 GiB |
117
 
118
- **Denoising is 4.2Γ— faster than the 8-step bar** β€” two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.
119
 
120
- Prompt encoding and VAE decode don't change with step count, so end to end sits below 4.2Γ— and rises toward it as the render grows. Denoise times at every trained resolution (7.0 s at 512Γ—512 to 39.6 s at 1440Γ—1440) are in the [Detailed Model Card](DETAILED-README.md).
121
 
122
  ---
123
 
@@ -138,20 +138,20 @@ At two steps the dial scales the adapter's whole job β€” turning two coarse call
138
 
139
  ## Current Checkpoint
140
 
141
- **`chk00031600`** (25 Sep 2026) replaces `chk00017464` (14 Sep 2026). It is 14,136 training samples later, aimed at what the previous checkpoint still got wrong β€” grain and excess texture at the largest sizes, colour under the teacher's there, small faces, stylised prompts β€” through a colour band at the teacher's level, nine critics taking turns, detail held to the teacher region by region, more training at the large sizes and a guarded running average.
142
 
143
- | axis | `chk00017464` | `chk00031600` |
144
- | ------------------------------------------------------ | ------------- | --------------- |
145
- | fine texture vs the teacher, 1280Β² / 1440Β² | 1.09 / 1.17 | **0.95 / 1.08** |
146
- | 16-px grid band, 1280Β² / 1440Β² | 1.08 / 1.08 | **0.95 / 1.05** |
147
- | grain in flat areas, sweep median | 1.26Γ— | **1.08Γ—** |
148
- | saturation vs the teacher, sweep mean | 0.93Γ— | **0.98Γ—** |
149
- | small faces keeping their structure (the teacher: 64%) | 53% | **63%** |
150
- | judge prefers the teacher (of 45) | 11 | 11 |
151
- | blind adherence rubric, points missing of 240 | 1 | 1 |
152
- | distance to the teacher, sweep mean | **0.406** | 0.413 |
153
 
154
- The excess fine energy at large sizes is gone and colour is at the teacher's level at every size; small faces keep their structure nearly as often as the teacher's, and prompt-following held level. [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) always names the checkpoint in the weight files; the full comparison is in the [Detailed Model Card](DETAILED-README.md).
155
 
156
  ---
157
 
@@ -176,16 +176,18 @@ Every one is being worked on; none is hidden in the sweeps or the examples.
176
 
177
  The student makes two calls, at Οƒ = 1.0 and 0.7595 β€” the first and fifth points of the teacher's 8-step grid at ΞΌ = 1.15 β€” with stock Euler between them. Euler's first step lands **exactly** on the flow-matching interpolant at Οƒ = 0.7595, so the first call's output is a legitimate image prediction and is judged as one.
178
 
179
- - **Distribution term** β€” the frozen teacher and a _fake-score_ adapter (rank 32, trained online on the student's current output, 4 updates per student step, discarded at the end) each denoise a freshly noised copy of the student's image; where they disagree is the direction toward the teacher's work. Averaging is never rewarded, so the student commits
180
  - **Trajectory anchor** β€” regression on the recorded chords at half weight keeps the student on the teacher's two-step grid
181
 
182
  **On top, each capped relative to the distribution term:**
183
 
184
  - **Spectral match** β€” student and teacher images compared through radial power spectra, on the whole latent and a decoded 256-px window, two-sided β€” the term that reached the 16/8-px grid grain at large sizes
185
- - **Three critics taking turns** β€” artefact, photo (half real photographs) and face heads on the frozen base's mid-network features; one pushes per step, filtered to structure finer than 32 px (photo critic: 24 px)
186
- - **Four detail terms** β€” anchor counts fine-detail error twice; one-sided photo floor at 3–10 px; the teacher's finish of the student's first call as the second call's target; smoothness limit on the fake adapter
 
 
187
 
188
- Shipped adapter is the **running (EMA) average** of the weights, not the last live state.
189
 
190
  ---
191
 
@@ -219,7 +221,7 @@ Only the photo critic's head sees the photos; the student and the fake adapter r
219
  | 1024Γ—1024 | 960Γ—1280 | 1280Γ—960 |
220
  | 1280Γ—1280 | 1440Γ—1280 | 1440Γ—1440 |
221
 
222
- Same 12 buckets as the 4-step adapter, interleaved in proportion to their remaining samples.
223
 
224
  ---
225
 
@@ -231,7 +233,7 @@ Trained on a **single RTX 3090 (24 GB)**. Frozen base **weight-only int8**. Ever
231
  - Student, fake adapter, spectral and detail terms, critic and teacher finish each build and free their own graph β€” peaks never overlap
232
  - Hard memory ceiling below the driver's paging threshold, so a step that doesn't fit fails loudly
233
 
234
- A full step with every term live reserves ~22.4 GB at 1440Γ—1440. Throughput is **~107 samples/hour** against the 4-step recipe's 470 β€” roughly a dozen model runs per sample instead of two. That is the objective's cost, not teacher generation: the trajectories were recorded once and are read from disk.
235
 
236
  Released LoRA is **bf16**.
237
 
@@ -281,13 +283,13 @@ Every sheet below: base model (8 steps), base at 2 steps **without** LoRA, base
281
 
282
  ---
283
 
284
- ### (12 more examples, and a two-panel sheet against the teacher for every prompt, in the [full version](DETAILED-README.md) β€” see [assets/](assets/) for all 15 prompts)
285
 
286
  ---
287
 
288
  ## Resolution Sweeps
289
 
290
- [`assets/resolution_sweeps/`](assets/resolution_sweeps/) β€” **this LoRA at every trained resolution for all 15 test prompts** (same prompts, seed, 2 steps, strength 1.0). Nothing cherry-picked.
291
 
292
  [`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step/) β€” official Krea 2 Turbo 8-step reference renders for the same prompts/seeds/resolutions. [`_turbo-base-NO-LoRA-2step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-2step/) β€” stock Turbo at 2 steps, the floor.
293
 
@@ -315,7 +317,7 @@ Every published checkpoint and its resolution sweep under [`_archive/`](_archive
315
 
316
  Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
317
 
318
- Next, aimed at the [Known Issues](#known-issues): small subjects in wide scenes; the finest edges and skin texture at the largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two plausible poses of the same subject can meet; the tactile surface of stylised materials such as clay β€” none of it allowed to cost the colour and the clean large sizes this checkpoint gained.
319
 
320
  A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
321
 
 
17
 
18
  # Krea 2 Turbo β€” 2-Step Distillation LoRA
19
 
20
+ **A quarter of the steps Β· 4Γ— faster denoising Β· fine detail at 0.97–1.09Γ— the teacher's across all 12 trained resolutions Β· colour at the teacher's level Β· 2 points missing of 352 on a blind prompt-adherence rubric Β· teacher preferred on 12 of 66 judged renders Β· 41,320 training samples on the 4-step project's recorded trajectories Β· 25 days on one RTX 3090 Β· still in training.**
21
 
22
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β€” Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
23
 
24
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
25
+ - ⏱️ **4Γ— faster denoising** β€” 76.4 s β†’ 19.3 s at 1024Γ—1024; the adapter's own cost per call is within measurement noise
26
+ - 🎯 **Fine detail at the teacher's level** β€” **0.97–1.09Γ—** the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: **0.40–0.65Γ—**); from 1 megapixel up, closer to the teacher than the 4-step adapter
27
  - πŸ“Š **Distribution matching, not imitation** β€” matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
28
+ - πŸ—£οΈ **Prompt-conditioned throughout** β€” teacher and fake scores both read each prompt's conditioning; a blind rubric finds **2 points missing of 352** (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on **12 of 66** (4-step adapter: 6 of 45, on the original 15 prompts), mostly on style
29
  - πŸ“ **12 trained resolutions** β€” multi-aspect from 512Γ—512 up to 1440Γ—1440
30
  - πŸ”Œ **Drop-in, no exceptions** β€” plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
31
  - 🧬 **Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β€” rank 64 on the same 228 modules
32
  - 🎲 **13,750 recorded teacher trajectories** from the 4-step project, reused β€” not one new teacher run
33
+ - πŸ”’ **41,320 training samples** in the 2-step stages, on top of the 4-step LoRA's 78,000
34
+ - πŸ“… **25 days** from the first 2-step launch to this checkpoint, on a single RTX 3090 β€” training continues
35
  - πŸ” **More than forty recipe adjustments** across two methods β€” each kept only when the renders did not get worse
36
 
37
  > πŸ§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β€” a portrait, a single figure, an object up close β€” hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
38
  >
39
  > πŸ“ **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ—2048 β€” Krea's published maximum recommended resolution, beyond this adapter's largest trained size β€” come within easy reach. Past 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
40
 
41
+ [![The 22 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
42
 
43
  ---
44
 
 
111
 
112
  | | denoise | per model call | GPU peak |
113
  | ----------------------------- | ---------- | -------------- | -------- |
114
+ | Turbo 8 steps (quality bar) | 76.4 s | 9.5 s | 25.2 GiB |
115
+ | Turbo 2 steps, no LoRA | 19.4 s | 9.7 s | 25.2 GiB |
116
+ | **Turbo 2 steps + this LoRA** | **19.3 s** | 9.6 s | 25.2 GiB |
117
 
118
+ **Denoising is 4.0Γ— faster than the 8-step bar** β€” two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.
119
 
120
+ Prompt encoding and VAE decode don't change with step count, so end to end sits below 4Γ— and rises toward it as the render grows. Denoise times at every trained resolution (6.5 s at 512Γ—512 to 38.3 s at 1440Γ—1440) are in the [Detailed Model Card](DETAILED-README.md).
121
 
122
  ---
123
 
 
138
 
139
  ## Current Checkpoint
140
 
141
+ **`chk00041320`** (2 Oct 2026) replaces `chk00031600` (25 Sep 2026). It is 9,720 training samples later, aimed at what matters first in a picture β€” subjects drawn twice or two poses blended into one at the large sizes, and faces in busy scenes β€” through the teacher's own layout held at the large sizes, eight critics, and every update to the running average checked for doubled subjects. It is also the first checkpoint measured on the full 22-prompt sweep.
142
 
143
+ | axis | `chk00031600` | `chk00041320` |
144
+ | ------------------------------------------------------------------ | ------------- | ------------- |
145
+ | structure renders with a fault the teacher's does not have (of 21) | 1 | **0** |
146
+ | ghosting, blur-invariant index, sweep mean | 1.24Γ— | **1.21Γ—** |
147
+ | market faces at 1280Β², layout closeness to the teacher | 0.71 | **0.79** |
148
+ | grain in flat areas, sweep mean | 1.07Γ— | **1.03Γ—** |
149
+ | fine texture vs the teacher, 1440Β² | 1.07 | **1.03** |
150
+ | judge prefers the teacher (of 66) | 15 | **12** |
151
+ | blind adherence rubric, points missing of 352 | 2 | 2 |
152
+ | saturation vs the teacher, sweep mean | 0.98Γ— | **0.99Γ—** |
153
 
154
+ Fewer doubled subjects and less ghosting at the large sizes, whole faces in the busy market scene, and prompt-following a step closer to the teacher's, while colour and fine detail stay at the teacher's level. [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) always names the checkpoint in the weight files; the full comparison is in the [Detailed Model Card](DETAILED-README.md#chk00041320-vs-chk00031600).
155
 
156
  ---
157
 
 
176
 
177
  The student makes two calls, at Οƒ = 1.0 and 0.7595 β€” the first and fifth points of the teacher's 8-step grid at ΞΌ = 1.15 β€” with stock Euler between them. Euler's first step lands **exactly** on the flow-matching interpolant at Οƒ = 0.7595, so the first call's output is a legitimate image prediction and is judged as one.
178
 
179
+ - **Distribution term** β€” the frozen teacher and a _fake-score_ adapter (rank 32, trained online on the student's current output, 3 updates per student step, discarded at the end) each denoise a freshly noised copy of the student's image; where they disagree is the direction toward the teacher's work. Averaging is never rewarded, so the student commits
180
  - **Trajectory anchor** β€” regression on the recorded chords at half weight keeps the student on the teacher's two-step grid
181
 
182
  **On top, each capped relative to the distribution term:**
183
 
184
  - **Spectral match** β€” student and teacher images compared through radial power spectra, on the whole latent and a decoded 256-px window, two-sided β€” the term that reached the 16/8-px grid grain at large sizes
185
+ - **Eight critics taking turns** β€” artefact, photo (half real photographs), large faces, the first call's layout against the teacher's own state, content, prompt, text, and the teacher's own finish, all heads on the frozen base's mid-network features; one pushes per step, each filtered and capped
186
+ - **Detail held to the teacher region by region** β€” the anchor counts fine-detail error twice (3.5Γ— where the teacher is most detailed); a one-sided photo floor at 3–10 px and a per-tile ceiling at 1.3Γ—; a direction-aware term; windows on eyes, nose and lips; the teacher's finish of the student's first call as the second call's target; a smoothness limit on the fake adapter
187
+ - **Colour band** β€” saturation held between the teacher's own level and 8% above it
188
+ - **Layout at the large sizes** β€” at every size with a 1280- or 1440-px side the distribution term pushes at half strength while the anchor keeps full weight, and the anchor counts the second call's coarse layout (64 px and up) twice
189
 
190
+ Shipped adapter is a **guarded running (EMA) average** of the weights, not the last live state: short excursions are kept out, and every update is checked for doubled subjects before it goes in.
191
 
192
  ---
193
 
 
221
  | 1024Γ—1024 | 960Γ—1280 | 1280Γ—960 |
222
  | 1280Γ—1280 | 1440Γ—1280 | 1440Γ—1440 |
223
 
224
+ Same 12 buckets as the 4-step adapter, with the larger sizes and stylised prompts drawn more often: 1440Γ—1440 about 6% of the samples, 1024Γ—1024 and 768Γ—1024 12% each, stylised prompts 12%.
225
 
226
  ---
227
 
 
233
  - Student, fake adapter, spectral and detail terms, critic and teacher finish each build and free their own graph β€” peaks never overlap
234
  - Hard memory ceiling below the driver's paging threshold, so a step that doesn't fit fails loudly
235
 
236
+ A full step with every term live reserves ~22.4 GB at 1440Γ—1440. Throughput is **~58 samples/hour** against the 4-step recipe's 470 β€” roughly a dozen model runs per sample instead of two. That is the objective's cost, not teacher generation: the trajectories were recorded once and are read from disk.
237
 
238
  Released LoRA is **bf16**.
239
 
 
283
 
284
  ---
285
 
286
+ ### (19 more examples, and a two-panel sheet against the teacher for every prompt, in the [full version](DETAILED-README.md) β€” see [assets/](assets/) for all 22 prompts)
287
 
288
  ---
289
 
290
  ## Resolution Sweeps
291
 
292
+ [`assets/resolution_sweeps/`](assets/resolution_sweeps/) β€” **this LoRA at every trained resolution for all 22 test prompts** (same prompts, seed, 2 steps, strength 1.0). Nothing cherry-picked.
293
 
294
  [`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step/) β€” official Krea 2 Turbo 8-step reference renders for the same prompts/seeds/resolutions. [`_turbo-base-NO-LoRA-2step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-2step/) β€” stock Turbo at 2 steps, the floor.
295
 
 
317
 
318
  Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
319
 
320
+ Next, aimed at the [Known Issues](#known-issues): small subjects in wide scenes; the finest edges and skin texture at the largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two plausible poses of the same subject can meet; the fine detail of close-up faces β€” eyes and skin texture; the tactile surface of stylised materials such as clay β€” none of it allowed to cost the structure, the colour and the clean large sizes this checkpoint gained.
321
 
322
  A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
323
 
krea2_turbo_2step_rank_64_lora.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:1df55a05e4ed3cd367e3cc06c84279d9dd2430394119b81038b9ad30c3b17952
3
- size 438161624
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c827d21150f96245fe421f75a51a374ac494006d1c7dac7d3166b57a486f5d5c
3
+ size 438161704
krea2_turbo_2step_rank_64_lora_checkpoint_info.md CHANGED
@@ -1,13 +1,13 @@
1
  # Which checkpoint is this?
2
 
3
  `krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
4
- **chk00031600** β€” the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00031600.safetensors` (and its
5
  `_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
6
  keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
7
 
8
  | file | SHA-256 | size |
9
  | --- | --- | --- |
10
- | `krea2_turbo_2step_rank_64_lora.safetensors` | `1df55a05e4ed3cd367e3cc06c84279d9dd2430394119b81038b9ad30c3b17952` | 418M |
11
- | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `e017280bc21c5279156adb0daf6654238896ca1ee0c12c8459b66b882edcfca4` | 418M |
12
 
13
- Updated: 25 Sep 2026
 
1
  # Which checkpoint is this?
2
 
3
  `krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
4
+ **chk00041320** β€” the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00041320.safetensors` (and its
5
  `_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
6
  keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
7
 
8
  | file | SHA-256 | size |
9
  | --- | --- | --- |
10
+ | `krea2_turbo_2step_rank_64_lora.safetensors` | `c827d21150f96245fe421f75a51a374ac494006d1c7dac7d3166b57a486f5d5c` | 418M |
11
+ | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `69c9d4e55530f26764ff0cfed39e230ba20c9d1b8bcf8c581beef26eb9dd9726` | 418M |
12
 
13
+ Updated: 2 Oct 2026
krea2_turbo_2step_rank_64_lora_comfyui.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e017280bc21c5279156adb0daf6654238896ca1ee0c12c8459b66b882edcfca4
3
- size 438142160
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:69c9d4e55530f26764ff0cfed39e230ba20c9d1b8bcf8c581beef26eb9dd9726
3
+ size 438142240