lvladikov commited on
Commit
a2bb052
Β·
verified Β·
1 Parent(s): 56b23db

Checkpoint 31600 Release

Browse files
DETAILED-README.md CHANGED
@@ -33,9 +33,6 @@ pipeline_tag: text-to-image
33
  > within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past
34
  > 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects β€” a property of the base model, with or without this
35
  > adapter.
36
- >
37
- > πŸ”€ **Also compatible with Krea 2 Raw.** With some prompts it works very well on Krea 2 Raw too, at 7 steps+ β€” see
38
- > [Using it on Raw](#using-it-on-raw) and the [dedicated experiment](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md).
39
 
40
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
41
  down to 2** β€” Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at
@@ -67,12 +64,12 @@ recommendation for quality renders.
67
  only and never ships.
68
  - 🎲 **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
69
  re-run.
70
- - πŸ”’ **17,464 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β€” all of them
71
  drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
72
  training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
73
  4-step project, read again at the two sigmas this schedule uses.
74
- - πŸ“… **7 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β€” and the project continues.
75
- - πŸ” **26 recipe adjustments** across two methods so far β€” seven of trajectory distillation before the switch, nineteen of distribution matching since β€” each kept only when the renders did not get worse.
76
  - πŸ–₯️ **One RTX 3090**, and a recipe shaped by its 24 GB.
77
 
78
  [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
@@ -88,7 +85,6 @@ comparisons with the 8-step teacher are in [Examples](#examples)._
88
  | `krea2_turbo_2step_rank_64_lora.safetensors` | the LoRA in diffusers key format β€” see [Inference with diffusers](#inference-with-diffusers); also for MLX or anything that reads safetensors |
89
  | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | the same weights under ComfyUI's key names β€” see [ComfyUI](#comfyui) |
90
  | `krea2_turbo_2step_lora_t2i.json` | a ready ComfyUI workflow, stock nodes only |
91
- | `krea2_raw_7step_lora_experiment_t2i.json` | the Krea 2 Raw 7-step experiment's ComfyUI workflow ([Using it on Raw](#using-it-on-raw)) |
92
  | [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **the quick place to check which checkpoint the two weight files are based on.** The pair above keeps its names and is updated in place as better checkpoints ship; this file always says what they are today. Every published checkpoint also sits in [`_archive/checkpoints/`](_archive/checkpoints) under its number |
93
  | `LICENSE.pdf` | the Krea 2 Community License Agreement, which covers this adapter β€” see [License](#license) |
94
  | `NOTICE.txt` | the attribution notice the license requires of a derivative |
@@ -103,13 +99,13 @@ place to check which checkpoint the current files are based on.
103
 
104
  | | |
105
  | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
106
- | lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β†’ 2-step trajectory distillation β†’ distribution matching β†’ a spectral match against the teacher's own images β†’ an artefact critic and detail terms β†’ three critics taking turns |
107
  | this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
108
- | what it gives | usable two-step renders at every trained resolution: fine detail at or just above the teacher's β€” from 1 megapixel up, closer to the teacher than the 4-step adapter β€” with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 1 point missing out of 240). A judge asked which render follows the prompt better still prefers the 8-step teacher on 11 of 45, against 6 for the 4-step adapter, mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
109
 
110
  ### Known issues
111
 
112
- The usual costs of two steps, in this order of how often they show. **Small subjects are the weak spot, people and objects alike**: a portrait-sized face or an object seen up close holds up, while small or distant subjects β€” faces in a crowd, a figure in a wide scene, the machines at the back of a room β€” can come out ghosted, smeared or misshapen, since at that size a whole subject is only a few of the blocks the model works in. Fine structure can come out soft or a few pixels out of register β€” feathers, hair strands, signage, the surface of a distant object β€” most at 1280Γ—1280 and above, and a faint doubled contour can show on limbs. On stylised prompts, **how the prompt says the image should look is followed less faithfully than what should be in it**: crisp anime linework, energetic brush strokes, the fingerprints in clay or a matte-painting finish come out closer to a generic rendering than the teacher's. On busy action or crowd scenes the composition can repeat itself β€” an extra hand or held object, a figure duplicated in a crowd β€” where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. Skin reads slightly smoother and less saturated than the teacher's, and colour overall runs a little under the teacher's at the largest sizes; freckles tend to gather into clusters rather than separate dots. A fine grain remains on the most textured subjects at the largest sizes, lighter than in the previous checkpoint. Every one of these is being worked on; none is hidden in the sweeps or the examples.
113
 
114
  ## How I got here
115
 
@@ -166,11 +162,102 @@ The recipe adjustments so far, each made on the measurement of the one before:
166
  11. **three critics taking turns** β€” the artefact critic joined by one whose real examples are half real photographs and one
167
  weighted toward faces, one of them pushing on each step while the others keep training in between, because the three did
168
  not fit in memory side by side
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
169
 
170
  Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
171
  a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
172
  not place exactly where the teacher had it.
173
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
174
  ## chk00017464 vs chk00013663
175
 
176
  `chk00013663` (published 12 Sep 2026) was the first public checkpoint: distribution matching with the spectral match, and nothing
@@ -238,38 +325,38 @@ teacher):
238
 
239
  | bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
240
  | --- | --- | --- | --- | --- | --- |
241
- | 512Γ—512 | 1.02 | 1.02 | 1.05 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
242
- | 768Γ—1024 | 1.02 | 0.98 | 1.07 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
243
- | 1024Γ—1024 | 1.06 | 1.07 | 1.12 | 1.11 Β· 1.17 Β· 1.16 | 0.39 |
244
- | 1280Γ—1280 | 1.09 | 1.08 | 1.13 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
245
- | 1440Γ—1440 | 1.17 | 1.08 | 1.21 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
246
 
247
  Two steps without the adapter carry 0.57Γ— the teacher's fine detail at 512Γ—512 and **less than half** (0.39–0.42Γ—) at the four larger sizes. With it, the detail
248
- sits at or just above the teacher's everywhere β€” and from 1 megapixel up it is closer to the teacher than the 4-step
249
- adapter, which carries more excess fine energy there.
250
 
251
  **Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
252
  in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
253
 
254
  | bucket | wins | ties | losses |
255
  | --- | --- | --- | --- |
256
- | 512Γ—512 | 1 | 10 | 4 |
257
- | 1280Γ—1280 | 1 | 10 | 4 |
258
- | 1440Γ—1440 | 0 | 12 | 3 |
259
 
260
  Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
261
  stylised prompts and on how a prompt says the picture should look β€” crisp linework, energetic brush strokes, the texture of
262
  clay β€” more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
263
  attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
264
- training prompt bank and never rendered before, the judge returned 0 wins, 10 ties, 5 losses; on another 15 fresh prompts, 1 win,
265
- 13 ties, 1 loss.
266
 
267
- **Checked for the damage this kind of training can do.** Saturation sits at 0.99Γ— the teacher's at 768Γ—1024 and 0.89Γ—
268
- at the larger sizes; edge detail 0.93–0.96Γ—; skin texture inside detected faces 1.07Γ— at 768Γ—1024 and 0.85Γ— at
269
- 1440Γ—1440, with skin saturation 0.92Γ— and 0.77–0.85Γ— at the larger sizes. The honest reading of those last two: **skin
270
- is the softest and least saturated part of this adapter's output at large sizes, and colour overall runs a little under
271
- the teacher's there.** Fine detail in flat regions β€” skies, walls, out-of-focus backgrounds β€” runs 1.53–1.85Γ— the
272
- teacher's, which is where two steps put grain that eight steps do not.
273
 
274
  **Distance to the teacher**, as a plain pixel measure, is 0.39–0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
275
  0.30–0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
@@ -292,33 +379,7 @@ of the 4-step grid, so the model is evaluated at two points it already knows.
292
 
293
  **This LoRA is trained on Krea 2 Turbo, against Turbo as its own teacher, and for Turbo.** Every layer it targets also
294
  exists in Krea 2 Raw, so it will load there without complaint β€” but that is a side effect of the shared architecture, not
295
- a supported mode.
296
-
297
- It still does something useful there. On Turbo the adapter runs a quarter of the teacher's steps β€” 2 of its 8 β€” so on
298
- Raw the same quarter of its usual 28 steps, **7**, is the natural place to start, and for most prompts it is a good one.
299
- For a quick preview, some prompts hold together at even 4–5 steps β€” not as good as 7 steps or more, but not broken
300
- either. Keep the guidance **light**: `guidance_scale=1.0` in diffusers, which is **cfg 2.0 in ComfyUI**. Heavier
301
- guidance, such as 4.5, crushes most images into near-black frames at 7 steps, and with no guidance at all the pictures
302
- come out flat.
303
-
304
- [![Krea 2 Raw + this LoRA at 7 steps β€” portrait](assets/resolution_sweeps/raw-LoRA-7steps-experiment/thumbs/portrait.jpg)](assets/resolution_sweeps/raw-LoRA-7steps-experiment/1024x768/portrait.jpg)
305
- [![Krea 2 Raw + this LoRA at 7 steps β€” pizza](assets/resolution_sweeps/raw-LoRA-7steps-experiment/thumbs/pizza.jpg)](assets/resolution_sweeps/raw-LoRA-7steps-experiment/1024x768/pizza.jpg)
306
-
307
- _Krea 2 Raw + this LoRA, 7 steps, light guidance, empty negative prompt, seed 4242, 1024Γ—768. Click for full size._
308
-
309
- Results are **mixed and subject-dependent**. Some prompts come through as finished pictures; others do not β€” an
310
- underexposed night street, a cityscape with less detail than Turbo gives at 2 steps β€” and for those a few more steps may
311
- help. The full experiment, with all 15 test prompts at 1024Γ—768, what worked and what did not, and a ComfyUI workflow,
312
- is in its [dedicated README](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md), in
313
- [`assets/resolution_sweeps/raw-LoRA-7steps-experiment/`](assets/resolution_sweeps/raw-LoRA-7steps-experiment). The same
314
- prompts on stock Raw at the same 7 steps and light guidance, without the LoRA, are in
315
- [`_raw-base-NO-LoRA-7step-cfg1/`](assets/resolution_sweeps/raw-LoRA-7steps-experiment/_raw-base-NO-LoRA-7step-cfg1).
316
-
317
- [![ComfyUI workflow β€” Krea 2 Raw + this LoRA, 7 steps](assets/resolution_sweeps/raw-LoRA-7steps-experiment/thumbs/workflow_preview_raw.jpg)](assets/resolution_sweeps/raw-LoRA-7steps-experiment/workflow_preview_raw.jpg)
318
-
319
- _The experiment's ComfyUI workflow,
320
- [`krea2_raw_7step_lora_experiment_t2i.json`](krea2_raw_7step_lora_experiment_t2i.json): Krea 2 Raw + this LoRA, 7 steps,
321
- cfg 2.0. Click for full size._
322
 
323
  ## Inference with diffusers
324
 
@@ -449,14 +510,15 @@ the latent, writing the file β€” is the same whether you run two steps or eight,
449
  end-to-end figure on your machine will sit below 4.2Γ— and rise toward it as the render gets larger.
450
 
451
  **By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
452
- size; sweep renders run one at a time, not the controlled measurement above):
 
453
 
454
  | resolution | denoising (2 calls) | resolution | denoising (2 calls) |
455
  | --- | --- | --- | --- |
456
- | 512Γ—512 | 6.4 s | 1024Γ—1024 | 20.4 s |
457
- | 512Γ—768 / 768Γ—512 | 8.9 s / 9.0 s | 1280Γ—960 / 960Γ—1280 | 23.2 s / 24.0 s |
458
- | 768Γ—768 | 11.9 s | 1280Γ—1280 | 32.2 s |
459
- | 768Γ—1024 / 1024Γ—768 | 14.9 s / 15.1 s | 1440Γ—1280 / 1440Γ—1440 | 34.8 s / 39.8 s |
460
 
461
  ## LoRA strength
462
 
@@ -545,8 +607,9 @@ freshly noised copy: the frozen teacher, and a second small adapter on the same
545
  is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
546
  the image more like the teacher's work and less like the student's habits, and the student is pushed that way
547
  (the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
548
- rank 32, starts as an exact copy of the teacher, updates four times per student step β€” often enough to keep up with a
549
- student that is still changing β€” and is discarded at the end.
 
550
 
551
  **The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
552
  student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
@@ -563,18 +626,18 @@ first distribution-matching checkpoint on the same prompts):
563
 
564
  | bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
565
  | --------- | --------------------------- | --------------- | -------------- | ----------------------- |
566
- | 512x512 | 1.02 (1.16) | 1.02 (1.23) | 1.05 (1.23) | 0.41 (0.44) |
567
- | 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.06 (1.33) | 0.39 (0.41) |
568
- | 768x512 | 1.02 (1.22) | 1.08 (1.37) | 1.04 (1.26) | 0.44 (0.47) |
569
- | 768x768 | 1.09 (1.43) | 1.05 (1.46) | 1.14 (1.50) | 0.40 (0.43) |
570
- | 768x1024 | 1.02 (1.36) | 0.98 (1.38) | 1.07 (1.42) | 0.39 (0.43) |
571
- | 1024x768 | 1.01 (1.35) | 0.98 (1.40) | 1.04 (1.38) | 0.43 (0.46) |
572
- | 1024x1024 | 1.06 (1.47) | 1.07 (1.55) | 1.12 (1.56) | 0.41 (0.44) |
573
- | 1280x960 | 1.03 (1.50) | 1.04 (1.59) | 1.07 (1.56) | 0.40 (0.42) |
574
- | 960x1280 | 1.10 (1.60) | 1.05 (1.63) | 1.19 (1.70) | 0.40 (0.43) |
575
- | 1280x1280 | 1.09 (1.65) | 1.08 (1.68) | 1.13 (1.69) | 0.41 (0.44) |
576
- | 1440x1280 | 1.08 (1.59) | 1.08 (1.66) | 1.14 (1.66) | 0.40 (0.42) |
577
- | 1440x1440 | 1.17 (1.81) | 1.08 (1.79) | 1.21 (1.89) | 0.39 (0.41) |
578
 
579
  ### The spectral match
580
 
@@ -593,7 +656,9 @@ judged too. The comparison is two-sided, so too much fine energy and too little
593
  does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
594
  refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
595
  without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too. The whole-latent comparison trains at that
596
- strength; the decoded window was later brought down to a quarter of it, which kept the detail and removed some grain.
 
 
597
 
598
  ### The artefact critic
599
 
@@ -615,31 +680,53 @@ teacher's finishing pass below take turns instead of sharing a step.
615
 
616
  ### Critics in turn
617
 
618
- One critic holds one idea of what is wrong. The artefact critic was joined by two more on the same frozen mid-network
619
- features and under the same rules β€” lightly re-noised inputs, an empty prompt, a push that is filtered and capped β€” each
620
- aimed at a different fault:
621
 
622
  - **A photo critic.** Half of its real examples are real photographs and half the teacher's finished images, so it learns
623
- what fine texture looks like in a photograph as well as in the teacher's rendering of one. Its push is filtered to
624
- periods finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher
625
- does not.
626
- - **A face critic.** Its real examples are the teacher's finished images of prompts with faces, the face regions counted at
627
- full weight and the rest at half, so its push concentrates on what small and mid-sized faces lose first.
628
-
629
- All three on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
630
- others train so none goes stale before its turn comes back. The two new heads started from the artefact critic's weights and
631
- trained on their own before they were allowed to push.
 
 
 
 
 
 
 
 
 
 
632
 
633
  ### Detail terms
634
 
635
- Four smaller terms sit on top, each capped relative to the distribution term so none of them can take over:
636
 
637
  - **A detail-weighted anchor.** The trajectory regression counts the fine-detail part of its error β€” everything finer
638
- than 32 pixels β€” twice, so the anchor stops tolerating softness it used to average away.
639
- - **A one-sided photo floor.** On the decoded window, the student's energy at periods of 3–10 pixels may not fall below
640
- the teacher's plus the margin real photographs carry over it at those scales. That margin is measured once from a
641
- pool of real photographs and clamped, and the term only ever pushes upward to that floor, never past it β€” so it
642
- lifts detail that is missing without adding grain that is not.
 
 
 
 
 
 
 
 
 
 
 
 
643
  - **The teacher's finish as a target.** Every second step, the teacher itself runs its remaining steps starting from
644
  the student's own first-call output. The result is a finished image that shares the student's layout, and the second
645
  call is pulled gently toward it β€” a target that lines up with what the student actually drew, where the recorded
@@ -647,6 +734,20 @@ Four smaller terms sit on top, each capped relative to the distribution term so
647
  - **A smoothness limit on the fake adapter.** The fake adapter's fine-detail energy is kept below the teacher's at the
648
  point where the distribution term is measured, so the difference between the two keeps pointing toward detail.
649
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
650
  ## What the LoRA touches
651
 
652
  Rank **64**, alpha = rank, bf16, the same **228 modules** as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA): the 8 attention and feed-forward linears
@@ -658,14 +759,17 @@ of all 28 transformer blocks, plus the four global linears β€” `time_embed.linea
658
  The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β€” Krea 2 Turbo's own 8-step run at mu = 1.15 and
659
  guidance 0.0, every latent and velocity stored β€” serve unchanged: a 2-step chord is two of the 4-step chords end to
660
  end. 203 held-out prompts measure the student–teacher gap on unseen prompts and never receive a gradient. The spectral
661
- match and the artefact critic read the teacher's finals for the training prompts. The 43,044 real-photo crops of the
 
662
  [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) enter in two places: as one precomputed statistic β€” how much fine-detail energy they carry at 3–10
663
  pixels relative to the teacher, clamped β€” which sets the photo floor, and as half of the photo critic's real examples. Only that
664
  critic's head sees them; the student and the fake adapter never do, and receive only its filtered, capped push.
665
 
666
  ## Resolutions
667
 
668
- The same 12 buckets as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA), interleaved in proportion to their remaining samples:
 
 
669
 
670
  | | | |
671
  | --------- | --------- | --------- |
@@ -695,9 +799,6 @@ resolution, one image per prompt, so any image can be compared 1:1 with its twin
695
  checkpoint published here, replaced whenever a better one ships. Which checkpoint that is today is in
696
  [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md), and every
697
  published checkpoint's tree is kept under its own number in [`_archive/resolution_sweeps/`](_archive/resolution_sweeps).
698
- - [`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step) and
699
- [`1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme) β€” the out-of-spec single-step material of the
700
- [bonus section](#bonus-the-1-step-extreme-test) further down.
701
 
702
  ```
703
  assets/resolution_sweeps/
@@ -713,9 +814,7 @@ assets/resolution_sweeps/
713
  β”‚ β”œβ”€β”€ 1440x1280/ …and the remaining buckets
714
  β”‚ └── 1440x1440/
715
  β”œβ”€β”€ _turbo-base-NO-LoRA-2step/ the same tree, stock Turbo at 2 steps β€” the floor
716
- β”œβ”€β”€ 2step-LoRA/ the same tree, rendered with this LoRA at 2 steps β€” the published checkpoint
717
- β”œβ”€β”€ _turbo-base-NO-LoRA-1step/ stock Turbo at ONE step (bonus section)
718
- └── 1step-LoRA-extreme/ this LoRA at ONE step, with side-by-side strips (bonus section)
719
  ```
720
 
721
  Two ways to read them, both useful:
@@ -759,8 +858,9 @@ visibly better than the last again.
759
  ## How it is judged
760
 
761
  At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
762
- 15 fixed prompts across four resolutions (512Γ—512, 768Γ—1024, 1280Γ—1280, 1440Γ—1440); milestone checkpoints get the same
763
- render at all 12 buckets, which is where the per-resolution table above comes from. Every image is measured against the
 
764
  teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
765
  and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
766
  seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
@@ -1085,52 +1185,14 @@ individual renders β€” click any image for full size.
1085
  sigmas are anchored to that grid.
1086
  - πŸ§ͺ **Not a finished adapter.** Usable at 2 steps for previews and drafts; not the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s quality, which remains the recommendation for quality renders. Training continues, and a later checkpoint replaces this file only when the sweeps and I visually agree it is better.
1087
 
1088
- ## Bonus: the 1-step extreme test
1089
-
1090
- > **This is an extreme, out-of-spec experiment β€” not recommended for any use.** This LoRA is trained for 2 steps; at
1091
- > 1 step it is half its trained step count and an eighth of the teacher's.
1092
-
1093
- Stock Krea 2 Turbo and Turbo + this LoRA, each run at just **one step** β€” a single call β€” on the same 15 prompts, the
1094
- same seed, at all 12 resolutions of the sweep, the LoRA at its normal strength (1.0). Stock Turbo returns a smear at one
1095
- step: a colour field with a ghost of the subject in it. With the LoRA the same single call returns a coherent picture β€”
1096
- the subject, the composition, the lighting and the colours are all there. What is missing is the fine detail the second
1097
- step adds: skin is soft, hair and fur come out streaked rather than in strands, and the finest structure (feathers,
1098
- falling snow, small text) is largely absent. That makes one step a rough **preview of composition and colour** at an
1099
- eighth of the teacher's cost, and nothing more.
1100
-
1101
- The full set is in [`assets/resolution_sweeps/1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme), one
1102
- folder per resolution: the LoRA's single-step render as `<prompt>.jpg` and the side-by-side strip as
1103
- `<prompt>_comparison.jpg` (native on the left, LoRA on the right); stock Turbo's single-step renders are in
1104
- [`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step). A few samples at 768Γ—1024:
1105
-
1106
- [![native vs LoRA at 1 step β€” portrait](assets/resolution_sweeps/1step-LoRA-extreme/thumbs/portrait.jpg)](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/portrait_comparison.jpg)
1107
- [![native vs LoRA at 1 step β€” inventor](assets/resolution_sweeps/1step-LoRA-extreme/thumbs/inventor.jpg)](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/inventor_comparison.jpg)
1108
- [![native vs LoRA at 1 step β€” mecha](assets/resolution_sweeps/1step-LoRA-extreme/thumbs/mecha.jpg)](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/mecha_comparison.jpg)
1109
- [![native vs LoRA at 1 step β€” pizza](assets/resolution_sweeps/1step-LoRA-extreme/thumbs/pizza.jpg)](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/pizza_comparison.jpg)
1110
- [![native vs LoRA at 1 step β€” racecar](assets/resolution_sweeps/1step-LoRA-extreme/thumbs/racecar.jpg)](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/racecar_comparison.jpg)
1111
- [![native vs LoRA at 1 step β€” sorceress](assets/resolution_sweeps/1step-LoRA-extreme/thumbs/sorceress.jpg)](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/sorceress_comparison.jpg)
1112
-
1113
- _Still out-of-spec and still edge case preview-only β€” this is one step, not the 2-step regime the rest of this page measures._
1114
-
1115
- > πŸ”­ **A 1-step adapter is a possible follow-on project.** The picture above is why: that a single call already holds
1116
- > together with an adapter trained for two suggests a dedicated one-step adapter is worth attempting once this one
1117
- > ships β€” as a booster on top of this LoRA rather than a replacement, trained by distribution matching alone (at one
1118
- > step there is no trajectory left to regress), and judged on seed variety as much as on detail, since one-step students
1119
- > are the ones that collapse to a favourite. Expectations set accordingly: a usable preview at an eighth of the
1120
- > teacher's cost, not the quality bar.
1121
-
1122
  ## What's next
1123
 
1124
  Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any
1125
- resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). The next step
1126
- is already running, and it goes after what this checkpoint still gets wrong: a critic that judges the first call against the
1127
- teacher's own intermediate state from the same noise, so a first call that blends two layouts is caught where the blend
1128
- happens; a critic weighted toward wherever the teacher put fine detail; the photo critic and the photo floor limited to
1129
- photographic prompts, so illustration, anime and 3D renders are no longer pulled toward photographic grain; detail held to the
1130
- teacher region by region, with a ceiling as well as a floor; a focus on the eyes, nose and lips of faces so they sharpen while
1131
- skin stays the teacher's; a colour floor, so colour at large sizes stops falling below the teacher's; and a more even mix of
1132
- resolutions. After it comes prompt adherence β€” a critic that learns whether an image belongs to its own prompt, and stylised
1133
- prompts drawn more often β€” held to the rule that none of it may cost the sharpness this checkpoint gained. A better checkpoint
1134
  replaces this one when the sweeps and I visually agree, the same discipline as the
1135
  [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
1136
  one is the fast preview.
 
33
  > within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past
34
  > 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects β€” a property of the base model, with or without this
35
  > adapter.
 
 
 
36
 
37
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
38
  down to 2** β€” Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at
 
64
  only and never ships.
65
  - 🎲 **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
66
  re-run.
67
+ - πŸ”’ **31,600 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β€” all of them
68
  drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
69
  training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
70
  4-step project, read again at the two sigmas this schedule uses.
71
+ - πŸ“… **18 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β€” and the project continues.
72
+ - πŸ” **More than forty recipe adjustments** across two methods so far β€” seven of trajectory distillation before the switch, the rest of distribution matching since β€” each kept only when the renders did not get worse.
73
  - πŸ–₯️ **One RTX 3090**, and a recipe shaped by its 24 GB.
74
 
75
  [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
 
85
  | `krea2_turbo_2step_rank_64_lora.safetensors` | the LoRA in diffusers key format β€” see [Inference with diffusers](#inference-with-diffusers); also for MLX or anything that reads safetensors |
86
  | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | the same weights under ComfyUI's key names β€” see [ComfyUI](#comfyui) |
87
  | `krea2_turbo_2step_lora_t2i.json` | a ready ComfyUI workflow, stock nodes only |
 
88
  | [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **the quick place to check which checkpoint the two weight files are based on.** The pair above keeps its names and is updated in place as better checkpoints ship; this file always says what they are today. Every published checkpoint also sits in [`_archive/checkpoints/`](_archive/checkpoints) under its number |
89
  | `LICENSE.pdf` | the Krea 2 Community License Agreement, which covers this adapter β€” see [License](#license) |
90
  | `NOTICE.txt` | the attribution notice the license requires of a derivative |
 
99
 
100
  | | |
101
  | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
102
+ | lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β†’ 2-step trajectory distillation β†’ distribution matching β†’ a spectral match against the teacher's own images β†’ an artefact critic and detail terms β†’ three critics taking turns β†’ nine critics, detail and colour held to the teacher region by region, a guarded running average |
103
  | this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
104
+ | what it gives | usable two-step renders at every trained resolution: fine detail and colour at the teacher's level β€” from 1 megapixel up, closer to the teacher than the 4-step adapter β€” with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 1 point missing out of 240). A judge asked which render follows the prompt better still prefers the 8-step teacher on 11 of 45, against 6 for the 4-step adapter, mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
105
 
106
  ### Known issues
107
 
108
+ The usual costs of two steps, in this order of how often they show. **Small subjects are the weak spot, people and objects alike**: a portrait-sized face or an object seen up close holds up, while small or distant subjects β€” faces in a crowd, a figure in a wide scene, the machines at the back of a room β€” can come out ghosted, smeared or misshapen, since at that size a whole subject is only a few of the blocks the model works in. Fine structure can come out soft or a few pixels out of register β€” feathers, hair strands, signage, the surface of a distant object β€” most at 1280Γ—1280 and above, and a faint doubled contour can show on limbs. On stylised prompts, **how the prompt says the image should look is followed less faithfully than what should be in it**: crisp anime linework, energetic brush strokes, the fingerprints in clay or a matte-painting finish come out closer to a generic rendering than the teacher's. On busy action or crowd scenes the composition can repeat itself β€” an extra hand or held object, a figure duplicated in a crowd β€” where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. Skin reads smoother than the teacher's β€” its finest texture, the pores, sits under the teacher's at the larger sizes, though its colour is now close to the teacher's β€” and freckles come out as dots, softer than the teacher's. Every one of these is being worked on; none is hidden in the sweeps or the examples.
109
 
110
  ## How I got here
111
 
 
162
  11. **three critics taking turns** β€” the artefact critic joined by one whose real examples are half real photographs and one
163
  weighted toward faces, one of them pushing on each step while the others keep training in between, because the three did
164
  not fit in memory side by side
165
+ 12. **a colour band** β€” a floor under the saturation of the whole image and of the decoded window at the teacher's own level,
166
+ and a ceiling 8% above it, so colour can neither fall under the teacher's at large sizes nor climb past it; grey and
167
+ black-and-white references are left alone
168
+ 13. **more critics, each with one job** β€” nine in all, still taking turns: the photo critic confined to photographic prompts; a
169
+ critic for large faces that pushes on every step, with the face critic kept beside it for photographs; and five new ones β€”
170
+ the first call's layout judged against the teacher's own intermediate state from the same noise, content weighted toward
171
+ wherever the teacher put detail, the prompt (the critic sees it, and is shown a mismatched one as a negative), text, and
172
+ the teacher's own finish of the student's second-call state
173
+ 14. **detail held to the teacher region by region** β€” the decoded-window spectral pull aimed at the teacher's most detailed
174
+ tenth of the image, with the anchor counting fine error there three and a half times; a per-area ceiling at 1.3Γ— the
175
+ teacher's detail; a direction-aware spectral term; on photographs, half the decoded windows centred on eyes, nose or
176
+ lips; the photo floor confined to photographic prompts, with floors at the teacher's own level for stylised prompts and
177
+ for flat areas
178
+ 15. the second call's losses also reaching the first call, at a fifth of their strength, so the first call is shaped for the
179
+ finish it feeds; small faces weighted up in the distribution term; and the fake-score adapter back to three updates per
180
+ student step, its lag behind the student measured inside the range it had at four, which bought back training pace
181
+ 16. **more training at large sizes and on stylised prompts** β€” 1440Γ—1440 from under 2% of the samples to about 6%, 1024Γ—1024
182
+ and 768Γ—1024 to 12% each; stylised prompts from 4% to 12%
183
+ 17. **a guarded running average** β€” the published weights are a running average of training, and averaging two layouts of
184
+ the same prompt had produced doubled subjects: a short-lived excursion of the training weights is now kept out of the
185
+ average and a lasting change taken in whole, and the average was restarted once, after a layout change it had blended
186
 
187
  Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
188
  a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
189
  not place exactly where the teacher had it.
190
 
191
+ ## chk00031600 vs chk00017464
192
+
193
+ `chk00017464` (published 14 Sep 2026) was the first checkpoint with critics and detail terms. `chk00031600` (25 Sep 2026) is
194
+ 14,136 training samples later, and those samples went to the faults its page listed β€” grain and excess texture at the largest
195
+ sizes, colour running under the teacher's there, small faces, and stylised prompts drifting toward a generic look β€” through the
196
+ changes numbered 12 to 17 under [How I got here](#how-i-got-here), each kept only after its own look at the renders:
197
+
198
+ 1. a colour band at the teacher's level
199
+ 2. nine critics taking turns, each with one job
200
+ 3. detail held to the teacher region by region, with a ceiling as well as floors
201
+ 4. the second call's losses reaching the first call, and small faces weighted up in the distribution term
202
+ 5. more training at the large sizes and on stylised prompts
203
+ 6. a guarded running average
204
+
205
+ **Grain and texture at large sizes β€” the headline.** Every one of the 12 sweep resolutions Γ— 15 prompts measured against the
206
+ 8-step teacher, as in [Measured against the teacher](#measured-against-the-teacher). The excess fine energy two steps used to put
207
+ into large images is gone: fine texture and both grid bands sit within 5% of the teacher's at 1280Γ—1280 and 1440Γ—1280, and within
208
+ 13% at 1440Γ—1440:
209
+
210
+ | 1.00 = the teacher | fine texture | 16-px band | 8-px band |
211
+ | --- | --- | --- | --- |
212
+ | 1280Γ—1280 | 1.09 β†’ **0.95** | 1.08 β†’ **0.95** | 1.13 β†’ **0.96** |
213
+ | 1440Γ—1280 | 1.08 β†’ **0.95** | 1.09 β†’ **0.97** | 1.14 β†’ **0.99** |
214
+ | 1440Γ—1440 | 1.17 β†’ **1.08** | 1.08 β†’ **1.05** | 1.21 β†’ **1.13** |
215
+
216
+ The 8-px band comes closer to the teacher at all 12 resolutions, the 16-px band at 9 and fine texture at 7, and the grain in flat
217
+ areas β€” skies, walls, out-of-focus backgrounds β€” at 10 of 12: 1.08Γ— the teacher's across the sweep, from 1.26Γ— (1440Γ—1440: 1.57Γ—
218
+ β†’ 1.14Γ—, median of the 15 prompts). The two ghosting indexes are about level (closer at 7 and at 5 of the 12), and the distance to
219
+ the teacher β€” a plain pixel measure this page reports but training does not optimise β€” is within a hundredth of the previous
220
+ checkpoint's.
221
+
222
+ **Colour, now at the teacher's level.** Saturation across the sweep rises to 0.98Γ— the teacher's from 0.93Γ—, closer at 10 of the 12
223
+ sizes; at 1440Γ—1440 it goes from 0.86Γ— to 0.96Γ—, at 1280Γ—1280 from 0.87Γ— to 0.93Γ—. Skin inside detected faces follows: its
224
+ saturation from 0.92Γ— to 0.97Γ— at 768Γ—1024 and from 0.85Γ— to 0.94Γ— at 1440Γ—1440.
225
+
226
+ **Small faces.** A crowd probe β€” 15 prompts full of small faces at 1280Γ—1280 β€” checks every frontal face for eyes, nose and mouth in
227
+ place, with the check calibrated on the teacher's own faces. **63% of the faces keep that structure, up from 53%, against the
228
+ teacher's 64%**; the blur that brings the teacher's own faces down to the same pass rate falls from 1.7 pixels to 1.0. The gain is largest on
229
+ faces 32–48 pixels tall (56% β†’ 69%) and 64–96 pixels tall (73% β†’ 87%). Small subjects stay the part of the image two steps find
230
+ hardest β€” see [Known issues](#known-issues).
231
+
232
+ **Faces up close.** On the test portrait at 1:1 the freckles come out as separate dots rather than the clusters of the previous
233
+ checkpoint, and the eyes stay clean, irises and catchlights in place; the skin between the freckles is smoother than the teacher's.
234
+
235
+ **Prompt adherence β€” held.** The judge that asks which of two renders follows the prompt better prefers the teacher on 11 of 45, as
236
+ before; the blind rubric finds the same 1 point missing out of 240; and on 15 prompts drawn fresh from the prompt bank for this
237
+ checkpoint, plus 5 black-and-white ones, both checkpoints come out the same: 1 win, 9 ties, 5 losses, and 0 Β· 4 Β· 1 in black and
238
+ white.
239
+
240
+ **Speed β€” unchanged.** The same adapter shape at the same cost: measured again at 1024Γ—1024, two steps with this LoRA took 19.8 and
241
+ 20.2 s against 84.3 s for the teacher's eight β€” the same 4.2Γ—.
242
+
243
+ **What stays a limit of two steps.** Small subjects in wide scenes, and the tactile surface of stylised materials such as clay,
244
+ are still where two steps fall furthest short of eight. Both were worked on across these samples β€” face critics of several kinds,
245
+ a small-face curriculum, small faces weighted up in the distribution term, a critic on the teacher's own finish β€” and both remain
246
+ the focus of what comes next. Fine edges and skin texture at the largest sizes sit a little under the teacher's; the figures are
247
+ in [Measured against the teacher](#measured-against-the-teacher).
248
+
249
+ | axis | `chk00017464` | `chk00031600` |
250
+ | --- | --- | --- |
251
+ | fine texture vs the teacher, 1280Β² / 1440Β² | 1.09 / 1.17 | **0.95 / 1.08** |
252
+ | 16-px grid band, 1280Β² / 1440Β² | 1.08 / 1.08 | **0.95 / 1.05** |
253
+ | grain in flat areas, sweep median | 1.26Γ— | **1.08Γ—** |
254
+ | saturation vs the teacher, sweep mean | 0.93Γ— | **0.98Γ—** |
255
+ | small faces keeping their structure (the teacher: 64%) | 53% | **63%** |
256
+ | judge prefers the teacher (of 45) | 11 | 11 |
257
+ | blind adherence rubric, points missing of 240 | 1 | 1 |
258
+ | distance to the teacher, sweep mean | **0.406** | 0.413 |
259
+ | training samples in the 2-step stages | 17,464 | 31,600 |
260
+
261
  ## chk00017464 vs chk00013663
262
 
263
  `chk00013663` (published 12 Sep 2026) was the first public checkpoint: distribution matching with the spectral match, and nothing
 
325
 
326
  | bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
327
  | --- | --- | --- | --- | --- | --- |
328
+ | 512Γ—512 | 1.00 | 1.03 | 1.04 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
329
+ | 768Γ—1024 | 0.95 | 0.93 | 0.98 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
330
+ | 1024Γ—1024 | 1.00 | 1.03 | 1.05 | 1.11 Β· 1.17 Β· 1.16 | 0.39 |
331
+ | 1280Γ—1280 | 0.95 | 0.95 | 0.96 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
332
+ | 1440Γ—1440 | 1.08 | 1.05 | 1.13 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
333
 
334
  Two steps without the adapter carry 0.57Γ— the teacher's fine detail at 512Γ—512 and **less than half** (0.39–0.42Γ—) at the four larger sizes. With it, the detail
335
+ sits at the teacher's level everywhere β€” between 0.93Γ— and 1.13Γ— of it β€” and from 1 megapixel up it is closer to the teacher
336
+ than the 4-step adapter, which carries more excess fine energy there.
337
 
338
  **Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
339
  in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
340
 
341
  | bucket | wins | ties | losses |
342
  | --- | --- | --- | --- |
343
+ | 512Γ—512 | 0 | 12 | 3 |
344
+ | 1280Γ—1280 | 0 | 11 | 4 |
345
+ | 1440Γ—1440 | 0 | 11 | 4 |
346
 
347
  Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
348
  stylised prompts and on how a prompt says the picture should look β€” crisp linework, energetic brush strokes, the texture of
349
  clay β€” more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
350
  attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
351
+ training prompt bank for this checkpoint and never rendered before, the judge returned 1 win, 9 ties, 5 losses, and on five fresh
352
+ black-and-white prompts 0 wins, 4 ties, 1 loss β€” with no colour cast in any of the five.
353
 
354
+ **Checked for the damage this kind of training can do.** Saturation sits at 1.04Γ— the teacher's at 768Γ—1024 and 0.94–0.97Γ—
355
+ at the larger sizes; edge detail 0.90–0.91Γ—; skin texture inside detected faces 1.04Γ— at 768Γ—1024 and 0.87Γ— / 0.81Γ— at
356
+ 1280Γ—1280 / 1440Γ—1440, with skin saturation 0.97Γ— and 0.80–0.94Γ— at the larger sizes. The honest reading: **colour is now at
357
+ the teacher's level at every size, and skin remains the softest part of this adapter's output at large sizes.** Fine detail
358
+ in flat regions β€” skies, walls, out-of-focus backgrounds β€” runs 1.33–1.69Γ— the teacher's on this measure, which is where two
359
+ steps put grain that eight steps do not.
360
 
361
  **Distance to the teacher**, as a plain pixel measure, is 0.39–0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
362
  0.30–0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
 
379
 
380
  **This LoRA is trained on Krea 2 Turbo, against Turbo as its own teacher, and for Turbo.** Every layer it targets also
381
  exists in Krea 2 Raw, so it will load there without complaint β€” but that is a side effect of the shared architecture, not
382
+ a supported mode: it is neither trained nor tuned for Raw's weights, steps or guidance.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
383
 
384
  ## Inference with diffusers
385
 
 
510
  end-to-end figure on your machine will sit below 4.2Γ— and rise toward it as the render gets larger.
511
 
512
  **By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
513
+ size; sweep renders run one at a time, not the controlled measurement above, so a second or two either way between one sweep and
514
+ the next is run-to-run variation β€” the adapter's shape, and so its cost, is the same at every checkpoint):
515
 
516
  | resolution | denoising (2 calls) | resolution | denoising (2 calls) |
517
  | --- | --- | --- | --- |
518
+ | 512Γ—512 | 7.0 s | 1024Γ—1024 | 19.6 s |
519
+ | 512Γ—768 / 768Γ—512 | 11.5 s / 10.9 s | 1280Γ—960 / 960Γ—1280 | 22.9 s / 24.2 s |
520
+ | 768Γ—768 | 12.6 s | 1280Γ—1280 | 31.5 s |
521
+ | 768Γ—1024 / 1024Γ—768 | 15.1 s / 15.1 s | 1440Γ—1280 / 1440Γ—1440 | 34.4 s / 39.6 s |
522
 
523
  ## LoRA strength
524
 
 
607
  is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
608
  the image more like the teacher's work and less like the student's habits, and the student is pushed that way
609
  (the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
610
+ rank 32, starts as an exact copy of the teacher, updates three times per student step β€” often enough to keep up with a
611
+ student that is still changing β€” and is discarded at the end. On prompts with small faces, the push on the face tokens is
612
+ weighted up, because small faces are what two steps get wrong most often.
613
 
614
  **The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
615
  student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
 
626
 
627
  | bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
628
  | --------- | --------------------------- | --------------- | -------------- | ----------------------- |
629
+ | 512x512 | 1.00 (1.16) | 1.03 (1.23) | 1.04 (1.23) | 0.42 (0.44) |
630
+ | 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.04 (1.33) | 0.40 (0.41) |
631
+ | 768x512 | 0.98 (1.22) | 1.03 (1.37) | 0.99 (1.26) | 0.44 (0.47) |
632
+ | 768x768 | 1.03 (1.43) | 1.05 (1.46) | 1.05 (1.50) | 0.41 (0.43) |
633
+ | 768x1024 | 0.95 (1.36) | 0.93 (1.38) | 0.98 (1.42) | 0.40 (0.43) |
634
+ | 1024x768 | 1.01 (1.35) | 1.00 (1.40) | 1.02 (1.38) | 0.42 (0.46) |
635
+ | 1024x1024 | 1.00 (1.47) | 1.03 (1.55) | 1.05 (1.56) | 0.42 (0.44) |
636
+ | 1280x960 | 0.95 (1.50) | 0.98 (1.59) | 1.00 (1.56) | 0.42 (0.42) |
637
+ | 960x1280 | 1.02 (1.60) | 0.99 (1.63) | 1.07 (1.70) | 0.42 (0.43) |
638
+ | 1280x1280 | 0.95 (1.65) | 0.95 (1.68) | 0.96 (1.69) | 0.42 (0.44) |
639
+ | 1440x1280 | 0.95 (1.59) | 0.97 (1.66) | 0.99 (1.66) | 0.40 (0.42) |
640
+ | 1440x1440 | 1.08 (1.81) | 1.05 (1.79) | 1.13 (1.89) | 0.39 (0.41) |
641
 
642
  ### The spectral match
643
 
 
656
  does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
657
  refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
658
  without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too. The whole-latent comparison trains at that
659
+ strength; the decoded window was later brought down to a quarter of it, which kept the detail and removed some grain. Later
660
+ still its pull was aimed: 0.4Γ— the distribution term on the tenth of the image where the teacher has the most fine detail, and
661
+ 0.05Γ— everywhere else, so detail is matched where the teacher has it and flat areas are left alone.
662
 
663
  ### The artefact critic
664
 
 
680
 
681
  ### Critics in turn
682
 
683
+ One critic holds one idea of what is wrong. The artefact critic is joined by eight more on the same frozen mid-network
684
+ features and under the same rules β€” lightly re-noised inputs, a push that is filtered and capped β€” each aimed at a
685
+ different fault:
686
 
687
  - **A photo critic.** Half of its real examples are real photographs and half the teacher's finished images, so it learns
688
+ what fine texture looks like in a photograph as well as in the teacher's rendering of one. It judges photographic prompts
689
+ only, so illustration, anime and 3D renders are not pulled toward photographic grain, and its push is filtered to periods
690
+ finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher does not.
691
+ - **Two face critics, on photographs.** One reads faces of every size, the face regions counted at full weight and the rest at
692
+ half, so its push concentrates on what small and mid-sized faces lose first; the other reads only large faces, 192 pixels and
693
+ up, and pushes on every step.
694
+ - **A structure critic.** It judges the first call β€” the layout, before any detail β€” against the teacher's own intermediate
695
+ state from the same noise, at the high noise levels where layout is decided and at periods of 32 pixels and coarser only,
696
+ so a first call that blends two layouts is caught where the blend happens.
697
+ - **A content critic.** It pushes on every step, weighted toward wherever the teacher put fine detail.
698
+ - **A prompt critic.** Unlike the others it sees the prompt: it learns whether an image belongs to its own prompt, and is
699
+ shown a mismatched prompt as a negative.
700
+ - **A text critic.** It trains on the prompts that ask for lettering, and takes priority on them.
701
+ - **A rollout critic.** Its real examples are the teacher's own finish from the student's second-call starting point, so the
702
+ second call is judged against what the teacher would have made from the same start.
703
+
704
+ All nine on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
705
+ others train so none goes stale before its turn comes back. Each new head started from the artefact critic's weights and
706
+ trained on its own before it was allowed to push.
707
 
708
  ### Detail terms
709
 
710
+ Smaller terms sit on top, each capped relative to the distribution term so none of them can take over:
711
 
712
  - **A detail-weighted anchor.** The trajectory regression counts the fine-detail part of its error β€” everything finer
713
+ than 32 pixels β€” twice, and three and a half times on the tenth of the image where the teacher has the most fine detail,
714
+ so the anchor stops tolerating softness it used to average away.
715
+ - **A one-sided photo floor.** On the decoded window of a photographic prompt, the student's energy at periods of 3–10
716
+ pixels may not fall below the teacher's plus the margin real photographs carry over it at those scales, judged tile by tile
717
+ on the window's textured tiles. That margin is measured once from a pool of real photographs and clamped, and the term
718
+ only ever pushes upward to that floor, never past it β€” so it lifts detail that is missing without adding grain that is not.
719
+ - **Floors at the teacher's own level.** Stylised prompts, and the flat tiles of every image, get a floor at the teacher's
720
+ own 3–10-pixel energy instead, so a clay surface or a painted sky cannot fade below the teacher's and nothing is added
721
+ above it.
722
+ - **A ceiling.** Tile by tile, detail at 3–16 pixels may not climb past 1.3Γ— the teacher's β€” the counterpart of the floors,
723
+ and what keeps grain from building up at large sizes.
724
+ - **A direction-aware term.** The spectral comparison is also made orientation by orientation, so the student's fine detail
725
+ runs in the same directions as the teacher's.
726
+ - **Windows on features.** On photographs, half of the decoded windows are centred on an eye, the nose or the lips of a face,
727
+ so the detail terms look hardest where a face is read first.
728
+ - **The second call reaching the first.** The second call's losses also flow back into the first call, at a fifth of their
729
+ strength, so the first call is shaped for the finish it feeds.
730
  - **The teacher's finish as a target.** Every second step, the teacher itself runs its remaining steps starting from
731
  the student's own first-call output. The result is a finished image that shares the student's layout, and the second
732
  call is pulled gently toward it β€” a target that lines up with what the student actually drew, where the recorded
 
734
  - **A smoothness limit on the fake adapter.** The fake adapter's fine-detail energy is kept below the teacher's at the
735
  point where the distribution term is measured, so the difference between the two keeps pointing toward detail.
736
 
737
+ ### Colour
738
+
739
+ A floor holds saturation at the teacher's own level β€” on the decoded window and on the whole image, every step β€” and a
740
+ ceiling 8% above it keeps it from climbing past. References that are grey or black-and-white are left alone, so a
741
+ monochrome prompt is never pushed toward colour.
742
+
743
+ ### A guarded running average
744
+
745
+ The published weights are a running average of training (decay 0.999), which smooths out the noise of single steps.
746
+ Averaging has one failure: while the training weights move between two layouts of the same prompt, their average draws
747
+ both β€” a doubled subject. A guard watches a fixed set of layout probes every ten steps; a move away from the trend that
748
+ comes back within a few rounds is kept out of the average, and a lasting move is taken in whole β€” the guard never resets
749
+ the average. The average itself was restarted once, by hand, after a layout change it had blended.
750
+
751
  ## What the LoRA touches
752
 
753
  Rank **64**, alpha = rank, bf16, the same **228 modules** as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA): the 8 attention and feed-forward linears
 
759
  The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β€” Krea 2 Turbo's own 8-step run at mu = 1.15 and
760
  guidance 0.0, every latent and velocity stored β€” serve unchanged: a 2-step chord is two of the 4-step chords end to
761
  end. 203 held-out prompts measure the student–teacher gap on unseen prompts and never receive a gradient. The spectral
762
+ match, the detail terms and the critics read the teacher's finals for the training prompts, and the face critics and the
763
+ windows on features use face masks computed once on those finals. The 43,044 real-photo crops of the
764
  [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) enter in two places: as one precomputed statistic β€” how much fine-detail energy they carry at 3–10
765
  pixels relative to the teacher, clamped β€” which sets the photo floor, and as half of the photo critic's real examples. Only that
766
  critic's head sees them; the student and the fake adapter never do, and receive only its filtered, capped push.
767
 
768
  ## Resolutions
769
 
770
+ The same 12 buckets as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Since `chk00017464` the draw leans toward the larger sizes, where two steps
771
+ had put the most grain β€” 1440Γ—1440 from under 2% of the samples to about 6%, 1024Γ—1024 and 768Γ—1024 to 12% each β€” and stylised
772
+ prompts are drawn three times as often as their share of the pool, 12% of the samples instead of 4%:
773
 
774
  | | | |
775
  | --------- | --------- | --------- |
 
799
  checkpoint published here, replaced whenever a better one ships. Which checkpoint that is today is in
800
  [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md), and every
801
  published checkpoint's tree is kept under its own number in [`_archive/resolution_sweeps/`](_archive/resolution_sweeps).
 
 
 
802
 
803
  ```
804
  assets/resolution_sweeps/
 
814
  β”‚ β”œβ”€β”€ 1440x1280/ …and the remaining buckets
815
  β”‚ └── 1440x1440/
816
  β”œβ”€β”€ _turbo-base-NO-LoRA-2step/ the same tree, stock Turbo at 2 steps β€” the floor
817
+ └── 2step-LoRA/ the same tree, rendered with this LoRA at 2 steps β€” the published checkpoint
 
 
818
  ```
819
 
820
  Two ways to read them, both useful:
 
858
  ## How it is judged
859
 
860
  At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
861
+ 15 fixed prompts across four resolutions (512Γ—512, 1280Γ—1280, 1440Γ—1440, 1440Γ—1280), after a layout check at both 1440
862
+ sizes that has to pass first; milestone checkpoints get the same render at all 12 buckets, which is where the
863
+ per-resolution table above comes from. Every image is measured against the
864
  teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
865
  and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
866
  seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
 
1185
  sigmas are anchored to that grid.
1186
  - πŸ§ͺ **Not a finished adapter.** Usable at 2 steps for previews and drafts; not the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s quality, which remains the recommendation for quality renders. Training continues, and a later checkpoint replaces this file only when the sweeps and I visually agree it is better.
1187
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1188
  ## What's next
1189
 
1190
  Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any
1191
+ resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Next, aimed
1192
+ at what this checkpoint still gets wrong: small subjects in wide scenes; the finest edges and the texture of skin at the
1193
+ largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two
1194
+ plausible poses of the same subject can meet; and the tactile surface of stylised materials such as clay β€” each held to the
1195
+ rule that none of it may cost the colour and the clean large sizes this checkpoint gained. A better checkpoint
 
 
 
 
1196
  replaces this one when the sweeps and I visually agree, the same discipline as the
1197
  [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
1198
  one is the fast preview.
README.md CHANGED
@@ -17,28 +17,26 @@ pipeline_tag: text-to-image
17
 
18
  # Krea 2 Turbo β€” 2-Step Distillation LoRA
19
 
20
- **A quarter of the steps Β· 4.2Γ— faster denoising Β· fine detail at 1.01–1.17Γ— the teacher's across all 12 trained resolutions Β· 1 point missing of 240 on a blind prompt-adherence rubric Β· teacher preferred on 11 of 45 judged renders Β· 17,464 training samples on the 4-step project's recorded trajectories Β· 7 days on one RTX 3090 Β· still in training.**
21
 
22
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β€” Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
23
 
24
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
25
  - ⏱️ **4.2Γ— faster denoising** β€” 81.4 s β†’ 19.5 s at 1024Γ—1024; the adapter's own cost per call is within measurement noise
26
- - 🎯 **Fine detail at or just above the teacher's** β€” **1.01–1.17Γ—** the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: **0.39–0.57Γ—**); from 1 megapixel up, closer to the teacher than the 4-step adapter
27
  - πŸ“Š **Distribution matching, not imitation** β€” matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
28
  - πŸ—£οΈ **Prompt-conditioned throughout** β€” teacher and fake scores both read each prompt's conditioning; a blind rubric finds **1 point missing of 240** (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on **11 of 45** (4-step adapter: 6), mostly on style
29
  - πŸ“ **12 trained resolutions** β€” multi-aspect from 512Γ—512 up to 1440Γ—1440
30
  - πŸ”Œ **Drop-in, no exceptions** β€” plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
31
  - 🧬 **Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β€” rank 64 on the same 228 modules
32
  - 🎲 **13,750 recorded teacher trajectories** from the 4-step project, reused β€” not one new teacher run
33
- - πŸ”’ **17,464 training samples** in the 2-step stages, on top of the 4-step LoRA's 78,000
34
- - πŸ“… **7 days** from the first 2-step launch to this checkpoint, on a single RTX 3090 β€” training continues
35
- - πŸ” **26 recipe adjustments** across two methods β€” each kept only when the renders did not get worse
36
 
37
  > πŸ§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β€” a portrait, a single figure, an object up close β€” hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
38
  >
39
  > πŸ“ **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ—2048 β€” Krea's published maximum recommended resolution, beyond this adapter's largest trained size β€” come within easy reach. Past 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
40
- >
41
- > πŸ”€ **Also compatible with Krea 2 Raw** β€” with some prompts, at 7+ steps and light guidance. See [Using it on Raw](#using-it-on-raw).
42
 
43
  [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
44
 
@@ -51,7 +49,6 @@ A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that tak
51
  | `krea2_turbo_2step_rank_64_lora.safetensors` | LoRA in diffusers key format β€” see [diffusers](#diffusers) |
52
  | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | Same weights under ComfyUI key names β€” see [ComfyUI](#comfyui) |
53
  | `krea2_turbo_2step_lora_t2i.json` | Ready ComfyUI workflow, stock nodes only |
54
- | `krea2_raw_7step_lora_experiment_t2i.json` | ComfyUI workflow of the Krea 2 Raw 7-step experiment β€” see [Using it on Raw](#using-it-on-raw) |
55
  | [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **Which checkpoint the two weight files are** β€” updated with every release |
56
  | `LICENSE.pdf` | Krea 2 Community License Agreement |
57
  | `NOTICE.txt` | Required attribution notice |
@@ -106,9 +103,7 @@ Load [`krea2_turbo_2step_lora_t2i.json`](krea2_turbo_2step_lora_t2i.json). Full
106
 
107
  ### Using it on Raw
108
 
109
- Trained on Turbo, for Turbo β€” it loads on **Krea 2 Raw** because the architecture is shared, a side effect rather than a supported mode. The same quarter of Raw's usual 28 steps, **7**, is the place to start; some prompts hold at 4–5 for a quick preview. Keep guidance **light**: `guidance_scale=1.0` in diffusers, **cfg 2.0** in ComfyUI β€” 4.5 crushes most images to near-black at 7 steps, and no guidance leaves them flat.
110
-
111
- Results are mixed and subject-dependent. All 15 prompts at 1024Γ—768, what worked and what didn't, and the workflow ([`krea2_raw_7step_lora_experiment_t2i.json`](krea2_raw_7step_lora_experiment_t2i.json)) are in the [experiment's README](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md).
112
 
113
  ---
114
 
@@ -122,7 +117,7 @@ Results are mixed and subject-dependent. All 15 prompts at 1024Γ—768, what worke
122
 
123
  **Denoising is 4.2Γ— faster than the 8-step bar** β€” two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.
124
 
125
- Prompt encoding and VAE decode don't change with step count, so end to end sits below 4.2Γ— and rises toward it as the render grows. Denoise times at every trained resolution (6.4 s at 512Γ—512 to 39.8 s at 1440Γ—1440) are in the [Detailed Model Card](DETAILED-README.md).
126
 
127
  ---
128
 
@@ -143,19 +138,20 @@ At two steps the dial scales the adapter's whole job β€” turning two coarse call
143
 
144
  ## Current Checkpoint
145
 
146
- **`chk00017464`** (14 Sep 2026) replaces `chk00013663` (12 Sep 2026). It is 3,801 training samples later, all aimed at what distribution matching leaves behind β€” grain and grid pattern at large sizes, small faces, dense detail β€” through five recipe changes, among them the artefact, photo and face critics taking turns and four detail terms.
147
 
148
- | axis | `chk00013663` | `chk00017464` |
149
- | --------------------------------------------- | ------------- | --------------- |
150
- | fine texture vs the teacher, 1280Β² / 1440Β² | 1.20 / 1.34 | **1.09 / 1.17** |
151
- | 16-px grid band, 1280Β² / 1440Β² | 1.16 / 1.25 | **1.08 / 1.08** |
152
- | grain in flat areas, sweep median | 1.36Γ— | **1.26Γ—** |
153
- | distance to the teacher, sweep mean | 0.413 | **0.406** |
154
- | judge prefers the teacher (of 45) | **6** | 11 |
155
- | blind adherence rubric, points missing of 240 | **0** | 1 |
156
- | saturation vs the teacher, sweep mean | **0.94Γ—** | 0.93Γ— |
 
157
 
158
- Fine texture and both grid bands came closer to the teacher at 10 of 12 resolutions. Prompt-following on stylised prompts and colour did not improve β€” both are what the next recipe changes target. [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) always names the checkpoint in the weight files; the full comparison is in the [Detailed Model Card](DETAILED-README.md).
159
 
160
  ---
161
 
@@ -168,8 +164,7 @@ The usual costs of two steps, in order of how often they show:
168
  - **Style** β€” on stylised prompts, _how_ the picture should look (crisp linework, brush strokes, fingerprints in clay, a matte-painting finish) is followed less faithfully than _what_ should be in it
169
  - **Repeats** β€” on busy action or crowd scenes the composition can repeat itself: an extra hand or held object, a figure duplicated in a crowd
170
  - **Different composition** β€” two steps is a shorter path from the same noise, so framing, pose or arrangement can differ from the 8-step render at the same seed. Treat the teacher's render as a quality reference, not the picture two steps will reproduce
171
- - **Skin and colour** β€” skin slightly smoother and less saturated than the teacher's, colour a little under it at the largest sizes; freckles gather into clusters rather than separate dots
172
- - **Grain** β€” a fine grain remains on the most textured subjects at the largest sizes, lighter than in the previous checkpoint
173
 
174
  Every one is being worked on; none is hidden in the sweeps or the examples.
175
 
@@ -305,26 +300,11 @@ assets/resolution_sweeps/
305
  β”‚ β”œβ”€β”€ 1024x1024/
306
  β”‚ └── ... (all 12 buckets)
307
  β”œβ”€β”€ _turbo-base-NO-LoRA-2step/ Same tree, stock Turbo at 2 steps (the floor)
308
- β”œβ”€β”€ 2step-LoRA/ Same tree, this LoRA at 2 steps (the published checkpoint)
309
- β”œβ”€β”€ _turbo-base-NO-LoRA-1step/ Stock Turbo at 1 step (bonus section)
310
- β”œβ”€β”€ 1step-LoRA-extreme/ This LoRA at 1 step, with side-by-side strips (bonus section)
311
- └── raw-LoRA-7steps-experiment/ Krea 2 Raw + this LoRA at 7 steps (Using it on Raw)
312
  ```
313
 
314
  ---
315
 
316
- ## Bonus: 1-Step Extreme Test
317
-
318
- > **Out-of-spec experiment β€” not recommended for any use.** Trained for 2 steps; at 1 step it runs half its trained count and an eighth of the teacher's.
319
-
320
- Stock Turbo returns a smear at one step β€” a colour field with a ghost of the subject. With the LoRA the same single call is a coherent picture: subject, composition, lighting and colours all there. Missing is the detail the second step adds β€” soft skin, streaked hair and fur, little of the finest structure (feathers, falling snow, small text). A rough **preview of composition and colour** at an eighth of the teacher's cost, nothing more.
321
-
322
- Full set at all 12 resolutions, render + side-by-side strip per prompt: [`assets/resolution_sweeps/1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme/). Stock Turbo at 1 step: [`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step/)
323
-
324
- > πŸ”­ **A 1-step adapter is a possible follow-on** β€” a booster on top of this LoRA, trained by distribution matching alone and judged on seed variety as much as detail. Expectation: a usable preview at an eighth of the teacher's cost, not the quality bar.
325
-
326
- ---
327
-
328
  ## Archive
329
 
330
  Every published checkpoint and its resolution sweep under [`_archive/`](_archive/) β€” [`checkpoints/`](_archive/checkpoints/) and [`resolution_sweeps/`](_archive/resolution_sweeps/), each under its number. Superseded, not maintained.
@@ -335,7 +315,7 @@ Every published checkpoint and its resolution sweep under [`_archive/`](_archive
335
 
336
  Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
337
 
338
- Next, aimed at the [Known Issues](#known-issues): a first call that blends two layouts caught where the blend happens; detail held to the teacher region by region, with a ceiling as well as a floor; photographic grain kept off illustration, anime and 3D; sharper eyes, nose and lips with skin kept the teacher's; colour at large sizes held up to the teacher's; a more even mix of resolutions. After that, prompt adherence on stylised prompts β€” none of it allowed to cost the sharpness this checkpoint gained.
339
 
340
  A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
341
 
 
17
 
18
  # Krea 2 Turbo β€” 2-Step Distillation LoRA
19
 
20
+ **A quarter of the steps Β· 4.2Γ— faster denoising Β· fine detail at 0.95–1.08Γ— the teacher's across all 12 trained resolutions Β· colour at the teacher's level Β· 1 point missing of 240 on a blind prompt-adherence rubric Β· teacher preferred on 11 of 45 judged renders Β· 31,600 training samples on the 4-step project's recorded trajectories Β· 18 days on one RTX 3090 Β· still in training.**
21
 
22
  A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β€” Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
23
 
24
  - ⚑ **A quarter of the steps** β€” 8 β†’ 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
25
  - ⏱️ **4.2Γ— faster denoising** β€” 81.4 s β†’ 19.5 s at 1024Γ—1024; the adapter's own cost per call is within measurement noise
26
+ - 🎯 **Fine detail at the teacher's level** β€” **0.95–1.08Γ—** the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: **0.39–0.57Γ—**); from 1 megapixel up, closer to the teacher than the 4-step adapter
27
  - πŸ“Š **Distribution matching, not imitation** β€” matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
28
  - πŸ—£οΈ **Prompt-conditioned throughout** β€” teacher and fake scores both read each prompt's conditioning; a blind rubric finds **1 point missing of 240** (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on **11 of 45** (4-step adapter: 6), mostly on style
29
  - πŸ“ **12 trained resolutions** β€” multi-aspect from 512Γ—512 up to 1440Γ—1440
30
  - πŸ”Œ **Drop-in, no exceptions** β€” plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
31
  - 🧬 **Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β€” rank 64 on the same 228 modules
32
  - 🎲 **13,750 recorded teacher trajectories** from the 4-step project, reused β€” not one new teacher run
33
+ - πŸ”’ **31,600 training samples** in the 2-step stages, on top of the 4-step LoRA's 78,000
34
+ - πŸ“… **18 days** from the first 2-step launch to this checkpoint, on a single RTX 3090 β€” training continues
35
+ - πŸ” **More than forty recipe adjustments** across two methods β€” each kept only when the renders did not get worse
36
 
37
  > πŸ§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β€” a portrait, a single figure, an object up close β€” hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
38
  >
39
  > πŸ“ **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ—2048 β€” Krea's published maximum recommended resolution, beyond this adapter's largest trained size β€” come within easy reach. Past 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
 
 
40
 
41
  [![The 15 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps](assets/thumbs/poster.jpg)](assets/poster.jpg)
42
 
 
49
  | `krea2_turbo_2step_rank_64_lora.safetensors` | LoRA in diffusers key format β€” see [diffusers](#diffusers) |
50
  | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | Same weights under ComfyUI key names β€” see [ComfyUI](#comfyui) |
51
  | `krea2_turbo_2step_lora_t2i.json` | Ready ComfyUI workflow, stock nodes only |
 
52
  | [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **Which checkpoint the two weight files are** β€” updated with every release |
53
  | `LICENSE.pdf` | Krea 2 Community License Agreement |
54
  | `NOTICE.txt` | Required attribution notice |
 
103
 
104
  ### Using it on Raw
105
 
106
+ Trained on Turbo, for Turbo β€” it loads on **Krea 2 Raw** because the architecture is shared, a side effect rather than a supported mode: it is neither trained nor tuned for Raw's weights, steps or guidance.
 
 
107
 
108
  ---
109
 
 
117
 
118
  **Denoising is 4.2Γ— faster than the 8-step bar** β€” two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.
119
 
120
+ Prompt encoding and VAE decode don't change with step count, so end to end sits below 4.2Γ— and rises toward it as the render grows. Denoise times at every trained resolution (7.0 s at 512Γ—512 to 39.6 s at 1440Γ—1440) are in the [Detailed Model Card](DETAILED-README.md).
121
 
122
  ---
123
 
 
138
 
139
  ## Current Checkpoint
140
 
141
+ **`chk00031600`** (25 Sep 2026) replaces `chk00017464` (14 Sep 2026). It is 14,136 training samples later, aimed at what the previous checkpoint still got wrong β€” grain and excess texture at the largest sizes, colour under the teacher's there, small faces, stylised prompts β€” through a colour band at the teacher's level, nine critics taking turns, detail held to the teacher region by region, more training at the large sizes and a guarded running average.
142
 
143
+ | axis | `chk00017464` | `chk00031600` |
144
+ | ------------------------------------------------------ | ------------- | --------------- |
145
+ | fine texture vs the teacher, 1280Β² / 1440Β² | 1.09 / 1.17 | **0.95 / 1.08** |
146
+ | 16-px grid band, 1280Β² / 1440Β² | 1.08 / 1.08 | **0.95 / 1.05** |
147
+ | grain in flat areas, sweep median | 1.26Γ— | **1.08Γ—** |
148
+ | saturation vs the teacher, sweep mean | 0.93Γ— | **0.98Γ—** |
149
+ | small faces keeping their structure (the teacher: 64%) | 53% | **63%** |
150
+ | judge prefers the teacher (of 45) | 11 | 11 |
151
+ | blind adherence rubric, points missing of 240 | 1 | 1 |
152
+ | distance to the teacher, sweep mean | **0.406** | 0.413 |
153
 
154
+ The excess fine energy at large sizes is gone and colour is at the teacher's level at every size; small faces keep their structure nearly as often as the teacher's, and prompt-following held level. [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) always names the checkpoint in the weight files; the full comparison is in the [Detailed Model Card](DETAILED-README.md).
155
 
156
  ---
157
 
 
164
  - **Style** β€” on stylised prompts, _how_ the picture should look (crisp linework, brush strokes, fingerprints in clay, a matte-painting finish) is followed less faithfully than _what_ should be in it
165
  - **Repeats** β€” on busy action or crowd scenes the composition can repeat itself: an extra hand or held object, a figure duplicated in a crowd
166
  - **Different composition** β€” two steps is a shorter path from the same noise, so framing, pose or arrangement can differ from the 8-step render at the same seed. Treat the teacher's render as a quality reference, not the picture two steps will reproduce
167
+ - **Skin** β€” smoother than the teacher's at the larger sizes, the pores softer, though its colour is now close to the teacher's; freckles come out as dots, softer than the teacher's
 
168
 
169
  Every one is being worked on; none is hidden in the sweeps or the examples.
170
 
 
300
  β”‚ β”œβ”€β”€ 1024x1024/
301
  β”‚ └── ... (all 12 buckets)
302
  β”œβ”€β”€ _turbo-base-NO-LoRA-2step/ Same tree, stock Turbo at 2 steps (the floor)
303
+ └── 2step-LoRA/ Same tree, this LoRA at 2 steps (the published checkpoint)
 
 
 
304
  ```
305
 
306
  ---
307
 
 
 
 
 
 
 
 
 
 
 
 
 
308
  ## Archive
309
 
310
  Every published checkpoint and its resolution sweep under [`_archive/`](_archive/) β€” [`checkpoints/`](_archive/checkpoints/) and [`resolution_sweeps/`](_archive/resolution_sweeps/), each under its number. Superseded, not maintained.
 
315
 
316
  Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β€” aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
317
 
318
+ Next, aimed at the [Known Issues](#known-issues): small subjects in wide scenes; the finest edges and skin texture at the largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two plausible poses of the same subject can meet; the tactile surface of stylised materials such as clay β€” none of it allowed to cost the colour and the clean large sizes this checkpoint gained.
319
 
320
  A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
321
 
krea2_turbo_2step_rank_64_lora.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:cc7a7d65ca04070b2aa1dbf60b3cd0cf210fde2a93d4d7d4fc18bdadf23060d3
3
- size 438161144
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1df55a05e4ed3cd367e3cc06c84279d9dd2430394119b81038b9ad30c3b17952
3
+ size 438161624
krea2_turbo_2step_rank_64_lora_checkpoint_info.md CHANGED
@@ -1,13 +1,13 @@
1
  # Which checkpoint is this?
2
 
3
  `krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
4
- **chk00017464** β€” the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00017464.safetensors` (and its
5
  `_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
6
  keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
7
 
8
  | file | SHA-256 | size |
9
  | --- | --- | --- |
10
- | `krea2_turbo_2step_rank_64_lora.safetensors` | `cc7a7d65ca04070b2aa1dbf60b3cd0cf210fde2a93d4d7d4fc18bdadf23060d3` | 418M |
11
- | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `5c02cac5de27dcb40ea98caee5fd5668ac1a928a105a1fb29c02eb9b0b1e5862` | 418M |
12
 
13
- Updated: 14 Sep 2026
 
1
  # Which checkpoint is this?
2
 
3
  `krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
4
+ **chk00031600** β€” the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00031600.safetensors` (and its
5
  `_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
6
  keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
7
 
8
  | file | SHA-256 | size |
9
  | --- | --- | --- |
10
+ | `krea2_turbo_2step_rank_64_lora.safetensors` | `1df55a05e4ed3cd367e3cc06c84279d9dd2430394119b81038b9ad30c3b17952` | 418M |
11
+ | `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `e017280bc21c5279156adb0daf6654238896ca1ee0c12c8459b66b882edcfca4` | 418M |
12
 
13
+ Updated: 25 Sep 2026
krea2_turbo_2step_rank_64_lora_comfyui.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:5c02cac5de27dcb40ea98caee5fd5668ac1a928a105a1fb29c02eb9b0b1e5862
3
- size 438141680
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:e017280bc21c5279156adb0daf6654238896ca1ee0c12c8459b66b882edcfca4
3
+ size 438142160