Instructions to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("krea/Krea-2-Turbo", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("lvladikov/Krea2-Turbo-Distill-2step-LoRA") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Checkpoint 31600 Release
Browse files
DETAILED-README.md
CHANGED
|
@@ -33,9 +33,6 @@ pipeline_tag: text-to-image
|
|
| 33 |
> within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past
|
| 34 |
> 2048Γ2048, stock Krea 2 itself begins to duplicate subjects β a property of the base model, with or without this
|
| 35 |
> adapter.
|
| 36 |
-
>
|
| 37 |
-
> π **Also compatible with Krea 2 Raw.** With some prompts it works very well on Krea 2 Raw too, at 7 steps+ β see
|
| 38 |
-
> [Using it on Raw](#using-it-on-raw) and the [dedicated experiment](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md).
|
| 39 |
|
| 40 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
|
| 41 |
down to 2** β Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β aiming at
|
|
@@ -67,12 +64,12 @@ recommendation for quality renders.
|
|
| 67 |
only and never ships.
|
| 68 |
- π² **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
|
| 69 |
re-run.
|
| 70 |
-
- π’ **
|
| 71 |
drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
|
| 72 |
training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
|
| 73 |
4-step project, read again at the two sigmas this schedule uses.
|
| 74 |
-
- π
**
|
| 75 |
-
- π **
|
| 76 |
- π₯οΈ **One RTX 3090**, and a recipe shaped by its 24 GB.
|
| 77 |
|
| 78 |
[](assets/poster.jpg)
|
|
@@ -88,7 +85,6 @@ comparisons with the 8-step teacher are in [Examples](#examples)._
|
|
| 88 |
| `krea2_turbo_2step_rank_64_lora.safetensors` | the LoRA in diffusers key format β see [Inference with diffusers](#inference-with-diffusers); also for MLX or anything that reads safetensors |
|
| 89 |
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | the same weights under ComfyUI's key names β see [ComfyUI](#comfyui) |
|
| 90 |
| `krea2_turbo_2step_lora_t2i.json` | a ready ComfyUI workflow, stock nodes only |
|
| 91 |
-
| `krea2_raw_7step_lora_experiment_t2i.json` | the Krea 2 Raw 7-step experiment's ComfyUI workflow ([Using it on Raw](#using-it-on-raw)) |
|
| 92 |
| [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **the quick place to check which checkpoint the two weight files are based on.** The pair above keeps its names and is updated in place as better checkpoints ship; this file always says what they are today. Every published checkpoint also sits in [`_archive/checkpoints/`](_archive/checkpoints) under its number |
|
| 93 |
| `LICENSE.pdf` | the Krea 2 Community License Agreement, which covers this adapter β see [License](#license) |
|
| 94 |
| `NOTICE.txt` | the attribution notice the license requires of a derivative |
|
|
@@ -103,13 +99,13 @@ place to check which checkpoint the current files are based on.
|
|
| 103 |
|
| 104 |
| | |
|
| 105 |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 106 |
-
| lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β 2-step trajectory distillation β distribution matching β a spectral match against the teacher's own images β an artefact critic and detail terms β three critics taking turns
|
| 107 |
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
|
| 108 |
-
| what it gives | usable two-step renders at every trained resolution: fine detail
|
| 109 |
|
| 110 |
### Known issues
|
| 111 |
|
| 112 |
-
The usual costs of two steps, in this order of how often they show. **Small subjects are the weak spot, people and objects alike**: a portrait-sized face or an object seen up close holds up, while small or distant subjects β faces in a crowd, a figure in a wide scene, the machines at the back of a room β can come out ghosted, smeared or misshapen, since at that size a whole subject is only a few of the blocks the model works in. Fine structure can come out soft or a few pixels out of register β feathers, hair strands, signage, the surface of a distant object β most at 1280Γ1280 and above, and a faint doubled contour can show on limbs. On stylised prompts, **how the prompt says the image should look is followed less faithfully than what should be in it**: crisp anime linework, energetic brush strokes, the fingerprints in clay or a matte-painting finish come out closer to a generic rendering than the teacher's. On busy action or crowd scenes the composition can repeat itself β an extra hand or held object, a figure duplicated in a crowd β where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. Skin reads
|
| 113 |
|
| 114 |
## How I got here
|
| 115 |
|
|
@@ -166,11 +162,102 @@ The recipe adjustments so far, each made on the measurement of the one before:
|
|
| 166 |
11. **three critics taking turns** β the artefact critic joined by one whose real examples are half real photographs and one
|
| 167 |
weighted toward faces, one of them pushing on each step while the others keep training in between, because the three did
|
| 168 |
not fit in memory side by side
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 169 |
|
| 170 |
Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
|
| 171 |
a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
|
| 172 |
not place exactly where the teacher had it.
|
| 173 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 174 |
## chk00017464 vs chk00013663
|
| 175 |
|
| 176 |
`chk00013663` (published 12 Sep 2026) was the first public checkpoint: distribution matching with the spectral match, and nothing
|
|
@@ -238,38 +325,38 @@ teacher):
|
|
| 238 |
|
| 239 |
| bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
|
| 240 |
| --- | --- | --- | --- | --- | --- |
|
| 241 |
-
| 512Γ512 | 1.
|
| 242 |
-
| 768Γ1024 |
|
| 243 |
-
| 1024Γ1024 | 1.
|
| 244 |
-
| 1280Γ1280 |
|
| 245 |
-
| 1440Γ1440 | 1.
|
| 246 |
|
| 247 |
Two steps without the adapter carry 0.57Γ the teacher's fine detail at 512Γ512 and **less than half** (0.39β0.42Γ) at the four larger sizes. With it, the detail
|
| 248 |
-
sits at
|
| 249 |
-
adapter, which carries more excess fine energy there.
|
| 250 |
|
| 251 |
**Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
|
| 252 |
in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
|
| 253 |
|
| 254 |
| bucket | wins | ties | losses |
|
| 255 |
| --- | --- | --- | --- |
|
| 256 |
-
| 512Γ512 |
|
| 257 |
-
| 1280Γ1280 |
|
| 258 |
-
| 1440Γ1440 | 0 |
|
| 259 |
|
| 260 |
Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
|
| 261 |
stylised prompts and on how a prompt says the picture should look β crisp linework, energetic brush strokes, the texture of
|
| 262 |
clay β more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
|
| 263 |
attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
|
| 264 |
-
training prompt bank and never rendered before, the judge returned
|
| 265 |
-
|
| 266 |
|
| 267 |
-
**Checked for the damage this kind of training can do.** Saturation sits at
|
| 268 |
-
at the larger sizes; edge detail 0.
|
| 269 |
-
1440Γ1440, with skin saturation 0.
|
| 270 |
-
|
| 271 |
-
|
| 272 |
-
|
| 273 |
|
| 274 |
**Distance to the teacher**, as a plain pixel measure, is 0.39β0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
|
| 275 |
0.30β0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
|
|
@@ -292,33 +379,7 @@ of the 4-step grid, so the model is evaluated at two points it already knows.
|
|
| 292 |
|
| 293 |
**This LoRA is trained on Krea 2 Turbo, against Turbo as its own teacher, and for Turbo.** Every layer it targets also
|
| 294 |
exists in Krea 2 Raw, so it will load there without complaint β but that is a side effect of the shared architecture, not
|
| 295 |
-
a supported mode.
|
| 296 |
-
|
| 297 |
-
It still does something useful there. On Turbo the adapter runs a quarter of the teacher's steps β 2 of its 8 β so on
|
| 298 |
-
Raw the same quarter of its usual 28 steps, **7**, is the natural place to start, and for most prompts it is a good one.
|
| 299 |
-
For a quick preview, some prompts hold together at even 4β5 steps β not as good as 7 steps or more, but not broken
|
| 300 |
-
either. Keep the guidance **light**: `guidance_scale=1.0` in diffusers, which is **cfg 2.0 in ComfyUI**. Heavier
|
| 301 |
-
guidance, such as 4.5, crushes most images into near-black frames at 7 steps, and with no guidance at all the pictures
|
| 302 |
-
come out flat.
|
| 303 |
-
|
| 304 |
-
[](assets/resolution_sweeps/raw-LoRA-7steps-experiment/1024x768/portrait.jpg)
|
| 305 |
-
[](assets/resolution_sweeps/raw-LoRA-7steps-experiment/1024x768/pizza.jpg)
|
| 306 |
-
|
| 307 |
-
_Krea 2 Raw + this LoRA, 7 steps, light guidance, empty negative prompt, seed 4242, 1024Γ768. Click for full size._
|
| 308 |
-
|
| 309 |
-
Results are **mixed and subject-dependent**. Some prompts come through as finished pictures; others do not β an
|
| 310 |
-
underexposed night street, a cityscape with less detail than Turbo gives at 2 steps β and for those a few more steps may
|
| 311 |
-
help. The full experiment, with all 15 test prompts at 1024Γ768, what worked and what did not, and a ComfyUI workflow,
|
| 312 |
-
is in its [dedicated README](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md), in
|
| 313 |
-
[`assets/resolution_sweeps/raw-LoRA-7steps-experiment/`](assets/resolution_sweeps/raw-LoRA-7steps-experiment). The same
|
| 314 |
-
prompts on stock Raw at the same 7 steps and light guidance, without the LoRA, are in
|
| 315 |
-
[`_raw-base-NO-LoRA-7step-cfg1/`](assets/resolution_sweeps/raw-LoRA-7steps-experiment/_raw-base-NO-LoRA-7step-cfg1).
|
| 316 |
-
|
| 317 |
-
[](assets/resolution_sweeps/raw-LoRA-7steps-experiment/workflow_preview_raw.jpg)
|
| 318 |
-
|
| 319 |
-
_The experiment's ComfyUI workflow,
|
| 320 |
-
[`krea2_raw_7step_lora_experiment_t2i.json`](krea2_raw_7step_lora_experiment_t2i.json): Krea 2 Raw + this LoRA, 7 steps,
|
| 321 |
-
cfg 2.0. Click for full size._
|
| 322 |
|
| 323 |
## Inference with diffusers
|
| 324 |
|
|
@@ -449,14 +510,15 @@ the latent, writing the file β is the same whether you run two steps or eight,
|
|
| 449 |
end-to-end figure on your machine will sit below 4.2Γ and rise toward it as the render gets larger.
|
| 450 |
|
| 451 |
**By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
|
| 452 |
-
size; sweep renders run one at a time, not the controlled measurement above
|
|
|
|
| 453 |
|
| 454 |
| resolution | denoising (2 calls) | resolution | denoising (2 calls) |
|
| 455 |
| --- | --- | --- | --- |
|
| 456 |
-
| 512Γ512 |
|
| 457 |
-
| 512Γ768 / 768Γ512 |
|
| 458 |
-
| 768Γ768 |
|
| 459 |
-
| 768Γ1024 / 1024Γ768 |
|
| 460 |
|
| 461 |
## LoRA strength
|
| 462 |
|
|
@@ -545,8 +607,9 @@ freshly noised copy: the frozen teacher, and a second small adapter on the same
|
|
| 545 |
is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
|
| 546 |
the image more like the teacher's work and less like the student's habits, and the student is pushed that way
|
| 547 |
(the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
|
| 548 |
-
rank 32, starts as an exact copy of the teacher, updates
|
| 549 |
-
student that is still changing β and is discarded at the end.
|
|
|
|
| 550 |
|
| 551 |
**The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
|
| 552 |
student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
|
|
@@ -563,18 +626,18 @@ first distribution-matching checkpoint on the same prompts):
|
|
| 563 |
|
| 564 |
| bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
|
| 565 |
| --------- | --------------------------- | --------------- | -------------- | ----------------------- |
|
| 566 |
-
| 512x512 | 1.
|
| 567 |
-
| 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.
|
| 568 |
-
| 768x512 |
|
| 569 |
-
| 768x768 | 1.
|
| 570 |
-
| 768x1024 |
|
| 571 |
-
| 1024x768 | 1.01 (1.35) |
|
| 572 |
-
| 1024x1024 | 1.
|
| 573 |
-
| 1280x960 |
|
| 574 |
-
| 960x1280 | 1.
|
| 575 |
-
| 1280x1280 |
|
| 576 |
-
| 1440x1280 |
|
| 577 |
-
| 1440x1440 | 1.
|
| 578 |
|
| 579 |
### The spectral match
|
| 580 |
|
|
@@ -593,7 +656,9 @@ judged too. The comparison is two-sided, so too much fine energy and too little
|
|
| 593 |
does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
|
| 594 |
refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
|
| 595 |
without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too. The whole-latent comparison trains at that
|
| 596 |
-
strength; the decoded window was later brought down to a quarter of it, which kept the detail and removed some grain.
|
|
|
|
|
|
|
| 597 |
|
| 598 |
### The artefact critic
|
| 599 |
|
|
@@ -615,31 +680,53 @@ teacher's finishing pass below take turns instead of sharing a step.
|
|
| 615 |
|
| 616 |
### Critics in turn
|
| 617 |
|
| 618 |
-
One critic holds one idea of what is wrong. The artefact critic
|
| 619 |
-
features and under the same rules β lightly re-noised inputs,
|
| 620 |
-
|
| 621 |
|
| 622 |
- **A photo critic.** Half of its real examples are real photographs and half the teacher's finished images, so it learns
|
| 623 |
-
what fine texture looks like in a photograph as well as in the teacher's rendering of one.
|
| 624 |
-
|
| 625 |
-
does not.
|
| 626 |
-
- **
|
| 627 |
-
|
| 628 |
-
|
| 629 |
-
|
| 630 |
-
|
| 631 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 632 |
|
| 633 |
### Detail terms
|
| 634 |
|
| 635 |
-
|
| 636 |
|
| 637 |
- **A detail-weighted anchor.** The trajectory regression counts the fine-detail part of its error β everything finer
|
| 638 |
-
than 32 pixels β twice,
|
| 639 |
-
|
| 640 |
-
|
| 641 |
-
|
| 642 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 643 |
- **The teacher's finish as a target.** Every second step, the teacher itself runs its remaining steps starting from
|
| 644 |
the student's own first-call output. The result is a finished image that shares the student's layout, and the second
|
| 645 |
call is pulled gently toward it β a target that lines up with what the student actually drew, where the recorded
|
|
@@ -647,6 +734,20 @@ Four smaller terms sit on top, each capped relative to the distribution term so
|
|
| 647 |
- **A smoothness limit on the fake adapter.** The fake adapter's fine-detail energy is kept below the teacher's at the
|
| 648 |
point where the distribution term is measured, so the difference between the two keeps pointing toward detail.
|
| 649 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 650 |
## What the LoRA touches
|
| 651 |
|
| 652 |
Rank **64**, alpha = rank, bf16, the same **228 modules** as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA): the 8 attention and feed-forward linears
|
|
@@ -658,14 +759,17 @@ of all 28 transformer blocks, plus the four global linears β `time_embed.linea
|
|
| 658 |
The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β Krea 2 Turbo's own 8-step run at mu = 1.15 and
|
| 659 |
guidance 0.0, every latent and velocity stored β serve unchanged: a 2-step chord is two of the 4-step chords end to
|
| 660 |
end. 203 held-out prompts measure the studentβteacher gap on unseen prompts and never receive a gradient. The spectral
|
| 661 |
-
match and the
|
|
|
|
| 662 |
[4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) enter in two places: as one precomputed statistic β how much fine-detail energy they carry at 3β10
|
| 663 |
pixels relative to the teacher, clamped β which sets the photo floor, and as half of the photo critic's real examples. Only that
|
| 664 |
critic's head sees them; the student and the fake adapter never do, and receive only its filtered, capped push.
|
| 665 |
|
| 666 |
## Resolutions
|
| 667 |
|
| 668 |
-
The same 12 buckets as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)
|
|
|
|
|
|
|
| 669 |
|
| 670 |
| | | |
|
| 671 |
| --------- | --------- | --------- |
|
|
@@ -695,9 +799,6 @@ resolution, one image per prompt, so any image can be compared 1:1 with its twin
|
|
| 695 |
checkpoint published here, replaced whenever a better one ships. Which checkpoint that is today is in
|
| 696 |
[`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md), and every
|
| 697 |
published checkpoint's tree is kept under its own number in [`_archive/resolution_sweeps/`](_archive/resolution_sweeps).
|
| 698 |
-
- [`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step) and
|
| 699 |
-
[`1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme) β the out-of-spec single-step material of the
|
| 700 |
-
[bonus section](#bonus-the-1-step-extreme-test) further down.
|
| 701 |
|
| 702 |
```
|
| 703 |
assets/resolution_sweeps/
|
|
@@ -713,9 +814,7 @@ assets/resolution_sweeps/
|
|
| 713 |
β βββ 1440x1280/ β¦and the remaining buckets
|
| 714 |
β βββ 1440x1440/
|
| 715 |
βββ _turbo-base-NO-LoRA-2step/ the same tree, stock Turbo at 2 steps β the floor
|
| 716 |
-
|
| 717 |
-
βββ _turbo-base-NO-LoRA-1step/ stock Turbo at ONE step (bonus section)
|
| 718 |
-
βββ 1step-LoRA-extreme/ this LoRA at ONE step, with side-by-side strips (bonus section)
|
| 719 |
```
|
| 720 |
|
| 721 |
Two ways to read them, both useful:
|
|
@@ -759,8 +858,9 @@ visibly better than the last again.
|
|
| 759 |
## How it is judged
|
| 760 |
|
| 761 |
At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
|
| 762 |
-
15 fixed prompts across four resolutions (512Γ512,
|
| 763 |
-
|
|
|
|
| 764 |
teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
|
| 765 |
and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
|
| 766 |
seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
|
|
@@ -1085,52 +1185,14 @@ individual renders β click any image for full size.
|
|
| 1085 |
sigmas are anchored to that grid.
|
| 1086 |
- π§ͺ **Not a finished adapter.** Usable at 2 steps for previews and drafts; not the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s quality, which remains the recommendation for quality renders. Training continues, and a later checkpoint replaces this file only when the sweeps and I visually agree it is better.
|
| 1087 |
|
| 1088 |
-
## Bonus: the 1-step extreme test
|
| 1089 |
-
|
| 1090 |
-
> **This is an extreme, out-of-spec experiment β not recommended for any use.** This LoRA is trained for 2 steps; at
|
| 1091 |
-
> 1 step it is half its trained step count and an eighth of the teacher's.
|
| 1092 |
-
|
| 1093 |
-
Stock Krea 2 Turbo and Turbo + this LoRA, each run at just **one step** β a single call β on the same 15 prompts, the
|
| 1094 |
-
same seed, at all 12 resolutions of the sweep, the LoRA at its normal strength (1.0). Stock Turbo returns a smear at one
|
| 1095 |
-
step: a colour field with a ghost of the subject in it. With the LoRA the same single call returns a coherent picture β
|
| 1096 |
-
the subject, the composition, the lighting and the colours are all there. What is missing is the fine detail the second
|
| 1097 |
-
step adds: skin is soft, hair and fur come out streaked rather than in strands, and the finest structure (feathers,
|
| 1098 |
-
falling snow, small text) is largely absent. That makes one step a rough **preview of composition and colour** at an
|
| 1099 |
-
eighth of the teacher's cost, and nothing more.
|
| 1100 |
-
|
| 1101 |
-
The full set is in [`assets/resolution_sweeps/1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme), one
|
| 1102 |
-
folder per resolution: the LoRA's single-step render as `<prompt>.jpg` and the side-by-side strip as
|
| 1103 |
-
`<prompt>_comparison.jpg` (native on the left, LoRA on the right); stock Turbo's single-step renders are in
|
| 1104 |
-
[`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step). A few samples at 768Γ1024:
|
| 1105 |
-
|
| 1106 |
-
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/portrait_comparison.jpg)
|
| 1107 |
-
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/inventor_comparison.jpg)
|
| 1108 |
-
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/mecha_comparison.jpg)
|
| 1109 |
-
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/pizza_comparison.jpg)
|
| 1110 |
-
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/racecar_comparison.jpg)
|
| 1111 |
-
[](assets/resolution_sweeps/1step-LoRA-extreme/768x1024/sorceress_comparison.jpg)
|
| 1112 |
-
|
| 1113 |
-
_Still out-of-spec and still edge case preview-only β this is one step, not the 2-step regime the rest of this page measures._
|
| 1114 |
-
|
| 1115 |
-
> π **A 1-step adapter is a possible follow-on project.** The picture above is why: that a single call already holds
|
| 1116 |
-
> together with an adapter trained for two suggests a dedicated one-step adapter is worth attempting once this one
|
| 1117 |
-
> ships β as a booster on top of this LoRA rather than a replacement, trained by distribution matching alone (at one
|
| 1118 |
-
> step there is no trajectory left to regress), and judged on seed variety as much as on detail, since one-step students
|
| 1119 |
-
> are the ones that collapse to a favourite. Expectations set accordingly: a usable preview at an eighth of the
|
| 1120 |
-
> teacher's cost, not the quality bar.
|
| 1121 |
-
|
| 1122 |
## What's next
|
| 1123 |
|
| 1124 |
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any
|
| 1125 |
-
resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
|
| 1126 |
-
|
| 1127 |
-
|
| 1128 |
-
|
| 1129 |
-
|
| 1130 |
-
teacher region by region, with a ceiling as well as a floor; a focus on the eyes, nose and lips of faces so they sharpen while
|
| 1131 |
-
skin stays the teacher's; a colour floor, so colour at large sizes stops falling below the teacher's; and a more even mix of
|
| 1132 |
-
resolutions. After it comes prompt adherence β a critic that learns whether an image belongs to its own prompt, and stylised
|
| 1133 |
-
prompts drawn more often β held to the rule that none of it may cost the sharpness this checkpoint gained. A better checkpoint
|
| 1134 |
replaces this one when the sweeps and I visually agree, the same discipline as the
|
| 1135 |
[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
|
| 1136 |
one is the fast preview.
|
|
|
|
| 33 |
> within easy reach. That makes the adapter a stepping stone to high-resolution renders as well as a fast preview. Past
|
| 34 |
> 2048Γ2048, stock Krea 2 itself begins to duplicate subjects β a property of the base model, with or without this
|
| 35 |
> adapter.
|
|
|
|
|
|
|
|
|
|
| 36 |
|
| 37 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
|
| 38 |
down to 2** β Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β aiming at
|
|
|
|
| 64 |
only and never ships.
|
| 65 |
- π² **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
|
| 66 |
re-run.
|
| 67 |
+
- π’ **31,600 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β all of them
|
| 68 |
drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
|
| 69 |
training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
|
| 70 |
4-step project, read again at the two sigmas this schedule uses.
|
| 71 |
+
- π
**18 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β and the project continues.
|
| 72 |
+
- π **More than forty recipe adjustments** across two methods so far β seven of trajectory distillation before the switch, the rest of distribution matching since β each kept only when the renders did not get worse.
|
| 73 |
- π₯οΈ **One RTX 3090**, and a recipe shaped by its 24 GB.
|
| 74 |
|
| 75 |
[](assets/poster.jpg)
|
|
|
|
| 85 |
| `krea2_turbo_2step_rank_64_lora.safetensors` | the LoRA in diffusers key format β see [Inference with diffusers](#inference-with-diffusers); also for MLX or anything that reads safetensors |
|
| 86 |
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | the same weights under ComfyUI's key names β see [ComfyUI](#comfyui) |
|
| 87 |
| `krea2_turbo_2step_lora_t2i.json` | a ready ComfyUI workflow, stock nodes only |
|
|
|
|
| 88 |
| [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **the quick place to check which checkpoint the two weight files are based on.** The pair above keeps its names and is updated in place as better checkpoints ship; this file always says what they are today. Every published checkpoint also sits in [`_archive/checkpoints/`](_archive/checkpoints) under its number |
|
| 89 |
| `LICENSE.pdf` | the Krea 2 Community License Agreement, which covers this adapter β see [License](#license) |
|
| 90 |
| `NOTICE.txt` | the attribution notice the license requires of a derivative |
|
|
|
|
| 99 |
|
| 100 |
| | |
|
| 101 |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 102 |
+
| lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β 2-step trajectory distillation β distribution matching β a spectral match against the teacher's own images β an artefact critic and detail terms β three critics taking turns β nine critics, detail and colour held to the teacher region by region, a guarded running average |
|
| 103 |
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
|
| 104 |
+
| what it gives | usable two-step renders at every trained resolution: fine detail and colour at the teacher's level β from 1 megapixel up, closer to the teacher than the 4-step adapter β with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 1 point missing out of 240). A judge asked which render follows the prompt better still prefers the 8-step teacher on 11 of 45, against 6 for the 4-step adapter, mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
|
| 105 |
|
| 106 |
### Known issues
|
| 107 |
|
| 108 |
+
The usual costs of two steps, in this order of how often they show. **Small subjects are the weak spot, people and objects alike**: a portrait-sized face or an object seen up close holds up, while small or distant subjects β faces in a crowd, a figure in a wide scene, the machines at the back of a room β can come out ghosted, smeared or misshapen, since at that size a whole subject is only a few of the blocks the model works in. Fine structure can come out soft or a few pixels out of register β feathers, hair strands, signage, the surface of a distant object β most at 1280Γ1280 and above, and a faint doubled contour can show on limbs. On stylised prompts, **how the prompt says the image should look is followed less faithfully than what should be in it**: crisp anime linework, energetic brush strokes, the fingerprints in clay or a matte-painting finish come out closer to a generic rendering than the teacher's. On busy action or crowd scenes the composition can repeat itself β an extra hand or held object, a figure duplicated in a crowd β where the 8-step and 4-step renders commit to one. On some prompts the composition itself differs from the 8-step render at the same seed: two steps is a shorter path from the same starting noise, so the image can settle on a different framing, pose or arrangement rather than a degraded version of the teacher's. Treat the teacher's render as a reference for quality, not as the picture two steps will reproduce. Skin reads smoother than the teacher's β its finest texture, the pores, sits under the teacher's at the larger sizes, though its colour is now close to the teacher's β and freckles come out as dots, softer than the teacher's. Every one of these is being worked on; none is hidden in the sweeps or the examples.
|
| 109 |
|
| 110 |
## How I got here
|
| 111 |
|
|
|
|
| 162 |
11. **three critics taking turns** β the artefact critic joined by one whose real examples are half real photographs and one
|
| 163 |
weighted toward faces, one of them pushing on each step while the others keep training in between, because the three did
|
| 164 |
not fit in memory side by side
|
| 165 |
+
12. **a colour band** β a floor under the saturation of the whole image and of the decoded window at the teacher's own level,
|
| 166 |
+
and a ceiling 8% above it, so colour can neither fall under the teacher's at large sizes nor climb past it; grey and
|
| 167 |
+
black-and-white references are left alone
|
| 168 |
+
13. **more critics, each with one job** β nine in all, still taking turns: the photo critic confined to photographic prompts; a
|
| 169 |
+
critic for large faces that pushes on every step, with the face critic kept beside it for photographs; and five new ones β
|
| 170 |
+
the first call's layout judged against the teacher's own intermediate state from the same noise, content weighted toward
|
| 171 |
+
wherever the teacher put detail, the prompt (the critic sees it, and is shown a mismatched one as a negative), text, and
|
| 172 |
+
the teacher's own finish of the student's second-call state
|
| 173 |
+
14. **detail held to the teacher region by region** β the decoded-window spectral pull aimed at the teacher's most detailed
|
| 174 |
+
tenth of the image, with the anchor counting fine error there three and a half times; a per-area ceiling at 1.3Γ the
|
| 175 |
+
teacher's detail; a direction-aware spectral term; on photographs, half the decoded windows centred on eyes, nose or
|
| 176 |
+
lips; the photo floor confined to photographic prompts, with floors at the teacher's own level for stylised prompts and
|
| 177 |
+
for flat areas
|
| 178 |
+
15. the second call's losses also reaching the first call, at a fifth of their strength, so the first call is shaped for the
|
| 179 |
+
finish it feeds; small faces weighted up in the distribution term; and the fake-score adapter back to three updates per
|
| 180 |
+
student step, its lag behind the student measured inside the range it had at four, which bought back training pace
|
| 181 |
+
16. **more training at large sizes and on stylised prompts** β 1440Γ1440 from under 2% of the samples to about 6%, 1024Γ1024
|
| 182 |
+
and 768Γ1024 to 12% each; stylised prompts from 4% to 12%
|
| 183 |
+
17. **a guarded running average** β the published weights are a running average of training, and averaging two layouts of
|
| 184 |
+
the same prompt had produced doubled subjects: a short-lived excursion of the training weights is now kept out of the
|
| 185 |
+
average and a lasting change taken in whole, and the average was restarted once, after a layout change it had blended
|
| 186 |
|
| 187 |
Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
|
| 188 |
a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
|
| 189 |
not place exactly where the teacher had it.
|
| 190 |
|
| 191 |
+
## chk00031600 vs chk00017464
|
| 192 |
+
|
| 193 |
+
`chk00017464` (published 14 Sep 2026) was the first checkpoint with critics and detail terms. `chk00031600` (25 Sep 2026) is
|
| 194 |
+
14,136 training samples later, and those samples went to the faults its page listed β grain and excess texture at the largest
|
| 195 |
+
sizes, colour running under the teacher's there, small faces, and stylised prompts drifting toward a generic look β through the
|
| 196 |
+
changes numbered 12 to 17 under [How I got here](#how-i-got-here), each kept only after its own look at the renders:
|
| 197 |
+
|
| 198 |
+
1. a colour band at the teacher's level
|
| 199 |
+
2. nine critics taking turns, each with one job
|
| 200 |
+
3. detail held to the teacher region by region, with a ceiling as well as floors
|
| 201 |
+
4. the second call's losses reaching the first call, and small faces weighted up in the distribution term
|
| 202 |
+
5. more training at the large sizes and on stylised prompts
|
| 203 |
+
6. a guarded running average
|
| 204 |
+
|
| 205 |
+
**Grain and texture at large sizes β the headline.** Every one of the 12 sweep resolutions Γ 15 prompts measured against the
|
| 206 |
+
8-step teacher, as in [Measured against the teacher](#measured-against-the-teacher). The excess fine energy two steps used to put
|
| 207 |
+
into large images is gone: fine texture and both grid bands sit within 5% of the teacher's at 1280Γ1280 and 1440Γ1280, and within
|
| 208 |
+
13% at 1440Γ1440:
|
| 209 |
+
|
| 210 |
+
| 1.00 = the teacher | fine texture | 16-px band | 8-px band |
|
| 211 |
+
| --- | --- | --- | --- |
|
| 212 |
+
| 1280Γ1280 | 1.09 β **0.95** | 1.08 β **0.95** | 1.13 β **0.96** |
|
| 213 |
+
| 1440Γ1280 | 1.08 β **0.95** | 1.09 β **0.97** | 1.14 β **0.99** |
|
| 214 |
+
| 1440Γ1440 | 1.17 β **1.08** | 1.08 β **1.05** | 1.21 β **1.13** |
|
| 215 |
+
|
| 216 |
+
The 8-px band comes closer to the teacher at all 12 resolutions, the 16-px band at 9 and fine texture at 7, and the grain in flat
|
| 217 |
+
areas β skies, walls, out-of-focus backgrounds β at 10 of 12: 1.08Γ the teacher's across the sweep, from 1.26Γ (1440Γ1440: 1.57Γ
|
| 218 |
+
β 1.14Γ, median of the 15 prompts). The two ghosting indexes are about level (closer at 7 and at 5 of the 12), and the distance to
|
| 219 |
+
the teacher β a plain pixel measure this page reports but training does not optimise β is within a hundredth of the previous
|
| 220 |
+
checkpoint's.
|
| 221 |
+
|
| 222 |
+
**Colour, now at the teacher's level.** Saturation across the sweep rises to 0.98Γ the teacher's from 0.93Γ, closer at 10 of the 12
|
| 223 |
+
sizes; at 1440Γ1440 it goes from 0.86Γ to 0.96Γ, at 1280Γ1280 from 0.87Γ to 0.93Γ. Skin inside detected faces follows: its
|
| 224 |
+
saturation from 0.92Γ to 0.97Γ at 768Γ1024 and from 0.85Γ to 0.94Γ at 1440Γ1440.
|
| 225 |
+
|
| 226 |
+
**Small faces.** A crowd probe β 15 prompts full of small faces at 1280Γ1280 β checks every frontal face for eyes, nose and mouth in
|
| 227 |
+
place, with the check calibrated on the teacher's own faces. **63% of the faces keep that structure, up from 53%, against the
|
| 228 |
+
teacher's 64%**; the blur that brings the teacher's own faces down to the same pass rate falls from 1.7 pixels to 1.0. The gain is largest on
|
| 229 |
+
faces 32β48 pixels tall (56% β 69%) and 64β96 pixels tall (73% β 87%). Small subjects stay the part of the image two steps find
|
| 230 |
+
hardest β see [Known issues](#known-issues).
|
| 231 |
+
|
| 232 |
+
**Faces up close.** On the test portrait at 1:1 the freckles come out as separate dots rather than the clusters of the previous
|
| 233 |
+
checkpoint, and the eyes stay clean, irises and catchlights in place; the skin between the freckles is smoother than the teacher's.
|
| 234 |
+
|
| 235 |
+
**Prompt adherence β held.** The judge that asks which of two renders follows the prompt better prefers the teacher on 11 of 45, as
|
| 236 |
+
before; the blind rubric finds the same 1 point missing out of 240; and on 15 prompts drawn fresh from the prompt bank for this
|
| 237 |
+
checkpoint, plus 5 black-and-white ones, both checkpoints come out the same: 1 win, 9 ties, 5 losses, and 0 Β· 4 Β· 1 in black and
|
| 238 |
+
white.
|
| 239 |
+
|
| 240 |
+
**Speed β unchanged.** The same adapter shape at the same cost: measured again at 1024Γ1024, two steps with this LoRA took 19.8 and
|
| 241 |
+
20.2 s against 84.3 s for the teacher's eight β the same 4.2Γ.
|
| 242 |
+
|
| 243 |
+
**What stays a limit of two steps.** Small subjects in wide scenes, and the tactile surface of stylised materials such as clay,
|
| 244 |
+
are still where two steps fall furthest short of eight. Both were worked on across these samples β face critics of several kinds,
|
| 245 |
+
a small-face curriculum, small faces weighted up in the distribution term, a critic on the teacher's own finish β and both remain
|
| 246 |
+
the focus of what comes next. Fine edges and skin texture at the largest sizes sit a little under the teacher's; the figures are
|
| 247 |
+
in [Measured against the teacher](#measured-against-the-teacher).
|
| 248 |
+
|
| 249 |
+
| axis | `chk00017464` | `chk00031600` |
|
| 250 |
+
| --- | --- | --- |
|
| 251 |
+
| fine texture vs the teacher, 1280Β² / 1440Β² | 1.09 / 1.17 | **0.95 / 1.08** |
|
| 252 |
+
| 16-px grid band, 1280Β² / 1440Β² | 1.08 / 1.08 | **0.95 / 1.05** |
|
| 253 |
+
| grain in flat areas, sweep median | 1.26Γ | **1.08Γ** |
|
| 254 |
+
| saturation vs the teacher, sweep mean | 0.93Γ | **0.98Γ** |
|
| 255 |
+
| small faces keeping their structure (the teacher: 64%) | 53% | **63%** |
|
| 256 |
+
| judge prefers the teacher (of 45) | 11 | 11 |
|
| 257 |
+
| blind adherence rubric, points missing of 240 | 1 | 1 |
|
| 258 |
+
| distance to the teacher, sweep mean | **0.406** | 0.413 |
|
| 259 |
+
| training samples in the 2-step stages | 17,464 | 31,600 |
|
| 260 |
+
|
| 261 |
## chk00017464 vs chk00013663
|
| 262 |
|
| 263 |
`chk00013663` (published 12 Sep 2026) was the first public checkpoint: distribution matching with the spectral match, and nothing
|
|
|
|
| 325 |
|
| 326 |
| bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
|
| 327 |
| --- | --- | --- | --- | --- | --- |
|
| 328 |
+
| 512Γ512 | 1.00 | 1.03 | 1.04 | 1.05 Β· 1.05 Β· 1.07 | 0.57 |
|
| 329 |
+
| 768Γ1024 | 0.95 | 0.93 | 0.98 | 1.05 Β· 1.05 Β· 1.08 | 0.39 |
|
| 330 |
+
| 1024Γ1024 | 1.00 | 1.03 | 1.05 | 1.11 Β· 1.17 Β· 1.16 | 0.39 |
|
| 331 |
+
| 1280Γ1280 | 0.95 | 0.95 | 0.96 | 1.20 Β· 1.18 Β· 1.22 | 0.41 |
|
| 332 |
+
| 1440Γ1440 | 1.08 | 1.05 | 1.13 | 1.24 Β· 1.23 Β· 1.32 | 0.42 |
|
| 333 |
|
| 334 |
Two steps without the adapter carry 0.57Γ the teacher's fine detail at 512Γ512 and **less than half** (0.39β0.42Γ) at the four larger sizes. With it, the detail
|
| 335 |
+
sits at the teacher's level everywhere β between 0.93Γ and 1.13Γ of it β and from 1 megapixel up it is closer to the teacher
|
| 336 |
+
than the 4-step adapter, which carries more excess fine energy there.
|
| 337 |
|
| 338 |
**Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
|
| 339 |
in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
|
| 340 |
|
| 341 |
| bucket | wins | ties | losses |
|
| 342 |
| --- | --- | --- | --- |
|
| 343 |
+
| 512Γ512 | 0 | 12 | 3 |
|
| 344 |
+
| 1280Γ1280 | 0 | 11 | 4 |
|
| 345 |
+
| 1440Γ1440 | 0 | 11 | 4 |
|
| 346 |
|
| 347 |
Eleven losses out of 45, where the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six against the same teacher. They gather on
|
| 348 |
stylised prompts and on how a prompt says the picture should look β crisp linework, energetic brush strokes, the texture of
|
| 349 |
clay β more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
|
| 350 |
attributes and relations, with the teacher scored identically, finds 1 point missing out of 240. On 15 prompts drawn fresh from the
|
| 351 |
+
training prompt bank for this checkpoint and never rendered before, the judge returned 1 win, 9 ties, 5 losses, and on five fresh
|
| 352 |
+
black-and-white prompts 0 wins, 4 ties, 1 loss β with no colour cast in any of the five.
|
| 353 |
|
| 354 |
+
**Checked for the damage this kind of training can do.** Saturation sits at 1.04Γ the teacher's at 768Γ1024 and 0.94β0.97Γ
|
| 355 |
+
at the larger sizes; edge detail 0.90β0.91Γ; skin texture inside detected faces 1.04Γ at 768Γ1024 and 0.87Γ / 0.81Γ at
|
| 356 |
+
1280Γ1280 / 1440Γ1440, with skin saturation 0.97Γ and 0.80β0.94Γ at the larger sizes. The honest reading: **colour is now at
|
| 357 |
+
the teacher's level at every size, and skin remains the softest part of this adapter's output at large sizes.** Fine detail
|
| 358 |
+
in flat regions β skies, walls, out-of-focus backgrounds β runs 1.33β1.69Γ the teacher's on this measure, which is where two
|
| 359 |
+
steps put grain that eight steps do not.
|
| 360 |
|
| 361 |
**Distance to the teacher**, as a plain pixel measure, is 0.39β0.44 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
|
| 362 |
0.30β0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
|
|
|
|
| 379 |
|
| 380 |
**This LoRA is trained on Krea 2 Turbo, against Turbo as its own teacher, and for Turbo.** Every layer it targets also
|
| 381 |
exists in Krea 2 Raw, so it will load there without complaint β but that is a side effect of the shared architecture, not
|
| 382 |
+
a supported mode: it is neither trained nor tuned for Raw's weights, steps or guidance.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 383 |
|
| 384 |
## Inference with diffusers
|
| 385 |
|
|
|
|
| 510 |
end-to-end figure on your machine will sit below 4.2Γ and rise toward it as the render gets larger.
|
| 511 |
|
| 512 |
**By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 15 test prompts per
|
| 513 |
+
size; sweep renders run one at a time, not the controlled measurement above, so a second or two either way between one sweep and
|
| 514 |
+
the next is run-to-run variation β the adapter's shape, and so its cost, is the same at every checkpoint):
|
| 515 |
|
| 516 |
| resolution | denoising (2 calls) | resolution | denoising (2 calls) |
|
| 517 |
| --- | --- | --- | --- |
|
| 518 |
+
| 512Γ512 | 7.0 s | 1024Γ1024 | 19.6 s |
|
| 519 |
+
| 512Γ768 / 768Γ512 | 11.5 s / 10.9 s | 1280Γ960 / 960Γ1280 | 22.9 s / 24.2 s |
|
| 520 |
+
| 768Γ768 | 12.6 s | 1280Γ1280 | 31.5 s |
|
| 521 |
+
| 768Γ1024 / 1024Γ768 | 15.1 s / 15.1 s | 1440Γ1280 / 1440Γ1440 | 34.4 s / 39.6 s |
|
| 522 |
|
| 523 |
## LoRA strength
|
| 524 |
|
|
|
|
| 607 |
is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
|
| 608 |
the image more like the teacher's work and less like the student's habits, and the student is pushed that way
|
| 609 |
(the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
|
| 610 |
+
rank 32, starts as an exact copy of the teacher, updates three times per student step β often enough to keep up with a
|
| 611 |
+
student that is still changing β and is discarded at the end. On prompts with small faces, the push on the face tokens is
|
| 612 |
+
weighted up, because small faces are what two steps get wrong most often.
|
| 613 |
|
| 614 |
**The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
|
| 615 |
student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
|
|
|
|
| 626 |
|
| 627 |
| bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
|
| 628 |
| --------- | --------------------------- | --------------- | -------------- | ----------------------- |
|
| 629 |
+
| 512x512 | 1.00 (1.16) | 1.03 (1.23) | 1.04 (1.23) | 0.42 (0.44) |
|
| 630 |
+
| 512x768 | 1.01 (1.23) | 1.06 (1.37) | 1.04 (1.33) | 0.40 (0.41) |
|
| 631 |
+
| 768x512 | 0.98 (1.22) | 1.03 (1.37) | 0.99 (1.26) | 0.44 (0.47) |
|
| 632 |
+
| 768x768 | 1.03 (1.43) | 1.05 (1.46) | 1.05 (1.50) | 0.41 (0.43) |
|
| 633 |
+
| 768x1024 | 0.95 (1.36) | 0.93 (1.38) | 0.98 (1.42) | 0.40 (0.43) |
|
| 634 |
+
| 1024x768 | 1.01 (1.35) | 1.00 (1.40) | 1.02 (1.38) | 0.42 (0.46) |
|
| 635 |
+
| 1024x1024 | 1.00 (1.47) | 1.03 (1.55) | 1.05 (1.56) | 0.42 (0.44) |
|
| 636 |
+
| 1280x960 | 0.95 (1.50) | 0.98 (1.59) | 1.00 (1.56) | 0.42 (0.42) |
|
| 637 |
+
| 960x1280 | 1.02 (1.60) | 0.99 (1.63) | 1.07 (1.70) | 0.42 (0.43) |
|
| 638 |
+
| 1280x1280 | 0.95 (1.65) | 0.95 (1.68) | 0.96 (1.69) | 0.42 (0.44) |
|
| 639 |
+
| 1440x1280 | 0.95 (1.59) | 0.97 (1.66) | 0.99 (1.66) | 0.40 (0.42) |
|
| 640 |
+
| 1440x1440 | 1.08 (1.81) | 1.05 (1.79) | 1.13 (1.89) | 0.39 (0.41) |
|
| 641 |
|
| 642 |
### The spectral match
|
| 643 |
|
|
|
|
| 656 |
does not satisfy it. Its gradient is added to the distribution push and capped per sample as a fraction of it, so it
|
| 657 |
refines rather than takes over. At its first strength it brought every resolution closer to the teacher's spectrum
|
| 658 |
without touching adherence, layout or variety; raised, it began closing the grain on the hardest subjects too. The whole-latent comparison trains at that
|
| 659 |
+
strength; the decoded window was later brought down to a quarter of it, which kept the detail and removed some grain. Later
|
| 660 |
+
still its pull was aimed: 0.4Γ the distribution term on the tenth of the image where the teacher has the most fine detail, and
|
| 661 |
+
0.05Γ everywhere else, so detail is matched where the teacher has it and flat areas are left alone.
|
| 662 |
|
| 663 |
### The artefact critic
|
| 664 |
|
|
|
|
| 680 |
|
| 681 |
### Critics in turn
|
| 682 |
|
| 683 |
+
One critic holds one idea of what is wrong. The artefact critic is joined by eight more on the same frozen mid-network
|
| 684 |
+
features and under the same rules β lightly re-noised inputs, a push that is filtered and capped β each aimed at a
|
| 685 |
+
different fault:
|
| 686 |
|
| 687 |
- **A photo critic.** Half of its real examples are real photographs and half the teacher's finished images, so it learns
|
| 688 |
+
what fine texture looks like in a photograph as well as in the teacher's rendering of one. It judges photographic prompts
|
| 689 |
+
only, so illustration, anime and 3D renders are not pulled toward photographic grain, and its push is filtered to periods
|
| 690 |
+
finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher does not.
|
| 691 |
+
- **Two face critics, on photographs.** One reads faces of every size, the face regions counted at full weight and the rest at
|
| 692 |
+
half, so its push concentrates on what small and mid-sized faces lose first; the other reads only large faces, 192 pixels and
|
| 693 |
+
up, and pushes on every step.
|
| 694 |
+
- **A structure critic.** It judges the first call β the layout, before any detail β against the teacher's own intermediate
|
| 695 |
+
state from the same noise, at the high noise levels where layout is decided and at periods of 32 pixels and coarser only,
|
| 696 |
+
so a first call that blends two layouts is caught where the blend happens.
|
| 697 |
+
- **A content critic.** It pushes on every step, weighted toward wherever the teacher put fine detail.
|
| 698 |
+
- **A prompt critic.** Unlike the others it sees the prompt: it learns whether an image belongs to its own prompt, and is
|
| 699 |
+
shown a mismatched prompt as a negative.
|
| 700 |
+
- **A text critic.** It trains on the prompts that ask for lettering, and takes priority on them.
|
| 701 |
+
- **A rollout critic.** Its real examples are the teacher's own finish from the student's second-call starting point, so the
|
| 702 |
+
second call is judged against what the teacher would have made from the same start.
|
| 703 |
+
|
| 704 |
+
All nine on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
|
| 705 |
+
others train so none goes stale before its turn comes back. Each new head started from the artefact critic's weights and
|
| 706 |
+
trained on its own before it was allowed to push.
|
| 707 |
|
| 708 |
### Detail terms
|
| 709 |
|
| 710 |
+
Smaller terms sit on top, each capped relative to the distribution term so none of them can take over:
|
| 711 |
|
| 712 |
- **A detail-weighted anchor.** The trajectory regression counts the fine-detail part of its error β everything finer
|
| 713 |
+
than 32 pixels β twice, and three and a half times on the tenth of the image where the teacher has the most fine detail,
|
| 714 |
+
so the anchor stops tolerating softness it used to average away.
|
| 715 |
+
- **A one-sided photo floor.** On the decoded window of a photographic prompt, the student's energy at periods of 3β10
|
| 716 |
+
pixels may not fall below the teacher's plus the margin real photographs carry over it at those scales, judged tile by tile
|
| 717 |
+
on the window's textured tiles. That margin is measured once from a pool of real photographs and clamped, and the term
|
| 718 |
+
only ever pushes upward to that floor, never past it β so it lifts detail that is missing without adding grain that is not.
|
| 719 |
+
- **Floors at the teacher's own level.** Stylised prompts, and the flat tiles of every image, get a floor at the teacher's
|
| 720 |
+
own 3β10-pixel energy instead, so a clay surface or a painted sky cannot fade below the teacher's and nothing is added
|
| 721 |
+
above it.
|
| 722 |
+
- **A ceiling.** Tile by tile, detail at 3β16 pixels may not climb past 1.3Γ the teacher's β the counterpart of the floors,
|
| 723 |
+
and what keeps grain from building up at large sizes.
|
| 724 |
+
- **A direction-aware term.** The spectral comparison is also made orientation by orientation, so the student's fine detail
|
| 725 |
+
runs in the same directions as the teacher's.
|
| 726 |
+
- **Windows on features.** On photographs, half of the decoded windows are centred on an eye, the nose or the lips of a face,
|
| 727 |
+
so the detail terms look hardest where a face is read first.
|
| 728 |
+
- **The second call reaching the first.** The second call's losses also flow back into the first call, at a fifth of their
|
| 729 |
+
strength, so the first call is shaped for the finish it feeds.
|
| 730 |
- **The teacher's finish as a target.** Every second step, the teacher itself runs its remaining steps starting from
|
| 731 |
the student's own first-call output. The result is a finished image that shares the student's layout, and the second
|
| 732 |
call is pulled gently toward it β a target that lines up with what the student actually drew, where the recorded
|
|
|
|
| 734 |
- **A smoothness limit on the fake adapter.** The fake adapter's fine-detail energy is kept below the teacher's at the
|
| 735 |
point where the distribution term is measured, so the difference between the two keeps pointing toward detail.
|
| 736 |
|
| 737 |
+
### Colour
|
| 738 |
+
|
| 739 |
+
A floor holds saturation at the teacher's own level β on the decoded window and on the whole image, every step β and a
|
| 740 |
+
ceiling 8% above it keeps it from climbing past. References that are grey or black-and-white are left alone, so a
|
| 741 |
+
monochrome prompt is never pushed toward colour.
|
| 742 |
+
|
| 743 |
+
### A guarded running average
|
| 744 |
+
|
| 745 |
+
The published weights are a running average of training (decay 0.999), which smooths out the noise of single steps.
|
| 746 |
+
Averaging has one failure: while the training weights move between two layouts of the same prompt, their average draws
|
| 747 |
+
both β a doubled subject. A guard watches a fixed set of layout probes every ten steps; a move away from the trend that
|
| 748 |
+
comes back within a few rounds is kept out of the average, and a lasting move is taken in whole β the guard never resets
|
| 749 |
+
the average. The average itself was restarted once, by hand, after a layout change it had blended.
|
| 750 |
+
|
| 751 |
## What the LoRA touches
|
| 752 |
|
| 753 |
Rank **64**, alpha = rank, bf16, the same **228 modules** as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA): the 8 attention and feed-forward linears
|
|
|
|
| 759 |
The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β Krea 2 Turbo's own 8-step run at mu = 1.15 and
|
| 760 |
guidance 0.0, every latent and velocity stored β serve unchanged: a 2-step chord is two of the 4-step chords end to
|
| 761 |
end. 203 held-out prompts measure the studentβteacher gap on unseen prompts and never receive a gradient. The spectral
|
| 762 |
+
match, the detail terms and the critics read the teacher's finals for the training prompts, and the face critics and the
|
| 763 |
+
windows on features use face masks computed once on those finals. The 43,044 real-photo crops of the
|
| 764 |
[4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) enter in two places: as one precomputed statistic β how much fine-detail energy they carry at 3β10
|
| 765 |
pixels relative to the teacher, clamped β which sets the photo floor, and as half of the photo critic's real examples. Only that
|
| 766 |
critic's head sees them; the student and the fake adapter never do, and receive only its filtered, capped push.
|
| 767 |
|
| 768 |
## Resolutions
|
| 769 |
|
| 770 |
+
The same 12 buckets as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Since `chk00017464` the draw leans toward the larger sizes, where two steps
|
| 771 |
+
had put the most grain β 1440Γ1440 from under 2% of the samples to about 6%, 1024Γ1024 and 768Γ1024 to 12% each β and stylised
|
| 772 |
+
prompts are drawn three times as often as their share of the pool, 12% of the samples instead of 4%:
|
| 773 |
|
| 774 |
| | | |
|
| 775 |
| --------- | --------- | --------- |
|
|
|
|
| 799 |
checkpoint published here, replaced whenever a better one ships. Which checkpoint that is today is in
|
| 800 |
[`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md), and every
|
| 801 |
published checkpoint's tree is kept under its own number in [`_archive/resolution_sweeps/`](_archive/resolution_sweeps).
|
|
|
|
|
|
|
|
|
|
| 802 |
|
| 803 |
```
|
| 804 |
assets/resolution_sweeps/
|
|
|
|
| 814 |
β βββ 1440x1280/ β¦and the remaining buckets
|
| 815 |
β βββ 1440x1440/
|
| 816 |
βββ _turbo-base-NO-LoRA-2step/ the same tree, stock Turbo at 2 steps β the floor
|
| 817 |
+
βββ 2step-LoRA/ the same tree, rendered with this LoRA at 2 steps β the published checkpoint
|
|
|
|
|
|
|
| 818 |
```
|
| 819 |
|
| 820 |
Two ways to read them, both useful:
|
|
|
|
| 858 |
## How it is judged
|
| 859 |
|
| 860 |
At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
|
| 861 |
+
15 fixed prompts across four resolutions (512Γ512, 1280Γ1280, 1440Γ1440, 1440Γ1280), after a layout check at both 1440
|
| 862 |
+
sizes that has to pass first; milestone checkpoints get the same render at all 12 buckets, which is where the
|
| 863 |
+
per-resolution table above comes from. Every image is measured against the
|
| 864 |
teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
|
| 865 |
and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
|
| 866 |
seen through their windows), straight-line artefacts, a graded judge, a pairwise preference against the teacher, and a blind rubric that
|
|
|
|
| 1185 |
sigmas are anchored to that grid.
|
| 1186 |
- π§ͺ **Not a finished adapter.** Usable at 2 steps for previews and drafts; not the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s quality, which remains the recommendation for quality renders. Training continues, and a later checkpoint replaces this file only when the sweeps and I visually agree it is better.
|
| 1187 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1188 |
## What's next
|
| 1189 |
|
| 1190 |
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any
|
| 1191 |
+
resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Next, aimed
|
| 1192 |
+
at what this checkpoint still gets wrong: small subjects in wide scenes; the finest edges and the texture of skin at the
|
| 1193 |
+
largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two
|
| 1194 |
+
plausible poses of the same subject can meet; and the tactile surface of stylised materials such as clay β each held to the
|
| 1195 |
+
rule that none of it may cost the colour and the clean large sizes this checkpoint gained. A better checkpoint
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1196 |
replaces this one when the sweeps and I visually agree, the same discipline as the
|
| 1197 |
[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
|
| 1198 |
one is the fast preview.
|
README.md
CHANGED
|
@@ -17,28 +17,26 @@ pipeline_tag: text-to-image
|
|
| 17 |
|
| 18 |
# Krea 2 Turbo β 2-Step Distillation LoRA
|
| 19 |
|
| 20 |
-
**A quarter of the steps Β· 4.2Γ faster denoising Β· fine detail at
|
| 21 |
|
| 22 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
|
| 23 |
|
| 24 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
|
| 25 |
- β±οΈ **4.2Γ faster denoising** β 81.4 s β 19.5 s at 1024Γ1024; the adapter's own cost per call is within measurement noise
|
| 26 |
-
- π― **Fine detail at
|
| 27 |
- π **Distribution matching, not imitation** β matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
|
| 28 |
- π£οΈ **Prompt-conditioned throughout** β teacher and fake scores both read each prompt's conditioning; a blind rubric finds **1 point missing of 240** (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on **11 of 45** (4-step adapter: 6), mostly on style
|
| 29 |
- π **12 trained resolutions** β multi-aspect from 512Γ512 up to 1440Γ1440
|
| 30 |
- π **Drop-in, no exceptions** β plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
|
| 31 |
- 𧬠**Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β rank 64 on the same 228 modules
|
| 32 |
- π² **13,750 recorded teacher trajectories** from the 4-step project, reused β not one new teacher run
|
| 33 |
-
- π’ **
|
| 34 |
-
- π
**
|
| 35 |
-
- π **
|
| 36 |
|
| 37 |
> π§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β a portrait, a single figure, an object up close β hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
|
| 38 |
>
|
| 39 |
> π **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ2048 β Krea's published maximum recommended resolution, beyond this adapter's largest trained size β come within easy reach. Past 2048Γ2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
|
| 40 |
-
>
|
| 41 |
-
> π **Also compatible with Krea 2 Raw** β with some prompts, at 7+ steps and light guidance. See [Using it on Raw](#using-it-on-raw).
|
| 42 |
|
| 43 |
[](assets/poster.jpg)
|
| 44 |
|
|
@@ -51,7 +49,6 @@ A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that tak
|
|
| 51 |
| `krea2_turbo_2step_rank_64_lora.safetensors` | LoRA in diffusers key format β see [diffusers](#diffusers) |
|
| 52 |
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | Same weights under ComfyUI key names β see [ComfyUI](#comfyui) |
|
| 53 |
| `krea2_turbo_2step_lora_t2i.json` | Ready ComfyUI workflow, stock nodes only |
|
| 54 |
-
| `krea2_raw_7step_lora_experiment_t2i.json` | ComfyUI workflow of the Krea 2 Raw 7-step experiment β see [Using it on Raw](#using-it-on-raw) |
|
| 55 |
| [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **Which checkpoint the two weight files are** β updated with every release |
|
| 56 |
| `LICENSE.pdf` | Krea 2 Community License Agreement |
|
| 57 |
| `NOTICE.txt` | Required attribution notice |
|
|
@@ -106,9 +103,7 @@ Load [`krea2_turbo_2step_lora_t2i.json`](krea2_turbo_2step_lora_t2i.json). Full
|
|
| 106 |
|
| 107 |
### Using it on Raw
|
| 108 |
|
| 109 |
-
Trained on Turbo, for Turbo β it loads on **Krea 2 Raw** because the architecture is shared, a side effect rather than a supported mode
|
| 110 |
-
|
| 111 |
-
Results are mixed and subject-dependent. All 15 prompts at 1024Γ768, what worked and what didn't, and the workflow ([`krea2_raw_7step_lora_experiment_t2i.json`](krea2_raw_7step_lora_experiment_t2i.json)) are in the [experiment's README](assets/resolution_sweeps/raw-LoRA-7steps-experiment/README.md).
|
| 112 |
|
| 113 |
---
|
| 114 |
|
|
@@ -122,7 +117,7 @@ Results are mixed and subject-dependent. All 15 prompts at 1024Γ768, what worke
|
|
| 122 |
|
| 123 |
**Denoising is 4.2Γ faster than the 8-step bar** β two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.
|
| 124 |
|
| 125 |
-
Prompt encoding and VAE decode don't change with step count, so end to end sits below 4.2Γ and rises toward it as the render grows. Denoise times at every trained resolution (
|
| 126 |
|
| 127 |
---
|
| 128 |
|
|
@@ -143,19 +138,20 @@ At two steps the dial scales the adapter's whole job β turning two coarse call
|
|
| 143 |
|
| 144 |
## Current Checkpoint
|
| 145 |
|
| 146 |
-
**`
|
| 147 |
|
| 148 |
-
| axis
|
| 149 |
-
| --------------------------------------------- | ------------- | --------------- |
|
| 150 |
-
| fine texture vs the teacher, 1280Β² / 1440Β²
|
| 151 |
-
| 16-px grid band, 1280Β² / 1440Β²
|
| 152 |
-
| grain in flat areas, sweep median
|
| 153 |
-
|
|
| 154 |
-
|
|
| 155 |
-
|
|
| 156 |
-
|
|
|
|
|
| 157 |
|
| 158 |
-
|
| 159 |
|
| 160 |
---
|
| 161 |
|
|
@@ -168,8 +164,7 @@ The usual costs of two steps, in order of how often they show:
|
|
| 168 |
- **Style** β on stylised prompts, _how_ the picture should look (crisp linework, brush strokes, fingerprints in clay, a matte-painting finish) is followed less faithfully than _what_ should be in it
|
| 169 |
- **Repeats** β on busy action or crowd scenes the composition can repeat itself: an extra hand or held object, a figure duplicated in a crowd
|
| 170 |
- **Different composition** β two steps is a shorter path from the same noise, so framing, pose or arrangement can differ from the 8-step render at the same seed. Treat the teacher's render as a quality reference, not the picture two steps will reproduce
|
| 171 |
-
- **Skin
|
| 172 |
-
- **Grain** β a fine grain remains on the most textured subjects at the largest sizes, lighter than in the previous checkpoint
|
| 173 |
|
| 174 |
Every one is being worked on; none is hidden in the sweeps or the examples.
|
| 175 |
|
|
@@ -305,26 +300,11 @@ assets/resolution_sweeps/
|
|
| 305 |
β βββ 1024x1024/
|
| 306 |
β βββ ... (all 12 buckets)
|
| 307 |
βββ _turbo-base-NO-LoRA-2step/ Same tree, stock Turbo at 2 steps (the floor)
|
| 308 |
-
|
| 309 |
-
βββ _turbo-base-NO-LoRA-1step/ Stock Turbo at 1 step (bonus section)
|
| 310 |
-
βββ 1step-LoRA-extreme/ This LoRA at 1 step, with side-by-side strips (bonus section)
|
| 311 |
-
βββ raw-LoRA-7steps-experiment/ Krea 2 Raw + this LoRA at 7 steps (Using it on Raw)
|
| 312 |
```
|
| 313 |
|
| 314 |
---
|
| 315 |
|
| 316 |
-
## Bonus: 1-Step Extreme Test
|
| 317 |
-
|
| 318 |
-
> **Out-of-spec experiment β not recommended for any use.** Trained for 2 steps; at 1 step it runs half its trained count and an eighth of the teacher's.
|
| 319 |
-
|
| 320 |
-
Stock Turbo returns a smear at one step β a colour field with a ghost of the subject. With the LoRA the same single call is a coherent picture: subject, composition, lighting and colours all there. Missing is the detail the second step adds β soft skin, streaked hair and fur, little of the finest structure (feathers, falling snow, small text). A rough **preview of composition and colour** at an eighth of the teacher's cost, nothing more.
|
| 321 |
-
|
| 322 |
-
Full set at all 12 resolutions, render + side-by-side strip per prompt: [`assets/resolution_sweeps/1step-LoRA-extreme/`](assets/resolution_sweeps/1step-LoRA-extreme/). Stock Turbo at 1 step: [`_turbo-base-NO-LoRA-1step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-1step/)
|
| 323 |
-
|
| 324 |
-
> π **A 1-step adapter is a possible follow-on** β a booster on top of this LoRA, trained by distribution matching alone and judged on seed variety as much as detail. Expectation: a usable preview at an eighth of the teacher's cost, not the quality bar.
|
| 325 |
-
|
| 326 |
-
---
|
| 327 |
-
|
| 328 |
## Archive
|
| 329 |
|
| 330 |
Every published checkpoint and its resolution sweep under [`_archive/`](_archive/) β [`checkpoints/`](_archive/checkpoints/) and [`resolution_sweeps/`](_archive/resolution_sweeps/), each under its number. Superseded, not maintained.
|
|
@@ -335,7 +315,7 @@ Every published checkpoint and its resolution sweep under [`_archive/`](_archive
|
|
| 335 |
|
| 336 |
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
|
| 337 |
|
| 338 |
-
Next, aimed at the [Known Issues](#known-issues):
|
| 339 |
|
| 340 |
A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
|
| 341 |
|
|
|
|
| 17 |
|
| 18 |
# Krea 2 Turbo β 2-Step Distillation LoRA
|
| 19 |
|
| 20 |
+
**A quarter of the steps Β· 4.2Γ faster denoising Β· fine detail at 0.95β1.08Γ the teacher's across all 12 trained resolutions Β· colour at the teacher's level Β· 1 point missing of 240 on a blind prompt-adherence rubric Β· teacher preferred on 11 of 45 judged renders Β· 31,600 training samples on the 4-step project's recorded trajectories Β· 18 days on one RTX 3090 Β· still in training.**
|
| 21 |
|
| 22 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
|
| 23 |
|
| 24 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
|
| 25 |
- β±οΈ **4.2Γ faster denoising** β 81.4 s β 19.5 s at 1024Γ1024; the adapter's own cost per call is within measurement noise
|
| 26 |
+
- π― **Fine detail at the teacher's level** β **0.95β1.08Γ** the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: **0.39β0.57Γ**); from 1 megapixel up, closer to the teacher than the 4-step adapter
|
| 27 |
- π **Distribution matching, not imitation** β matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
|
| 28 |
- π£οΈ **Prompt-conditioned throughout** β teacher and fake scores both read each prompt's conditioning; a blind rubric finds **1 point missing of 240** (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on **11 of 45** (4-step adapter: 6), mostly on style
|
| 29 |
- π **12 trained resolutions** β multi-aspect from 512Γ512 up to 1440Γ1440
|
| 30 |
- π **Drop-in, no exceptions** β plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
|
| 31 |
- 𧬠**Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β rank 64 on the same 228 modules
|
| 32 |
- π² **13,750 recorded teacher trajectories** from the 4-step project, reused β not one new teacher run
|
| 33 |
+
- π’ **31,600 training samples** in the 2-step stages, on top of the 4-step LoRA's 78,000
|
| 34 |
+
- π
**18 days** from the first 2-step launch to this checkpoint, on a single RTX 3090 β training continues
|
| 35 |
+
- π **More than forty recipe adjustments** across two methods β each kept only when the renders did not get worse
|
| 36 |
|
| 37 |
> π§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β a portrait, a single figure, an object up close β hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
|
| 38 |
>
|
| 39 |
> π **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ2048 β Krea's published maximum recommended resolution, beyond this adapter's largest trained size β come within easy reach. Past 2048Γ2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
|
|
|
|
|
|
|
| 40 |
|
| 41 |
[](assets/poster.jpg)
|
| 42 |
|
|
|
|
| 49 |
| `krea2_turbo_2step_rank_64_lora.safetensors` | LoRA in diffusers key format β see [diffusers](#diffusers) |
|
| 50 |
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | Same weights under ComfyUI key names β see [ComfyUI](#comfyui) |
|
| 51 |
| `krea2_turbo_2step_lora_t2i.json` | Ready ComfyUI workflow, stock nodes only |
|
|
|
|
| 52 |
| [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) | **Which checkpoint the two weight files are** β updated with every release |
|
| 53 |
| `LICENSE.pdf` | Krea 2 Community License Agreement |
|
| 54 |
| `NOTICE.txt` | Required attribution notice |
|
|
|
|
| 103 |
|
| 104 |
### Using it on Raw
|
| 105 |
|
| 106 |
+
Trained on Turbo, for Turbo β it loads on **Krea 2 Raw** because the architecture is shared, a side effect rather than a supported mode: it is neither trained nor tuned for Raw's weights, steps or guidance.
|
|
|
|
|
|
|
| 107 |
|
| 108 |
---
|
| 109 |
|
|
|
|
| 117 |
|
| 118 |
**Denoising is 4.2Γ faster than the 8-step bar** β two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.
|
| 119 |
|
| 120 |
+
Prompt encoding and VAE decode don't change with step count, so end to end sits below 4.2Γ and rises toward it as the render grows. Denoise times at every trained resolution (7.0 s at 512Γ512 to 39.6 s at 1440Γ1440) are in the [Detailed Model Card](DETAILED-README.md).
|
| 121 |
|
| 122 |
---
|
| 123 |
|
|
|
|
| 138 |
|
| 139 |
## Current Checkpoint
|
| 140 |
|
| 141 |
+
**`chk00031600`** (25 Sep 2026) replaces `chk00017464` (14 Sep 2026). It is 14,136 training samples later, aimed at what the previous checkpoint still got wrong β grain and excess texture at the largest sizes, colour under the teacher's there, small faces, stylised prompts β through a colour band at the teacher's level, nine critics taking turns, detail held to the teacher region by region, more training at the large sizes and a guarded running average.
|
| 142 |
|
| 143 |
+
| axis | `chk00017464` | `chk00031600` |
|
| 144 |
+
| ------------------------------------------------------ | ------------- | --------------- |
|
| 145 |
+
| fine texture vs the teacher, 1280Β² / 1440Β² | 1.09 / 1.17 | **0.95 / 1.08** |
|
| 146 |
+
| 16-px grid band, 1280Β² / 1440Β² | 1.08 / 1.08 | **0.95 / 1.05** |
|
| 147 |
+
| grain in flat areas, sweep median | 1.26Γ | **1.08Γ** |
|
| 148 |
+
| saturation vs the teacher, sweep mean | 0.93Γ | **0.98Γ** |
|
| 149 |
+
| small faces keeping their structure (the teacher: 64%) | 53% | **63%** |
|
| 150 |
+
| judge prefers the teacher (of 45) | 11 | 11 |
|
| 151 |
+
| blind adherence rubric, points missing of 240 | 1 | 1 |
|
| 152 |
+
| distance to the teacher, sweep mean | **0.406** | 0.413 |
|
| 153 |
|
| 154 |
+
The excess fine energy at large sizes is gone and colour is at the teacher's level at every size; small faces keep their structure nearly as often as the teacher's, and prompt-following held level. [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) always names the checkpoint in the weight files; the full comparison is in the [Detailed Model Card](DETAILED-README.md).
|
| 155 |
|
| 156 |
---
|
| 157 |
|
|
|
|
| 164 |
- **Style** β on stylised prompts, _how_ the picture should look (crisp linework, brush strokes, fingerprints in clay, a matte-painting finish) is followed less faithfully than _what_ should be in it
|
| 165 |
- **Repeats** β on busy action or crowd scenes the composition can repeat itself: an extra hand or held object, a figure duplicated in a crowd
|
| 166 |
- **Different composition** β two steps is a shorter path from the same noise, so framing, pose or arrangement can differ from the 8-step render at the same seed. Treat the teacher's render as a quality reference, not the picture two steps will reproduce
|
| 167 |
+
- **Skin** β smoother than the teacher's at the larger sizes, the pores softer, though its colour is now close to the teacher's; freckles come out as dots, softer than the teacher's
|
|
|
|
| 168 |
|
| 169 |
Every one is being worked on; none is hidden in the sweeps or the examples.
|
| 170 |
|
|
|
|
| 300 |
β βββ 1024x1024/
|
| 301 |
β βββ ... (all 12 buckets)
|
| 302 |
βββ _turbo-base-NO-LoRA-2step/ Same tree, stock Turbo at 2 steps (the floor)
|
| 303 |
+
βββ 2step-LoRA/ Same tree, this LoRA at 2 steps (the published checkpoint)
|
|
|
|
|
|
|
|
|
|
| 304 |
```
|
| 305 |
|
| 306 |
---
|
| 307 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 308 |
## Archive
|
| 309 |
|
| 310 |
Every published checkpoint and its resolution sweep under [`_archive/`](_archive/) β [`checkpoints/`](_archive/checkpoints/) and [`resolution_sweeps/`](_archive/resolution_sweeps/), each under its number. Superseded, not maintained.
|
|
|
|
| 315 |
|
| 316 |
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
|
| 317 |
|
| 318 |
+
Next, aimed at the [Known Issues](#known-issues): small subjects in wide scenes; the finest edges and skin texture at the largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two plausible poses of the same subject can meet; the tactile surface of stylised materials such as clay β none of it allowed to cost the colour and the clean large sizes this checkpoint gained.
|
| 319 |
|
| 320 |
A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
|
| 321 |
|
krea2_turbo_2step_rank_64_lora.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:1df55a05e4ed3cd367e3cc06c84279d9dd2430394119b81038b9ad30c3b17952
|
| 3 |
+
size 438161624
|
krea2_turbo_2step_rank_64_lora_checkpoint_info.md
CHANGED
|
@@ -1,13 +1,13 @@
|
|
| 1 |
# Which checkpoint is this?
|
| 2 |
|
| 3 |
`krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
|
| 4 |
-
**
|
| 5 |
`_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
|
| 6 |
keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
|
| 7 |
|
| 8 |
| file | SHA-256 | size |
|
| 9 |
| --- | --- | --- |
|
| 10 |
-
| `krea2_turbo_2step_rank_64_lora.safetensors` | `
|
| 11 |
-
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `
|
| 12 |
|
| 13 |
-
Updated:
|
|
|
|
| 1 |
# Which checkpoint is this?
|
| 2 |
|
| 3 |
`krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
|
| 4 |
+
**chk00031600** β the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00031600.safetensors` (and its
|
| 5 |
`_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
|
| 6 |
keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
|
| 7 |
|
| 8 |
| file | SHA-256 | size |
|
| 9 |
| --- | --- | --- |
|
| 10 |
+
| `krea2_turbo_2step_rank_64_lora.safetensors` | `1df55a05e4ed3cd367e3cc06c84279d9dd2430394119b81038b9ad30c3b17952` | 418M |
|
| 11 |
+
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `e017280bc21c5279156adb0daf6654238896ca1ee0c12c8459b66b882edcfca4` | 418M |
|
| 12 |
|
| 13 |
+
Updated: 25 Sep 2026
|
krea2_turbo_2step_rank_64_lora_comfyui.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:e017280bc21c5279156adb0daf6654238896ca1ee0c12c8459b66b882edcfca4
|
| 3 |
+
size 438142160
|