Instructions to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use lvladikov/Krea2-Turbo-Distill-2step-LoRA with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("krea/Krea-2-Turbo", dtype=torch.bfloat16, device_map="cuda") pipe.load_lora_weights("lvladikov/Krea2-Turbo-Distill-2step-LoRA") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
chk41320 release
Browse files
DETAILED-README.md
CHANGED
|
@@ -43,8 +43,8 @@ recommendation for quality renders.
|
|
| 43 |
- π― **The aim** β the best two-step quality this base can give, at every one of the same 12 resolutions, measured
|
| 44 |
against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
|
| 45 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas.
|
| 46 |
-
- β±οΈ **4
|
| 47 |
-
changes: **
|
| 48 |
measurement noise. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
|
| 49 |
[Performance](#performance).
|
| 50 |
- π **Distribution matching, not imitation** β the training objective that got the renders improving again after the
|
|
@@ -64,17 +64,17 @@ recommendation for quality renders.
|
|
| 64 |
only and never ships.
|
| 65 |
- π² **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
|
| 66 |
re-run.
|
| 67 |
-
- π’ **
|
| 68 |
drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
|
| 69 |
training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
|
| 70 |
4-step project, read again at the two sigmas this schedule uses.
|
| 71 |
-
- π
**
|
| 72 |
- π **More than forty recipe adjustments** across two methods so far β seven of trajectory distillation before the switch, the rest of distribution matching since β each kept only when the renders did not get worse.
|
| 73 |
- π₯οΈ **One RTX 3090**, and a recipe shaped by its 24 GB.
|
| 74 |
|
| 75 |
-
[._
|
| 80 |
|
|
@@ -99,9 +99,9 @@ place to check which checkpoint the current files are based on.
|
|
| 99 |
|
| 100 |
| | |
|
| 101 |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 102 |
-
| lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β 2-step trajectory distillation β distribution matching β a spectral match against the teacher's own images β an artefact critic and detail terms β three critics taking turns β nine critics, detail and colour held to the teacher region by region, a guarded running average |
|
| 103 |
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
|
| 104 |
-
| what it gives | usable two-step renders at every trained resolution: fine detail and colour at the teacher's level β from 1 megapixel up, closer to the teacher than the 4-step adapter β with the prompt's objects, counts, attributes and relations in place (a blind rubric finds
|
| 105 |
|
| 106 |
### Known issues
|
| 107 |
|
|
@@ -183,11 +183,96 @@ The recipe adjustments so far, each made on the measurement of the one before:
|
|
| 183 |
17. **a guarded running average** β the published weights are a running average of training, and averaging two layouts of
|
| 184 |
the same prompt had produced doubled subjects: a short-lived excursion of the training weights is now kept out of the
|
| 185 |
average and a lasting change taken in whole, and the average was restarted once, after a layout change it had blended
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 186 |
|
| 187 |
Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
|
| 188 |
a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
|
| 189 |
not place exactly where the teacher had it.
|
| 190 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 191 |
## chk00031600 vs chk00017464
|
| 192 |
|
| 193 |
`chk00017464` (published 14 Sep 2026) was the first checkpoint with critics and detail terms. `chk00031600` (25 Sep 2026) is
|
|
@@ -315,51 +400,51 @@ point missing out of 240, and on the same 15 fresh prompts the previous checkpoi
|
|
| 315 |
## Measured against the teacher
|
| 316 |
|
| 317 |
Every number here compares a render of this LoRA at 2 steps with the 8-step reference render of the **same prompt at the
|
| 318 |
-
same seed**, across the
|
| 319 |
the figures have a floor and a ceiling around them: **stock Krea 2 Turbo at 2 steps**, which is what the base model does
|
| 320 |
without the adapter, and the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) at its own 4 steps, which is the better tool and the thing worth
|
| 321 |
-
being compared against.
|
| 322 |
|
| 323 |
**Detail, band by band.** Fine-texture energy and the two grid bands, as a ratio to the teacher's own (1.00 = the
|
| 324 |
teacher):
|
| 325 |
|
| 326 |
-
| bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three) | stock 2-step, fine texture |
|
| 327 |
| --- | --- | --- | --- | --- | --- |
|
| 328 |
-
| 512Γ512 | 1.00 | 1.
|
| 329 |
-
| 768Γ1024 |
|
| 330 |
-
| 1024Γ1024 | 1.
|
| 331 |
-
| 1280Γ1280 |
|
| 332 |
-
| 1440Γ1440 | 1.
|
| 333 |
|
| 334 |
-
Two steps without the adapter carry 0.
|
| 335 |
-
sits at the teacher's level everywhere β between 0.
|
| 336 |
-
than the 4-step adapter, which carries more excess fine energy there.
|
| 337 |
|
| 338 |
**Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
|
| 339 |
in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
|
| 340 |
|
| 341 |
| bucket | wins | ties | losses |
|
| 342 |
| --- | --- | --- | --- |
|
| 343 |
-
| 512Γ512 |
|
| 344 |
-
| 1280Γ1280 | 0 |
|
| 345 |
-
| 1440Γ1440 | 0 |
|
| 346 |
-
|
| 347 |
-
|
| 348 |
-
stylised prompts and on how a prompt says the picture should look β crisp linework, energetic brush strokes, the texture
|
| 349 |
-
clay β more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
|
| 350 |
-
attributes and relations, with the teacher scored identically, finds
|
| 351 |
-
training prompt bank for this checkpoint and never rendered before, the judge returned
|
| 352 |
-
black-and-white prompts 0 wins, 4 ties, 1 loss β with no colour cast in any of the five.
|
| 353 |
-
|
| 354 |
-
**Checked for the damage this kind of training can do.** Saturation sits at 1.04Γ the teacher's at 768Γ1024 and 0.
|
| 355 |
-
|
| 356 |
-
1280Γ1280 / 1440Γ1440, with skin saturation 0.
|
| 357 |
-
|
| 358 |
-
|
| 359 |
-
|
| 360 |
-
|
| 361 |
-
**Distance to the teacher**, as a plain pixel measure, is 0.39β0.
|
| 362 |
-
0.30β0.37. That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
|
| 363 |
not the teacher's render of it β see [Known issues](#known-issues).
|
| 364 |
|
| 365 |
## Usage
|
|
@@ -497,28 +582,28 @@ the model already loaded, so the numbers are the render itself and not a model l
|
|
| 497 |
|
| 498 |
| configuration | denoising | per model call | GPU peak |
|
| 499 |
| --- | --- | --- | --- |
|
| 500 |
-
| Krea 2 Turbo β 8 steps (the reference) | **
|
| 501 |
-
| Krea 2 Turbo β 2 steps, no LoRA |
|
| 502 |
-
| **Krea 2 Turbo β 2 steps + this LoRA** | **19.
|
| 503 |
|
| 504 |
-
**Denoising is 4.
|
| 505 |
not show up in this measurement: the runs with it came in marginally faster than those without, which is measurement noise, not a
|
| 506 |
speed-up. A rank-64 low-rank product is small beside the transformer it is added to, and it adds no measurable memory.
|
| 507 |
|
| 508 |
Denoising is the part the step count changes. What a complete render costs on top of it β encoding the prompt, decoding
|
| 509 |
the latent, writing the file β is the same whether you run two steps or eight, and it depends on your pipeline, so the
|
| 510 |
-
end-to-end figure on your machine will sit below 4
|
| 511 |
|
| 512 |
-
**By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the
|
| 513 |
size; sweep renders run one at a time, not the controlled measurement above, so a second or two either way between one sweep and
|
| 514 |
the next is run-to-run variation β the adapter's shape, and so its cost, is the same at every checkpoint):
|
| 515 |
|
| 516 |
| resolution | denoising (2 calls) | resolution | denoising (2 calls) |
|
| 517 |
| --- | --- | --- | --- |
|
| 518 |
-
| 512Γ512 |
|
| 519 |
-
| 512Γ768 / 768Γ512 |
|
| 520 |
-
| 768Γ768 | 12.
|
| 521 |
-
| 768Γ1024 / 1024Γ768 | 15.1 s / 15.
|
| 522 |
|
| 523 |
## LoRA strength
|
| 524 |
|
|
@@ -621,23 +706,23 @@ and under-fits fine structure there, and a per-pixel normaliser lands harder as
|
|
| 621 |
sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
|
| 622 |
1.4β1.8Γ the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
|
| 623 |
not hold, and the spectral match replaced it. The same sweep of the published checkpoint, its running-average weights, fixed
|
| 624 |
-
seed,
|
| 625 |
-
first distribution-matching checkpoint on the
|
| 626 |
|
| 627 |
| bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
|
| 628 |
| --------- | --------------------------- | --------------- | -------------- | ----------------------- |
|
| 629 |
-
| 512x512 | 1.00 (1.16) | 1.
|
| 630 |
-
| 512x768 |
|
| 631 |
-
| 768x512 |
|
| 632 |
-
| 768x768 | 1.
|
| 633 |
-
| 768x1024 |
|
| 634 |
-
| 1024x768 |
|
| 635 |
-
| 1024x1024 | 1.
|
| 636 |
-
| 1280x960 | 0.
|
| 637 |
-
| 960x1280 | 1.
|
| 638 |
-
| 1280x1280 |
|
| 639 |
-
| 1440x1280 | 0.
|
| 640 |
-
| 1440x1440 | 1.
|
| 641 |
|
| 642 |
### The spectral match
|
| 643 |
|
|
@@ -680,7 +765,7 @@ teacher's finishing pass below take turns instead of sharing a step.
|
|
| 680 |
|
| 681 |
### Critics in turn
|
| 682 |
|
| 683 |
-
One critic holds one idea of what is wrong. The artefact critic is joined by
|
| 684 |
features and under the same rules β lightly re-noised inputs, a push that is filtered and capped β each aimed at a
|
| 685 |
different fault:
|
| 686 |
|
|
@@ -688,9 +773,9 @@ different fault:
|
|
| 688 |
what fine texture looks like in a photograph as well as in the teacher's rendering of one. It judges photographic prompts
|
| 689 |
only, so illustration, anime and 3D renders are not pulled toward photographic grain, and its push is filtered to periods
|
| 690 |
finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher does not.
|
| 691 |
-
- **
|
| 692 |
-
|
| 693 |
-
|
| 694 |
- **A structure critic.** It judges the first call β the layout, before any detail β against the teacher's own intermediate
|
| 695 |
state from the same noise, at the high noise levels where layout is decided and at periods of 32 pixels and coarser only,
|
| 696 |
so a first call that blends two layouts is caught where the blend happens.
|
|
@@ -701,7 +786,7 @@ different fault:
|
|
| 701 |
- **A rollout critic.** Its real examples are the teacher's own finish from the student's second-call starting point, so the
|
| 702 |
second call is judged against what the teacher would have made from the same start.
|
| 703 |
|
| 704 |
-
All
|
| 705 |
others train so none goes stale before its turn comes back. Each new head started from the artefact critic's weights and
|
| 706 |
trained on its own before it was allowed to push.
|
| 707 |
|
|
@@ -716,11 +801,8 @@ Smaller terms sit on top, each capped relative to the distribution term so none
|
|
| 716 |
pixels may not fall below the teacher's plus the margin real photographs carry over it at those scales, judged tile by tile
|
| 717 |
on the window's textured tiles. That margin is measured once from a pool of real photographs and clamped, and the term
|
| 718 |
only ever pushes upward to that floor, never past it β so it lifts detail that is missing without adding grain that is not.
|
| 719 |
-
- **
|
| 720 |
-
|
| 721 |
-
above it.
|
| 722 |
-
- **A ceiling.** Tile by tile, detail at 3β16 pixels may not climb past 1.3Γ the teacher's β the counterpart of the floors,
|
| 723 |
-
and what keeps grain from building up at large sizes.
|
| 724 |
- **A direction-aware term.** The spectral comparison is also made orientation by orientation, so the student's fine detail
|
| 725 |
runs in the same directions as the teacher's.
|
| 726 |
- **Windows on features.** On photographs, half of the decoded windows are centred on an eye, the nose or the lips of a face,
|
|
@@ -745,8 +827,21 @@ monochrome prompt is never pushed toward colour.
|
|
| 745 |
The published weights are a running average of training (decay 0.999), which smooths out the noise of single steps.
|
| 746 |
Averaging has one failure: while the training weights move between two layouts of the same prompt, their average draws
|
| 747 |
both β a doubled subject. A guard watches a fixed set of layout probes every ten steps; a move away from the trend that
|
| 748 |
-
comes back within
|
| 749 |
-
the average.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 750 |
|
| 751 |
## What the LoRA touches
|
| 752 |
|
|
@@ -786,8 +881,9 @@ strong at the size it saw most while quietly softer at the ones it barely saw. O
|
|
| 786 |
which is why every cut of this run is rendered across the buckets and why the per-resolution table under
|
| 787 |
[Method](#method) above exists.
|
| 788 |
|
| 789 |
-
[`assets/resolution_sweeps/`](assets/resolution_sweeps) holds the evidence β same
|
| 790 |
-
|
|
|
|
| 791 |
|
| 792 |
- [`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step) β the **official Krea 2 Turbo 8-step reference renders**:
|
| 793 |
the stock model, no LoRA, at its native settings (8 steps, guidance 0.0). The teacher this LoRA is distilled from and
|
|
@@ -806,8 +902,8 @@ assets/resolution_sweeps/
|
|
| 806 |
β βββ 512x512/ one folder per resolution
|
| 807 |
β β βββ portrait.jpg
|
| 808 |
β β βββ kingfisher.jpg
|
| 809 |
-
β β βββ β¦
|
| 810 |
-
β β βββ
|
| 811 |
β βββ 768x1024/
|
| 812 |
β βββ 1024x1024/
|
| 813 |
β βββ 1280x1280/
|
|
@@ -834,15 +930,15 @@ One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's che
|
|
| 834 |
host memory above 0.3 megapixels; the student, the fake adapter, the spectral and detail terms, the critic and the teacher's finishing pass each
|
| 835 |
build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
|
| 836 |
does not fit fails loudly. A full step with every term live and every critic pushing reserves about 22.4 GB at 1440Γ1440,
|
| 837 |
-
of 24. The price of the objective is throughput: **about
|
| 838 |
-
|
| 839 |
-
fraction of the samples.
|
| 840 |
|
| 841 |
**Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
|
| 842 |
The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
|
| 843 |
disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
|
| 844 |
of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
|
| 845 |
-
alongside to imitate the student β and that second adapter takes
|
| 846 |
The spectral and detail terms decode part of the image out of the latent to compare its texture with the teacher's,
|
| 847 |
each critic reads half the network twice more, and every second step the teacher finishes the image from the student's
|
| 848 |
first call.
|
|
@@ -858,12 +954,14 @@ visibly better than the last again.
|
|
| 858 |
## How it is judged
|
| 859 |
|
| 860 |
At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
|
| 861 |
-
|
| 862 |
sizes that has to pass first; milestone checkpoints get the same render at all 12 buckets, which is where the
|
| 863 |
per-resolution table above comes from. Every image is measured against the
|
| 864 |
teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
|
| 865 |
and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
|
| 866 |
-
seen through their windows), straight-line artefacts, a
|
|
|
|
|
|
|
| 867 |
scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
|
| 868 |
identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
|
| 869 |
checkpoint. Latent distances β the held-out chord gap and the two-step rollout
|
|
@@ -1179,6 +1277,139 @@ individual renders β click any image for full size.
|
|
| 1179 |
|
| 1180 |
---
|
| 1181 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1182 |
## Notes and limitations
|
| 1183 |
|
| 1184 |
- π― **Krea 2 Turbo only**, at **2 steps**, guidance **0.0** (cfg 1.0 in ComfyUI), **mu = 1.15** β the two training
|
|
@@ -1191,8 +1422,9 @@ Training continues from this checkpoint, one recipe change at a time, each kept
|
|
| 1191 |
resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Next, aimed
|
| 1192 |
at what this checkpoint still gets wrong: small subjects in wide scenes; the finest edges and the texture of skin at the
|
| 1193 |
largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two
|
| 1194 |
-
plausible poses of the same subject can meet;
|
| 1195 |
-
rule that none of it may cost the colour and the clean large
|
|
|
|
| 1196 |
replaces this one when the sweeps and I visually agree, the same discipline as the
|
| 1197 |
[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
|
| 1198 |
one is the fast preview.
|
|
|
|
| 43 |
- π― **The aim** β the best two-step quality this base can give, at every one of the same 12 resolutions, measured
|
| 44 |
against the 8-step teacher and against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) as the reference. Not a claim to reach either.
|
| 45 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas.
|
| 46 |
+
- β±οΈ **4Γ faster denoising** β the model runs twice instead of eight times, and denoising is the part this adapter
|
| 47 |
+
changes: **76.4 s β 19.3 s** measured at 1024Γ1024 on the same prompts, the adapter's own cost per call within
|
| 48 |
measurement noise. What a whole render costs on top of that is unchanged by the LoRA and depends on your pipeline; see
|
| 49 |
[Performance](#performance).
|
| 50 |
- π **Distribution matching, not imitation** β the training objective that got the renders improving again after the
|
|
|
|
| 64 |
only and never ships.
|
| 65 |
- π² **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
|
| 66 |
re-run.
|
| 67 |
+
- π’ **41,320 training samples** in the 2-step stages, on top of the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 78,000 β all of them
|
| 68 |
drawn from the **same recorded material**: no new prompts, no new text embeddings and not one new teacher run. A
|
| 69 |
training sample is one pass over a prompt that was already encoded and already traced by the teacher for the
|
| 70 |
4-step project, read again at the two sigmas this schedule uses.
|
| 71 |
+
- π
**25 days** from the first 2-step training launch to this checkpoint, on a single RTX 3090 β and the project continues.
|
| 72 |
- π **More than forty recipe adjustments** across two methods so far β seven of trajectory distillation before the switch, the rest of distribution matching since β each kept only when the renders did not get worse.
|
| 73 |
- π₯οΈ **One RTX 3090**, and a recipe shaped by its 24 GB.
|
| 74 |
|
| 75 |
+
[](assets/poster.jpg)
|
| 76 |
|
| 77 |
+
_All of the above were created with this LoRA at 2 steps: the 22 test prompts, Krea 2 Turbo + the
|
| 78 |
LoRA, seed 4242, each at one of its trained resolutions. Click for full size. The side-by-side
|
| 79 |
comparisons with the 8-step teacher are in [Examples](#examples)._
|
| 80 |
|
|
|
|
| 99 |
|
| 100 |
| | |
|
| 101 |
| ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
| 102 |
+
| lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β 2-step trajectory distillation β distribution matching β a spectral match against the teacher's own images β an artefact critic and detail terms β three critics taking turns β nine critics, detail and colour held to the teacher region by region, a guarded running average β the teacher's own layout held at the large sizes, eight critics, every update to the average checked for doubled subjects |
|
| 103 |
| this release | the current run's latest probed checkpoint, chosen by the 12-bucket sweep and by my own look at the renders; the run continues from it one recipe change at a time |
|
| 104 |
+
| what it gives | usable two-step renders at every trained resolution: fine detail and colour at the teacher's level β from 1 megapixel up, closer to the teacher than the 4-step adapter β with the prompt's objects, counts, attributes and relations in place (a blind rubric finds 2 points missing out of 352, the teacher's own renders 1). A judge asked which render follows the prompt better still prefers the 8-step teacher on 12 of 66 (the 4-step adapter: 6 of 45, on the original 15 prompts), mostly on how a stylised prompt says things should look. What it does not give is the teacher's own picture: see [Known issues](#known-issues) and [Measured against the teacher](#measured-against-the-teacher) |
|
| 105 |
|
| 106 |
### Known issues
|
| 107 |
|
|
|
|
| 183 |
17. **a guarded running average** β the published weights are a running average of training, and averaging two layouts of
|
| 184 |
the same prompt had produced doubled subjects: a short-lived excursion of the training weights is now kept out of the
|
| 185 |
average and a lasting change taken in whole, and the average was restarted once, after a layout change it had blended
|
| 186 |
+
18. **three of the newest additions taken back out** β the floors at the teacher's level for stylised prompts and for flat areas,
|
| 187 |
+
and a face critic for photographs, all of which had trained the last stretch of the previous checkpoint: a render of the
|
| 188 |
+
training weights drew one animal as two joined bodies, and the recipe went back to the one that had trained the weeks before,
|
| 189 |
+
with eight critics
|
| 190 |
+
19. **a stricter guard on the running average** β a change in the training weights is taken into the average only once it has held
|
| 191 |
+
for six rounds instead of three; shorter swings are left out
|
| 192 |
+
20. **the second call held to the teacher's layout** β the anchor to the teacher's recorded trajectory counts the coarse layout of
|
| 193 |
+
the second call, everything 64 pixels and up, twice
|
| 194 |
+
21. **every update to the running average checked for doubled subjects first** β before a stretch of training goes into the
|
| 195 |
+
average, the average it would make renders a fixed probe at the large sizes, and a stretch that would put a second head on the
|
| 196 |
+
subject is held out
|
| 197 |
+
22. **the distribution term halved at the large sizes** β at every size with a 1280- or 1440-pixel side, while the anchor to the
|
| 198 |
+
teacher's trajectory keeps its full weight there, so at those sizes the teacher's own layout carries twice the share it did;
|
| 199 |
+
the doubled subjects had formed at exactly the sizes trained most
|
| 200 |
+
23. the running average rebuilt once more, from the training weights since the halving, after the guard had held it still through
|
| 201 |
+
a long change of pose
|
| 202 |
|
| 203 |
Two further ideas were tried and taken back out: confining the distribution term to the second call's noise range, and
|
| 204 |
a detail pyramid compared pixel by pixel against the teacher, which on inspection rewarded fading any detail it could
|
| 205 |
not place exactly where the teacher had it.
|
| 206 |
|
| 207 |
+
## chk00041320 vs chk00031600
|
| 208 |
+
|
| 209 |
+
`chk00031600` (published 25 Sep 2026) was the first checkpoint with a colour band, nine critics and a guarded running average.
|
| 210 |
+
`chk00041320` (2 Oct 2026) is 9,720 training samples later, and those samples went to what matters first in a picture β subjects
|
| 211 |
+
drawn twice, or two poses blended into one, at the large sizes β and to faces in busy scenes, through the changes numbered 18 to 23
|
| 212 |
+
under [How I got here](#how-i-got-here), each kept only after its own look at the renders:
|
| 213 |
+
|
| 214 |
+
1. three of the newest additions taken back out, eight critics
|
| 215 |
+
2. a stricter guard on the running average
|
| 216 |
+
3. the second call held to the teacher's layout
|
| 217 |
+
4. every update to the running average checked for doubled subjects
|
| 218 |
+
5. the distribution term halved at the large sizes
|
| 219 |
+
6. the running average rebuilt from the recent training
|
| 220 |
+
|
| 221 |
+
It is also the first checkpoint measured on the full 22-prompt sweep: the original 15, plus seven scenes of people, animals and
|
| 222 |
+
action added on 29 Sep 2026 because structure is what they test β five friends on a beach, a family at a table seen from above, three
|
| 223 |
+
kittens in a basket, two dogs in a tug of war, a show jumper, a pianist's hands, a flock of flamingos. Every figure below compares the
|
| 224 |
+
two checkpoints on those 22, each against the 8-step teacher.
|
| 225 |
+
|
| 226 |
+
**Structure β the headline.** A judge compares each of the 21 structure renders β the seven scenes at 1280Γ1280, 1440Γ1440 and
|
| 227 |
+
1440Γ1280 β with the teacher's for missing, extra or merged body parts and subjects, and every flag is checked by eye at 1:1. The
|
| 228 |
+
previous checkpoint has one real fault among them β a ghost saddle pad and boot behind the show jumper's neck at 1280Γ1280 β and this
|
| 229 |
+
one has none, and the snow leopard the guard watches has come out as one animal at every large size in every checkpoint since the
|
| 230 |
+
average was rebuilt. Doubling at the level of fine structure falls too: the two ghosting indexes come closer to the teacher at 11 and
|
| 231 |
+
10 of the 12 sizes, most at the large ones β at 1440Γ1440 from 1.17Γ the teacher's to 1.06Γ.
|
| 232 |
+
|
| 233 |
+
**Faces in busy scenes.** The market scene at 1280Γ1280 β vendors leaning over a stall β now comes out with whole faces where the
|
| 234 |
+
previous checkpoint drew a broken one, and its layout sits much closer to the teacher's (a correlation of the stall side with the
|
| 235 |
+
teacher's render: 0.71 β 0.79). Inside the faces the sweep detects, skin texture comes closer to the teacher's at 1280Γ1280 and
|
| 236 |
+
1440Γ1440 (0.89Γ β 0.90Γ, 0.83Γ β 0.87Γ), and skin colour at 1280Γ1280 rises from 0.86Γ of the teacher's to 0.95Γ.
|
| 237 |
+
|
| 238 |
+
**Detail and colour.** At the two largest sizes the excess energy two steps put into the grid bands comes down toward the
|
| 239 |
+
teacher's level:
|
| 240 |
+
|
| 241 |
+
| 1.00 = the teacher | fine texture | 16-px band | 8-px band |
|
| 242 |
+
| --- | --- | --- | --- |
|
| 243 |
+
| 1440Γ1280 | 1.01 β 0.99 | 1.03 β **1.00** | 1.10 β **1.09** |
|
| 244 |
+
| 1440Γ1440 | 1.07 β **1.03** | 1.08 β **1.05** | 1.13 β **1.09** |
|
| 245 |
+
|
| 246 |
+
At 1280Γ1280 and below the fine texture stays within a few percent of the teacher's, and the grain in flat areas β skies, walls,
|
| 247 |
+
out-of-focus backgrounds β comes closer to the teacher's at 8 of the 12 sizes: 1.03Γ the teacher's across the sweep, from 1.07Γ.
|
| 248 |
+
Edge detail rises at 768Γ1024 and 1280Γ1280 (0.91Γ β 0.94Γ, 0.91Γ β 0.93Γ), and colour stays at the teacher's level: 0.99Γ across
|
| 249 |
+
the sweep, from 0.98Γ.
|
| 250 |
+
|
| 251 |
+
**Prompt adherence.** The judge that asks which of two renders follows the prompt better prefers the teacher on 12 of 66 β the 22
|
| 252 |
+
prompts at 512Γ512, 1280Γ1280 and 1440Γ1440 β against 15 for the previous checkpoint, and gives this one the only outright win. The
|
| 253 |
+
blind rubric, now 22 prompts at four sizes, finds the same 2 points missing out of 352 for both; the teacher's own renders lose 1. On
|
| 254 |
+
15 prompts drawn fresh from the prompt bank for this checkpoint, plus 5 black-and-white ones: 0 wins, 14 ties, 1 loss, and 0 Β· 4 Β· 1
|
| 255 |
+
in black and white.
|
| 256 |
+
|
| 257 |
+
**Speed β unchanged.** The same adapter shape at the same cost: measured again at 1024Γ1024, two steps with this LoRA took 19.3
|
| 258 |
+
and 19.2 s against 76.4 and 76.3 s for the teacher's eight β 4.0Γ, two model calls instead of eight.
|
| 259 |
+
|
| 260 |
+
**What stays a limit of two steps.** Small subjects in wide scenes and the tactile surface of stylised materials such as clay stay
|
| 261 |
+
where two steps fall furthest short of eight. The figures are in [Measured against the teacher](#measured-against-the-teacher).
|
| 262 |
+
|
| 263 |
+
| axis | `chk00031600` | `chk00041320` |
|
| 264 |
+
| --- | --- | --- |
|
| 265 |
+
| structure renders with a fault the teacher's does not have (of 21, checked at 1:1) | 1 | **0** |
|
| 266 |
+
| ghosting, blur-invariant index, sweep mean | 1.24Γ | **1.21Γ** |
|
| 267 |
+
| market faces at 1280Β², layout closeness to the teacher | 0.71 | **0.79** |
|
| 268 |
+
| grain in flat areas, sweep mean | 1.07Γ | **1.03Γ** |
|
| 269 |
+
| fine texture vs the teacher, 1440Β² | 1.07 | **1.03** |
|
| 270 |
+
| judge prefers the teacher (of 66) | 15 | **12** |
|
| 271 |
+
| blind adherence rubric, points missing of 352 | 2 | 2 |
|
| 272 |
+
| saturation vs the teacher, sweep mean | 0.98Γ | **0.99Γ** |
|
| 273 |
+
| distance to the teacher, sweep mean | 0.41 | 0.41 |
|
| 274 |
+
| training samples in the 2-step stages | 31,600 | 41,320 |
|
| 275 |
+
|
| 276 |
## chk00031600 vs chk00017464
|
| 277 |
|
| 278 |
`chk00017464` (published 14 Sep 2026) was the first checkpoint with critics and detail terms. `chk00031600` (25 Sep 2026) is
|
|
|
|
| 400 |
## Measured against the teacher
|
| 401 |
|
| 402 |
Every number here compares a render of this LoRA at 2 steps with the 8-step reference render of the **same prompt at the
|
| 403 |
+
same seed**, across the 22 test prompts and the 12 trained resolutions. Two other columns are measured the same way, so
|
| 404 |
the figures have a floor and a ceiling around them: **stock Krea 2 Turbo at 2 steps**, which is what the base model does
|
| 405 |
without the adapter, and the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) at its own 4 steps, which is the better tool and the thing worth
|
| 406 |
+
being compared against β its renders exist for the original 15 prompts, so its figures are on those 15.
|
| 407 |
|
| 408 |
**Detail, band by band.** Fine-texture energy and the two grid bands, as a ratio to the teacher's own (1.00 = the
|
| 409 |
teacher):
|
| 410 |
|
| 411 |
+
| bucket | fine texture | 16-px band | 8-px band | 4-step LoRA (same three, original 15) | stock 2-step, fine texture |
|
| 412 |
| --- | --- | --- | --- | --- | --- |
|
| 413 |
+
| 512Γ512 | 1.00 | 1.01 | 1.09 | 1.05 Β· 1.05 Β· 1.07 | 0.65 |
|
| 414 |
+
| 768Γ1024 | 1.00 | 0.98 | 1.02 | 1.05 Β· 1.05 Β· 1.08 | 0.44 |
|
| 415 |
+
| 1024Γ1024 | 1.06 | 1.10 | 1.13 | 1.11 Β· 1.17 Β· 1.16 | 0.42 |
|
| 416 |
+
| 1280Γ1280 | 1.02 | 1.02 | 1.03 | 1.20 Β· 1.18 Β· 1.22 | 0.44 |
|
| 417 |
+
| 1440Γ1440 | 1.03 | 1.05 | 1.09 | 1.24 Β· 1.23 Β· 1.32 | 0.52 |
|
| 418 |
|
| 419 |
+
Two steps without the adapter carry 0.65Γ the teacher's fine detail at 512Γ512 and **about half or less** (0.42β0.52Γ) at the four
|
| 420 |
+
larger sizes. With it, the fine texture sits at the teacher's level everywhere β between 0.97Γ and 1.09Γ of it across all 12
|
| 421 |
+
sizes β and from 1 megapixel up it is closer to the teacher than the 4-step adapter, which carries more excess fine energy there.
|
| 422 |
|
| 423 |
**Prompt adherence, judged.** A vision-language judge is shown the teacher's render and this LoRA's for the same prompt,
|
| 424 |
in both orders, and asked which follows the prompt better; a loss means the teacher was preferred both times:
|
| 425 |
|
| 426 |
| bucket | wins | ties | losses |
|
| 427 |
| --- | --- | --- | --- |
|
| 428 |
+
| 512Γ512 | 1 | 15 | 6 |
|
| 429 |
+
| 1280Γ1280 | 0 | 18 | 4 |
|
| 430 |
+
| 1440Γ1440 | 0 | 20 | 2 |
|
| 431 |
+
|
| 432 |
+
Twelve losses out of 66 (the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) scores six of 45 against the same teacher on the original 15). They
|
| 433 |
+
gather on stylised prompts and on how a prompt says the picture should look β crisp linework, energetic brush strokes, the texture
|
| 434 |
+
of clay β more than on what should be in it: a blind rubric that scores each render on its own against the prompt's objects, counts,
|
| 435 |
+
attributes and relations, with the teacher scored identically, finds 2 points missing out of 352 (the teacher's own renders lose 1).
|
| 436 |
+
On 15 prompts drawn fresh from the training prompt bank for this checkpoint and never rendered before, the judge returned 0 wins,
|
| 437 |
+
14 ties and 1 loss, and on five fresh black-and-white prompts 0 wins, 4 ties, 1 loss β with no colour cast in any of the five.
|
| 438 |
+
|
| 439 |
+
**Checked for the damage this kind of training can do.** Saturation sits at 1.04Γ the teacher's at 768Γ1024 and 0.98Γ at
|
| 440 |
+
the larger sizes; edge detail 0.92β0.94Γ; skin texture inside detected faces 1.02Γ at 768Γ1024 and 0.90Γ / 0.87Γ at
|
| 441 |
+
1280Γ1280 / 1440Γ1440, with skin saturation 0.98Γ and 0.95Γ at the larger sizes. The honest reading: **colour is at the
|
| 442 |
+
teacher's level at every size, and skin remains the softest part of this adapter's output at large sizes.** Fine detail in flat
|
| 443 |
+
regions β skies, walls, out-of-focus backgrounds β runs 1.32β1.57Γ the teacher's on this measure, which is where two steps put
|
| 444 |
+
grain that eight steps do not.
|
| 445 |
+
|
| 446 |
+
**Distance to the teacher**, as a plain pixel measure, is 0.39β0.43 at every size against the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s
|
| 447 |
+
0.30β0.37 (on the original 15). That gap is what two model calls cost instead of four: the image is a good render of the prompt, but it is
|
| 448 |
not the teacher's render of it β see [Known issues](#known-issues).
|
| 449 |
|
| 450 |
## Usage
|
|
|
|
| 582 |
|
| 583 |
| configuration | denoising | per model call | GPU peak |
|
| 584 |
| --- | --- | --- | --- |
|
| 585 |
+
| Krea 2 Turbo β 8 steps (the reference) | **76.4 s** | 9.5 s | 25.2 GiB |
|
| 586 |
+
| Krea 2 Turbo β 2 steps, no LoRA | 19.4 s | 9.7 s | 25.2 GiB |
|
| 587 |
+
| **Krea 2 Turbo β 2 steps + this LoRA** | **19.3 s** | 9.6 s | 25.2 GiB |
|
| 588 |
|
| 589 |
+
**Denoising is 4.0Γ faster than the 8-step reference** β two model calls instead of eight. The adapter's own cost per call did
|
| 590 |
not show up in this measurement: the runs with it came in marginally faster than those without, which is measurement noise, not a
|
| 591 |
speed-up. A rank-64 low-rank product is small beside the transformer it is added to, and it adds no measurable memory.
|
| 592 |
|
| 593 |
Denoising is the part the step count changes. What a complete render costs on top of it β encoding the prompt, decoding
|
| 594 |
the latent, writing the file β is the same whether you run two steps or eight, and it depends on your pipeline, so the
|
| 595 |
+
end-to-end figure on your machine will sit below 4Γ and rise toward it as the render gets larger.
|
| 596 |
|
| 597 |
+
**By resolution.** The two model calls of this LoRA's own sweep renders on the same machine (median of the 22 test prompts per
|
| 598 |
size; sweep renders run one at a time, not the controlled measurement above, so a second or two either way between one sweep and
|
| 599 |
the next is run-to-run variation β the adapter's shape, and so its cost, is the same at every checkpoint):
|
| 600 |
|
| 601 |
| resolution | denoising (2 calls) | resolution | denoising (2 calls) |
|
| 602 |
| --- | --- | --- | --- |
|
| 603 |
+
| 512Γ512 | 6.5 s | 1024Γ1024 | 19.5 s |
|
| 604 |
+
| 512Γ768 / 768Γ512 | 10.7 s / 11.3 s | 1280Γ960 / 960Γ1280 | 22.8 s / 22.7 s |
|
| 605 |
+
| 768Γ768 | 12.7 s | 1280Γ1280 | 32.0 s |
|
| 606 |
+
| 768Γ1024 / 1024Γ768 | 15.1 s / 15.2 s | 1440Γ1280 / 1440Γ1440 | 33.3 s / 38.3 s |
|
| 607 |
|
| 608 |
## LoRA strength
|
| 609 |
|
|
|
|
| 706 |
sweep of the first distribution-matching checkpoint located the problem at 1 megapixel and above (fine-texture energy
|
| 707 |
1.4β1.8Γ the teacher's at the five largest buckets); scaling the push per bucket from that measurement was tried and did
|
| 708 |
not hold, and the spectral match replaced it. The same sweep of the published checkpoint, its running-average weights, fixed
|
| 709 |
+
seed, 22 prompts per bucket, every image measured against the teacher's render of the same prompt and seed (in brackets: the
|
| 710 |
+
first distribution-matching checkpoint, on the original 15 prompts):
|
| 711 |
|
| 712 |
| bucket | fine texture vs the teacher | 16-px grid band | 8-px grid band | distance to the teacher |
|
| 713 |
| --------- | --------------------------- | --------------- | -------------- | ----------------------- |
|
| 714 |
+
| 512x512 | 1.00 (1.16) | 1.01 (1.23) | 1.09 (1.23) | 0.43 (0.44) |
|
| 715 |
+
| 512x768 | 0.97 (1.23) | 0.99 (1.37) | 1.02 (1.33) | 0.41 (0.41) |
|
| 716 |
+
| 768x512 | 1.01 (1.22) | 1.07 (1.37) | 1.02 (1.26) | 0.42 (0.47) |
|
| 717 |
+
| 768x768 | 1.05 (1.43) | 1.07 (1.46) | 1.09 (1.50) | 0.40 (0.43) |
|
| 718 |
+
| 768x1024 | 1.00 (1.36) | 0.98 (1.38) | 1.02 (1.42) | 0.39 (0.43) |
|
| 719 |
+
| 1024x768 | 0.98 (1.35) | 0.98 (1.40) | 0.98 (1.38) | 0.40 (0.46) |
|
| 720 |
+
| 1024x1024 | 1.06 (1.47) | 1.10 (1.55) | 1.13 (1.56) | 0.42 (0.44) |
|
| 721 |
+
| 1280x960 | 0.97 (1.50) | 0.99 (1.59) | 0.98 (1.56) | 0.43 (0.42) |
|
| 722 |
+
| 960x1280 | 1.09 (1.60) | 1.02 (1.63) | 1.13 (1.70) | 0.42 (0.43) |
|
| 723 |
+
| 1280x1280 | 1.02 (1.65) | 1.02 (1.68) | 1.03 (1.69) | 0.43 (0.44) |
|
| 724 |
+
| 1440x1280 | 0.99 (1.59) | 1.00 (1.66) | 1.09 (1.66) | 0.40 (0.42) |
|
| 725 |
+
| 1440x1440 | 1.03 (1.81) | 1.05 (1.79) | 1.09 (1.89) | 0.40 (0.41) |
|
| 726 |
|
| 727 |
### The spectral match
|
| 728 |
|
|
|
|
| 765 |
|
| 766 |
### Critics in turn
|
| 767 |
|
| 768 |
+
One critic holds one idea of what is wrong. The artefact critic is joined by seven more on the same frozen mid-network
|
| 769 |
features and under the same rules β lightly re-noised inputs, a push that is filtered and capped β each aimed at a
|
| 770 |
different fault:
|
| 771 |
|
|
|
|
| 773 |
what fine texture looks like in a photograph as well as in the teacher's rendering of one. It judges photographic prompts
|
| 774 |
only, so illustration, anime and 3D renders are not pulled toward photographic grain, and its push is filtered to periods
|
| 775 |
finer than 24 pixels and held lower than the artefact critic's, because photographs carry grain the teacher does not.
|
| 776 |
+
- **A large-face critic, on photographs.** It reads only large faces, 192 pixels and up, and pushes on every step. A second
|
| 777 |
+
face critic, reading faces of every size, trained the previous checkpoint and was taken back out (item 18 under
|
| 778 |
+
[How I got here](#how-i-got-here)).
|
| 779 |
- **A structure critic.** It judges the first call β the layout, before any detail β against the teacher's own intermediate
|
| 780 |
state from the same noise, at the high noise levels where layout is decided and at periods of 32 pixels and coarser only,
|
| 781 |
so a first call that blends two layouts is caught where the blend happens.
|
|
|
|
| 786 |
- **A rollout critic.** Its real examples are the teacher's own finish from the student's second-call starting point, so the
|
| 787 |
second call is judged against what the teacher would have made from the same start.
|
| 788 |
|
| 789 |
+
All eight on every step do not fit in 24 GB, so they take turns: on each step one critic pushes, and on alternate steps the
|
| 790 |
others train so none goes stale before its turn comes back. Each new head started from the artefact critic's weights and
|
| 791 |
trained on its own before it was allowed to push.
|
| 792 |
|
|
|
|
| 801 |
pixels may not fall below the teacher's plus the margin real photographs carry over it at those scales, judged tile by tile
|
| 802 |
on the window's textured tiles. That margin is measured once from a pool of real photographs and clamped, and the term
|
| 803 |
only ever pushes upward to that floor, never past it β so it lifts detail that is missing without adding grain that is not.
|
| 804 |
+
- **A ceiling.** Tile by tile, detail at 3β16 pixels may not climb past 1.3Γ the teacher's β the counterpart of the photo
|
| 805 |
+
floor, and what keeps grain from building up at large sizes.
|
|
|
|
|
|
|
|
|
|
| 806 |
- **A direction-aware term.** The spectral comparison is also made orientation by orientation, so the student's fine detail
|
| 807 |
runs in the same directions as the teacher's.
|
| 808 |
- **Windows on features.** On photographs, half of the decoded windows are centred on an eye, the nose or the lips of a face,
|
|
|
|
| 827 |
The published weights are a running average of training (decay 0.999), which smooths out the noise of single steps.
|
| 828 |
Averaging has one failure: while the training weights move between two layouts of the same prompt, their average draws
|
| 829 |
both β a doubled subject. A guard watches a fixed set of layout probes every ten steps; a move away from the trend that
|
| 830 |
+
comes back within six rounds is kept out of the average, and a lasting move is taken in whole β the guard never resets
|
| 831 |
+
the average. Before any stretch of training goes in, the average it would make renders a fixed probe at the large sizes, and a
|
| 832 |
+
stretch that would put a second head on the subject is held out, even when the training weights themselves drew a single one
|
| 833 |
+
throughout. The average itself
|
| 834 |
+
has been rebuilt twice, by hand: once after a layout change it had blended, and once from the training since the change below,
|
| 835 |
+
after the guard had held it still through a long change of pose.
|
| 836 |
+
|
| 837 |
+
### Layout at the large sizes
|
| 838 |
+
|
| 839 |
+
The doubled subjects that two steps can draw β a second head, two bodies joined, two poses blended β formed at the sizes
|
| 840 |
+
trained most, 1280 pixels and up, and the distribution term is the one that pushes the student toward whatever the teacher would
|
| 841 |
+
plausibly draw, which at a two-step jump can be more than one layout at once. So at every size with a 1280- or 1440-pixel side the
|
| 842 |
+
distribution term pushes at half strength, while the anchor to the teacher's recorded trajectory keeps its full weight there: at
|
| 843 |
+
those sizes the teacher's own layout for the same noise carries twice the share it did. The anchor also counts the coarse layout of
|
| 844 |
+
the second call β everything 64 pixels and up β twice, since that is the call in which a second head appears.
|
| 845 |
|
| 846 |
## What the LoRA touches
|
| 847 |
|
|
|
|
| 881 |
which is why every cut of this run is rendered across the buckets and why the per-resolution table under
|
| 882 |
[Method](#method) above exists.
|
| 883 |
|
| 884 |
+
[`assets/resolution_sweeps/`](assets/resolution_sweeps) holds the evidence β same 22 prompts (the original 15, plus seven
|
| 885 |
+
scenes of people, animals and action added on 29 Sep 2026 to check structure), same seed, one folder per resolution, one
|
| 886 |
+
image per prompt, so any image can be compared 1:1 with its twin in the next tree:
|
| 887 |
|
| 888 |
- [`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step) β the **official Krea 2 Turbo 8-step reference renders**:
|
| 889 |
the stock model, no LoRA, at its native settings (8 steps, guidance 0.0). The teacher this LoRA is distilled from and
|
|
|
|
| 902 |
β βββ 512x512/ one folder per resolution
|
| 903 |
β β βββ portrait.jpg
|
| 904 |
β β βββ kingfisher.jpg
|
| 905 |
+
β β βββ β¦ 19 more, one per test prompt
|
| 906 |
+
β β βββ flamingos.jpg
|
| 907 |
β βββ 768x1024/
|
| 908 |
β βββ 1024x1024/
|
| 909 |
β βββ 1280x1280/
|
|
|
|
| 930 |
host memory above 0.3 megapixels; the student, the fake adapter, the spectral and detail terms, the critic and the teacher's finishing pass each
|
| 931 |
build and free their own graph in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
|
| 932 |
does not fit fails loudly. A full step with every term live and every critic pushing reserves about 22.4 GB at 1440Γ1440,
|
| 933 |
+
of 24. The price of the objective is throughput: **about 58 training samples an hour** with every critic and detail term in
|
| 934 |
+
the recipe, against the [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β about eight times the cost per sample, and so far a
|
| 935 |
+
small fraction of the samples.
|
| 936 |
|
| 937 |
**Where that cost comes from.** Distribution matching is simply a heavier objective than trajectory distillation.
|
| 938 |
The 4-step project's recipe compared the student's own output with a teacher state that had already been recorded to
|
| 939 |
disk, so a training step was one student pass plus a small adversarial head. Here every step also needs the *score*
|
| 940 |
of two models at a freshly noised point: the frozen teacher's, and a second adapter's that is being trained
|
| 941 |
+
alongside to imitate the student β and that second adapter takes three optimiser steps of its own per student step.
|
| 942 |
The spectral and detail terms decode part of the image out of the latent to compare its texture with the teacher's,
|
| 943 |
each critic reads half the network twice more, and every second step the teacher finishes the image from the student's
|
| 944 |
first call.
|
|
|
|
| 954 |
## How it is judged
|
| 955 |
|
| 956 |
At regular intervals, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
|
| 957 |
+
22 fixed prompts across four resolutions (512Γ512, 1280Γ1280, 1440Γ1440, 1440Γ1280), after a layout check at three large
|
| 958 |
sizes that has to pass first; milestone checkpoints get the same render at all 12 buckets, which is where the
|
| 959 |
per-resolution table above comes from. Every image is measured against the
|
| 960 |
teacher's render of the same prompt and seed: distance, fine-texture energy, the 16-pixel and 8-pixel grid bands, skin
|
| 961 |
and flat-region grain, saturation, faces cut out at 1:1, fixed content windows (small faces in a crowd, shop interiors
|
| 962 |
+
seen through their windows), straight-line artefacts, a structure judge that compares the seven scenes of people, animals
|
| 963 |
+
and action with the teacher's for missing, extra or merged parts (every flag checked by eye at 1:1), a graded judge, a pairwise
|
| 964 |
+
preference against the teacher, and a blind rubric that
|
| 965 |
scores each render on its own against the prompt's objects, counts, attributes and relations, with the teacher scored
|
| 966 |
identically. Fifteen prompts drawn fresh from the prompt bank, never rendered before, are judged the same way at every
|
| 967 |
checkpoint. Latent distances β the held-out chord gap and the two-step rollout
|
|
|
|
| 1277 |
|
| 1278 |
---
|
| 1279 |
|
| 1280 |
+
### A group photo of five friends standing side by side on a sunny beach, arms around each other's shoulders, all smiling at the camera; five clearly different people, each with a unique face, no twins or lookalikes: a tall bearded man in his forties, a young woman with curly red hair and freckles, an older East Asian man with grey hair and glasses, a Black woman with short natural hair, and a teenage boy with messy blond hair, photograph, sharp detail
|
| 1281 |
+
|
| 1282 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1283 |
+
|
| 1284 |
+
[](assets/group5_compare_turbo.jpg)
|
| 1285 |
+
|
| 1286 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1287 |
+
|
| 1288 |
+
[](assets/group5_compare_teacher.jpg)
|
| 1289 |
+
|
| 1290 |
+
**Individual frames** β click any panel to open that render full size
|
| 1291 |
+
|
| 1292 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1293 |
+
| --------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
|
| 1294 |
+
| [](assets/group5_turbo_8step.jpg) | [](assets/group5_turbo_2step.jpg) | [](assets/group5_turbo_2step_lora.jpg) |
|
| 1295 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1296 |
+
|
| 1297 |
+
---
|
| 1298 |
+
|
| 1299 |
+
### A family of four having dinner at a round wooden table, seen from directly above, plates, glasses and bowls of food, warm evening light, photograph
|
| 1300 |
+
|
| 1301 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1302 |
+
|
| 1303 |
+
[](assets/family4_compare_turbo.jpg)
|
| 1304 |
+
|
| 1305 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1306 |
+
|
| 1307 |
+
[](assets/family4_compare_teacher.jpg)
|
| 1308 |
+
|
| 1309 |
+
**Individual frames** β click any panel to open that render full size
|
| 1310 |
+
|
| 1311 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1312 |
+
| ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
| 1313 |
+
| [](assets/family4_turbo_8step.jpg) | [](assets/family4_turbo_2step.jpg) | [](assets/family4_turbo_2step_lora.jpg) |
|
| 1314 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1315 |
+
|
| 1316 |
+
---
|
| 1317 |
+
|
| 1318 |
+
### Three kittens sitting side by side in a wicker basket with a tall arched handle over them, all looking at the camera, soft natural light, photograph, sharp detail
|
| 1319 |
+
|
| 1320 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1321 |
+
|
| 1322 |
+
[](assets/kittens3_compare_turbo.jpg)
|
| 1323 |
+
|
| 1324 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1325 |
+
|
| 1326 |
+
[](assets/kittens3_compare_teacher.jpg)
|
| 1327 |
+
|
| 1328 |
+
**Individual frames** β click any panel to open that render full size
|
| 1329 |
+
|
| 1330 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1331 |
+
| ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
|
| 1332 |
+
| [](assets/kittens3_turbo_8step.jpg) | [](assets/kittens3_turbo_2step.jpg) | [](assets/kittens3_turbo_2step_lora.jpg) |
|
| 1333 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1334 |
+
|
| 1335 |
+
---
|
| 1336 |
+
|
| 1337 |
+
### Two golden retrievers playing tug of war with a red rope on a green lawn, action shot, photograph, sharp detail
|
| 1338 |
+
|
| 1339 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1340 |
+
|
| 1341 |
+
[](assets/dogs2_compare_turbo.jpg)
|
| 1342 |
+
|
| 1343 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1344 |
+
|
| 1345 |
+
[](assets/dogs2_compare_teacher.jpg)
|
| 1346 |
+
|
| 1347 |
+
**Individual frames** β click any panel to open that render full size
|
| 1348 |
+
|
| 1349 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1350 |
+
| ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ |
|
| 1351 |
+
| [](assets/dogs2_turbo_8step.jpg) | [](assets/dogs2_turbo_2step.jpg) | [](assets/dogs2_turbo_2step_lora.jpg) |
|
| 1352 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1353 |
+
|
| 1354 |
+
---
|
| 1355 |
+
|
| 1356 |
+
### A horse and rider jumping over a wooden fence at a show jumping event, side view, sports photography, sharp detail
|
| 1357 |
+
|
| 1358 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1359 |
+
|
| 1360 |
+
[](assets/horserider_compare_turbo.jpg)
|
| 1361 |
+
|
| 1362 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1363 |
+
|
| 1364 |
+
[](assets/horserider_compare_teacher.jpg)
|
| 1365 |
+
|
| 1366 |
+
**Individual frames** β click any panel to open that render full size
|
| 1367 |
+
|
| 1368 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1369 |
+
| ----------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
|
| 1370 |
+
| [](assets/horserider_turbo_8step.jpg) | [](assets/horserider_turbo_2step.jpg) | [](assets/horserider_turbo_2step_lora.jpg) |
|
| 1371 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1372 |
+
|
| 1373 |
+
---
|
| 1374 |
+
|
| 1375 |
+
### Close-up of a pianist's two hands playing the keys of a grand piano, dramatic side light, photograph, sharp detail
|
| 1376 |
+
|
| 1377 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1378 |
+
|
| 1379 |
+
[](assets/pianist_compare_turbo.jpg)
|
| 1380 |
+
|
| 1381 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1382 |
+
|
| 1383 |
+
[](assets/pianist_compare_teacher.jpg)
|
| 1384 |
+
|
| 1385 |
+
**Individual frames** β click any panel to open that render full size
|
| 1386 |
+
|
| 1387 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1388 |
+
| ----------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- |
|
| 1389 |
+
| [](assets/pianist_turbo_8step.jpg) | [](assets/pianist_turbo_2step.jpg) | [](assets/pianist_turbo_2step_lora.jpg) |
|
| 1390 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1391 |
+
|
| 1392 |
+
---
|
| 1393 |
+
|
| 1394 |
+
### A flock of pink flamingos standing in shallow turquoise water with their reflections, wildlife photography, sharp detail
|
| 1395 |
+
|
| 1396 |
+
**3-way comparison** β one image, all three renders side by side
|
| 1397 |
+
|
| 1398 |
+
[](assets/flamingos_compare_turbo.jpg)
|
| 1399 |
+
|
| 1400 |
+
**Against the teacher** β this LoRA at 2 steps beside the 8-step render
|
| 1401 |
+
|
| 1402 |
+
[](assets/flamingos_compare_teacher.jpg)
|
| 1403 |
+
|
| 1404 |
+
**Individual frames** β click any panel to open that render full size
|
| 1405 |
+
|
| 1406 |
+
| Turbo β 8 steps | Turbo β 2 steps, no LoRA | **Turbo β 2 steps + this LoRA** |
|
| 1407 |
+
| --------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------- |
|
| 1408 |
+
| [](assets/flamingos_turbo_8step.jpg) | [](assets/flamingos_turbo_2step.jpg) | [](assets/flamingos_turbo_2step_lora.jpg) |
|
| 1409 |
+
| as shipped Β· 8 NFE | the deficit this closes | strength 1.0 Β· 2 NFE |
|
| 1410 |
+
|
| 1411 |
+
---
|
| 1412 |
+
|
| 1413 |
## Notes and limitations
|
| 1414 |
|
| 1415 |
- π― **Krea 2 Turbo only**, at **2 steps**, guidance **0.0** (cfg 1.0 in ComfyUI), **mu = 1.15** β the two training
|
|
|
|
| 1422 |
resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). Next, aimed
|
| 1423 |
at what this checkpoint still gets wrong: small subjects in wide scenes; the finest edges and the texture of skin at the
|
| 1424 |
largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two
|
| 1425 |
+
plausible poses of the same subject can meet; the fine detail of close-up faces β eyes and skin texture; and the tactile surface of
|
| 1426 |
+
stylised materials such as clay β each held to the rule that none of it may cost the structure, the colour and the clean large
|
| 1427 |
+
sizes this checkpoint gained. A better checkpoint
|
| 1428 |
replaces this one when the sweeps and I visually agree, the same discipline as the
|
| 1429 |
[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA); until then the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) remains the recommendation for quality renders, and this
|
| 1430 |
one is the fast preview.
|
README.md
CHANGED
|
@@ -17,28 +17,28 @@ pipeline_tag: text-to-image
|
|
| 17 |
|
| 18 |
# Krea 2 Turbo β 2-Step Distillation LoRA
|
| 19 |
|
| 20 |
-
**A quarter of the steps Β· 4
|
| 21 |
|
| 22 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
|
| 23 |
|
| 24 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
|
| 25 |
-
- β±οΈ **4
|
| 26 |
-
- π― **Fine detail at the teacher's level** β **0.
|
| 27 |
- π **Distribution matching, not imitation** β matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
|
| 28 |
-
- π£οΈ **Prompt-conditioned throughout** β teacher and fake scores both read each prompt's conditioning; a blind rubric finds **
|
| 29 |
- π **12 trained resolutions** β multi-aspect from 512Γ512 up to 1440Γ1440
|
| 30 |
- π **Drop-in, no exceptions** β plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
|
| 31 |
- 𧬠**Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β rank 64 on the same 228 modules
|
| 32 |
- π² **13,750 recorded teacher trajectories** from the 4-step project, reused β not one new teacher run
|
| 33 |
-
- π’ **
|
| 34 |
-
- π
**
|
| 35 |
- π **More than forty recipe adjustments** across two methods β each kept only when the renders did not get worse
|
| 36 |
|
| 37 |
> π§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β a portrait, a single figure, an object up close β hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
|
| 38 |
>
|
| 39 |
> π **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ2048 β Krea's published maximum recommended resolution, beyond this adapter's largest trained size β come within easy reach. Past 2048Γ2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
|
| 40 |
|
| 41 |
-
[ β **this LoRA at every trained resolution for all
|
| 291 |
|
| 292 |
[`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step/) β official Krea 2 Turbo 8-step reference renders for the same prompts/seeds/resolutions. [`_turbo-base-NO-LoRA-2step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-2step/) β stock Turbo at 2 steps, the floor.
|
| 293 |
|
|
@@ -315,7 +317,7 @@ Every published checkpoint and its resolution sweep under [`_archive/`](_archive
|
|
| 315 |
|
| 316 |
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
|
| 317 |
|
| 318 |
-
Next, aimed at the [Known Issues](#known-issues): small subjects in wide scenes; the finest edges and skin texture at the largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two plausible poses of the same subject can meet; the tactile surface of stylised materials such as clay β none of it allowed to cost the colour and the clean large sizes this checkpoint gained.
|
| 319 |
|
| 320 |
A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
|
| 321 |
|
|
|
|
| 17 |
|
| 18 |
# Krea 2 Turbo β 2-Step Distillation LoRA
|
| 19 |
|
| 20 |
+
**A quarter of the steps Β· 4Γ faster denoising Β· fine detail at 0.97β1.09Γ the teacher's across all 12 trained resolutions Β· colour at the teacher's level Β· 2 points missing of 352 on a blind prompt-adherence rubric Β· teacher preferred on 12 of 66 judged renders Β· 41,320 training samples on the 4-step project's recorded trajectories Β· 25 days on one RTX 3090 Β· still in training.**
|
| 21 |
|
| 22 |
A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from **8 steps down to 2** β Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β aiming at the best quality two steps can give. It is for **fast previews and drafts**; the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** remains the recommendation for quality renders.
|
| 23 |
|
| 24 |
- β‘ **A quarter of the steps** β 8 β 2, on Turbo's own deployment sigmas `[1.0, 0.7595]`
|
| 25 |
+
- β±οΈ **4Γ faster denoising** β 76.4 s β 19.3 s at 1024Γ1024; the adapter's own cost per call is within measurement noise
|
| 26 |
+
- π― **Fine detail at the teacher's level** β **0.97β1.09Γ** the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: **0.40β0.65Γ**); from 1 megapixel up, closer to the teacher than the 4-step adapter
|
| 27 |
- π **Distribution matching, not imitation** β matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
|
| 28 |
+
- π£οΈ **Prompt-conditioned throughout** β teacher and fake scores both read each prompt's conditioning; a blind rubric finds **2 points missing of 352** (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on **12 of 66** (4-step adapter: 6 of 45, on the original 15 prompts), mostly on style
|
| 29 |
- π **12 trained resolutions** β multi-aspect from 512Γ512 up to 1440Γ1440
|
| 30 |
- π **Drop-in, no exceptions** β plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
|
| 31 |
- 𧬠**Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β rank 64 on the same 228 modules
|
| 32 |
- π² **13,750 recorded teacher trajectories** from the 4-step project, reused β not one new teacher run
|
| 33 |
+
- π’ **41,320 training samples** in the 2-step stages, on top of the 4-step LoRA's 78,000
|
| 34 |
+
- π
**25 days** from the first 2-step launch to this checkpoint, on a single RTX 3090 β training continues
|
| 35 |
- π **More than forty recipe adjustments** across two methods β each kept only when the renders did not get worse
|
| 36 |
|
| 37 |
> π§ͺ **Fast-preview adapter, still in training.** Subjects that are close and fill a good part of the frame β a portrait, a single figure, an object up close β hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA). See [Known Issues](#known-issues).
|
| 38 |
>
|
| 39 |
> π **The saved steps can also go into resolution.** A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ2048 β Krea's published maximum recommended resolution, beyond this adapter's largest trained size β come within easy reach. Past 2048Γ2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.
|
| 40 |
|
| 41 |
+
[](assets/poster.jpg)
|
| 42 |
|
| 43 |
---
|
| 44 |
|
|
|
|
| 111 |
|
| 112 |
| | denoise | per model call | GPU peak |
|
| 113 |
| ----------------------------- | ---------- | -------------- | -------- |
|
| 114 |
+
| Turbo 8 steps (quality bar) | 76.4 s | 9.5 s | 25.2 GiB |
|
| 115 |
+
| Turbo 2 steps, no LoRA | 19.4 s | 9.7 s | 25.2 GiB |
|
| 116 |
+
| **Turbo 2 steps + this LoRA** | **19.3 s** | 9.6 s | 25.2 GiB |
|
| 117 |
|
| 118 |
+
**Denoising is 4.0Γ faster than the 8-step bar** β two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.
|
| 119 |
|
| 120 |
+
Prompt encoding and VAE decode don't change with step count, so end to end sits below 4Γ and rises toward it as the render grows. Denoise times at every trained resolution (6.5 s at 512Γ512 to 38.3 s at 1440Γ1440) are in the [Detailed Model Card](DETAILED-README.md).
|
| 121 |
|
| 122 |
---
|
| 123 |
|
|
|
|
| 138 |
|
| 139 |
## Current Checkpoint
|
| 140 |
|
| 141 |
+
**`chk00041320`** (2 Oct 2026) replaces `chk00031600` (25 Sep 2026). It is 9,720 training samples later, aimed at what matters first in a picture β subjects drawn twice or two poses blended into one at the large sizes, and faces in busy scenes β through the teacher's own layout held at the large sizes, eight critics, and every update to the running average checked for doubled subjects. It is also the first checkpoint measured on the full 22-prompt sweep.
|
| 142 |
|
| 143 |
+
| axis | `chk00031600` | `chk00041320` |
|
| 144 |
+
| ------------------------------------------------------------------ | ------------- | ------------- |
|
| 145 |
+
| structure renders with a fault the teacher's does not have (of 21) | 1 | **0** |
|
| 146 |
+
| ghosting, blur-invariant index, sweep mean | 1.24Γ | **1.21Γ** |
|
| 147 |
+
| market faces at 1280Β², layout closeness to the teacher | 0.71 | **0.79** |
|
| 148 |
+
| grain in flat areas, sweep mean | 1.07Γ | **1.03Γ** |
|
| 149 |
+
| fine texture vs the teacher, 1440Β² | 1.07 | **1.03** |
|
| 150 |
+
| judge prefers the teacher (of 66) | 15 | **12** |
|
| 151 |
+
| blind adherence rubric, points missing of 352 | 2 | 2 |
|
| 152 |
+
| saturation vs the teacher, sweep mean | 0.98Γ | **0.99Γ** |
|
| 153 |
|
| 154 |
+
Fewer doubled subjects and less ghosting at the large sizes, whole faces in the busy market scene, and prompt-following a step closer to the teacher's, while colour and fine detail stay at the teacher's level. [`krea2_turbo_2step_rank_64_lora_checkpoint_info.md`](krea2_turbo_2step_rank_64_lora_checkpoint_info.md) always names the checkpoint in the weight files; the full comparison is in the [Detailed Model Card](DETAILED-README.md#chk00041320-vs-chk00031600).
|
| 155 |
|
| 156 |
---
|
| 157 |
|
|
|
|
| 176 |
|
| 177 |
The student makes two calls, at Ο = 1.0 and 0.7595 β the first and fifth points of the teacher's 8-step grid at ΞΌ = 1.15 β with stock Euler between them. Euler's first step lands **exactly** on the flow-matching interpolant at Ο = 0.7595, so the first call's output is a legitimate image prediction and is judged as one.
|
| 178 |
|
| 179 |
+
- **Distribution term** β the frozen teacher and a _fake-score_ adapter (rank 32, trained online on the student's current output, 3 updates per student step, discarded at the end) each denoise a freshly noised copy of the student's image; where they disagree is the direction toward the teacher's work. Averaging is never rewarded, so the student commits
|
| 180 |
- **Trajectory anchor** β regression on the recorded chords at half weight keeps the student on the teacher's two-step grid
|
| 181 |
|
| 182 |
**On top, each capped relative to the distribution term:**
|
| 183 |
|
| 184 |
- **Spectral match** β student and teacher images compared through radial power spectra, on the whole latent and a decoded 256-px window, two-sided β the term that reached the 16/8-px grid grain at large sizes
|
| 185 |
+
- **Eight critics taking turns** β artefact, photo (half real photographs), large faces, the first call's layout against the teacher's own state, content, prompt, text, and the teacher's own finish, all heads on the frozen base's mid-network features; one pushes per step, each filtered and capped
|
| 186 |
+
- **Detail held to the teacher region by region** β the anchor counts fine-detail error twice (3.5Γ where the teacher is most detailed); a one-sided photo floor at 3β10 px and a per-tile ceiling at 1.3Γ; a direction-aware term; windows on eyes, nose and lips; the teacher's finish of the student's first call as the second call's target; a smoothness limit on the fake adapter
|
| 187 |
+
- **Colour band** β saturation held between the teacher's own level and 8% above it
|
| 188 |
+
- **Layout at the large sizes** β at every size with a 1280- or 1440-px side the distribution term pushes at half strength while the anchor keeps full weight, and the anchor counts the second call's coarse layout (64 px and up) twice
|
| 189 |
|
| 190 |
+
Shipped adapter is a **guarded running (EMA) average** of the weights, not the last live state: short excursions are kept out, and every update is checked for doubled subjects before it goes in.
|
| 191 |
|
| 192 |
---
|
| 193 |
|
|
|
|
| 221 |
| 1024Γ1024 | 960Γ1280 | 1280Γ960 |
|
| 222 |
| 1280Γ1280 | 1440Γ1280 | 1440Γ1440 |
|
| 223 |
|
| 224 |
+
Same 12 buckets as the 4-step adapter, with the larger sizes and stylised prompts drawn more often: 1440Γ1440 about 6% of the samples, 1024Γ1024 and 768Γ1024 12% each, stylised prompts 12%.
|
| 225 |
|
| 226 |
---
|
| 227 |
|
|
|
|
| 233 |
- Student, fake adapter, spectral and detail terms, critic and teacher finish each build and free their own graph β peaks never overlap
|
| 234 |
- Hard memory ceiling below the driver's paging threshold, so a step that doesn't fit fails loudly
|
| 235 |
|
| 236 |
+
A full step with every term live reserves ~22.4 GB at 1440Γ1440. Throughput is **~58 samples/hour** against the 4-step recipe's 470 β roughly a dozen model runs per sample instead of two. That is the objective's cost, not teacher generation: the trajectories were recorded once and are read from disk.
|
| 237 |
|
| 238 |
Released LoRA is **bf16**.
|
| 239 |
|
|
|
|
| 283 |
|
| 284 |
---
|
| 285 |
|
| 286 |
+
### (19 more examples, and a two-panel sheet against the teacher for every prompt, in the [full version](DETAILED-README.md) β see [assets/](assets/) for all 22 prompts)
|
| 287 |
|
| 288 |
---
|
| 289 |
|
| 290 |
## Resolution Sweeps
|
| 291 |
|
| 292 |
+
[`assets/resolution_sweeps/`](assets/resolution_sweeps/) β **this LoRA at every trained resolution for all 22 test prompts** (same prompts, seed, 2 steps, strength 1.0). Nothing cherry-picked.
|
| 293 |
|
| 294 |
[`_teacher-8step/`](assets/resolution_sweeps/_teacher-8step/) β official Krea 2 Turbo 8-step reference renders for the same prompts/seeds/resolutions. [`_turbo-base-NO-LoRA-2step/`](assets/resolution_sweeps/_turbo-base-NO-LoRA-2step/) β stock Turbo at 2 steps, the floor.
|
| 295 |
|
|
|
|
| 317 |
|
| 318 |
Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β aiming at the best quality two steps can give, not at matching the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
|
| 319 |
|
| 320 |
+
Next, aimed at the [Known Issues](#known-issues): small subjects in wide scenes; the finest edges and skin texture at the largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two plausible poses of the same subject can meet; the fine detail of close-up faces β eyes and skin texture; the tactile surface of stylised materials such as clay β none of it allowed to cost the structure, the colour and the clean large sizes this checkpoint gained.
|
| 321 |
|
| 322 |
A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.
|
| 323 |
|
krea2_turbo_2step_rank_64_lora.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:c827d21150f96245fe421f75a51a374ac494006d1c7dac7d3166b57a486f5d5c
|
| 3 |
+
size 438161704
|
krea2_turbo_2step_rank_64_lora_checkpoint_info.md
CHANGED
|
@@ -1,13 +1,13 @@
|
|
| 1 |
# Which checkpoint is this?
|
| 2 |
|
| 3 |
`krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
|
| 4 |
-
**
|
| 5 |
`_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
|
| 6 |
keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
|
| 7 |
|
| 8 |
| file | SHA-256 | size |
|
| 9 |
| --- | --- | --- |
|
| 10 |
-
| `krea2_turbo_2step_rank_64_lora.safetensors` | `
|
| 11 |
-
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `
|
| 12 |
|
| 13 |
-
Updated:
|
|
|
|
| 1 |
# Which checkpoint is this?
|
| 2 |
|
| 3 |
`krea2_turbo_2step_rank_64_lora.safetensors` and `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` in this folder are
|
| 4 |
+
**chk00041320** β the same weights as `_archive/checkpoints/krea2_turbo_2step_rank_64_lora_chk00041320.safetensors` (and its
|
| 5 |
`_comfyui` twin). The pair here is updated in place whenever a better checkpoint ships; the archive
|
| 6 |
keeps every one that did. The same checkpoint id is in each file's safetensors metadata (`checkpoint`).
|
| 7 |
|
| 8 |
| file | SHA-256 | size |
|
| 9 |
| --- | --- | --- |
|
| 10 |
+
| `krea2_turbo_2step_rank_64_lora.safetensors` | `c827d21150f96245fe421f75a51a374ac494006d1c7dac7d3166b57a486f5d5c` | 418M |
|
| 11 |
+
| `krea2_turbo_2step_rank_64_lora_comfyui.safetensors` | `69c9d4e55530f26764ff0cfed39e230ba20c9d1b8bcf8c581beef26eb9dd9726` | 418M |
|
| 12 |
|
| 13 |
+
Updated: 2 Oct 2026
|
krea2_turbo_2step_rank_64_lora_comfyui.safetensors
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:69c9d4e55530f26764ff0cfed39e230ba20c9d1b8bcf8c581beef26eb9dd9726
|
| 3 |
+
size 438142240
|