Krea 2 Turbo β€” 2-Step Distillation LoRA

A quarter of the steps Β· 4Γ— faster denoising Β· fine detail at 0.97–1.09Γ— the teacher's across all 12 trained resolutions Β· colour at the teacher's level Β· 2 points missing of 352 on a blind prompt-adherence rubric Β· teacher preferred on 12 of 66 judged renders Β· 41,320 training samples on the 4-step project's recorded trajectories Β· 25 days on one RTX 3090 Β· still in training.

A LoRA for Krea 2 Turbo that takes the model from 8 steps down to 2 β€” Turbo's own weights and sigmas, guidance 0.0, a quarter of the denoising passes β€” aiming at the best quality two steps can give. It is for fast previews and drafts; the 4-step LoRA remains the recommendation for quality renders.

  • ⚑ A quarter of the steps β€” 8 β†’ 2, on Turbo's own deployment sigmas [1.0, 0.7595]
  • ⏱️ 4Γ— faster denoising β€” 76.4 s β†’ 19.3 s at 1024Γ—1024; the adapter's own cost per call is within measurement noise
  • 🎯 Fine detail at the teacher's level β€” 0.97–1.09Γ— the teacher's fine-texture energy at every trained resolution (stock Turbo at 2 steps: 0.40–0.65Γ—); from 1 megapixel up, closer to the teacher than the 4-step adapter
  • πŸ“Š Distribution matching, not imitation β€” matches what the teacher would plausibly produce rather than its exact trajectory, so the student commits instead of averaging into blur and doubled edges
  • πŸ—£οΈ Prompt-conditioned throughout β€” teacher and fake scores both read each prompt's conditioning; a blind rubric finds 2 points missing of 352 (objects, counts, attributes, relations), and a judge prefers the 8-step teacher on 12 of 66 (4-step adapter: 6 of 45, on the original 15 prompts), mostly on style
  • πŸ“ 12 trained resolutions β€” multi-aspect from 512Γ—512 up to 1440Γ—1440
  • πŸ”Œ Drop-in, no exceptions β€” plain LoRA, stock Euler, diffusers / ComfyUI / MLX. No custom nodes, no custom sampler
  • 🧬 Same shape as the 4-step adapter β€” rank 64 on the same 228 modules
  • 🎲 13,750 recorded teacher trajectories from the 4-step project, reused β€” not one new teacher run
  • πŸ”’ 41,320 training samples in the 2-step stages, on top of the 4-step LoRA's 78,000
  • πŸ“… 25 days from the first 2-step launch to this checkpoint, on a single RTX 3090 β€” training continues
  • πŸ” More than forty recipe adjustments across two methods β€” each kept only when the renders did not get worse

πŸ§ͺ Fast-preview adapter, still in training. Subjects that are close and fill a good part of the frame β€” a portrait, a single figure, an object up close β€” hold up well at two steps. Small subjects are where it still falls short: faces in a crowd or figures in a wide scene can come out ghosted or smeared. For those, and whenever quality matters more than speed, use the 4-step LoRA. See Known Issues.

πŸ“ The saved steps can also go into resolution. A larger render makes a small subject bigger, and at a quarter of the teacher's steps, renders up to 2048Γ—2048 β€” Krea's published maximum recommended resolution, beyond this adapter's largest trained size β€” come within easy reach. Past 2048Γ—2048, stock Krea 2 itself begins to duplicate subjects, with or without this adapter.

The 22 test prompts, rendered by Krea 2 Turbo with this LoRA at 2 steps


Files

file what it is
krea2_turbo_2step_rank_64_lora.safetensors LoRA in diffusers key format β€” see diffusers
krea2_turbo_2step_rank_64_lora_comfyui.safetensors Same weights under ComfyUI key names β€” see ComfyUI
krea2_turbo_2step_lora_t2i.json Ready ComfyUI workflow, stock nodes only
krea2_turbo_2step_rank_64_lora_checkpoint_info.md Which checkpoint the two weight files are β€” updated with every release
LICENSE.pdf Krea 2 Community License Agreement
NOTICE.txt Required attribution notice

Both weight files are one adapter β€” only key names differ. Both carry training details in safetensors metadata. File names never change: a better checkpoint replaces both in place, and every published checkpoint stays in _archive/checkpoints/ under its number.


Quick Start

diffusers

pip install git+https://github.com/huggingface/diffusers.git
import torch
from diffusers import Krea2Pipeline
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file

pipe = Krea2Pipeline.from_pretrained("krea/Krea-2-Turbo", torch_dtype=torch.bfloat16).to("cuda")

lora = hf_hub_download("lvladikov/Krea2-Turbo-Distill-2step-LoRA", "krea2_turbo_2step_rank_64_lora.safetensors")
state = load_file(lora)
state = {f"transformer.{k}": v for k, v in state.items() if not k.endswith(".alpha")}
pipe.load_lora_weights(state, adapter_name="2step")

image = pipe("a fox in the snow", num_inference_steps=2, guidance_scale=0.0).images[0]
image.save("krea2_2step.png")
  • num_inference_steps=2 is the whole config. Pipeline applies Turbo's fixed timestep shift (ΞΌ = 1.15) and evaluates at Οƒ = 1.0, 0.7595 β€” exactly the two points the LoRA was trained on. Keep guidance_scale=0.0.
  • Strength: pipe.set_adapters(["2step"], adapter_weights=[0.75]). Stock Turbo for comparison: pipe.unload_lora_weights() + num_inference_steps=8.
  • Use the diffusers file, not _comfyui β€” diffusers reads Krea's own key naming, not ComfyUI's lora_down/lora_up.

ComfyUI

file put it in
krea2_turbo_2step_rank_64_lora_comfyui.safetensors ComfyUI/models/loras/
krea2_turbo_bf16.safetensors (from Comfy-Org/Krea-2) ComfyUI/models/diffusion_models/
qwen3vl_4b_bf16.safetensors (same repo) ComfyUI/models/text_encoders/
qwen_image_vae.safetensors (same repo) ComfyUI/models/vae/

Load krea2_turbo_2step_lora_t2i.json. Full bf16, no quantisation, runs on CUDA / Apple Silicon / CPU unchanged. Smaller Turbo builds work too β€” ComfyUI's loader applies the LoRA to any build β€” except fp8_scaled on Apple Silicon (MPS has no Float8_e4m3fn).

Settings: steps 2, cfg 1.0, sampler euler / simple, LoRA strength 1.0.

βš™οΈ cfg 1.0, not 0.0. ComfyUI expresses "no CFG" as 1.0 (one forward pass); diffusers uses 0.0. Setting 0.0 in ComfyUI is not the same thing.

Using it on Raw

Trained on Turbo, for Turbo β€” it loads on Krea 2 Raw because the architecture is shared, a side effect rather than a supported mode: it is neither trained nor tuned for Raw's weights, steps or guidance.


Performance (1024Γ—1024, Apple Silicon MLX bf16)

denoise per model call GPU peak
Turbo 8 steps (quality bar) 76.4 s 9.5 s 25.2 GiB
Turbo 2 steps, no LoRA 19.4 s 9.7 s 25.2 GiB
Turbo 2 steps + this LoRA 19.3 s 9.6 s 25.2 GiB

Denoising is 4.0Γ— faster than the 8-step bar β€” two model calls instead of eight. The adapter adds no measurable cost per call and no measurable memory; the runs with it came in marginally faster, which is noise, not a speed-up.

Prompt encoding and VAE decode don't change with step count, so end to end sits below 4Γ— and rises toward it as the render grows. Denoise times at every trained resolution (6.5 s at 512Γ—512 to 38.3 s at 1440Γ—1440) are in the Detailed Model Card.


LoRA Strength

strength what happens
below 1.0 Correction only partly applied β€” softer skin and hair, less fine structure, closer to 2 steps without the adapter
1.0 Trained point, recommended
1.0–1.5 Extrapolation β€” texture denser than the subject warrants, fine structure reads wiry rather than sharp. Usable per prompt
above 1.5 Not recommended, not measured

At two steps the dial scales the adapter's whole job β€” turning two coarse calls into a finished image β€” so there is less reason to go below 1.0 than with a 4-step adapter. Reach for steps before strength: when quality matters more than speed, the 4-step adapter is the better tool.

LoRA strength comparison


Current Checkpoint

chk00041320 (2 Oct 2026) replaces chk00031600 (25 Sep 2026). It is 9,720 training samples later, aimed at what matters first in a picture β€” subjects drawn twice or two poses blended into one at the large sizes, and faces in busy scenes β€” through the teacher's own layout held at the large sizes, eight critics, and every update to the running average checked for doubled subjects. It is also the first checkpoint measured on the full 22-prompt sweep.

axis chk00031600 chk00041320
structure renders with a fault the teacher's does not have (of 21) 1 0
ghosting, blur-invariant index, sweep mean 1.24Γ— 1.21Γ—
market faces at 1280Β², layout closeness to the teacher 0.71 0.79
grain in flat areas, sweep mean 1.07Γ— 1.03Γ—
fine texture vs the teacher, 1440Β² 1.07 1.03
judge prefers the teacher (of 66) 15 12
blind adherence rubric, points missing of 352 2 2
saturation vs the teacher, sweep mean 0.98Γ— 0.99Γ—

Fewer doubled subjects and less ghosting at the large sizes, whole faces in the busy market scene, and prompt-following a step closer to the teacher's, while colour and fine detail stay at the teacher's level. krea2_turbo_2step_rank_64_lora_checkpoint_info.md always names the checkpoint in the weight files; the full comparison is in the Detailed Model Card.


Known Issues

The usual costs of two steps, in order of how often they show:

  • Small subjects β€” the weak spot, people and objects alike. Faces in a crowd, a figure in a wide scene, the machines at the back of a room can come out ghosted, smeared or misshapen; a portrait-sized face or an object up close holds up
  • Fine structure β€” feathers, hair strands and signage can be soft or a few pixels out of register, most at 1280Γ—1280 and above; a faint doubled contour can show on limbs
  • Style β€” on stylised prompts, how the picture should look (crisp linework, brush strokes, fingerprints in clay, a matte-painting finish) is followed less faithfully than what should be in it
  • Repeats β€” on busy action or crowd scenes the composition can repeat itself: an extra hand or held object, a figure duplicated in a crowd
  • Different composition β€” two steps is a shorter path from the same noise, so framing, pose or arrangement can differ from the 8-step render at the same seed. Treat the teacher's render as a quality reference, not the picture two steps will reproduce
  • Skin β€” smoother than the teacher's at the larger sizes, the pores softer, though its colour is now close to the teacher's; freckles come out as dots, softer than the teacher's

Every one is being worked on; none is hidden in the sweeps or the examples.


Method

Distribution matching (DMD2 family) with a trajectory anchor, Krea 2 Turbo as its own teacher, on the recorded 8-step trajectories.

The student makes two calls, at Οƒ = 1.0 and 0.7595 β€” the first and fifth points of the teacher's 8-step grid at ΞΌ = 1.15 β€” with stock Euler between them. Euler's first step lands exactly on the flow-matching interpolant at Οƒ = 0.7595, so the first call's output is a legitimate image prediction and is judged as one.

  • Distribution term β€” the frozen teacher and a fake-score adapter (rank 32, trained online on the student's current output, 3 updates per student step, discarded at the end) each denoise a freshly noised copy of the student's image; where they disagree is the direction toward the teacher's work. Averaging is never rewarded, so the student commits
  • Trajectory anchor β€” regression on the recorded chords at half weight keeps the student on the teacher's two-step grid

On top, each capped relative to the distribution term:

  • Spectral match β€” student and teacher images compared through radial power spectra, on the whole latent and a decoded 256-px window, two-sided β€” the term that reached the 16/8-px grid grain at large sizes
  • Eight critics taking turns β€” artefact, photo (half real photographs), large faces, the first call's layout against the teacher's own state, content, prompt, text, and the teacher's own finish, all heads on the frozen base's mid-network features; one pushes per step, each filtered and capped
  • Detail held to the teacher region by region β€” the anchor counts fine-detail error twice (3.5Γ— where the teacher is most detailed); a one-sided photo floor at 3–10 px and a per-tile ceiling at 1.3Γ—; a direction-aware term; windows on eyes, nose and lips; the teacher's finish of the student's first call as the second call's target; a smoothness limit on the fake adapter
  • Colour band β€” saturation held between the teacher's own level and 8% above it
  • Layout at the large sizes β€” at every size with a 1280- or 1440-px side the distribution term pushes at half strength while the anchor keeps full weight, and the anchor counts the second call's coarse layout (64 px and up) twice

Shipped adapter is a guarded running (EMA) average of the weights, not the last live state: short excursions are kept out, and every update is checked for doubled subjects before it goes in.


What the LoRA Touches

Rank 64, alpha = rank (scale 1.0), bf16 β€” the same 228 modules as the 4-step adapter:

  • 224 block linears β€” all 28 transformer blocks: to_q, to_k, to_v, to_gate, to_out.0, ff.gate, ff.up, ff.down
  • 4 global linears β€” time_embed.linear_1, time_embed.linear_2, time_mod_proj, final_layer.linear β€” the ones a step-count change needs most

Nothing about the base model changes.


Training Data

  • 13,750 recorded teacher trajectories from the 4-step project, prompts from Lakonik/t2i-prompts-3m β€” Turbo's own 8 steps at ΞΌ = 1.15, guidance 0.0, every latent and velocity stored; a 2-step chord is two 4-step chords end to end
  • 203 held-out prompts β€” measure the student–teacher gap, never receive a gradient
  • 43,044 real-photo crops from the 4-step project β€” set the photo floor (one precomputed statistic) and make up half of the photo critic's real examples

Only the photo critic's head sees the photos; the student and the fake adapter receive only its filtered, capped push. No new prompts, text embeddings or teacher runs.


Resolutions (12 buckets)

512Γ—512 512Γ—768 768Γ—512
768Γ—768 768Γ—1024 1024Γ—768
1024Γ—1024 960Γ—1280 1280Γ—960
1280Γ—1280 1440Γ—1280 1440Γ—1440

Same 12 buckets as the 4-step adapter, with the larger sizes and stylised prompts drawn more often: 1440Γ—1440 about 6% of the samples, 1024Γ—1024 and 768Γ—1024 12% each, stylised prompts 12%.


Hardware

Trained on a single RTX 3090 (24 GB). Frozen base weight-only int8. Every term fits up to 1440Γ—1440 via:

  • Checkpointed block inputs staged to pinned host memory above 0.3 MP
  • Student, fake adapter, spectral and detail terms, critic and teacher finish each build and free their own graph β€” peaks never overlap
  • Hard memory ceiling below the driver's paging threshold, so a step that doesn't fit fails loudly

A full step with every term live reserves 22.4 GB at 1440Γ—1440. Throughput is **58 samples/hour** against the 4-step recipe's 470 β€” roughly a dozen model runs per sample instead of two. That is the objective's cost, not teacher generation: the trajectories were recorded once and are read from disk.

Released LoRA is bf16.


Usage Notes

  • 🎯 Krea 2 Turbo only β€” trained against Turbo's weights and schedule; on Raw it is a side effect (see Using it on Raw)
  • 🚫 Keep guidance at 0.0 (ComfyUI: cfg 1.0, not 0.0)
  • πŸ“ Keep ΞΌ = 1.15 β€” the two training sigmas are anchored to that grid
  • πŸ§ͺ Not a finished adapter β€” for previews and drafts; a later checkpoint replaces this one only when the sweeps and I visually agree it is better
  • πŸ”¬ Training used an int8 base; released LoRA is bf16

Examples

Every sheet below: base model (8 steps), base at 2 steps without LoRA, base at 2 steps with LoRA β€” same seed throughout. Compare panels 2 vs 3 to isolate LoRA effect; panel 1 is the quality bar, not a pixel target. NFE = steps (Turbo is CFG-free).

Portrait of a young woman with freckles and windswept auburn hair...

portrait comparison

Turbo 8 steps Turbo 2 steps, no LoRA Turbo 2 steps + LoRA
8 steps 2 steps 2 steps + LoRA

Kingfisher bursting out of water...

kingfisher comparison

Turbo 8 steps Turbo 2 steps, no LoRA Turbo 2 steps + LoRA
8 steps 2 steps 2 steps + LoRA

Rainy night city street with glowing neon signs...

neonstreet comparison

Turbo 8 steps Turbo 2 steps, no LoRA Turbo 2 steps + LoRA
8 steps 2 steps 2 steps + LoRA

(19 more examples, and a two-panel sheet against the teacher for every prompt, in the full version β€” see assets/ for all 22 prompts)


Resolution Sweeps

assets/resolution_sweeps/ β€” this LoRA at every trained resolution for all 22 test prompts (same prompts, seed, 2 steps, strength 1.0). Nothing cherry-picked.

_teacher-8step/ β€” official Krea 2 Turbo 8-step reference renders for the same prompts/seeds/resolutions. _turbo-base-NO-LoRA-2step/ β€” stock Turbo at 2 steps, the floor.

Layout:

assets/resolution_sweeps/
β”œβ”€β”€ _teacher-8step/              8-step stock Turbo reference
β”‚   β”œβ”€β”€ 512x512/
β”‚   β”œβ”€β”€ 1024x1024/
β”‚   └── ... (all 12 buckets)
β”œβ”€β”€ _turbo-base-NO-LoRA-2step/   Same tree, stock Turbo at 2 steps (the floor)
└── 2step-LoRA/                  Same tree, this LoRA at 2 steps (the published checkpoint)

Archive

Every published checkpoint and its resolution sweep under _archive/ β€” checkpoints/ and resolution_sweeps/, each under its number. Superseded, not maintained.


What's next

Training continues from this checkpoint, one recipe change at a time, each kept only if the pictures do not degrade at any resolution β€” aiming at the best quality two steps can give, not at matching the 4-step LoRA.

Next, aimed at the Known Issues: small subjects in wide scenes; the finest edges and skin texture at the largest sizes, lifted to the teacher's level without bringing the grain back; the layout at the largest sizes, where two plausible poses of the same subject can meet; the fine detail of close-up faces β€” eyes and skin texture; the tactile surface of stylised materials such as clay β€” none of it allowed to cost the structure, the colour and the clean large sizes this checkpoint gained.

A better checkpoint replaces this one when the sweeps and I visually agree; until then the 4-step adapter remains the recommendation for quality renders, and this one is the fast preview.


Detailed Model Card

For more details, if interested, have a look at the Detailed Model Card.


License

This adapter is a Derivative of Krea 2 Turbo under the Krea 2 Community License Agreement. Everything the agreement says about Krea 2 Turbo applies to this LoRA: Acceptable Use Policy, revenue threshold for commercial use, content-filtering duty for deployments.

Copy of agreement: LICENSE.pdf, required notice: NOTICE.txt. See krea.ai/krea-2-licensing.

Not an official Krea product, not endorsed by Krea. Base model is Krea's; adapter weights and everything else in this repo are mine.

Downloads last month
5,045
Inference Providers NEW

Model tree for lvladikov/Krea2-Turbo-Distill-2step-LoRA

Base model

krea/Krea-2-Raw
Adapter
(1694)
this model

Spaces using lvladikov/Krea2-Turbo-Distill-2step-LoRA 2