lvladikov commited on
Commit
f0fdea0
Β·
verified Β·
1 Parent(s): 9bb040d

Upload 3 files

Browse files
Files changed (4) hide show
  1. .gitattributes +1 -0
  2. LICENSE.pdf +3 -0
  3. NOTICE.txt +5 -0
  4. README.md +191 -0
.gitattributes CHANGED
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
 
 
33
  *.zip filter=lfs diff=lfs merge=lfs -text
34
  *.zst filter=lfs diff=lfs merge=lfs -text
35
  *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ LICENSE.pdf filter=lfs diff=lfs merge=lfs -text
LICENSE.pdf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b82a2805162bde714a4eb27b9063c4fc3345d08a30be055134a6160e5430ba74
3
+ size 137711
NOTICE.txt ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ Krea 2 is licensed under the Krea 2 Community License Agreement. For more information, visit https://krea.ai/krea-2-licensing.
2
+
3
+ This repository distributes a Derivative of the Krea 2 Turbo model, as defined in that agreement: a LoRA adapter produced by distilling Krea 2 Turbo to a 2-step schedule. The Krea Model has been modified. This adapter is not an official Krea product and is not endorsed by Krea.ai, Inc.
4
+
5
+ Use of this adapter is subject to the Krea 2 Community License Agreement (LICENSE.pdf in this repository) and the Krea Acceptable Use Policy (https://krea.ai/krea-2-use-policy), exactly as use of Krea 2 Turbo itself is.
README.md ADDED
@@ -0,0 +1,191 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ base_model: krea/Krea-2-Turbo
3
+ base_model_relation: adapter
4
+ license: other
5
+ license_name: krea-2-community-license
6
+ license_link: https://huggingface.co/lvladikov/Krea2-Turbo-Distill-2step-LoRA/blob/main/LICENSE.pdf
7
+ library_name: diffusers
8
+ tags:
9
+ - lora
10
+ - text-to-image
11
+ - distillation
12
+ - step-distillation
13
+ - distribution-matching
14
+ - krea-2
15
+ pipeline_tag: text-to-image
16
+ ---
17
+
18
+ # Krea 2 Turbo β€” 2-Step Distillation LoRA
19
+
20
+ > 🚧 **Work in progress.** This page is the live account of the project β€” what the adapter is meant to do, how it is
21
+ > being trained, and where it stands. It is updated as the recipe changes. There are no shipped files yet; the samples
22
+ > and side-by-side comparisons will be added when there is a candidate worth showing.
23
+
24
+ A LoRA for **[Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo)** that takes the model from its usual **8 steps
25
+ down to 2** β€” Turbo's own weights and its own two sigmas, guidance 0.0, a quarter of the denoising passes β€” with the
26
+ aim of matching what the **[4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** delivers
27
+ today, at every resolution it delivers it at.
28
+
29
+ - 🎯 **The bar** β€” the [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s output quality, across the same 12 resolutions, in two steps.
30
+ - πŸ”Œ **Drop-in, no exceptions** β€” a plain LoRA sampled by stock Euler at sigmas `[1.0, 0.5128]` in diffusers, ComfyUI
31
+ or MLX. No custom sampler, no policy head, no per-step tricks. If the quality needs a special sampler it is not this
32
+ project.
33
+ - 🧬 **Same shape as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)** β€” rank 64 on the same 228 modules; a second adapter exists during training
34
+ only and never ships.
35
+ - πŸ“Š **Distribution matching, not imitation** β€” the training objective that finally moved this (see [Method](#method)).
36
+ - 🎲 **The same 13,750 recorded teacher trajectories** the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) trained on, reused without a single teacher
37
+ re-run.
38
+ - πŸ–₯️ **One RTX 3090**, and a recipe shaped by its 24 GB.
39
+
40
+ ## Where it stands
41
+
42
+ | | |
43
+ | --- | --- |
44
+ | lineage | [4-step LoRA](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) (78,000 samples) β†’ 9,000 samples of 2-step trajectory distillation β†’ distribution matching from there |
45
+ | current run | **run 2**: distribution matching + a critic judged against the teacher's own finals, per-resolution weights measured from a full 12-bucket sweep |
46
+ | best result so far | run 1 at 1,000 samples: at 512Γ—512 and 768Γ—1024 the sharpest, most coherent 2-step renders of the project β€” crowds resolved into people, faces intact at 768Γ—1024; at 1280Γ—1280 and above the same weights over-render (fine texture at 1.42Γ— the teacher's), which run 2 addresses |
47
+ | known gaps | faces at 512Γ—512 still carry a faint doubled contour; fine structure at 1280Γ—1280+ renders as fragments rather than the teacher's objects β€” the two things run 2 exists to fix |
48
+
49
+ ## How I got here
50
+
51
+ The [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) closed its page with a promise: a 2-step LoRA as the next project, and a guess at the lever it would
52
+ need β€” matching the teacher's *distribution* rather than its trajectory. That guess turned out to be the whole story.
53
+
54
+ The project began where the [4-step one](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) ended, from its final weights, and ran the same recipe at two steps:
55
+ **progressive distillation** on the recorded teacher trajectories, each student call covering four teacher steps, with
56
+ the LADD-style critic as the finisher. Nine thousand samples in, every number had stopped moving and the pictures had a
57
+ signature the numbers could not see: doubled contours on faces and limbs, soft fine texture, crowds averaged into
58
+ translucent overlaps. Several variations followed β€” the critic re-weighted, judged per token, a heavier hand on the
59
+ final call, the student's own first-step output fed into its second β€” and each traded one of those faults for another
60
+ without moving past them. A capacity probe ruled out adapter rank; a learning-rate shock ruled out the optimiser.
61
+
62
+ The reason is structural, and worth stating plainly because it decides the whole design. A regression loss asks the
63
+ student to land on the teacher's *specific* image for each prompt. When a two-step jump is wide enough that several
64
+ images are plausible, the answer that minimises the squared error is their average β€” and the average of two sharp
65
+ images is a blurred one with doubled edges. Every earlier recipe rewarded that average. Tuning its weights could not
66
+ change what it rewarded.
67
+
68
+ **Distribution matching** asks a different question: not "does your image match this one" but "would the teacher
69
+ plausibly have produced your image". The first run of that objective, on top of the 9,000-sample weights, produced
70
+ in a thousand samples what twenty thousand samples of the old recipe never had β€” and it did so while every latent
71
+ distance to the teacher *rose*, which is exactly what a mode-seeking objective predicts and what a mean-seeking metric
72
+ punishes. The distances are reported on this page; they are not optimised for, and they are not what decides a
73
+ checkpoint. Pictures are, at fixed seeds, at every resolution, with faces viewed at 1:1.
74
+
75
+ The recipe adjustments so far, each made on the measurement of the one before:
76
+
77
+ 1. progressive distillation at two steps from the [[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) weights, with the LADD critic β€” the baseline
78
+ 2. the critic made to judge the first call's endpoint against the teacher's mid-states, which removed the gross
79
+ ghosting and left the doubled contours
80
+ 3. the critic's weight, its per-token form and the student's own first-step output as the second call's input β€”
81
+ tried one at a time; the objective flip that mattered was not among them
82
+ 4. **distribution matching** (the DMD2 family) as the primary objective, trajectory regression demoted to an anchor at
83
+ half weight, the critic switched off to read the new term alone
84
+ 5. the running average of the weights restarted at the objective switch, after it was caught averaging two lineages
85
+ into composites
86
+ 6. per-resolution weights for the distribution term, measured from a full 12-bucket sweep rather than estimated
87
+ 7. the critic back in, in its DMD2 form β€” on the trained score network's features, judged against the teacher's own
88
+ finals, conditional on the prompt, its gradient capped against the distribution term's
89
+
90
+ ## Method
91
+
92
+ **Distribution matching with a trajectory anchor**, Krea 2 Turbo as its own teacher, on the recorded 8-step
93
+ trajectories.
94
+
95
+ The student makes two calls, at Οƒ = 1.0 and Οƒ = 0.5128 β€” the first and fifth points of the teacher's 8-step grid at
96
+ mu = 1.15 β€” and stock Euler carries it between them. That grid is what makes the objective a drop-in: Euler's first step
97
+ from pure noise lands *exactly* on the flow-matching interpolant at Οƒ = 0.5128 with the same noise and the student's
98
+ own clean-image prediction as the data point. So the student's first-call output is a legitimate image prediction that
99
+ can be judged as an image, and the second call is fed from it during training the way it will be at inference.
100
+
101
+ **The distribution term.** For an image the student produces, two denoisers estimate how it should be cleaned up from a
102
+ freshly noised copy: the frozen teacher, and a second small adapter on the same frozen base β€” the *fake score* β€” that
103
+ is trained online to denoise whatever the student currently makes. Where the two disagree is the direction that makes
104
+ the image more like the teacher's work and less like the student's habits, and the student is pushed that way
105
+ (the DMD2 gradient, per-sample normalised). Averaging is never rewarded, so the student commits. The fake adapter is
106
+ rank 32, starts as an exact copy of the teacher, updates twice per student step, and is discarded at the end.
107
+
108
+ **The anchor.** Plain trajectory regression on the teacher's recorded chords stays in at half weight. It keeps the
109
+ student on the teacher's two-step grid so the distribution term cannot wander into a different sampler behaviour, and
110
+ it is what the 9,000 earlier samples had already satisfied β€” which is why the first distribution-matching run moved
111
+ so far so fast.
112
+
113
+ **Per resolution.** The distribution term's push grows with resolution: the fake adapter sees few large-bucket samples
114
+ and under-fits fine structure there, and a per-pixel normaliser lands harder as pixel counts grow. Rather than guess a
115
+ taper, a full 12-bucket sweep of the run-1 weights measured the fine-texture ratio to the teacher at every resolution,
116
+ and the term's weight per bucket is set from that measurement so that each bucket is pushed toward the teacher's
117
+ texture rather than past it.
118
+
119
+ ### The critic (run 2)
120
+
121
+ The fake adapter cannot notice its own mistakes, and at large resolutions those mistakes render as fragments β€”
122
+ right amount of detail, wrong structure. A **discriminator** can: trained to tell the student's images from the
123
+ teacher's, it learns whatever systematic difference exists between them, fragments included, and its gradient points
124
+ back at the teacher. This is the second half of DMD2 and it differs from the [[4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) critic in two ways that
125
+ follow from what the earlier runs taught:
126
+
127
+ - it sits on the **trained fake adapter's features**, not the frozen teacher's, so it can learn a decision rather than
128
+ reweight what the teacher already computes;
129
+ - its real class is the **teacher's own finals**, conditional on the prompt β€” the bar is the teacher β€” and its gradient
130
+ is **capped per sample** at twice the distribution term's, so it can correct but never take over.
131
+
132
+ ## What the LoRA touches
133
+
134
+ Rank **64**, alpha = rank, bf16, the same **228 modules** as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA): the 8 attention and feed-forward linears
135
+ of all 28 transformer blocks, plus the four global linears β€” `time_embed.linear_1`, `time_embed.linear_2`,
136
+ `time_mod_proj`, `final_layer.linear` β€” that a step-count change needs most. Nothing about the base model changes.
137
+
138
+ ## Training data
139
+
140
+ The **13,750 recorded teacher trajectories** of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) β€” Krea 2 Turbo's own 8-step run at mu = 1.15 and
141
+ guidance 0.0, every latent and velocity stored β€” serve unchanged: a 2-step chord is two of the 4-step chords end to
142
+ end. 203 held-out prompts measure the student–teacher gap on unseen prompts and never receive a gradient. The critic's
143
+ real class is the teacher's finals for the training prompts; the 43,044 real-photo crops of the [4-step project](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) are
144
+ available to it as a second source and are not used in the current run.
145
+
146
+ ## Resolutions
147
+
148
+ The same 12 buckets as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA), interleaved in proportion to their remaining samples:
149
+
150
+ | | | |
151
+ | --------- | --------- | --------- |
152
+ | 512Γ—512 | 512Γ—768 | 768Γ—512 |
153
+ | 768Γ—768 | 768Γ—1024 | 1024Γ—768 |
154
+ | 1024Γ—1024 | 960Γ—1280 | 1280Γ—960 |
155
+ | 1280Γ—1280 | 1440Γ—1280 | 1440Γ—1440 |
156
+
157
+ ## Hardware
158
+
159
+ One **RTX 3090 (24 GB)**. The frozen base is weight-only int8; the student's checkpointed block inputs stage to pinned
160
+ host memory above 0.3 megapixels; the student, the fake adapter and the critic each build and free their own graph
161
+ in turn, so their peaks never overlap; a hard memory ceiling sits below the driver's paging threshold so a step that
162
+ does not fit fails loudly. A full step with every term live reserves about 21.5 GB at 1440Γ—1440. The price of the
163
+ objective is throughput: about 160 training samples an hour against the [4-step recipe](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA)'s 470 β€” three times the cost
164
+ per sample, and so far a small fraction of the samples.
165
+
166
+ ## How it is judged
167
+
168
+ Every 1,000 samples, both the live weights and their running average are pulled, merged and rendered at fixed seeds on
169
+ 15 fixed prompts across four resolutions, with faces cut out at 1:1, a fine-texture ratio and frequency bands against
170
+ the teacher, a graded judge, and a pairwise preference against the teacher. Latent distances β€” the held-out chord gap
171
+ and the two-step rollout error β€” are recorded but never used to keep or stop a run: this project's clearest lesson is
172
+ that they reward blur, and a run that improved them while its pictures collapsed was stopped by its pictures.
173
+
174
+ ## Notes and limitations
175
+
176
+ - 🎯 **Krea 2 Turbo only**, at **2 steps**, guidance **0.0** (cfg 1.0 in ComfyUI), **mu = 1.15** β€” the two training
177
+ sigmas are anchored to that grid.
178
+ - 🚧 Not released. Everything on this page is the current state of a run in progress.
179
+
180
+ ## What's next
181
+
182
+ Run 2 is training. Its first probe decides whether the critic removes the large-resolution fragments; if it does not,
183
+ the next lever is a perceptual-space term on decoded image crops, which judges structure rather than energy. The
184
+ shipping point, when there is one, is chosen by the sweep and the eye, not by distance run β€” the same discipline as the
185
+ [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA).
186
+
187
+ ## License
188
+
189
+ The adapter is a derivative of [Krea 2 Turbo](https://huggingface.co/krea/Krea-2-Turbo) and is covered by the
190
+ **Krea 2 Community License Agreement** (`LICENSE.pdf` in this repository), as the [4-step adapter](https://huggingface.co/lvladikov/Krea2-Turbo-Distill-4step-LoRA) is. The attribution
191
+ notice the license requires of a derivative ships with the files.