Instructions to use SG161222/SPARK.Chroma_v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use SG161222/SPARK.Chroma_v1 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("SG161222/SPARK.Chroma_v1", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Feedback
Hi everyone! Here you can post your feedback about this model. Your feedback will help me make future versions better!
Currently, I will be busy exploring models such as Ernie and Z-Image. Training Chroma proved to be challenging for me, but I will continue to conduct new tests with Chroma to achieve significantly better results on my hardware.
Hi, I haven't tested the current iteration yet, but I have been following the progress of Spark Chroma since anons started posting about it last year.
I appreciate your work, and I do not mean to "armchair train" here but the dataset you have, 2500 images, is probably more suited towards lora training than finetuning. Perhaps Z-Image and Ernie can benefit from it as loras? It would also allow much quicker experimentation to determine what clicks with these models.
Based on license and legal reasons I expect the answer to be "No", but I would like to still try my luck by asking if it is possible to share the dataset you have curated, so that other people attempting to train realism loras for Chroma and other models can benefit from it?
Thanks regardless and good luck with your next project.
Downloading the new model now, will publish some images here soon
I was already greatly satisfied with the preview, it was HD capable (my gens were 1300x1000 and it was mostly fine with it), and it stabilized the style 100% of my needs.
I even do not know if anything else was to wish :), I will not(stiked) ->now test this _1024 as a replacement, but will be nice to see your vision on what exactly improved?
Downloaded it already and firing first attempts....
well, right from the start there is a problem, I just replaced SPARK.Chroma_preview.safetensors with SPARK.Chroma_v1_1024.safetensors, everything else was super stable and according to your recommendations, and it immediately started to generate defected faces in a row, like, 5 out of 5 gens already (was strictly ZERO of this for 1000+ gens with Chroma or your SPARK preview).
That's with like "3 people in the street, facing camera, full height", like that.
All 100% of faces started ending this.. not good.
Checked once again everything. I have the following -
Chroma-DC-2K-BF16.gguf
chroma_v10HD.safetensors
SPARK.Chroma_preview.safetensors
SPARK.Chroma_v1_1024.safetensors
every version generates just fine, this is all 100% standard - 1024x1024, 40 steps of simple / dpmpp_2m, fp16 of T5, all rock solid and basic,
and only this new SPARK.Chroma_v1_1024.safetensors is generating 100% of broken faces. The rest, except faces, seems more or less fine - approximately like was in preview.
but with these faces it is 100% broken as it seems, cannot find the cause in my setup.
... tried your exactly wokflow with zero changes, your exactly negative prompt, all the same, faces broken
?? a pity
... but at the same time I should repeat once again how good SPARK.Chroma_preview.safetensors is, at least for me, it is doing the job 100% - removes this annoying Chroma style variations and stabilizes to photo style, just superb, thanks a lot for your efforts!
Yes, you're right. There's likely no issue with your workflowβthe problem lies with the model itself. As I mentioned above, Chroma is quite finicky to train, based on my experience with it.
The preview was trained at 768px, while V1 at 512/1024px. The preview also used a Timestep Shift of 1.0, which is great for details, but anatomy requires higher values. I did adjust it for V1, but it seems something went wrong in the process.
At this point, I'm not sure I want to keep pushing forward with Chroma when models like Ernie and Z-Image are available. They're far more predictable and stable during training (I've already run a quick test on Ernie with promising results).
I sincerely apologize for the time you've spent waiting for this model. I hope the pivot to a different base model will yield much better results.
I find Chroma's reality with your preview very nice. It is a pity if it was an accidental achivement that cannot be repeated and improved straightfowardly :),
but all this AI stuff is a very, very fresh air and we all are like in a vacuum of knowledge and verything is anew and no guides from the past. I know this because I am also developing around all this.
So lets' have "good luck and keep going"!
I run a facial fixing pass on my workflow, which greatly aids in alleviating the issue presented
well, running extra pass of something like WAN2 can even remove extra legs and arms :), but I do not like when things get this complicated
I also run a lot of extra on selected images, for upscaling at least, with a lot of denoise which can fix faces if in need, maybe it will also fix everything here, but back to the question,
what was wrong with preview then? for me it was like a solid release;
what's bad with v1 we discovered; but that supposed to be good? compared to preview?
what's bad with v1 we discovered; but that supposed to be good? compared to preview?
There are more details in the image, especially on skin, from my experience with preview and these releases
I find skin detail very easy to adjust in any model by just altering detailer steps with low denoise (of around 0.2 - 0.3). Just add steps and you will get any level of details. Unfortunately, this only plays reliably with the skin.
Anyway I appreciate moving to new models - at least I can surely say that no matter what, models based on modern LLMs input, are a lot better at prompt processing, this T5 is an issue in Chroma anyway; it is more keyword-style programming, which is in 2026 sort of annoying. Yes it can parse interactions to some extent, but any modern is way, way better.
I run overnight generation with prompt sweep, for different photos. Single person's face is also deformed way too often. With preview or originan Chrome, faces were always ok-class, never broken, can improve or just leave, to the taste, was OK. Here strong face correction is mandatory, to the point that you should just renoise and redraw them.
After playing around with this model, my conclusions are as follows:
1 - this seems like a sidegrade compared to 512 and a downgrade to preview
2 - It is better at following some specific prompt details, like film grain
3 - It is worse at replicating non-photographic styles (unable to do american comic books, worse at black and white manga for two examples)
4 - It seems worse at generating faces on a crowd, something I didn't felt with 512 model
5 - the preview model still is the best overall, making a good compromise between realism, details and rendering non-photographic styles
6 - Chroma is still the best for NSFW work, being the only one that is able to generate correct male genitalia (not just on the man that is penetrating, but also on males on a crowd or by themselves) and blood/gore (being able to render carcasses, body interiors, blood drips etc. much more coherently than other models)
Preview was a very good try. It also managed to extend usable generation to 1536x1024 resolution without messing things too much - a great deal. Any other version of Chroma tend to hallucinate way more at this resolution, or generate stacked multi-images, while preview was solid at 1536, with maybe +20% more broken images at most, which is good rate overall.
Spark.Chroma is what convinced me to even try Chroma out; before the release of it, I wrote it off as a complete wash
Currently, I will be busy exploring models such as Ernie and Z-Image. Training Chroma proved to be challenging for me, but I will continue to conduct new tests with Chroma to achieve significantly better results on my hardware.
Not ready yet, but I also have high hopes for https://huggingface.co/lodestones/Zeta-Chroma
preview was very good. It felt a bit biased towards dark and sometimes a bit low res. But it has a very realistic feel to it. That maybe naive (i train loras a lot but not finetune models) , but... do what you did there with a higher res version? :) I know it's a bit mean after all the effort to find the preview better. But i think it seems to be the honest feedback here. And except for the slightly low res look, there is indeed nothing wrong with the preview.
I'm gonna test the new one, but the faces thing looks not promising. i do run chroma and spark preview often with a Flux "Super-realism" lora, which adds to it, especially for base chroma. maybe it helps.
little update here: it is really mainly the face issue that is letting it down. It seems for example better at anatomy in multi people scene than preview.
Also with the conclusions in this chat, consider moving the preview version as a alternative into the files section of this model, instead of keeping it (only) on a different one.
the preview everyone is talking about: SG161222/SPARK.Chroma_preview
The advantage of Chroma as a base is that it is fully uncensored and nsfw capable. Let's face it, that is a big part of why Chroma is so popular. And not the "oh it looks kinda like a pen** if you look at it from an angle" nsfw capable as some modern models are supposed to be, but aren't. Ernie is supposed to be, but that is apocalyptically shit at anatomy in general, which isn't a good start. It's fast but it's crap.
Chroma is actually capable, really fully knowing, what it is doing. Because it was trained with it. And the preview inherited that and made it look better. But i strongly suspect the core capabilities came from Chroma, which can do this stuff. Hard to tell without knowing the dataset. There is nothing i'm aware of in that league.
Except for some older SDXL finetunes like Illustrious (2.0 last i think, still great SDXL finetune, chroma level nsfw capable, for anime AND OTHER art style up to semi realistic 2.5d images.). or lustify, also sdxl. Which is still good (except gay stuff). And as SDXL finetunes they are ridiculously fast on modern hardware. Better at anatomy than most sdxl, but still rubbish with text). But you can use a modern image edit model to do that later. Or to transform a illustrious artsy image to realism.
So over the past few days, I've been comparing Ernie Image and Flux Klein 9B, and honestly, Klein 9B is the clear winner when it comes to trainability. I tested both on SFW and NSFW content, and from what I've seen, Klein 9B picks up NSFW stuff way better than Ernie.
Right now I'm testing Timestep Shiftβjust 1 or 2 more training runs to lock in the final settings. After that, I'll put together a basic NSFW dataset and kick off training on a mixed SFW+NSFW set.
Really hoping switching to this model is the right move. What do you think?
Will it be distilled or full after the training? Distilleds like ZIT/Klein are annoying at generating absolutely the same person no matter what. The image and composition itself may be perfect, not talking about speed for sure, but lack of any diversity is showstopper in many of scenarios.
Will it be distilled or full after the training? Distilleds like ZIT/Klein are annoying at generating absolutely the same person no matter what. The image and composition itself may be perfect, not talking about speed for sure, but lack of any diversity is showstopper in many of scenarios.
I ran all my tests on the non-distilled versions of the models, and I'm planning to train only the non-distilled version of Klein 9B.
Some LoRAs work on the distilled versions even though they were trained for non-distilled versions.
Will you also try the 4B variant of FLUX or just the 9B?
Some LoRAs work on the distilled versions even though they were trained for non-distilled versions.
Will you also try the 4B variant of FLUX or just the 9B?
So far I've only tested the 9B versionβgonna try out the 4B model later.
Klein 9B has higher quality than Ernie, so unless Ernie responds exceptionally well to training (which doesn't seem to be the case) then there are no strong reasons to go with Ernie over Klein.
The distill is obviously preferred by most people for inference, but it would be fairly insane to do any serious finetuning on the distilled checkpoint while a more receptive and malleable CFG capable base alternative exists.
Once a great 9b-base checkpoint is created, I am sure there will be variety of options to make inference easier (quants, making a new distill, diffing base and distill and using that as distill lora over the new checkpoint, diffing new checkpoint and old 9b-base and using that as lora over distilled checkpoint, etc.)
Klein 9B has a license with certain restrictions though, you may want to give that a careful read before sinking a lot of time and money. If none of those are absolute deal breakers, then I would recommend going with 9b over 4b for training, it has significantly higher quality than the smaller counterpart. Also there is a case of a 4b finetuning attempt failing (Kaleidoscope), although this is not conclusive evidence that 4b can't be finetuned.
Lastly, I would prioritize curating a large number of high quality images over smaller dataset with most perfect hand-reviewed captions. If you can hit 10k mark perhaps, that could facilitate training a checkpoint without needing detrimental amount of regularization.
The flux 2 vae converges faster than flux 1 vae inside Chroma, so this number of images, and perhaps more, should still be possible to train with your setup under a sane time frame.
Best of luck in your endeavors.
Once a great 9b-base checkpoint is created, I am sure there will be variety of options to make inference easier (quants, making a new distill, diffing base and distill and using that as distill lora over the new checkpoint, diffing new checkpoint and old 9b-base and using that as lora over distilled checkpoint, etc.)
this one is damn good. pixel space vae so no encoder slamming gpu/cpu
https://huggingface.co/Lakonik/AsymFLUX.2-klein-9B
ive been doing testing and it has some tiling going on but im currently working on a node to remove it, just testing its quality.
Ernie is just crap. At least for people. I don't understand how it has so many likes. It doesn't need any complicated prompt or scene to mess anatomy up. A simple person has almost always something obviously wrong.
So you fight It's bad anatomy beside the nsfw stuff. Flux.2 has anatomy issues too in all models, but it gets worse, the smaller it gets. So 4b is by far the worst. There are some anatomy fixer loras that help a bit.
of course Klein 4b is also by far the easiest and fastest to train but that comes at too high a price. It's klein (small) and fast for a reason. 9b is the better choice.
Another important point to consider is prompt understanding. Viewstyle can be adjusted by re-gens quite easily, any detailer workflow will apply the style, including across models, just fine. At upscaler/refiner stage you almost always adjusting generation, you can as well use other or even custom models to adjust more, and it works good enough. Take some time but 100% solvable.
The harder thing is to get the composition right. And here we have a clear grade of something like, simplifed,
- SDXL keywords-based
- T5 encoders/embedders
- QWENs encoders/embedders
- QWEN models
many out of scope, but like that.
Any next step is a major upgrade. 2->3 is sometimes required to get the scene right: for example, with Chroma which is T2, it is kind of possible to make 2 person interaction and have some success with it, but 3 person is close to impossible no matter how you refine the prompt. Step up into type 3 area, and 3 persons will interact just fine as listed without any special measures.
These days I often compose with ZIT/Klein now, or even QWEN V models, and then refine with Chroma.
But this is not always possible when the model denies to compose.
Short TDLR - This is another important point to consider. Try complex prompts, not only complex, but with some interaction that you want to get EXACTLY. Not exactly the mood as above - the mood will be catched by the T5 perfectly. The action. This is important.
talking of complex prompts and interactions, i actually think Chroma is about as capable as qwen or Flux.2 (minus the editing of course) . Long prompts with multiple people, quite a few moving parts (THOSE and in general) and it is remarkably good at getting it right. Which is remarkable, because it is based on Flux.1 Schnell which was also catastrophically bad at anatomy.
It also proves that T5XXL gets a way too bad press compared with all the fancy (and different one !) llm's modern models use as clip. The only thing that feels outdated about Chrome (Spark and normal) is the resolution. In terms of image quality, Flux.2 is probably better, Qwen not really.
Sure, T5 based models understand the prompt well up to the whole T5 size. Which is a lot.
The problem is that after some complexity threshold you start getting "as written" in less than 10%, the rest is perfectly themed by the prompt, but important details went wrong.
This is not the case with newer embedders.
you say that. And i hear that. I use Chroma, Wan and some others still that use T5. And i think you and "they" are wrong. Frankly, i call bs on this new LLM-clip meta. And even the limits for a full size T5 are beyond what even i need. While for llm clips i still get recommended keeping it under 500 or whatever tokens. Which i ignore in both cases.
Of course you have to prompt differently, not with these flowery short stories that are so recommended these days, and often unnecessary in many non distilled models. So not only can it handle long prompts, most models do not actually NEED these book sized prompts that are so hyped up these days. I.E. the limited tokens of T5 are fine in the first place!
Some people write super long over the top detailed prompts that just suffocate the generation on more capable (non-distilled) models. (no, not "these people", but their OTHER LLM, they need to write for the first LLM, because as clip it has no bloody benefit. )
But even so, T5 (plus the model itself of course) figures it out as well.














