Text-to-Image
Diffusion Single File
English
anima
comfyui

This model seems to have forgotten some concepts learned from the ArtStation dataset

#21
by oblevdor - opened

I tested concepts in computer graphics or CG such as subsurface scattering, dust particles, contact shadows, edge wear, etc. I combined these concepts with common booru image types for generation, everything else being the same except the diffusion model. Compared to base 1.0, the output of 2.9B preview 1 noticeably lacks these concepts, making the images quite mediocre (or more booru style? Depends on subjective aesthetics).
The comparison sample image is a bit sensitive.

https://p.inari.site/guest/26-08/15/6a7ff64631092.png

Looks more like a CFG/sampler issue, you can see the right image has way more contrast, you can try euler ancestral cfg++ with 1.5/2 CFG.

I tested concepts in computer graphics or CG such as subsurface scattering, dust particles, contact shadows, edge wear, etc. I combined these concepts with common booru image types for generation, everything else being the same except the diffusion model. Compared to base 1.0, the output of 2.9B preview 1 noticeably lacks these concepts, making the images quite mediocre (or more booru style? Depends on subjective aesthetics).
The comparison sample image is a bit sensitive.

https://p.inari.site/guest/26-08/15/6a7ff64631092.png

This is of course a subjective opinion, but I like the 2.9b version better for the richness of the image

I tested concepts in computer graphics or CG such as subsurface scattering, dust particles, contact shadows, edge wear, etc. I combined these concepts with common booru image types for generation, everything else being the same except the diffusion model. Compared to base 1.0, the output of 2.9B preview 1 noticeably lacks these concepts, making the images quite mediocre (or more booru style? Depends on subjective aesthetics).

For those who have studied art theory, it's not really a subjective opinion. The left is objectively better. People who don't know much about art would prefer the image on the right, simply because it's sloppier and shinier. They don't understand the concepts you're talking about.

The base 1.0 model is actually not the best Anima model in terms of academic art theory either. Once training prioritized the characters knowledge, everything went down the drain. This is why I prefer the base+previews mixes.

Sign up or log in to comment