File size: 10,439 Bytes
28bffb6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
---

base_model:
- circlestone-labs/Anima
language:
- en
- zh
license: other
license_name: circlestone-labs-non-commercial-license
license_link: https://huggingface.co/circlestone-labs/Anima/resolve/main/LICENSE.md
library_name: diffusion-single-file
pipeline_tag: text-to-image
tags:
- text-to-image
- comfyui
---


<p align="center">
  <img src="banner.jpg" width="100%">
</p>

<p align="center">
English |
<a href="https://huggingface.co/Johnny-Z/Anima-Light-Lavender/blob/main/README-cn.md" target="_blank">δΈ­ζ–‡</a>
</p>

> **TL;DR**: A major boost to natural-language understanding at the 512-token scale; prompt with a long natural-language description by default (the `image_description` field). **We recommend following the "Quick Start" guide below and assembling the structured caption with comfyui-zako-pe**; if you don't want to write a long description, you can also provide just tags and let [zako-pe](https://huggingface.co/Johnny-Z/zako-pe) expand them.



## πŸ“Œ Overview



Anima-Light-Lavender is a version obtained by post-training on top of **Anima-Base v1.0** (`anima-base-v1.0.safetensors`) from the [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima) series. Compared with the base model, this release focuses on **improving the model's understanding of natural-language descriptions at the 512-token scale**, and leverages higher-quality training data to improve the overall generation quality.



This release makes **no architectural changes**: the layer count, parameter count, and file layout are all identical to the base model, so you can swap in the upstream weights directly and plug it into existing workflows.



| Item | Description |

| :--- | :--- |

| Base model | Anima-Base v1.0 (`anima-base-v1.0.safetensors`) from the [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima) series, built on [nvidia/Cosmos-Predict2-2B-Text2Image](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Text2Image) |

| Parameters / Architecture | Identical to the base model (~2B parameters, a 28-layer DiT); no added layers, no distillation |

| Task type | text-to-image: anime illustrations, characters, and stylized art |

| Text encoder | Qwen3-0.6B (`qwen_3_06b_base.safetensors`) |
| VAE | Qwen-Image VAE (`qwen_image_vae.safetensors`) |
| Weight files | `anima-light-lavender.safetensors` (single file, BF16); an MXFP8 version `anima-light-lavender_mxfp8.safetensors` is also available (single file, faster inference) |
| Training data | ~1.7M Danbooru images annotated and filtered by an in-house data pipeline; **all models used in the pipeline β€” annotation, filtering, preference judgment, and other steps β€” were trained independently by the author** |
| Training data cutoff | End of November 2025 |
| License | CircleStone Labs Non-Commercial License |

## 🎯 Use Cases

**Best for βœ…**

- **Anime illustrations and character art**: for generating images with a specific art style, composition, lighting, and atmosphere.
- **Natural-language driven**: for precise control over the image with a descriptive passage (rather than stacked tags).
- **Drop-in upgrade**: for users already running an Anima workflow β€” switch the checkpoint directly to improve generation quality.

**Not suitable for β›”**

- **Photo-realism**: photorealistic photography is not the goal of this series; results are limited.
- **Long-text rendering**: hard to faithfully reproduce text on signs, subtitles, or long multi-line sentences.

## πŸš€ Quick Start (ComfyUI)

The Anima series is **natively supported** in ComfyUI, and this release can directly replace the upstream weights.

### 1. Place the Model Files

| File | Target directory |
| :--- | :--- |
| `anima-light-lavender.safetensors` | `ComfyUI/models/diffusion_models` |
| `anima-light-lavender_mxfp8.safetensors` | `ComfyUI/models/diffusion_models` |
| `qwen_3_06b_base.safetensors` | `ComfyUI/models/text_encoders` |
| `qwen_image_vae.safetensors` | `ComfyUI/models/vae` |

Weights are provided in two versions, **BF16** and **MXFP8** β€” pick **either one** and place it into `diffusion_models`; the MXFP8 version speeds up inference.

(The last two files are the Qwen3-0.6B text encoder and Qwen-Image VAE used by upstream Anima; if you already run an Anima workflow, you can reuse them as-is.)

### 2. Install the Custom Nodes

Install the [comfyui-zako-pe](https://github.com/aa0525/comfyui-zako-pe) custom nodes: clone the repository into `ComfyUI/custom_nodes/` and restart ComfyUI:

```bash

cd ComfyUI/custom_nodes

git clone https://github.com/aa0525/comfyui-zako-pe.git

```

If git is not installed, you can also click **Code β†’ Download ZIP** on the repository page, extract it to `ComfyUI/custom_nodes/`, and restart ComfyUI for it to take effect.

The node pack provides `Danbooru Caption JSON` and `Danbooru Prompt Extend (OpenAI)`: the former assembles the structured caption, while the latter calls an OpenAI-compatible service to expand a tag-style `image_description` into natural language. For the expansion service, we recommend the companion model [zako-pe](https://huggingface.co/Johnny-Z/zako-pe) (ZAKO-V0.1): load its GGUF weights in LM Studio and start the local server (default address `http://127.0.0.1:1234`) to use it; see its model card for downloads and detailed usage.

### 3. Load the Workflows

This model repository ships with two workflow files; **drag them into the ComfyUI canvas** to load them β€” no manual wiring needed:

| Workflow file | Description |
| :--- | :--- |
| `anima.json` | Basic workflow: feeds the assembled structured caption directly into the positive prompt encoding; write `image_description` yourself as a long natural-language description |
| `anima-pe.json` | Adds a `Danbooru Prompt Extend (OpenAI)` node on top of the basic workflow: write only tags in `image_description`, and the expansion service rewrites them into natural language automatically |

Both workflows load the MXFP8 weights (`anima-light-lavender_mxfp8.safetensors`) by default; if you only downloaded the BF16 version, switch to `anima-light-lavender.safetensors` in the **Load Diffusion Model** node.

### 4. Generation Parameters (Recommended)

| Parameter | Recommended Value |
| :--- | :--- |
| Sampler | `euler` |
| Scheduler | `simple` |
| Steps | 25 |
| CFG | 4.0 |
| Resolution | Around 1280Γ—1280 pixels |
| Negative prompt | **Leave empty** |

## ✍️ Prompting

### Structured Caption

This release is trained on structured captions, so it can understand text blocks with a fixed field order:

```text

{

  "year": 2025,

  "preference_level": "best",

  "artist": [...],

  "copyright": [...],

  "character": [...],

  "image_description": "...",

  "extra_tags": [...]

}

```

| Field | Purpose and usage |
| :--- | :--- |
| `year` | Year anchor (integer, default `2025`) |
| `preference_level` | Preference level: `normal` / `high` / `very_high` / `best` (default `best`) |
| `artist` | Artist / style tags |
| `copyright` | Series / franchise tags |
| `character` | Character tags |
| `image_description` | **Image content description**: a natural-language description, as detailed as possible |
| `extra_tags` | Supplementary tags for the image content |

## 🧠 Training Details

### Data

- ~**1.7M Danbooru images**, all annotated and filtered by an in-house data pipeline; training data goes up to **the end of November 2025**.
- **All models involved in the data pipeline were trained independently by the author**: from image annotation and quality filtering to the preference judgment mentioned below, every model used was trained by the author rather than an off-the-shelf third-party model.
- Uses the structured caption shown above, in which `image_description` is **pure natural language**.
- `preference_level` tiers are determined by a preference judgment model trained on ~3M samples.

### Training Dynamics

| Item | Setting |
| :--- | :--- |
| Precision | Training runs in **BF16** overall; matrix multiplications in the MLP and Attention use **MXFP8 GEMM** |
| Optimizer | Composite optimizer: 2D parameters of the MLP and Attention are updated by **Muon** (momentum 0.95, `match_rms_adamw`); the remaining parameters stay in BF16 and are updated by **AdamW** (Ξ² = 0.9 / 0.95, Ξ΅ = 1e-8) |
| Batch size | 1024 |
| Learning rate | 4e-5 |
| Weight decay | 1e-3 |
| Gradient clipping | 1.0 |

### Training Focus

- Substantially enhances understanding of **natural-language descriptions at the 512-token scale**, making precise description β€” rather than rolling random tags β€” the default way to generate.
- The improved data quality also improves the aesthetics of generated images: even **without adding quality words or filling in a negative prompt**, the results are still clean and tidy enough.

## πŸ”Œ Compatibility

- The architecture is fully identical to the upstream base: the layer count, parameter count, and file format are all the same, so the checkpoint can be replaced directly.
- Uses the same text encoder and VAE (Qwen3-0.6B text encoder, Qwen-Image VAE) β€” no workflow changes needed.
- Single-file weights: put them into `models/diffusion_models` and ComfyUI can load them.

## ⚠️ Limitations

- **Not good at realism**: the model targets anime illustrations and non-photorealistic art; the photorealistic photography style is out of its scope.
- **Weak text rendering**: single words and short phrases usually render correctly, but the rendering of long text is unreliable.

## πŸ“œ License

The model weights and any derivatives are subject to the [CircleStone Labs Non-Commercial License](https://huggingface.co/circlestone-labs/Anima/resolve/main/LICENSE.md) and are **for non-commercial use only**. In addition, this model is a Derivative Model of Cosmos-Predict2-2B-Text2Image and is therefore also subject to the corresponding terms of the [NVIDIA Open Model License Agreement](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).

## πŸ™ Acknowledgements

- Base model: Anima-Base v1.0 (`anima-base-v1.0.safetensors`) from the [circlestone-labs/Anima](https://huggingface.co/circlestone-labs/Anima) series, built on [nvidia/Cosmos-Predict2-2B-Text2Image](https://huggingface.co/nvidia/Cosmos-Predict2-2B-Text2Image).