raubatz/loras-qwen / scat2512
19.7 GB
135 files
Updated 29 days ago
Name
Size
scat-qwen-image-2512
scat-qwen-image-2512-v1.safetensors295 MB
xet
origin.txt100 Bytes
xet
README.md2.09 kB
xet
README.md

Scat LoRA — Qwen-Image-2512

Model description

Scat LoRA. Use natural language, "shitting" was the preferred verb. Only trained on photos so far, style LoRAs are probably necessary for non-realism.

Training info

v1

Dataset

111 video screencaps, manually tagged with the actions and other things VLMs struggle with. Images (with their tags described in the text prompt) were then captioned using Qwen3-VL-32B-Thinking-heretic, with a custom system prompt, then manual corrections were sometimes applied.

Images were mostly very high quality, but some were phone quality with light compression artifacts. The lowest quality ones were passed through SeedVR2 which cleaned them up a bit.

Watermarks were left alone but captioned. Faces left in, but general appearance (approximate age, ethnicity, body type, hair) was captioned.

Hyperparameters

Parameter Value
Trainer musubi-tuner
Optimizer AdamW8Bit (defaults)
LR Schedule cosine_with_min_lr
LR 2e-4
Min LR ratio 0.1
Lora+ LR ratio 4
Warmup steps 300
Epochs 100 (11100 steps)
Timestep Sampling shift
Discrete Flow Shift 2.2

Evaluation

Generally works very well and is flexible in terms of content, but output style is quite rigidly restricted to amateur/phone-quality photos.

Shot distance is stuck close up, seems to mostly ignore a lot of camera prompts like "wide shot" (dataset leaned heavily towards front or back closeup shots but had a variety of vertical angles).

Skin tone seems to lean a bit darker than base model and more limited in range (dataset had a roughly 50% split of dark and light skinned).

Watermarks are generally never emitted but when combined with other LoRAs they can appear.

Download model

Download them in the Files & versions tab.

Total size
19.7 GB
Files
135
Last updated
Aug 26
Pre-warmed CDN
US EU US EU

Contributors