AetherArt β€” SDXL Ukiyo-e LoRA

Do not use the auto-generated "Use this model" snippet above. It loads in bfloat16 via device_map="cuda" and omits the required madebyollin/sdxl-vae-fp16-fix VAE. Copying it as-is has been measured to cause severe VRAM oversubscription on 8 GB-class GPUs β€” roughly 280 seconds per denoising step (over an hour for one image) instead of the ~10–20 s/image this adapter is designed for, and on some hardware it will produce black/NaN output instead. Use the verified code in Usage below.

A rank-8 LoRA adapter that steers Stable Diffusion XL toward Japanese ukiyo-e woodblock print style at 1024Γ—1024 resolution. Trained on 80 WikiArt Ukiyo-e images against stabilityai/stable-diffusion-xl-base-1.0. Activate the style with the trigger token ukyowood anywhere in the prompt.

This is the SDXL companion to gauravgandhi2411/aetherart-ukiyo-sd21, trained with identical rank, dataset, and step count for a controlled cross-resolution comparison. Both visual evaluations independently selected the checkpoint-1000 step count β€” at 512 for SD 2.1 and 1024 for SDXL.

Sample outputs (checkpoint-1000, seed 42)

Mount Fuji at sunset "ukyowood ukiyo-e woodblock print of Mount Fuji at sunset"

Crane over ocean waves "ukyowood ukiyo-e print of a crane over ocean waves"

Samurai in bamboo forest "ukyowood ukiyo-e woodblock print of a samurai in a bamboo forest"

Usage

Required: load the madebyollin/sdxl-vae-fp16-fix VAE alongside the base model. SDXL's default fp16 VAE produces black images without this fix.

from diffusers import AutoencoderKL, DPMSolverMultistepScheduler, StableDiffusionXLPipeline
import torch

vae = AutoencoderKL.from_pretrained(
    "madebyollin/sdxl-vae-fp16-fix",
    torch_dtype=torch.float16,
)
pipe = StableDiffusionXLPipeline.from_pretrained(
    "stabilityai/stable-diffusion-xl-base-1.0",
    vae=vae,
    torch_dtype=torch.float16,
    variant="fp16",
    use_safetensors=True,
)
pipe.scheduler = DPMSolverMultistepScheduler.from_config(pipe.scheduler.config)
pipe.enable_model_cpu_offload()  # required on GPUs < 16 GB VRAM

pipe.load_lora_weights("gauravgandhi2411/aetherart-ukiyo-sdxl")

img = pipe(
    "ukyowood ukiyo-e woodblock print of Mount Fuji at sunset",
    negative_prompt=(
        "text, watermark, calligraphy, writing, letters, words, signature, "
        "blurry, low quality, western art, photograph, 3d render"
    ),
    num_inference_steps=25,
    guidance_scale=7.5,
    height=1024,
    width=1024,
).images[0]
img.save("output.png")

Training details

Parameter Value
Base model stabilityai/stable-diffusion-xl-base-1.0
VAE (fp16-fix) madebyollin/sdxl-vae-fp16-fix
LoRA rank 8
Training images 80 (WikiArt Ukiyo-e)
Resolution 1024 Γ— 1024
Steps 1500
Precision fp16 mixed
Batch size 1 (gradient accumulation = 4, effective batch = 4)
Learning rate 1e-4
Seed 42
Trigger token ukyowood
Hardware GCP g2-standard-4 β€” NVIDIA L4 24 GB VRAM
Training time 4h 26m
Compute cost ~$3.50 (GCP on-demand, us-central1)

Checkpoint selection

Evaluated checkpoints 500, 1000, and 1500 against four fixed prompts at seed 42, 25 DPM-Solver++ steps, 1024Γ—1024.

Checkpoint Pixel mean (ref. prompt) Visual verdict
500 145.52 Over-calligraphed; samurai figure scale weak β€” not selected
1000 143.47 Deepest Hiroshige palette; figure preserved; cartouches integrated β€” SELECTED
1500 135.48 Samurai figure dropout; cherry-blossom colour regression β€” not selected

Checkpoint-1000 was selected because it is the only checkpoint that simultaneously delivers (1) strong ukiyo-e style lift above the SDXL pretraining baseline, (2) figure preservation across all four test prompts, and (3) calligraphy rendered as single-panel title cartouches rather than scattered characters.

The step-1500 loss bump (~0.008 β†’ ~0.08) correlates with a visible quality regression in the images β€” this is not microbatch noise. The 1000-step convergence point is consistent with the SD 2.1 companion run, which also selected checkpoint-1000.

What the LoRA adds above SDXL's baseline

SDXL's pretraining already contains strong ukiyo-e priors. The LoRA's value is measurable rather than obvious: it deepens the Hiroshige teal/blue water palette, adds warm coral-pink atmospheric haze to horizon gradients, and reinforces the characteristic Japanese flat-plane treatment of foliage. The baseline renders cleaner figures but lacks the palette weight and compositional integration of traditional woodblock prints. The LoRA earns its place.

Calligraphy artifact at 1024 (an evolution from 512)

The SD 2.1 companion adapter (trained at 512) produced scattered calligraphy characters in image margins β€” a training data artefact from WikiArt metadata captions embedded in source images. At 1024, the same signal manifests differently: single-panel title cartouches and red banner seals, compositionally placed where a real woodblock print would carry its title block. The artefact is still present, but at the higher resolution and with a stronger base prior, it reads as a learned stylistic element rather than noise. The default negative prompt suppresses most residual scatter.

Default negative prompt

Applied automatically by the AetherArt application whenever this adapter is active:

text, watermark, calligraphy, writing, letters, words, signature, blurry, low quality, western art, photograph, 3d render

Known limitations

  • Measured cost/benefit of this adapter (independent-axis VLM judge, n=90 paired, sdxl_base = no adapter as the reference point):
    • sdxl_base alone scores artifact_absence 0.9222 β€” cleaner than this adapter. Applying this LoRA measurably increases visible embedded text/calligraphy/cartouche marks relative to generating the same prompts with no adapter at all: published checkpoint βˆ’0.0500 (3.49Γ— the paired SEM), an unfiltered-training-set retrain βˆ’0.0422 (3.11Γ— SEM) β€” both are individually significant regressions, not noise. This is the entanglement between "ukiyo-e style" and the WikiArt source images' embedded captions/signatures/script that produced the style signal the adapter learned; it is a real, measured tradeoff of using this adapter, not fully "mitigated" by the default negative prompt.
    • What the adapter buys in exchange: a measured style-adherence lift over sdxl_base alone β€” a curated-training-set retrain lifts style_adherence +0.0100 over base (2.82Γ— SEM); the published checkpoint (unfiltered training set) lifts it +0.0056 (1.68Γ— SEM, not itself significant at this n). sdxl_base already scores 0.9389 on style_adherence for ukiyo-e-styled prompts from its own pretraining, so headroom for any adapter to add is small. A positive control confirms this rubric CAN distinguish real ukiyo-e art from off-style contrasts decisively (diff/SEM = +25.062 vs. real Pattachitra art, +25.580 vs. generic sdxl_base outputs) β€” these lift numbers are confirmed to be measuring a real style-adherence signal, not an instrument artifact.
    • Net read: this adapter's main value is a modest style lift over what sdxl_base already renders unassisted (confirmed, not provisional, though still modest β€” below the 2Γ—SEM bar for the unfiltered checkpoint), at the cost of a real, larger, and statistically significant increase in embedded-text artifacts. Whether that trade is worth it depends on the use case β€” for artifact-sensitive generations, consider sdxl_base alone with an explicit "ukiyo-e style" prompt, or add this adapter and screen outputs for text artifacts downstream. A follow-up retrain-and-eval attempt investigating whether more aggressive dataset curation or a different training recipe can close this gap is tracked in docs/NEXT_MODEL_SPEC.md, not yet completed.
    • Full methodology and numbers: docs/MODEL_VERDICT.md Β§4.6–§4.9 in the AetherArt GitHub repo.
  • CLIP scoring does not capture the quality improvements from this adapter. See the CLIP-blindness finding below β€” nine experiments showed CLIP delta <1 SE while LPIPS ranged 0.40–0.73; the underfitting paradox (smaller adapters score higher on CLIP) confirms CLIP optimises in the wrong direction for style transfer evaluation.
  • Evaluated at 1024Γ—1024. Results at other resolutions are untested.
  • enable_model_cpu_offload() is required on GPUs with less than ~16 GB VRAM; expect ~60–90 s/image under offload on an 8 GB card.

Links

Downloads last month
38
Inference Providers NEW

Model tree for gauravgandhi2411/aetherart-ukiyo-sdxl

Adapter
(9745)
this model