assembledchaos commited on
Commit
b17392c
路
verified 路
1 Parent(s): 9ddc8a3

Upload folder using huggingface_hub

Browse files
Files changed (3) hide show
  1. README.md +12 -5
  2. app.py +49 -12
  3. requirements.txt +1 -0
README.md CHANGED
@@ -1,5 +1,5 @@
1
  ---
2
- title: Qwen Image 2.1 Studio
3
  emoji: 馃帹
4
  colorFrom: indigo
5
  colorTo: purple
@@ -7,15 +7,16 @@ sdk: gradio
7
  sdk_version: 6.28.0
8
  python_version: "3.12"
9
  app_file: app.py
10
- short_description: Create, edit, and generate transparent images with Qwen
11
  startup_duration_timeout: 1h
12
  models:
 
13
  - Qwen/Qwen-Image-2.1
14
  ---
15
 
16
- # Qwen Image 2.1 Studio
17
 
18
- A Gradio demo of [Qwen/Qwen-Image-2.1](https://huggingface.co/Qwen/Qwen-Image-2.1), running the model directly on Hugging Face ZeroGPU.
19
 
20
  - **Create an image:** describe a scene or choose an example.
21
  - **Edit an image:** upload one reference image and describe the change.
@@ -26,7 +27,13 @@ The default is the model card's recommended 40 inference steps; lower this for q
26
 
27
  ## Hosting
28
 
29
- This Space requires `zero-a10g` hardware. Configure hardware in the Space settings; README metadata does not select hardware. No external inference API key is needed. The model weights and Diffusers revision are pinned in the source. The initial download is approximately 33 GB and can take several minutes.
 
 
 
 
 
 
30
 
31
  The model loads at startup and is registered with ZeroGPU. The Gradio queue runs one generation at a time. Visitors use Hugging Face's daily ZeroGPU quota. Uploaded and generated files are temporary; download results you want to keep. The app does not send images to an external API or publish a community gallery. Gradio cached files expire after 24 hours; cached examples can be reused across visitors.
32
 
 
1
  ---
2
+ title: Qwen Image 2.1 GGUF Studio
3
  emoji: 馃帹
4
  colorFrom: indigo
5
  colorTo: purple
 
7
  sdk_version: 6.28.0
8
  python_version: "3.12"
9
  app_file: app.py
10
+ short_description: Qwen Image 2.1 Uncensored GGUF demo with Q4_K_M weights
11
  startup_duration_timeout: 1h
12
  models:
13
+ - abenzerps/Qwen-Image-2.1-Uncensored-GGUF
14
  - Qwen/Qwen-Image-2.1
15
  ---
16
 
17
+ # Qwen Image 2.1 GGUF Studio
18
 
19
+ A Gradio demo of [abenzerps/Qwen-Image-2.1-Uncensored-GGUF](https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF), using its recommended **Q4_K_M** checkpoint on Hugging Face ZeroGPU. This replaces the original diffusion transformer in this Space.
20
 
21
  - **Create an image:** describe a scene or choose an example.
22
  - **Edit an image:** upload one reference image and describe the change.
 
27
 
28
  ## Hosting
29
 
30
+ This Space requires `zero-a10g` hardware. Configure hardware in the Space settings; README metadata does not select hardware. No external inference API key is needed. Model and library revisions are pinned in the source. The initial download is approximately 24 GB and can take several minutes.
31
+
32
+ The 4.60 GB `qwen-image-2.1-Q4_K_M.gguf` file is checked against the publisher's SHA256 checksum. All 297 tensors are loaded with strict key and shape validation. The GGUF values are expanded to BF16 once at startup so ZeroGPU can pack and stream ordinary PyTorch tensors; the model does not remain compressed in GPU memory. This trades the runtime memory saving for predictable inference latency within the 60-second GPU reservation.
33
+
34
+ The compatible text encoder, VAE, processor, scheduler, and transformer configuration come from the pinned upstream `Qwen/Qwen-Image-2.1` repository. Its diffusion transformer weights are neither downloaded nor used as a fallback. The publisher describes the GGUF as a quantization of the original upstream weights, not a separately trained model. The Q8_0 variant is not used because the model card reports a shape-mismatch issue.
35
+
36
+ Example caching uses a separate directory for this GGUF revision, preventing results from the previous model from being served. PNG metadata records the exact repository, revision, filename, checksum, and generation settings.
37
 
38
  The model loads at startup and is registered with ZeroGPU. The Gradio queue runs one generation at a time. Visitors use Hugging Face's daily ZeroGPU quota. Uploaded and generated files are temporary; download results you want to keep. The app does not send images to an external API or publish a community gallery. Gradio cached files expire after 24 hours; cached examples can be reused across visitors.
39
 
app.py CHANGED
@@ -1,21 +1,31 @@
1
  import os
2
 
3
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
 
4
 
5
  import spaces
6
  import gradio as gr
7
  import torch
 
8
  import json
9
  import random
10
  import tempfile
11
  import time
12
  from pathlib import Path
13
 
14
- from diffusers import QwenImage21Pipeline
 
 
 
 
15
  from PIL import Image, ImageOps, PngImagePlugin
16
 
17
- MODEL_ID = "Qwen/Qwen-Image-2.1"
18
- MODEL_REVISION = "790c92633540aa0cb11d9abf19eb46d861714758"
 
 
 
 
19
  MODES = ["Create an image", "Edit an image", "Transparent PNG"]
20
  SIZES = {
21
  "Square 路 1:1": (1024, 1024),
@@ -27,10 +37,34 @@ SIZES = {
27
  MAX_SEED = 2**31 - 1
28
 
29
  print(f"Loading {MODEL_ID} at {MODEL_REVISION}", flush=True)
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
30
  pipe = QwenImage21Pipeline.from_pretrained(
31
- MODEL_ID, revision=MODEL_REVISION, torch_dtype=torch.bfloat16
 
32
  ).to("cuda")
33
- print("Qwen-Image-2.1 pipeline loaded on CUDA via ZeroGPU.", flush=True)
34
 
35
 
36
  @spaces.GPU(duration=60)
@@ -44,7 +78,7 @@ def generate(
44
  randomize_seed: bool = True,
45
  progress: gr.Progress = gr.Progress(track_tqdm=True),
46
  ) -> tuple[str, str, int, str]:
47
- """Generate or edit an image with Qwen-Image-2.1 and return a PNG.
48
 
49
  Choose Create an image, Edit an image (requires a reference), or
50
  Transparent PNG. Disable randomize_seed to reuse a seed. The result
@@ -94,6 +128,9 @@ def generate(
94
  metadata = {
95
  "model": MODEL_ID,
96
  "revision": MODEL_REVISION,
 
 
 
97
  "prompt": prompt,
98
  "effective_prompt": effective_prompt,
99
  "mode": mode,
@@ -106,12 +143,12 @@ def generate(
106
  png_info = PngImagePlugin.PngInfo()
107
  png_info.add_text("generation", json.dumps(metadata, ensure_ascii=False))
108
  # Gradio owns its output cache and expires files after 24 hours.
109
- output_dir = Path(tempfile.mkdtemp(prefix="qwen21-", dir=gr.utils.get_upload_folder()))
110
- path = output_dir / f"qwen-image-2.1-{actual_seed}.png"
111
  result.save(path, pnginfo=png_info)
112
  details = (
113
  f"{result.width} 脳 {result.height} 路 {int(steps)} steps 路 "
114
- f"{elapsed:.1f}s 路 seed {actual_seed} 路 {result.mode} PNG"
115
  )
116
  if mode == "Transparent PNG" and (
117
  "A" not in result.getbands() or result.getchannel("A").getextrema()[0] == 255
@@ -135,11 +172,11 @@ CSS = """
135
  #generate { min-height: 48px; }
136
  """
137
 
138
- with gr.Blocks(title="Qwen Image 2.1 Studio", delete_cache=(3600, 86400)) as demo:
139
  gr.Markdown(
140
- "# Qwen Image 2.1 Studio\n"
141
  "Turn an idea into an image. Reimagine a photo. Create a transparent asset.\n\n"
142
- "[Model](https://huggingface.co/Qwen/Qwen-Image-2.1) 路 "
143
  "[Qwen Research License](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE)",
144
  elem_id="intro",
145
  )
 
1
  import os
2
 
3
  os.environ.setdefault("PYTORCH_CUDA_ALLOC_CONF", "expandable_segments:True")
4
+ os.environ["GRADIO_EXAMPLES_CACHE"] = ".gradio/gguf-q4-k-m-e92386d3-examples"
5
 
6
  import spaces
7
  import gradio as gr
8
  import torch
9
+ import hashlib
10
  import json
11
  import random
12
  import tempfile
13
  import time
14
  from pathlib import Path
15
 
16
+ from accelerate import init_empty_weights
17
+ from diffusers import QwenImage21Pipeline, QwenImage21Transformer2DModel
18
+ from diffusers.models.model_loading_utils import load_gguf_checkpoint
19
+ from diffusers.quantizers.gguf.utils import dequantize_gguf_tensor
20
+ from huggingface_hub import hf_hub_download
21
  from PIL import Image, ImageOps, PngImagePlugin
22
 
23
+ MODEL_ID = "abenzerps/Qwen-Image-2.1-Uncensored-GGUF"
24
+ MODEL_REVISION = "e92386d3b86fd77ae1650534815ab3f6b85d46ba"
25
+ CHECKPOINT = "qwen-image-2.1-Q4_K_M.gguf"
26
+ CHECKPOINT_SHA256 = "833439e91bc1152d28f37aa198c7f6f4218b7de95754c2f7a318a2422ab4b2f8"
27
+ COMPANION_ID = "Qwen/Qwen-Image-2.1"
28
+ COMPANION_REVISION = "790c92633540aa0cb11d9abf19eb46d861714758"
29
  MODES = ["Create an image", "Edit an image", "Transparent PNG"]
30
  SIZES = {
31
  "Square 路 1:1": (1024, 1024),
 
37
  MAX_SEED = 2**31 - 1
38
 
39
  print(f"Loading {MODEL_ID} at {MODEL_REVISION}", flush=True)
40
+ checkpoint_path = hf_hub_download(MODEL_ID, CHECKPOINT, revision=MODEL_REVISION)
41
+ with open(checkpoint_path, "rb") as checkpoint_file:
42
+ checksum = hashlib.file_digest(checkpoint_file, "sha256").hexdigest()
43
+ if checksum != CHECKPOINT_SHA256:
44
+ raise RuntimeError("The GGUF checkpoint checksum does not match the published SHA256SUMS.")
45
+
46
+ # Expand the requested GGUF once for ZeroGPU's eager tensor packing. This
47
+ # preserves the quantized checkpoint's values without per-step dequantization.
48
+ weights = load_gguf_checkpoint(checkpoint_path)
49
+ for name in weights:
50
+ weights[name] = dequantize_gguf_tensor(weights[name]).to(torch.bfloat16)
51
+ config = QwenImage21Transformer2DModel.load_config(
52
+ COMPANION_ID, subfolder="transformer", revision=COMPANION_REVISION
53
+ )
54
+ with init_empty_weights():
55
+ transformer = QwenImage21Transformer2DModel.from_config(config)
56
+ transformer.load_state_dict(weights, strict=True, assign=True)
57
+ transformer.eval().requires_grad_(False)
58
+ print(f"Verified {CHECKPOINT}: loaded all {len(weights)} tensors; sha256={checksum}", flush=True)
59
+ del weights
60
+
61
+ # Supplying the transformer prevents the upstream diffusion weights from
62
+ # being downloaded. Only its compatible text encoder, VAE and config are used.
63
  pipe = QwenImage21Pipeline.from_pretrained(
64
+ COMPANION_ID, revision=COMPANION_REVISION, transformer=transformer,
65
+ torch_dtype=torch.bfloat16,
66
  ).to("cuda")
67
+ print(f"{MODEL_ID} / {CHECKPOINT} loaded on CUDA via ZeroGPU.", flush=True)
68
 
69
 
70
  @spaces.GPU(duration=60)
 
78
  randomize_seed: bool = True,
79
  progress: gr.Progress = gr.Progress(track_tqdm=True),
80
  ) -> tuple[str, str, int, str]:
81
+ """Generate or edit an image with abenzerps' Qwen-Image-2.1 Q4_K_M GGUF.
82
 
83
  Choose Create an image, Edit an image (requires a reference), or
84
  Transparent PNG. Disable randomize_seed to reuse a seed. The result
 
128
  metadata = {
129
  "model": MODEL_ID,
130
  "revision": MODEL_REVISION,
131
+ "checkpoint": CHECKPOINT,
132
+ "checkpoint_sha256": CHECKPOINT_SHA256,
133
+ "runtime_dtype": "bfloat16",
134
  "prompt": prompt,
135
  "effective_prompt": effective_prompt,
136
  "mode": mode,
 
143
  png_info = PngImagePlugin.PngInfo()
144
  png_info.add_text("generation", json.dumps(metadata, ensure_ascii=False))
145
  # Gradio owns its output cache and expires files after 24 hours.
146
+ output_dir = Path(tempfile.mkdtemp(prefix="qwen21-gguf-", dir=gr.utils.get_upload_folder()))
147
+ path = output_dir / f"qwen-image-2.1-gguf-{actual_seed}.png"
148
  result.save(path, pnginfo=png_info)
149
  details = (
150
  f"{result.width} 脳 {result.height} 路 {int(steps)} steps 路 "
151
+ f"{elapsed:.1f}s 路 seed {actual_seed} 路 {result.mode} PNG 路 Q4_K_M"
152
  )
153
  if mode == "Transparent PNG" and (
154
  "A" not in result.getbands() or result.getchannel("A").getextrema()[0] == 255
 
172
  #generate { min-height: 48px; }
173
  """
174
 
175
+ with gr.Blocks(title="Qwen Image 2.1 GGUF Studio", delete_cache=(3600, 86400)) as demo:
176
  gr.Markdown(
177
+ "# Qwen Image 2.1 GGUF Studio\n"
178
  "Turn an idea into an image. Reimagine a photo. Create a transparent asset.\n\n"
179
+ "[abenzerps / Qwen-Image-2.1-Uncensored-GGUF](https://huggingface.co/abenzerps/Qwen-Image-2.1-Uncensored-GGUF) 路 **Q4_K_M** 路 "
180
  "[Qwen Research License](https://huggingface.co/Qwen/Qwen-Image-2.1/blob/main/LICENSE)",
181
  elem_id="intro",
182
  )
requirements.txt CHANGED
@@ -6,3 +6,4 @@ safetensors
6
  Pillow
7
  sentencepiece
8
  mcp
 
 
6
  Pillow
7
  sentencepiece
8
  mcp
9
+ gguf==0.19.0