--- license: apache-2.0 library_name: stable-diffusion.cpp pipeline_tag: text-to-image base_model: - Tongyi-MAI/Z-Image-Turbo tags: - gguf - stable-diffusion.cpp - text-to-image - cpu - cpu-only - no-gpu - on-device - edge - local - quantized - z-image - zimage - pocket - vidraft --- > ### πŸ†• **[POCKET-Qwen3.8-Flash-Next](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF)** β€” a **180B** model running on a **laptop with 8 GB VRAM + 32 GB RAM** Β· **4.17 tok/s measured**. > [![New](https://img.shields.io/badge/πŸ†•_NEW-POCKET--180B_on_a_laptop-6c3fd1)](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF) [![VRAM](https://img.shields.io/badge/VRAM-8_GB-1baf7a)]() [![RAM](https://img.shields.io/badge/RAM-32_GB-1baf7a)]() [![Speed](https://img.shields.io/badge/measured-4.17_tok%2Fs-243456)]() > ### πŸ“š Collections > **β–Ά [POCKET Models](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6)** β€” this family (on-device, no GPU) > [Darwin Family](https://huggingface.co/collections/FINAL-Bench/darwin-family-699987b1f652864af0122193) Β· [Aether Foundation](https://huggingface.co/collections/FINAL-Bench/aether-foundation-model-6a5c7f2fa1a4165c0414e53a) Β· [VKAE Accelerated](https://huggingface.co/collections/FINAL-Bench/vkae-accelerated-6a47231d7e7999dd8227675a) # POCKET-Zimage-CPU **Pick your build β†’** [![35B](https://img.shields.io/badge/POCKET--35B-GGUF-243456)](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) [![26B](https://img.shields.io/badge/POCKET--26B-GGUF-2a9d8f)](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) [![KR GGUF](https://img.shields.io/badge/POCKET--KR-GGUF-7A1F3D)](https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF) [![KR MLX](https://img.shields.io/badge/POCKET--KR-MLX_(iPhone)-0f6e56)](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) [![EN GGUF](https://img.shields.io/badge/POCKET--EN-GGUF-185fa5)](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) [![180B laptop](https://img.shields.io/badge/POCKET--180B-Qwen3.8_Flash--Next_Β·_laptop-6c3fd1)](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF) [![Image NF4](https://img.shields.io/badge/POCKET--Image-Z--Image_NF4-d35400)](https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage) [![Image CPU](https://img.shields.io/badge/POCKET--Zimage-CPU-ff6b35)](https://huggingface.co/FINAL-Bench/POCKET-Zimage-CPU) ### Photoreal images in **46 seconds** on a **CPU only**. No GPU. No CUDA. No Python. > πŸš€ **Try it live, no install β†’** [![Space](https://img.shields.io/badge/πŸ€—_Space-POCKET--Zimage_CPU_studio-ff6b35)](https://huggingface.co/spaces/FINAL-Bench/POCKET-Zimage-CPU) β€” generating on a **CPU-only** box. [![License](https://img.shields.io/badge/License-Apache_2.0-0f6e56)](https://www.apache.org/licenses/LICENSE-2.0) [![Runtime](https://img.shields.io/badge/runtime-stable--diffusion.cpp-f0992a)](https://github.com/leejet/stable-diffusion.cpp) [![No GPU](https://img.shields.io/badge/GPU-not_required-1baf7a)]() [![Base](https://img.shields.io/badge/base-Z--Image--Turbo-185fa5)](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) [![RAM](https://img.shields.io/badge/RAM-6.4_GB-7A1F3D)]() **The POCKET lineup β†’** [![Zimage CPU](https://img.shields.io/badge/POCKET--Zimage-CPU_(images)-ff6b35)](https://huggingface.co/FINAL-Bench/POCKET-Zimage-CPU) [![35B](https://img.shields.io/badge/POCKET--35B-GGUF-243456)](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) [![KR GGUF](https://img.shields.io/badge/POCKET--KR-GGUF-7A1F3D)](https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF) [![KR MLX](https://img.shields.io/badge/POCKET--KR-MLX_(iPhone)-0f6e56)](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) [![EN GGUF](https://img.shields.io/badge/POCKET--EN-GGUF-185fa5)](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) [![26B](https://img.shields.io/badge/POCKET--26B-GGUF-2a9d8f)](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) --- ## Why this exists Every image model assumes you have a GPU. Most machines don't. POCKET-Zimage-CPU is Z-Image-Turbo packaged so that **a plain office PC** β€” no graphics card, no CUDA, no Python environment β€” produces a photoreal 512Γ—512 image in **under a minute**. One binary, three files, done. ## Samples

3 steps Β· 48.6 s Β· the default we ship

4 steps Β· 60.7 s Β· no visible gain

Korean prompt Β· note it lost the count
Prompt: a red apple on a wooden table, photorealistic β€” Korean: λ‚˜λ¬΄ νƒμž μœ„μ— 놓인 λΉ¨κ°„ 사과, 사싀적인 사진. Same seed, CPU only. ## Measured, not claimed All numbers below are from our own runs. **GPU count used: zero.** | Resolution | Time | Sampling | VAE | Peak RAM | |---|---|---|---|---| | **512 Γ— 512** | **46.4 s** | 32.7 s | 12.5 s | **6.42 GB** | | 512 Γ— 512 (Korean prompt) | **45.3 s** | 32.1 s | 12.0 s | 6.42 GB | | 1024 Γ— 1024 | 192.7 s | 135.1 s | 55.8 s | 6.76 GB | Intel Xeon Gold 6526Y Γ—2 (32 cores / 64 threads), 48 threads, Q4_0, 3 steps, `--fa --vae-tiling`. Single run per row. **Korean prompts cost nothing extra** β€” 45.3 s vs 46.4 s. Language is not a speed penalty here. ## How it got 5.3Γ— faster We started at 244 seconds and ended at 46. Every step is measured: | Change | Time | Peak RAM | |---|---|---| | Default settings (20 steps) | 244 s | 8.16 GB | | β†’ 4 steps | 62.1 s | 8.16 GB | | β†’ `--fa` (flash attention) | 59.5 s | 8.18 GB | | β†’ 3 steps | 48.6 s | 8.00 GB | | β†’ `--vae-tiling` | **46.4 s** | **6.42 GB** | **The big one is step count.** Z-Image **Turbo** is distilled for few-step sampling, but the tool's default is 20. Using the default throws away 5Γ— for nothing. **3 steps is the floor.** At 4 and 3 we cannot tell the images apart. At 2 the surface collapses β€” water droplets and wood grain vanish and the texture turns cloth-like. **`--vae-tiling` is free.** It cuts VAE time 16% *and* peak RAM by 1.27 GB. At 1024Γ—1024 it is the difference between 13.3 GB and 6.76 GB. **Do not use every thread you have.** On a 32-core / 64-thread box, 48 threads took 59.5 s and **64 threads took 108.8 s** β€” 1.8Γ— slower. Hyper-threads fight each other for the same physical cores. ## Files | File | Size | What | |---|---|---| | `z_image_turbo-Q4_0-pocket.gguf` | 3.51 GB | Diffusion model, 4-bit (VIDRAFT CPU build) | | *(bring your own)* Qwen3-4B-Instruct-2507-Q4_K_M | 2.58 GB | Text encoder β€” [download](https://huggingface.co/lmstudio-community/Qwen3-4B-Instruct-2507-GGUF) | | *(bring your own)* `ae.safetensors` | 0.16 GB | VAE β€” [download](https://huggingface.co/Tongyi-MAI/Z-Image-Turbo) | | **Total** | **β‰ˆ 6.25 GB** | | The diffusion model here is VIDRAFT's own CPU build of Z-Image-Turbo: 3.51 GB, with no visible quality change. ## Run it Get a `stable-diffusion.cpp` binary ([releases](https://github.com/leejet/stable-diffusion.cpp/releases)), then: ```bash sd-cli \ --diffusion-model z_image_turbo-Q4_0-pocket.gguf \ --vae ae.safetensors \ --llm Qwen3-4B-Instruct-2507-Q4_K_M.gguf \ -p "a red apple on a wooden table, photorealistic" \ --cfg-scale 1.0 --steps 3 --fa --vae-tiling \ -t 8 -H 512 -W 512 -o out.png ``` Set `-t` to your physical core count β€” **not** your thread count. ## Honest limits - **It cannot render text.** Any words inside the image come out garbled, in every language. If you need accurate text in an image, this is the wrong tool. - **Korean prompts lose count.** "A red apple" gives one apple in English and five or six in Korean. Korean has no articles, so the singular signal is weak for the encoder. Reproduced at both 20 and 3 steps, so it is the encoder, not the step count. - **1024 Γ— 1024 takes 3 minutes** on the machine above. Slower CPUs scale accordingly. - **Measured on a server CPU.** Laptop and mini-PC numbers are not in yet. ## Credits and licensing | Component | License | Author | |---|---|---| | Z-Image-Turbo (diffusion) | Apache-2.0 | Tongyi-MAI / Hangzhou Tongyi Laboratory | | Qwen3-4B-Instruct (text encoder) | Apache-2.0 | Qwen, Alibaba | | GGUF conversion (upstream) | Apache-2.0 | leejet | | stable-diffusion.cpp (runtime) | MIT | leejet | This repository redistributes a re-quantized copy of Z-Image-Turbo and keeps the original copyright notices. **We did not train this model.** What is ours is the CPU packaging, the 3-step setting, the re-quantization, and the measurements on this page. ## Related | | Runs on | Strength | |---|---|---| | [POCKET-Image-Zimage](https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage) | GPU (Python) | Faster, renders Korean text via glyph-init | | **POCKET-Zimage-CPU** (this) | **CPU only** | **No GPU, single binary** | Different jobs. Use the first if you have a graphics card, this one if you don't. --- ## 🧩 The POCKET Family β€” On-device AI by VIDRAFT *Big models, small hardware. No GPU, no cloud.* **Models** - πŸ“¦ [POCKET-35B-GGUF](https://huggingface.co/FINAL-Bench/POCKET-35B-GGUF) β€” flagship, PC / server, no GPU - πŸ“¦ [POCKET-26B-GGUF](https://huggingface.co/FINAL-Bench/POCKET-26B-GGUF) β€” compact 26B - πŸ‡°πŸ‡· [POCKET-KR-GGUF](https://huggingface.co/FINAL-Bench/POCKET-KR-GGUF) β€” Korean, Android - 🍎 [POCKET-KR-MLX](https://huggingface.co/FINAL-Bench/POCKET-KR-MLX) β€” Korean, iPhone / Mac - 🌍 [POCKET-EN-GGUF](https://huggingface.co/FINAL-Bench/POCKET-EN-GGUF) β€” English, phone / PC - πŸ’» [POCKET-Qwen3.8-Flash-Next-GGUF](https://huggingface.co/FINAL-Bench/POCKET-Qwen3.8-Flash-Next-GGUF) β€” **180B on a laptop** (8 GB VRAM + 32 GB RAM) - πŸ–ΌοΈ [POCKET-Image-Zimage](https://huggingface.co/FINAL-Bench/POCKET-Image-Zimage) β€” character-perfect text in any image - πŸ–₯️ [POCKET-Zimage-CPU](https://huggingface.co/FINAL-Bench/POCKET-Zimage-CPU) β€” photoreal images on a CPU only **Demos & tools (Spaces)** - 🎨 [POCKET-Image Studio](https://huggingface.co/spaces/FINAL-Bench/POCKET-Image-Studio) β€” text-in-image, generate in-page - πŸ–₯️ [POCKET-35B-CPU](https://huggingface.co/spaces/FINAL-Bench/POCKET-35B-CPU) β€” 35B answering on a CPU - πŸ–₯️ [POCKET-26B-CPU](https://huggingface.co/spaces/FINAL-Bench/POCKET-26B-CPU) β€” 26B on a CPU - πŸ–ΌοΈ [POCKET-Zimage-CPU](https://huggingface.co/spaces/FINAL-Bench/POCKET-Zimage-CPU) β€” image generation on a CPU πŸ“š [Full POCKET collection](https://huggingface.co/collections/FINAL-Bench/pocket-models-6a618ee5d23eafb7e185a5c6)