Instructions to use Qwen/Qwen-Image-2.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use Qwen/Qwen-Image-2.1 with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("Qwen/Qwen-Image-2.1", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Qwen-Image-2.1 on Android in 200s — MNN + OpenCL (int4 DiT, Turbo 6 step & Tiny VAE)
Hi Qwen team and everyone,
Claude code and I have open-sourced libQwenImage21, an implementation enabling Alibaba's Qwen-Image-2.1 to run 100% locally and fully on-device on Android smartphones, powered by Alibaba's MNN framework with OpenCL GPU acceleration.
- GitHub Repository: https://github.com/scsonic/libQwenImage21
- Pre-converted MNN Models: https://huggingface.co/evankuo/Qwen-Image-2.1-MNN
The goal of this project is to run state-of-the-art Qwen-Image-2.1 locally on edge mobile hardware without needing servers, cloud APIs, CUDA, or Python runtimes.
Special Thanks to the Open Source Community: TAESD Makes Android Deployment Possible
On mobile devices, running the original VAE decoder was the single biggest roadblock: memory spikes during decoding frequently reached over 4 GB on CPU, putting the process at constant risk of being terminated by Android's Low Memory Killer (LMKD).
Huge thanks to the open-source community and the TAESD (Tiny AutoEncoder) project by @madebyollin (https://github.com/madebyollin/taesd). By adapting TAESD for Qwen-Image-2.1, running the full pipeline on a smartphone has finally become genuinely feasible:
| VAE Decoder | Device | Decode Time | Runtime Memory | Mobile Viability |
|---|---|---|---|---|
| Original VAE | CPU (fp16) | 19.54 s | 4326 MB | High risk of LMK kill |
| Tiny VAE (TAESD) | GPU (OpenCL) | 0.09 s | 187 MB | Fast and stable |
Thanks to TAESD, decoding time dropped from ~20 seconds down to sub-100 ms, while slashing runtime memory from 4.3 GB to 187 MB — completely eliminating the mobile OOM barrier.
Highlights & Open Source Foundations
- 100% Offline and On-Device on Android
- Mobile GPU Acceleration: Powered by Alibaba's MNN OpenCL backend (https://github.com/alibaba/MNN)
- Lossless int4 DiT Weights: Repacked from leejet/Qwen-Image-2.1-GGUF (https://huggingface.co/leejet/Qwen-Image-2.1-GGUF / https://github.com/leejet/stable-diffusion.cpp) into MNN asymmetric int4
- Block-Causal KV Prefix Caching: Prompt and condition tokens computed only once
- 6-Step Turbo Mode: Integrated Viggle-turbo LoRA (https://huggingface.co/Viggle/Qwen-Image-2.1-viggle-turbo) to achieve ~3.5 min total generation
- Ultra-Lightweight Decoding: Adapted from madebyollin/taesd (https://github.com/madebyollin/taesd) for near-zero latency and 187 MB footprint
Benchmarks (Snapdragon 8 Gen 2 / Adreno 740)
- Standard Mode (512x512, 20 steps):
- Text encoder (Qwen3-VL-8B int4, CPU): ~21 s
- DiT diffusion (OpenCL): 20 steps x ~19 s
- VAE Decode (Tiny VAE, OpenCL): 0.09 s (187 MB memory)
- Total time: ~400-450 s
- Turbo Mode (512x512, 6 steps):
- Total time:
200-220 s (3.5 minutes on-device)
- Total time:
Links
- Source Code: https://github.com/scsonic/libQwenImage21
- MNN Models: https://huggingface.co/evankuo/Qwen-Image-2.1-MNN
Feedback and benchmark reports across different mobile devices are warmly welcomed! If you find this project interesting, feel free to give it a star on GitHub!
@evankuo
could you explain how to use it on POCO F6? I have 12gb ram + 12gb extended RAM, but at the very start of generation it says "Out of memory before DiT: needs about 5526mb, only 3980 availble. I use tiny vae, and the smallest resolution ~320px. So how i it possible that ram is not enough.
So far using any AI on phone, i never had a feeling that those 12gb of extended RAM do anything, probably because AI doesn't use extended ram...
can desktop comfyui use this tine vae as well? or does it has other version?
@evankuo
could you explain how to use it on POCO F6? I have 12gb ram + 12gb extended RAM, but at the very start of generation it says "Out of memory before DiT: needs about 5526mb, only 3980 availble. I use tiny vae, and the smallest resolution ~320px. So how i it possible that ram is not enough.
So far using any AI on phone, i never had a feeling that those 12gb of extended RAM do anything, probably because AI doesn't use extended ram...
It looks like your other apps are taking up too much memory.
You were able to get past the Text Encoder, but you still can’t load the DiT model.
Try closing some background apps and restarting the device to see if you can free up more memory for the DiT model.
You might be able to try the 20-step DiT model, since it’s around 4.2 GB.
The 6-step Turbo version is around 5 GB because it combines the LoRA.
My 8 Gen 2 device has 16 GB of RAM, so it looks like I’ll have to find a way to integrate a 2-bit or 3-bit version.
can desktop comfyui use this tine vae as well? or does it has other version?
It looks like tiny vae hasn’t been integrated yet.
But you can modify ComfyUI based on the example code he provided. The instructions are here:
@evankuo
ok with 20 step dit, it seems to work with 320px size. Barely enough ram. Image Done in 255 seconds. Also tried 384px, surprisingly it also starts generating (350 seconds).
Even 512px works (620 s) and Real VAE too, how is it not too much i dont know 🤔
3 images ate 10% battery 😃
So, only Turbo DiT doesn't work with any resolution.
can desktop comfyui use this tine vae as well? or does it has other version?
@thrallemprize TAEQI2.1 support was merged to ComfyUI GitHub this week https://github.com/Comfy-Org/ComfyUI/pull/16552