Update README.md
Browse files
README.md
CHANGED
|
@@ -17,7 +17,7 @@ pipeline_tag: text-to-image
|
|
| 17 |
|
| 18 |
[Arxiv](https://arxiv.org/abs/2508.08098) | [Github](https://github.com/DruryXu/TBAC-UniImage)
|
| 19 |
|
| 20 |
-
LLM | GenEval | DPG-Bench |
|
| 33 |
| :--- | :--- | :--- | :--- |
|
|
@@ -47,8 +54,12 @@ Our model is composed of two components: the [Qwen2.5-VL-3B-Instruct](https://hu
|
|
| 47 |
|
| 48 |

|
| 49 |
|
| 50 |
-
##
|
| 51 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
|
| 53 |
### Few Prompts Used in Teaser
|
| 54 |
|
|
@@ -59,9 +70,6 @@ Please refer to [Github](https://github.com/DruryXu/TBAC-UniImage) for the infer
|
|
| 59 |
- Steampunk architecture in the forest, reactor, rusty green color scheme with Studio Ghibli style, lots of details, mechanical, green, forest, trees, moss, 8K, Unreal Engine, C4D rendering, Ultra HD details
|
| 60 |
- An astronaut holding a stop sign on the moon.
|
| 61 |
|
| 62 |
-
## Limitations
|
| 63 |
-
For a better experience, please use the text-to-image mode. The image-text-to-image capability is currently weaker (but you can still try it).
|
| 64 |
-
|
| 65 |
## Acknowledgements
|
| 66 |
The training and inference codes are modified from [MetaQuery](https://github.com/facebookresearch/metaquery). We thank them for their contribution!
|
| 67 |
|
|
|
|
| 17 |
|
| 18 |
[Arxiv](https://arxiv.org/abs/2508.08098) | [Github](https://github.com/DruryXu/TBAC-UniImage)
|
| 19 |
|
| 20 |
+

|
| 21 |
|
| 22 |
## Overview
|
| 23 |
This repository contains the official model checkpoints of **TBAC-UniImage-3B**, an unified understanding and generation model developed by Basic Algorithm Center, Platform and Content Group, Tencent.
|
|
|
|
| 28 |
|
| 29 |
## Performance
|
| 30 |
|
| 31 |
+
### Qualitative Results for Text-to-Image Task
|
| 32 |
+

|
| 33 |
+
|
| 34 |
+
### Qualitative Results for Image-Text-to-Image Task
|
| 35 |
+

|
| 36 |
+
**The input image is processed by the Qwen2.5-VL image encoder and then fed into the MLLM along with text and learnable queries. We use only the learnable queries, which have fused the multimodal information, as the generative condition, without directly incorporating any image VAE representations like other works. Despite this, the model still achieves promising multimodal understanding and consistency performance in Image-Text-to-Image tasks.**
|
| 37 |
+
|
| 38 |
### GenEval and DPG-Bench
|
| 39 |
| Method | Base (M)LLM | GenEval | DPG-Bench |
|
| 40 |
| :--- | :--- | :--- | :--- |
|
|
|
|
| 54 |
|
| 55 |

|
| 56 |
|
| 57 |
+
### ImgEdit
|
| 58 |
+
|
| 59 |
+

|
| 60 |
+
|
| 61 |
+
## Train and Inference
|
| 62 |
+
Please refer to [Github](https://github.com/DruryXu/TBAC-UniImage) for train and inference codes.
|
| 63 |
|
| 64 |
### Few Prompts Used in Teaser
|
| 65 |
|
|
|
|
| 70 |
- Steampunk architecture in the forest, reactor, rusty green color scheme with Studio Ghibli style, lots of details, mechanical, green, forest, trees, moss, 8K, Unreal Engine, C4D rendering, Ultra HD details
|
| 71 |
- An astronaut holding a stop sign on the moon.
|
| 72 |
|
|
|
|
|
|
|
|
|
|
| 73 |
## Acknowledgements
|
| 74 |
The training and inference codes are modified from [MetaQuery](https://github.com/facebookresearch/metaquery). We thank them for their contribution!
|
| 75 |
|