File size: 2,849 Bytes
1ca2e72
ea68aca
b1fefb8
57e3ede
 
1ca2e72
 
 
57e3ede
b1fefb8
ea68aca
 
b1fefb8
 
ea68aca
b1fefb8
ea68aca
 
b1fefb8
ea68aca
 
 
b1fefb8
ea68aca
 
 
b1fefb8
ea68aca
b1fefb8
ea68aca
 
 
 
 
 
 
b1fefb8
ea68aca
 
 
 
b1fefb8
ea68aca
b1fefb8
ea68aca
 
 
 
 
94cd091
 
 
b1fefb8
ea68aca
 
b1fefb8
ea68aca
b1fefb8
ea68aca
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
title: MiniMax-H3 Prompt Rewriter (Omni)
emoji: 🎬
colorFrom: red
colorTo: pink
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: "3.12"
startup_duration_timeout: 1h
pinned: false
short_description: Structured MiniMax-H3 audio-video prompts from short ideas
---

# MiniMax-H3 Prompt Rewriter · Qwen2.5-Omni LoRA

Turns a short request — plus optional image, video or audio references — into a
structured, production-ready **MiniMax-H3** audio-video prompt.

- LoRA adapter: [`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni`](https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni)
- Base model: [`Qwen/Qwen2.5-Omni-7B`](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) (Thinker, bf16)
- Target video model: [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3)

**Output is text only.** This is a prompt rewriter, not a video generator — feed
the rewritten prompt and the same reference assets into a MiniMax-H3 pipeline to
render.

## Tasks

| Task | Inputs | Output schema |
| --- | --- | --- |
| `T2AV` | text only | `integrated_multimodal_description`, `overall_soundscape`, `non_diegetic_music` |
| `I2AV` | text + 1 image (exact **first** frame) | same 3 fields |
| `L2AV` | text + 1 image (exact **last** frame) | same 3 fields |
| `FL2AV` | text + 2 ordered images (first & last frames) | same 3 fields |
| `Ref2AV` | text + ordered images / video / audio references | `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, `overall_soundscape`, `non_diegetic_music` |

Duration is requested in whole seconds (4–15) and snapped to MiniMax-H3's legal
`17 * n + 5` frame grid at 24 fps. Base tasks accept any aspect preset; `Ref2AV`
supports only `16:9` and `9:16`. `Ref2AV` prompts must mention every supplied
reference label (`<Picture 1>`, `<Video 1>`, `<Audio 1>`) and no others.

## Fidelity to the reference implementation

The system prompts (`system_prompt.py`) are copied verbatim from the adapter
repo, and message construction, the duration grid, reference labelling /
validation and the output schema check are ported 1:1 from the adapter repo's
`infer.py`. Encoding uses the same `ProcessorMixin.apply_chat_template` path
with `max_pixels` 301056 (images) / 100352 (video), `load_audio_from_video=False`
and `use_audio_in_video=False`. Those processing kwargs are passed through
transformers 5.x's `processor_kwargs` dict (transformers 4.x, which the
reference pins, is incompatible with `huggingface_hub` 1.x / Gradio 6).

Runs on ZeroGPU: the Thinker is loaded in bf16 with SDPA attention and the LoRA
is attached with PEFT at module scope.

## Credits

Example prompts and reference frames are the adapter authors' own, from
`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni` (`assets/examples/`), released
under the repo's Apache-2.0-named license.