Spaces:
Running on Zero
Running on Zero
File size: 2,849 Bytes
1ca2e72 ea68aca b1fefb8 57e3ede 1ca2e72 57e3ede b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca 94cd091 b1fefb8 ea68aca b1fefb8 ea68aca b1fefb8 ea68aca | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 | ---
title: MiniMax-H3 Prompt Rewriter (Omni)
emoji: 🎬
colorFrom: red
colorTo: pink
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: "3.12"
startup_duration_timeout: 1h
pinned: false
short_description: Structured MiniMax-H3 audio-video prompts from short ideas
---
# MiniMax-H3 Prompt Rewriter · Qwen2.5-Omni LoRA
Turns a short request — plus optional image, video or audio references — into a
structured, production-ready **MiniMax-H3** audio-video prompt.
- LoRA adapter: [`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni`](https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni)
- Base model: [`Qwen/Qwen2.5-Omni-7B`](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) (Thinker, bf16)
- Target video model: [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3)
**Output is text only.** This is a prompt rewriter, not a video generator — feed
the rewritten prompt and the same reference assets into a MiniMax-H3 pipeline to
render.
## Tasks
| Task | Inputs | Output schema |
| --- | --- | --- |
| `T2AV` | text only | `integrated_multimodal_description`, `overall_soundscape`, `non_diegetic_music` |
| `I2AV` | text + 1 image (exact **first** frame) | same 3 fields |
| `L2AV` | text + 1 image (exact **last** frame) | same 3 fields |
| `FL2AV` | text + 2 ordered images (first & last frames) | same 3 fields |
| `Ref2AV` | text + ordered images / video / audio references | `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, `overall_soundscape`, `non_diegetic_music` |
Duration is requested in whole seconds (4–15) and snapped to MiniMax-H3's legal
`17 * n + 5` frame grid at 24 fps. Base tasks accept any aspect preset; `Ref2AV`
supports only `16:9` and `9:16`. `Ref2AV` prompts must mention every supplied
reference label (`<Picture 1>`, `<Video 1>`, `<Audio 1>`) and no others.
## Fidelity to the reference implementation
The system prompts (`system_prompt.py`) are copied verbatim from the adapter
repo, and message construction, the duration grid, reference labelling /
validation and the output schema check are ported 1:1 from the adapter repo's
`infer.py`. Encoding uses the same `ProcessorMixin.apply_chat_template` path
with `max_pixels` 301056 (images) / 100352 (video), `load_audio_from_video=False`
and `use_audio_in_video=False`. Those processing kwargs are passed through
transformers 5.x's `processor_kwargs` dict (transformers 4.x, which the
reference pins, is incompatible with `huggingface_hub` 1.x / Gradio 6).
Runs on ZeroGPU: the Thinker is loaded in bf16 with SDPA attention and the LoRA
is attached with PEFT at module scope.
## Credits
Example prompts and reference frames are the adapter authors' own, from
`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni` (`assets/examples/`), released
under the repo's Apache-2.0-named license.
|