multimodalart's picture
multimodalart HF Staff
Port to transformers 5.x API (hub 1.x / gradio 6 compatible)
94cd091 verified
|
Raw History Blame Contribute Delete
2.85 kB
---
title: MiniMax-H3 Prompt Rewriter (Omni)
emoji: 🎬
colorFrom: red
colorTo: pink
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: "3.12"
startup_duration_timeout: 1h
pinned: false
short_description: Structured MiniMax-H3 audio-video prompts from short ideas
---
# MiniMax-H3 Prompt Rewriter Β· Qwen2.5-Omni LoRA
Turns a short request β€” plus optional image, video or audio references β€” into a
structured, production-ready **MiniMax-H3** audio-video prompt.
- LoRA adapter: [`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni`](https://huggingface.co/lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni)
- Base model: [`Qwen/Qwen2.5-Omni-7B`](https://huggingface.co/Qwen/Qwen2.5-Omni-7B) (Thinker, bf16)
- Target video model: [`MiniMaxAI/MiniMax-H3`](https://huggingface.co/MiniMaxAI/MiniMax-H3)
**Output is text only.** This is a prompt rewriter, not a video generator β€” feed
the rewritten prompt and the same reference assets into a MiniMax-H3 pipeline to
render.
## Tasks
| Task | Inputs | Output schema |
| --- | --- | --- |
| `T2AV` | text only | `integrated_multimodal_description`, `overall_soundscape`, `non_diegetic_music` |
| `I2AV` | text + 1 image (exact **first** frame) | same 3 fields |
| `L2AV` | text + 1 image (exact **last** frame) | same 3 fields |
| `FL2AV` | text + 2 ordered images (first & last frames) | same 3 fields |
| `Ref2AV` | text + ordered images / video / audio references | `subject_definitions`, `summary`, `retention_analysis`, `detailed_description`, `overall_soundscape`, `non_diegetic_music` |
Duration is requested in whole seconds (4–15) and snapped to MiniMax-H3's legal
`17 * n + 5` frame grid at 24 fps. Base tasks accept any aspect preset; `Ref2AV`
supports only `16:9` and `9:16`. `Ref2AV` prompts must mention every supplied
reference label (`<Picture 1>`, `<Video 1>`, `<Audio 1>`) and no others.
## Fidelity to the reference implementation
The system prompts (`system_prompt.py`) are copied verbatim from the adapter
repo, and message construction, the duration grid, reference labelling /
validation and the output schema check are ported 1:1 from the adapter repo's
`infer.py`. Encoding uses the same `ProcessorMixin.apply_chat_template` path
with `max_pixels` 301056 (images) / 100352 (video), `load_audio_from_video=False`
and `use_audio_in_video=False`. Those processing kwargs are passed through
transformers 5.x's `processor_kwargs` dict (transformers 4.x, which the
reference pins, is incompatible with `huggingface_hub` 1.x / Gradio 6).
Runs on ZeroGPU: the Thinker is loaded in bf16 with SDPA attention and the LoRA
is attached with PEFT at module scope.
## Credits
Example prompts and reference frames are the adapter authors' own, from
`lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni` (`assets/examples/`), released
under the repo's Apache-2.0-named license.