multimodalart's picture
multimodalart HF Staff
Port to transformers 5.x API (hub 1.x / gradio 6 compatible)
94cd091 verified
|
Raw History Blame Contribute Delete
2.85 kB

A newer version of the Gradio SDK is available: 6.29.0

Upgrade
metadata
title: MiniMax-H3 Prompt Rewriter (Omni)
emoji: 🎬
colorFrom: red
colorTo: pink
sdk: gradio
sdk_version: 6.26.0
app_file: app.py
python_version: '3.12'
startup_duration_timeout: 1h
pinned: false
short_description: Structured MiniMax-H3 audio-video prompts from short ideas

MiniMax-H3 Prompt Rewriter · Qwen2.5-Omni LoRA

Turns a short request — plus optional image, video or audio references — into a structured, production-ready MiniMax-H3 audio-video prompt.

Output is text only. This is a prompt rewriter, not a video generator — feed the rewritten prompt and the same reference assets into a MiniMax-H3 pipeline to render.

Tasks

Task Inputs Output schema
T2AV text only integrated_multimodal_description, overall_soundscape, non_diegetic_music
I2AV text + 1 image (exact first frame) same 3 fields
L2AV text + 1 image (exact last frame) same 3 fields
FL2AV text + 2 ordered images (first & last frames) same 3 fields
Ref2AV text + ordered images / video / audio references subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music

Duration is requested in whole seconds (4–15) and snapped to MiniMax-H3's legal 17 * n + 5 frame grid at 24 fps. Base tasks accept any aspect preset; Ref2AV supports only 16:9 and 9:16. Ref2AV prompts must mention every supplied reference label (<Picture 1>, <Video 1>, <Audio 1>) and no others.

Fidelity to the reference implementation

The system prompts (system_prompt.py) are copied verbatim from the adapter repo, and message construction, the duration grid, reference labelling / validation and the output schema check are ported 1:1 from the adapter repo's infer.py. Encoding uses the same ProcessorMixin.apply_chat_template path with max_pixels 301056 (images) / 100352 (video), load_audio_from_video=False and use_audio_in_video=False. Those processing kwargs are passed through transformers 5.x's processor_kwargs dict (transformers 4.x, which the reference pins, is incompatible with huggingface_hub 1.x / Gradio 6).

Runs on ZeroGPU: the Thinker is loaded in bf16 with SDPA attention and the LoRA is attached with PEFT at module scope.

Credits

Example prompts and reference frames are the adapter authors' own, from lightx2v/MiniMax-H3-Prompt-Rewriter-LoRA-Omni (assets/examples/), released under the repo's Apache-2.0-named license.