Instructions to use pablodawson/MiniMax-H3-360-Orbit-LoRA with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Inference
- Notebooks
- Google Colab
- Kaggle
MiniMax-H3 360° Orbit LoRA (first–last frame)
A LoRA for MiniMax-H3 (FL2VA variant) that turns one photo into a frozen-time 360° camera orbit, and lands back on the exact frame it started from.
Give the model the same image as the first and the last keyframe, and the LoRA orbits the camera all the way around the subject while the scene stays frozen. The clip closes on its own first frame, so orbits can be chained or stitched into longer shots without a visible seam.
Results
Each comparison shows three renders of the same input with the same seed, prompt, resolution and step count:
| left | middle | right |
|---|---|---|
| base model, first frame only | base model, first + last frame | base + this LoRA, first + last frame |




The previews are downscaled animated WebP. Full-resolution MP4s: comparisons skate · man2 · girl2 · man; LoRA-only clips skate · man2 · girl2 · man.
About the comparisons
- Base, first + last frame: when both keyframes are the same image, the base model reads the clip as a still and barely moves.
- Base, first frame only: the camera moves, but drifts away and never returns to the starting view.
- With the LoRA: the camera travels the full orbit and returns to the starting frame, with smooth motion and no cuts.
Why this LoRA?
Reference-to-video (Ref2VA) treats its input images as loose appearance references, so clips don't end on a known frame and can't be joined cleanly. First–last-frame generation (FL2VA) pins both ends, but out of the box it freezes when both ends are the same image. This LoRA teaches the model a real, geometry-consistent orbit, while FL2VA keyframe pinning at inference keeps both ends exact.
Prompt
Use this prompt verbatim.
One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
Recommended settings
| setting | value |
|---|---|
| base weights | MiniMax-H3 FL2VA, pruned (Comfy-Org/MiniMax-H3 int8 ConvRot repack) |
| keyframes | the same image as first and last frame, for a full 360° loop |
| resolution | 768 × 768 |
| frames | 73 (about 3 s at 24 fps) |
| steps | 28 |
| guidance | none. MiniMax-H3 is guidance-distilled, so there is no CFG or negative prompt |
| LoRA strength | 1.0 |
| audio | off |
Usage
These weights were trained and tested with ostris/ai-toolkit and its minimax_h3 model extension. The keys use ComfyUI naming (diffusion_model.*).
Training details
Data
Hand-selected high-quality human Gaussian splats collected from the internet, with an orbit video rendered around each one. Because the splats are static 3D scenes, every frame is geometrically consistent by construction. The dataset shows exactly the motion the LoRA should learn: the camera moves and nothing else does.
| clips | 28 orbit renders |
| clip format | 768 × 768, 73 frames, 24 fps |
| captions | the single prompt above, for every clip; caption dropout 0.05 |
| audio | none |
Hyperparameters
| trainer | ostris/ai-toolkit, arch minimax_h3, partition fl2va_pruned |
| base model | MiniMax-H3 FL2VA pruned, int8 ConvRot (convrot8); Qwen3-VL-32B text encoder in nvfp4 |
| training adapter | ostris/minimax_h3_training_adapter v1, active during training only and not needed for inference |
| conditioning | first frame of each clip as the keyframe (i2v) |
| LoRA | rank 16, alpha 16, all transformer linear layers except adaln_proj (208 modules) |
| steps | 3000 (≈107 epochs) |
| batch size | 1, no gradient accumulation |
| optimizer | AdamW 8-bit, lr 1e-4, weight decay 1e-4 |
| precision | bf16, gradient checkpointing |
| noise schedule | flow matching, shifted timesteps |
| loss | MSE, guidance loss target 3.5 |
| resolution buckets | 512, 768 |
| hardware | 1 × NVIDIA A100 80 GB, 10 h 24 min (about 12.5 s/step) |
Limitations
- The domain is narrow. The training data is 28 human-centric square clips of 73 frames. Other subjects, aspect ratios and clip lengths are untested.
- You may still get some subtle movements, blinks, etc.
Author
Pablo Dawson
- Downloads last month
- 1,429
Model tree for pablodawson/MiniMax-H3-360-Orbit-LoRA
Base model
MiniMaxAI/MiniMax-H3