Nothing to upload — the whole clip, picture and soundtrack, is invented from the prompt.
The uploaded image becomes the first frame; the model animates forward from it.
Give both ends and the model interpolates the shot between them. Either box may be left empty — supply only the last frame and it invents the run‑up that lands on it.
Drop reference material — subjects, styles, motion, voices — and the model carries it into a new clip. Limits: 9 images, 3 videos, 3 audio clips, and 12 files total.
0/9 images · 0/3 videos · 0/3 audio · 0/12 total
This mode needs the Ref2VA checkpoint, which is not loaded on the pod right now.
MiniMax H3 responds to an explicitly staged prompt far better than to a single sentence. Four sections, in this order:
- Look — one paragraph of subject, lens, lighting and setting. No timing here.
- Timeline — bracketed windows
[0s‑2s]. Put every camera move and action in a window. Dialogue goes inside a window as …and says clearly in English: "the line". The quoted text is what gets spoken. - Camera grammar — one continuous shot, or hard cuts. Saying "no zooms" actually suppresses zooms.
- Audio — the headline feature. Name the ambience and where it sits in the stereo field, the voice and its exact time window, and every music or sfx entrance by timestamp. This is not a TTS track bolted on afterwards — the model decodes picture and stereo sound from the same latent, so anything you don't specify it will invent for you.
Finish with negative constraints such as no on‑screen text, no subtitles. Eleven languages are supported for dialogue, but naming the language in the line ("says clearly in English") is what makes it reliable.
Output
Advanced dimensions
Both must be multiples of 32. Anything much above the 768 short side is untested on this checkpoint — the 2K pass is API‑only and not part of the local model.
The video VAE only accepts frame counts where length % 17 == 5,
so the slider snaps to the nearest legal clip. The server re‑snaps and reports the value it actually used.
Above 243 frames (10.1 s) is beyond the lengths validated on this pod — expect long runs and higher VRAM.
20 is the calibrated default. The checkpoint is CFG‑distilled, so there is no guidance scale — steps are the only quality dial, and time scales almost linearly with them.
Leave empty for a fresh random seed each run. The server returns the seed it used, so any clip can be reproduced.
LoRAs
No LoRAs found. Drop .safetensors files into
ComfyUI/models/loras/ on the pod, then hit Rescan.
Applied to the diffusion model only, in the order listed. Strength 1.0 is how the LoRA was trained; below ~0.3 the effect can vanish into the checkpoint's int8 quantisation.
Extrapolated from three measured runs on this exact card. Treat it as a ballpark — queue waits, a cold model load (~90 s) and VAE decode all move it.
Nothing rendered yet
Write a prompt with a timeline and an audio paragraph, then hit generate. A 5‑second 768p clip takes about six minutes on this card and comes back as one mp4 with 32 kHz stereo already muxed in.
History
Finished clips land here.