Launching here tomorrow, and the thing I keep going back and forth on is how much the audio step actually costs people.
Most video models hand back a silent MP4. The workflow I see most often is: render, then go find music, then cut sound effects, then if there is dialogue, dub it and hope the lips line up. Three tools and a timeline for something that started as one prompt.
Most video models hand back a silent MP4 that you then have to score. H3 generates dialogue, sound effects and music in the same forward pass as the picture, 32 kHz stereo, dialogue stable in eleven languages - so there is no silent version to fix later, and no mute toggle, because the sound is not a separate layer.
Four inputs, 4-15 seconds at any whole second, 24 fps, 768P or 2K, six aspect ratios. First clip runs with no account. Independent third party, not affiliated with MiniMax.