MiniMax H3
Video Generator · 20 credits / 768p 5s video
MiniMax H3 video generation from text, start and end frames, or multimodal references.
Create MiniMax H3 videos with native audio.
MiniMax H3 video capabilities
Use the MiniMax H3 video generator to combine text, image, video, and audio references, make precise multimodal edits, preserve typography and UI details, and generate synchronized native stereo audio in one workflow.
MiniMax H3 Reference to Video
Add up to 9 images, 3 video clips, and 3 audio clips to one MiniMax H3 reference-to-video generation. Use each reference to guide character identity, products, performance, camera movement, visual style, editing rhythm, voice, or sound—then carry those signals through one coherent video.
Precise MiniMax H3 Video Editing
Use natural-language MiniMax H3 video editing to replace a product, rewrite on-screen text, change dialogue, relight a scene, or add and remove objects. Fine-grained multimodal controls focus the edit on what you describe while preserving the surrounding character, composition, motion, pacing, and sound.
Text, UI & Brand Graphics
Generate product pages, game menus, HUDs, subtitles, title cards, and animated UI with clearer text and consistent visual structure. MiniMax H3 accepts prompts up to 7,000 characters, giving you room to describe complete shot lists, interface motion, branding, camera work, and timing in one AI video prompt.
MiniMax H3 Native Audio
Generate native 32 kHz stereo audio with the MiniMax H3 video: dialogue, ambience, foley, sound effects, and music timed to the scene. Add an audio reference to guide a voice or performance and keep sound aligned with the action from the first frame to the last.
MiniMax H3 Examples
Explore MiniMax H3 videos for brand films, consistent characters, stylized animation, gameplay, motion transfer, and precise AI video editing—all from the same multimodal video model.
First & Last Frame
“Use four reference images as sequential MiniMax H3 keyframes inside a fixed vintage binocular viewfinder. Begin slightly out of focus, push toward the installation, and rack focus as each product detail appears. Connect the frames with fast optical scans, restrained motion blur, film grain, soft halation, and precise red typography while keeping the twin-lens mask locked in place.”
Reference to Video
“Use the first image to lock the character identity, hairstyle, silver crown, layered blue costume, and accessories. Use the second reference for storyboard order and pacing. Generate a cinematic 16:9 wuxia sequence with expressive hand movement, controlled camera motion, consistent facial details, luminous fabric, and seamless transitions between every story beat.”
Video Style Reference
“Preserve the original buildings, pedestrians, timing, camera path, and natural lighting from the live-action video. Transform only the vehicles and trees into tactile voxel objects using the image as a style reference. Keep motion physically believable and make every block, shadow, reflection, and contact point feel integrated into the real environment.”
Text to Video
“Create a 15-second 16:9 late-night laundromat scene that blends live action with hand-drawn luminous animation. Use flickering fluorescent lights, turning washers, window reflections, quiet room tone, and a nostalgic atmosphere. Film it like a spontaneous one-handed phone recording with slight shake, exposure shifts, delayed autofocus, and an apparition moving through the space.”
Motion Reference
“Match the complete movement sequence and locked wide-camera timing from the reference video, but replace the performers with three photoreal capybaras. Preserve every drop, roll, jump, position change, and final pyramid pose. Integrate realistic fur, weight, floor contact, lighting, and shadows without changing the original choreography.”
AI Video Editing
“Remove the green-screen background from the source performance and replace it with a cinematic fairytale environment guided by the second video. Match the new scenery to the performers’ movement, camera perspective, and timing. Relight the characters, rebuild contact shadows, and balance color so the foreground and generated background feel like one continuous shot.”
MiniMax H3 online workflow
Start with a MiniMax H3 prompt or reference materials. Choose the matching generation mode, then set the duration, framing, and audio direction.
For MiniMax H3, write the subject, action, setting, camera movement, spoken lines, sound effects, and music in one prompt. Upload media when an element must remain recognizable.
Use MiniMax H3 text to video for a new scene, image to video for a fixed opening, keyframes for a controlled transition, or reference mode for mixed media inputs.
Generate with MiniMax H3 after setting aspect ratio, duration, resolution, and native audio. Review visual consistency and sound synchronization before continuing the workflow.
Choose the MiniMax H3 workflow closest to your intent. Create from text, animate an image, or combine references to control identity, motion, camera, style, and voice.

Use MiniMax H3 text to video to generate a complete clip and native stereo soundtrack from the scene, camera, dialogue, effects, and music in your prompt.
Open workflow
Use MiniMax H3 image to video to turn your upload into the opening frame, then guide motion, camera path, dialogue, and sound with a prompt.
Open workflow
Use MiniMax H3 reference mode to combine images, clips, and audio that guide character identity, products, style, movement, camera work, and voice.
Open workflowClear answers about MiniMax H3 video generation, supported inputs, native audio, output specifications, prompting, pricing, and the relationship between EZMiniMax and MiniMax.
MiniMax H3 is an omni-modal generative system. It understands text, images, video, and audio in one context and creates native stereo audio with dialogue, environmental sound, effects, and music.
MiniMax H3 supports text prompts, first and last frame images, and mixed reference inputs. Reference mode accepts up to 9 images, 3 video clips, and 3 audio clips, with no more than 12 files in one request.
Yes. MiniMax H3 uses an uploaded image as the first frame, then develops the subject, scene, and camera motion while generating synchronized dialogue, sound effects, ambience, and music.
Yes. MiniMax H3 generates 32 kHz stereo audio natively. Describe dialogue, sound effects, ambience, and music in the same prompt so their timing and relationship to the visuals are modeled together.
MiniMax H3 supports videos from 4 to 15 seconds at 24 FPS across common landscape, square, portrait, and cinematic aspect ratios. The base workflow uses a 768-pixel short edge, while the complete official regeneration workflow can produce output up to 2K.
No. EZMiniMax is an independent multi-model creative platform and is not affiliated with or endorsed by MiniMax or its affiliates. Third-party model and company names identify compatible workflows only. Model availability can change, so the in-product model selector is the source of truth.
Reference to video lets you mix images, videos, and audio in one request. You can tell MiniMax H3 which reference should control the character or product, visual style, motion, camera movement, editing rhythm, or voice in the generated clip.
Describe the complete audiovisual timeline: scene and subjects first, then actions, shot changes, camera movement, spoken lines, sound effects, ambience, and music. For reference generation, explicitly assign each image, video, or audio input to the role it should control.
Generation cost depends on the selected duration, resolution, and workflow. EZMiniMax uses credits and shows the current task cost in the product before generation. Review the pricing page for available plans and credit packages.
Commercial use depends on the current MiniMax H3 license and provider terms, your EZMiniMax plan, and the rights attached to prompts, uploads, references, brands, people, and generated content. Review the applicable terms and confirm that you hold the necessary rights before publishing.
Start from a prompt, an opening image, first and last frames, or a set of image, video, and audio references. Generate the video and synchronized soundtrack in one workflow.
Generate with MiniMax H3