Inputs
ReadyText, 1–9 images, Video, 2–9 references
Models
Omni-modal video creation
Create directed 4–15 second videos from text, a first-frame image, reference footage, or combined visual sources, with 768P or 2K output and a generated audio track.
DreamMotion AI provides an independent workspace for this model. Current controls and credits reflect the verified integration.
Current integration
The MiniMax H3 AI Video Generator supports text-to-video, first-frame image animation, reference-video transformation, and multi-image direction. Choose a 4–15 second duration, supported aspect ratio, and 768P or 2K output; the submitted task can return a synchronized audio track. Review the exact credit total before generation.
Inputs
ReadyText, 1–9 images, Video, 2–9 references
Resolution
Ready768p, 2K
Aspect ratios
Ready21:9, 16:9, 4:3, 1:1, 3:4, 9:16
Duration
Ready4s, 5s, 6s, 7s, 8s, 9s, 10s, 11s, 12s, 13s, 14s, 15s
Create on this page
The workspace is preselected for this model. Add a prompt or references, choose supported settings, and review the exact credit total before generating.
Choose an example
Prompt framework
Build stronger MiniMax H3 prompts by defining the subject, ordered action, environment, camera movement, shot sequence, continuity, lighting, audio, and final output constraints.
Define the opening image, ordered action, camera progression, transitions, audio moments, and closing frame. A short sequence of explicit beats gives H3 a usable timeline without repeating the model name or filling the prompt with generic quality terms.
State whether each source controls character identity, product detail, environment, movement, camera language, or sound. Compatible roles make the result easier to direct and reduce contradictions between visual references.
Quote required visible wording and list the facial features, product geometry, wardrobe, props, and spatial relationships that must remain unchanged. Then describe only the movement and atmosphere that should evolve.
Reusable example
Create a [4–15 second] [aspect ratio] video featuring [subject] completing [ordered action] in [environment]. Begin with [opening frame], move through [camera path and shot beats], and finish on [defined final frame]. Preserve [identity, product, typography, or reference details]. Coordinate [dialogue, ambience, music, and effects] with specific visual moments, avoid [continuity errors], and deliver at [768P or 2K] with coherent motion and a synchronized audio track.
Model overview
MiniMax H3 is a multimodal video creation model for generating a new shot from text, animating a still image, transforming existing footage, or combining several visual references. These four starting points let you choose between open-ended creation and tighter control over identity, composition, movement, environment, and style.
The MiniMax H3 AI Video Generator brings MiniMax’s omni-modal video model into the DreamMotion AI workspace. It supports a written brief, first-frame animation, reference footage, and coordinated multimodal sources for 4–15 second MP4 videos at 768P or 2K. Picture and an audio track can be planned together for dialogue-led scenes, product campaigns, dynamic posters, game interfaces, and cinematic stories. The workspace exposes the active input mode, source limits, aspect ratio, duration, resolution, and exact credit estimate before submission. Prompt moderation runs before credits are reserved or APIMart receives a generation task. The same workspace keeps references, prompts, generation status, recent results, and failed-task refunds together, so the creative direction and production settings remain visible throughout the task.
How to use
Follow the same four-step path from the first input to the final credit review, using only controls available in the current DreamMotion workspace.
Use text for a new scene, one image for first-frame animation, reference footage for motion direction, or multiple images when identity, product, setting, and style need separate controls.
Describe the subject, action order, camera, lighting, continuity, dialogue, ambience, effects, and final frame as one production brief.
Choose a 4–15 second duration, supported aspect ratio, and either 768P or 2K output for the actual destination.
Confirm the displayed total before generation. Moderation runs first, and confirmed failed generation tasks return reserved credits automatically.
Workflow guide
Start from the asset that already contains the most important creative decisions. The right input mode reduces unnecessary prompt instructions and gives the model a clearer definition of what may change and what must stay consistent.
Build a complete 4–15 second scene from a production brief. Define the opening composition, ordered actions, camera path, lighting changes, dialogue, ambience, effects, and closing frame so MiniMax H3 can coordinate picture and an audio track in one task.
Best for: Dialogue scenes, cinematic concepts, product stories, and sound-aware social video.
Start from an approved first frame when identity, product geometry, typography, or composition must remain recognizable. Describe the motion, camera development, evolving atmosphere, sound cues, and final beat without restating visual details that are already clear in the source.
Best for: Portrait performance, product reveals, dynamic posters, and designed campaign assets.
Use reference footage to establish action, timing, or camera rhythm, then specify the treatment MiniMax H3 should change. Separate protected motion beats from new character, environment, material, lighting, and audio direction to keep the transformation readable.
Best for: Motion transfer, visual restyling, alternate campaign treatments, and shot development.
Combine compatible images and video when different sources need to control identity, product detail, setting, motion, or style. Assign every reference one explicit role, resolve conflicts in the prompt, and describe the sound cues and details that must remain continuous.
Best for: Reference-led brand films, character continuity, ecommerce sequences, and multimodal concepts.
Capabilities and best use cases
Explore MiniMax H3 examples for 2K detail, structured scene direction, synchronized picture and sound, and reference-led production.
Higher-detail output
Use 2K output when faces, products, materials, or environmental texture need more room than an iteration preview.
Best for

Native audiovisual direction
Direct sound at the same visual beats as performance and camera motion, producing a connected audiovisual result.
Best for

Structured sequence direction
Define ordered actions, transitions, continuity, visible details, and an ending beat instead of relying on broad style words.
Best for

Visible design and product control
Coordinate product detail, visible text, interface motion, transitions, and sound for a campaign-ready concept.
Best for
Capability examples use official minimax h3 product media from the cited official source.
Open official sourceModel comparison
Choose by generation workflow, duration, creative control, output quality, and current credit range.
Text and image to video
Cinematic prompt-to-video shots
Longer cinematic sequences
2K cinematic video
Cinematic text-to-video scenes
Connected multishot story planning
Choose MiniMax H3 when the brief depends on 2K delivery, coordinated picture and sound, visible text, structured instructions, or combined image and video references. Choose Seedance 2.5 for a longer 4–30 second production window and a larger reference allowance. Choose Seedance 2.0 when output up to 4K is the priority. Compare each model by input type, maximum duration, final resolution, source limits, audio behavior, and current credit total.
Good to know
Clear answers about the current integration, credits, references, and generation behavior.
It is an omni-modal MiniMax video model available through DreamMotion AI and APIMart. The current integration supports text, first-frame image, reference video, and multi-image workflows for 4–15 second MP4 results at 768P or 2K.
MiniMax H3 can return video with an audio track. Include dialogue, ambience, effects, music, timing, and sound direction in the same prompt as the visible action. APIMart also documents reference-audio support at the model API level; DreamMotion AI displays the source types currently available in each workspace mode.
DreamMotion AI offers continuous durations from 4 through 15 seconds and 768P or 2K output. Concrete landscape, portrait, square, standard, and ultrawide aspect ratios are available.
Yes. Image to Video uses a first-frame image, References to Video accepts two to nine images, and Video to Video can use a reference clip. Assign each source a clear role and avoid mixing incompatible creative directions.
The total depends on duration and whether 768P or 2K is selected. DreamMotion AI calculates the requirement from the current controls and displays it before the task can be submitted.
A moderation denial stops before credits are consumed or a provider task begins. If an authorized provider task later reaches a confirmed failed state, DreamMotion AI returns its reserved credits.
Start inside the embedded workspace or open the full creator to keep more room for references, settings, and generated results.
10 free credits on sign-up · No credit card required