Quick answer
To use MiniMax H3 on minimaxh3.tv, follow six steps: choose Text to Video, Start & End Frames, or References; prepare the required input; describe the motion and camera behavior; check duration, aspect ratio, and the fixed 2K setting; review the credit quote; then generate and watch the entire clip. Text mode is the simplest place to start. Frame mode gives you control over the opening image and, optionally, the ending image. Reference mode accepts ordered images, videos, and audio when you need the model to follow a subject, movement, camera pattern, or sound cue. On this site, MiniMax H3 currently uses 5 to 15 second integer durations and native 2K output. Text and reference modes offer six aspect ratios, while frame mode follows the uploaded image. The official MiniMax API supports 4 to 15 seconds, so keep the official limit separate from this site's current interface.
Last verified: August 3, 2026. minimaxh3.tv is an independent service and is not owned, operated, authorized, or endorsed by MiniMax.

The first run is easier to debug when mode, input, motion, settings, generation, and review stay separate.
Choose text, image, or reference mode
Open the MiniMax H3 model page or go straight to one of the three generator pages. All three routes use MiniMax H3, but they ask for different inputs.
| Mode | Start with | Best fit | Do not choose it when |
|---|---|---|---|
| Text to Video | A written description | You can describe the whole shot without a fixed source image | The opening composition must match an existing image |
| Start & End Frames | A required start image and an optional end image | You need to preserve an opening frame or guide a transition toward an ending | You need several subject, motion, or audio references |
| References | One or more ordered reference assets | You need to guide subject identity, movement, camera behavior, style, voice, or rhythm | A short text prompt already describes the shot clearly |

Text starts from a description. Frame mode controls the opening or ending. Reference mode assigns jobs to ordered assets.
Prepare your inputs
Text to video
You only need a prompt. Write the shot as a sequence, not a pile of visual keywords. Name the subject, what changes, how the environment reacts, what the camera does, and how the shot ends.
A useful pattern is:
subject + action over time + environment response + camera movement + ending state
Keep the important action early. The current minimaxh3.tv prompt box shows a 2,500-character limit, while the official API documentation lists a 7,000-character prompt limit. The smaller site limit is the one that matters when you use this interface.
Start and end frames
The current image mode asks for a start image and lets you add an optional end image. Use a sharp source with a clear subject and enough room for the motion you want. Each image can be up to 30 MB on this site. JPG, JPEG, PNG, and WEBP are supported.
Do not spend the prompt repeating everything already visible. Describe what should move, what should remain stable, and how the camera should travel from the start frame toward the end frame. The output composition follows the uploaded frame, so this mode does not expose the normal aspect-ratio choice.
Image, video, and audio references
Reference mode accepts up to 9 images, 3 videos, and 3 audio files. Each reference video can be 2 to 15 seconds, with no more than 15 seconds of reference video in total. The same 15-second total applies to reference audio. Video files can be up to 50 MB, images up to 30 MB, and WAV or MP3 audio up to 15 MB per file.
At least one image or video is required. Audio cannot be the only reference. The official API also caps a mixed request at 12 assets, so do not assume that every per-type maximum can be used at the same time.
Put the files in a deliberate order, then refer to that order in the prompt: "Use Image 1 for the character, Video 2 for the camera move, and Audio 3 for the rhythm." This is more precise than asking the model to "use all references."
Write motion and camera instructions
Static descriptions tell the model what a frame looks like. A video prompt also needs time. Write what happens first, what follows, and where the shot ends.
This existing minimaxh3.tv text-to-video example used the prompt:
A cinematic sunrise over a calm alpine lake, soft mist drifting above the water, slow camera push forward.
The prompt names the subject, visible motion, and one camera instruction. For a first run, one clear camera move is usually easier to diagnose than several competing moves.
Select current settings
The public interface currently defaults to MiniMax H3, 5 seconds, 16:9, and 2K in text mode. The duration can be any whole number from 5 to 15 seconds. Text and reference modes support 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Frame mode follows the uploaded image.
| Setting | minimaxh3.tv interface | Official MiniMax API |
|---|---|---|
| Duration | 5 to 15 seconds, whole seconds | 4 to 15 seconds, whole seconds |
| Resolution | 2K | 768P or 2K |
| Text/reference ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16 | Common ratios or adaptive, depending on mode |
| Frame-mode ratio | Follows the input image | Adaptive |
| Prompt length | 2,500 characters in the current prompt box | Up to 7,000 characters |
These ranges were checked on August 3, 2026. Use the values shown in the generator when they change. The quote beside the Generate button is calculated from the selected duration and reference inputs. Check that quote before submitting because a submitted H3 job cannot be cancelled.
Generate and check the full clip
- Sign in, choose the mode, and confirm that MiniMax H3 is selected.
- Add the prompt and any required source files.
- Check the full credit quote, duration, ratio, and 2K setting.
- Submit once. Do not press Generate again because the first task looks slow.
- Watch the result from beginning to end before deciding what to change.
Check the opening composition, subject identity, action order, camera direction, and ending frame. If you supplied audio, listen for timing as well as sound quality. Small defects near the end are easy to miss when you only preview the first seconds.
Completed upstream results are available for 24 hours, so save a result you want to keep. A terminal provider failure restores the reserved credits, but a submitted task cannot be cancelled manually.
Common first-run problems
| Problem | Likely cause | What to change | Evidence |
|---|---|---|---|
| The clip looks static | The prompt describes appearance but not change over time | Add one action sequence and one camera instruction | Prompt structure and the verified text example above |
| The opening frame drifts | Text mode was used even though the first composition matters | Switch to Start & End Frames and upload the opening image | Current frame workflow |
| Reference mode will not submit | Only audio was added, or no image/video is present | Add at least one image or video | Current site and official input rules |
| An upload is rejected | The file type, size, duration, or total reference duration is outside the current limit | Convert or trim the asset before uploading | Current site limits checked August 3, 2026 |
| The result mixes reference roles | The prompt never says which file controls which part | Name the ordered asset and its job in one sentence | Official reference workflow |
| The final seconds break down | Too many actions compete inside a short clip | Reduce the number of beats or extend the duration within the 15-second limit | Practical prompt revision, not a measured success-rate claim |
Which workflow fits which task
| Task | Recommended workflow | Reason |
|---|---|---|
| Explore an idea from scratch | Text to Video | It needs no source asset and is the fastest mode to set up |
| Animate a product still or illustration | Start & End Frames | The source image fixes the opening composition |
| Build a controlled transition | Start & End Frames | An optional end frame gives the shot a destination |
| Follow a character, object, camera move, or audio cue | References | Each ordered asset can be assigned a specific job |
| Integrate H3 into an application | API | Use the MiniMax H3 API guide instead of the browser workflow |
If you are still checking availability, naming, and launch scope, read the MiniMax H3 release guide. For hands-on creation, open the MiniMax H3 generator and start with the simplest mode that preserves the control you need.
FAQ
Can I use MiniMax H3 online without the API?
Yes. minimaxh3.tv provides browser-based text, frame, and reference workflows. It is an independent service, not an official MiniMax website.
Which mode should a beginner use first?
Use Text to Video if you do not need to preserve an existing composition. It has the fewest required inputs, which makes prompt problems easier to isolate.
Can I upload only audio in Reference mode?
No. Reference mode requires at least one image or video. Audio can be added alongside those assets.
What is the current MiniMax H3 duration on this site?
The current minimaxh3.tv interface supports whole-second durations from 5 to 15 seconds. The official API documentation lists 4 to 15 seconds. These are different surfaces.
Does MiniMax H3 support an end frame?
Yes. The site's Start & End Frames workflow uses a required start image and an optional end image in its current interface.
Can I change the aspect ratio in frame mode?
Frame mode follows the uploaded image. Text and reference modes expose six ratio choices.
How much does a generation cost?
The generator displays a credit quote before submission. The quote changes with duration and reference inputs, so use the current on-screen value instead of a fixed number from a guide.
Can I cancel a submitted task?
No. Check the prompt, files, settings, and quote before you submit.
Sources and methodology
This guide was checked on August 3, 2026 against the current public MiniMax H3 model page, the three generator routes, one existing public minimaxh3.tv text-to-video example, the MiniMax H3 launch article, and the official MiniMax Video Generation guide. No new paid video task was created for this article. The public workflow and upload limits come from the current site. Official model and API limits come from MiniMax documentation. No external video or community claim is used.



