Image-to-Video Prompt Examples for MiniMax H3

MiniMax H3
|
Published on Aug 5, 2026

Quick answer

An image-to-video prompt should anchor the input frame and direct the motion. Tell MiniMax H3 what must stay fixed, what begins moving, how the action develops, and where the shot settles. Start with a clean frame, choose one subject action and one camera move, then add only the environmental motion that the image can support.

Use this working formula:

first-frame anchor + subject action + environmental motion + camera motion + continuity constraint + ending state

The examples below are writing templates, not claims of tested results. Replace the bracketed details with what is actually visible in your image. MiniMax's official H3 guide recommends developing an image-to-video shot from the first-frame state through the action onset, continuous movement, and a visible result or reaction. The current H3 API also requires a non-empty text prompt alongside the image.

Last verified: August 5, 2026. minimaxh3.tv is an independent service and is not owned, operated, authorized, or endorsed by MiniMax.

Six original image-to-video starting frames: portrait, product bottle, fox illustration, interior, dessert, and cafe

Six original sample inputs for planning motion. These synthetic stills are not MiniMax H3 output examples.

Image-to-video prompt formula

Read the input image before writing motion. Note the subject, pose, camera angle, lighting, empty space, and anything that must remain legible. Then write the prompt in this order:

  1. Anchor the visible frame: name the subject, framing, style, and fixed spatial relationships already present.
  2. Start the action: describe the first small movement that can plausibly grow from the still image.
  3. Develop the movement: state what changes over time and what stays fixed.
  4. Move the camera: use one clear move and add speed or range only when it matters.
  5. Land the shot: describe the final pose, composition, or reaction.

MiniMax's H3 prompt guide treats the input as the actual first frame. It recommends preserving identity, clothing, colors, key objects, and spatial relationships while the scene develops forward. For camera motion, it suggests a motion type plus meaningful speed or amplitude, written as a natural action inside the shot.

Four-panel motion plan for an unbranded product bottle, showing the input, camera arc, environmental movement, and final composition

Plan the shot in layers: preserve the product, move the camera, add compatible environmental motion, then define the final composition.

Prompt part What to write What to avoid
Input anchor Facts already visible in the frame Re-inventing clothes, objects, or layout
Subject action One observable action with a clear start Several unrelated actions at once
Environment Motion that fits the scene, such as curtains or steam New objects with no visual basis
Camera One move, direction, speed, and optional range A stack of conflicting camera terms
Continuity What must keep its shape, color, or position Guarantees such as "perfect consistency"
Ending A visible final pose or composition An abstract mood with no visual outcome

Portrait motion prompts

Subtle portrait turn

Input image condition: a fictional adult in side profile, shoulders visible, uncluttered background, soft light on the face.

Motion goal: create a restrained turn toward the light without changing identity or clothing.

Prompt:

Live-action, cinematic. Begin from the exact side-profile portrait in the input image, preserving the person's facial features, hairstyle, linen shirt, shoulder position, background, and soft window light. The person takes a quiet breath, blinks once, and slowly turns their eyes toward the brighter side of the frame before the chin follows slightly. A few loose strands of hair move gently. The camera pushes in with small amplitude at slow speed. End on a calm three-quarter profile with the original framing and color palette intact.

Checkpoint: watch the face, hairline, earrings, neckline, and shoulder position. If several features drift, remove the camera move and test only the eye and head motion.

Portrait with environmental motion

Input image condition: a seated portrait near a curtain or plant, with room around the subject.

Motion goal: keep the subject nearly still while the environment supplies the movement.

Prompt:

Preserve the seated person's appearance, pose, clothing, chair, and room layout from the input frame. The subject makes a small natural blink and shifts their gaze toward the window. The curtain lifts once in a light breeze and settles; nearby leaves sway with the same airflow. Hold a static camera. End with the subject in the original pose and the room quiet again.

Checkpoint: the curtain and leaves should move in the same direction. If the face changes, remove all subject motion and keep only the breeze.

Product photo prompts

Slow product arc

Input image condition: one unbranded product on a clean pedestal, no small print, enough space around the object.

Motion goal: add depth while preserving the product silhouette and material.

Prompt:

Product film, clean studio lighting. Preserve the exact bottle shape, cap, ceramic texture, muted colors, pedestal, and background from the input image. The camera arcs slowly from left to right with small amplitude while soft leaf shadows drift across the backdrop. Fine dust catches the side light. The bottle remains still and centered. End on a slightly more frontal angle with the silhouette and surface details unchanged.

Checkpoint: compare the cap, neck, label area, and base between the first and last moments. If geometry bends, switch to a static camera and animate only light or shadow.

Simple reveal for a packaged object

Input image condition: a box or container with no protected logo and no text that must remain exact.

Motion goal: create a short reveal without inventing a second product.

Prompt:

Start from the exact product arrangement in the image. A narrow band of warm light moves from left to right across the package, revealing the existing texture. The background darkens slightly as the camera pushes in at slow speed. Keep the package dimensions, edges, color, and placement fixed. End when the light reaches the right edge and the product fills slightly more of the frame.

Checkpoint: any readable packaging text may warp. Use a text-free asset or add final typography during editing instead of asking the generation to preserve small copy.

Illustration and anime-style prompts

Watercolor fox animation

Input image condition: an original illustrated fox among simple plants, clear silhouette, no copyrighted character design.

Motion goal: animate a small reaction while keeping the drawn medium.

Prompt:

Hand-painted watercolor animation. Preserve the fox's original proportions, orange-and-cream markings, ink outline, seated pose, and surrounding botanical composition. The fox's ears twitch, its eyes follow a drifting leaf, and the tail curls inward slightly. Two nearby stems bend in a gentle breeze. The camera holds still. Keep the paper texture, watercolor edges, and limited palette consistent. End as the leaf settles beside the fox's front paws.

Checkpoint: verify that the medium stays watercolor and the fox does not become a different character or shift into a 3D style.

Single-action character beat

Input image condition: one original character with visible hands and a simple prop.

Motion goal: turn the still into a readable one-beat action.

Prompt:

Preserve the original 2D illustration style, character design, outfit colors, hand shape, and prop. The character looks down, tightens their grip on the paper lantern, then raises it until warm light reaches their face. The camera pulls out with small amplitude at slow speed. Keep the background linework stable. End with the lantern at chest height and the character looking forward.

Checkpoint: if the hands or prop deform, shorten the movement and keep both within the silhouette already implied by the image.

Architecture and interior prompts

Sunlight across a quiet room

Input image condition: an interior with straight architectural lines, a visible window, and no people.

Motion goal: make the room feel alive without moving furniture.

Prompt:

Photoreal interior film. Preserve the room layout, walls, windows, sofa, tables, artwork, and all straight architectural lines from the input image. Morning sunlight moves slowly across the floor and sofa. The sheer curtain lifts gently, and plant leaves sway near the window. The camera pushes in with small amplitude at slow speed. Keep every piece of furniture fixed. End as the light reaches the coffee table.

Checkpoint: inspect vertical lines, table edges, and furniture count. If the room warps, use a static camera and animate only light, curtains, and plants.

Exterior facade with a controlled tilt

Input image condition: a centered building facade with clear verticals and open sky.

Motion goal: reveal height while protecting geometry.

Prompt:

Begin from the exact building facade and street composition in the input image. Clouds move slowly behind the roofline while a few tree leaves stir at the bottom of frame. The camera tilts upward at slow speed with small amplitude, keeping the facade centered and all windows, columns, and edges structurally unchanged. End with more sky visible above the building.

Checkpoint: if windows multiply or bend, remove the tilt and use only cloud and foliage motion.

Food and macro prompts

Dessert close-up

Input image condition: one plated dessert, clean surface, no hands or utensils intersecting the food.

Motion goal: add appetizing micro-motion without changing ingredients.

Prompt:

Macro food film. Preserve the tart shell, strawberry arrangement, cream shape, garnish, plate, and background from the input image. A thin ribbon of glaze catches the light as the camera pushes in with small amplitude at slow speed. A faint wisp of cool mist passes behind the dessert and fades. Keep every fruit piece and garnish in place. End on a tighter close-up of the existing surface texture.

Checkpoint: compare ingredient count and placement. If fruit appears or disappears, remove the glaze action and animate only the camera and light.

Steam without object drift

Input image condition: a hot drink or bowl photographed from a stable angle.

Motion goal: introduce steam and a small focus shift.

Prompt:

Preserve the cup, handle, saucer, liquid level, tabletop, and background exactly as shown. Thin steam rises in irregular curls and disperses naturally. The camera remains fixed while focus shifts slowly from the rim to the rising steam, then returns to the cup. End on the original composition.

Checkpoint: the vessel and liquid level should not change. If they do, keep the camera and focus fixed and request only steam.

Social vertical video prompts

Cafe establishing clip

Input image condition: a vertical cafe or storefront image with no readable brand signage and clear foreground space.

Motion goal: create a short opening shot for a social edit.

Prompt:

Vertical lifestyle film. Preserve the cafe layout, tables, chairs, plants, window frames, and warm interior light from the input image. Outside, one cyclist passes through the window reflection while leaves near the glass move gently. The camera pushes forward at slow speed toward the nearest empty table. Keep the furniture fixed and the room free of new people. End with the table centered and enough clean space for text added later in editing.

Checkpoint: generate the clean footage first. Add captions and logos in the editor, where spelling and placement stay controllable.

Product detail for a vertical cut

Input image condition: a close product crop with clear top and bottom space.

Motion goal: make a short insert that can sit between wider shots.

Prompt:

Preserve the product's shape, material, color, and exact position. A soft highlight travels across the surface while the camera trucks right with small amplitude at slow speed. Keep the background simple and unchanged. End after the highlight reaches the far edge, leaving clean negative space above the product.

Checkpoint: keep the action short and singular. If the product rotates while the camera also trucks, remove one of those movements.

Weak prompt vs better prompt

Weak:

Make this image cinematic and dynamic with cool camera movement.

This leaves the model to guess what moves, which details matter, and where the shot should end.

Better:

Preserve the exact ceramic bottle, cap, stone pedestal, and muted teal background from the input image. The camera arcs slowly from left to right with small amplitude while leaf shadows drift across the backdrop. Keep the bottle still and centered. End on a slightly more frontal view with the original silhouette and material unchanged.

The better version anchors visible facts, assigns one camera move, separates subject motion from environmental motion, and defines an ending.

Failure-to-fix table

Symptom Likely cause Next edit Evidence status
Face or product shape changes Too much subject and camera motion at once Keep the camera static and test one small action Practical troubleshooting template, not a measured success rate
The clip feels frozen Prompt repeats the image but gives no action onset Add one observable first movement Based on the official forward-development structure
Camera motion fights the subject Several camera commands compete Keep one move and state its speed Based on the official camera-motion guidance
Furniture, windows, or ingredients multiply Prompt asks for new scene content State what stays fixed and remove new objects Template guidance, not a guarantee
Text warps Small lettering is treated as image detail Use a text-free input and add copy in editing Common workflow precaution, not an H3-specific benchmark
First and last frames do not connect The prompt describes both frames but skips the path Write the observable intermediate changes in order Based on the official first-and-last-frame structure
Motion is too busy Too many actions share a short clip Keep one subject action and one environmental action Template guidance, not a measured limit

FAQ

Should an image-to-video prompt describe the whole image?

Only enough to anchor the frame. Name the subject, composition, style, and details that must remain stable, then spend most of the prompt on motion and the ending state. Rewriting every background detail can crowd out the action.

Should I write camera terms or plain English?

Plain English works well when it is precise. The official H3 guide recommends writing camera motion as a natural action, such as a slow small push toward the subject. Choose one move before combining anything more complex.

Can I use a first frame and a last frame?

Yes. Current H3 documentation supports first-frame, last-frame, and first-plus-last-frame image-to-video inputs. Describe the motion path between the two frames instead of repeating two static descriptions.

Does the image control the aspect ratio?

For the H3 API's first-frame or last-frame image-to-video mode, the ratio is adaptive and follows the input image. Crop the source intentionally before generation, especially for a vertical social clip.

Are these prompts tested outputs?

No. They are source-backed writing templates paired with original example inputs. Results vary by image and settings. The article does not report a controlled H3 test set or a success rate.

Where can I try a prompt?

Open the MiniMax H3 image-to-video generator, upload a rights-cleared image, and test one motion variable at a time. For the broader prompt framework, read the MiniMax H3 prompt guide. You can also browse the showcase for published examples, but do not assume a displayed result used the same input or prompt unless the page says so.

Sources and methodology

Checked on August 5, 2026:

The prompt examples were drafted from the official H3 structure and paired with original synthetic input illustrations. No paid generation was run for this article, so the examples are not presented as tested outputs. No external video or community claim was used.

#MiniMax H3#image to video#prompt examples#AI video
Related Posts
View all articles
MiniMax H3 Prompt Guide: Formula, Examples, and Fixes

MiniMax H3 Prompt Guide: Formula, Examples, and Fixes

Learn a practical MiniMax H3 prompt formula, copyable text, frame and reference templates, verified examples, camera terms, and single-variable fixes.

How to Use MiniMax H3: Text, Images, and Video References

How to Use MiniMax H3: Text, Images, and Video References

Learn how to use MiniMax H3 with text, start and end frames, or image, video, and audio references using current 2K settings and limits.

MiniMax H3 API Guide: Create, Poll, and Download Video

MiniMax H3 API Guide: Create, Poll, and Download Video

Use the official MiniMax H3 V2 API to create text, frame, or reference video tasks, poll status, download results, and estimate current pricing.

MiniMax H3 Open Source Status: Hugging Face, GitHub, Weights, and License

MiniMax H3 Open Source Status: Hugging Face, GitHub, Weights, and License

Check MiniMax H3 Hugging Face, GitHub, ModelScope, open weights, download, and license status, verified August 1, 2026.