MiniMaxDirector — Turbo
The director workflow with
larryvrh's Turbo LoRA
between the model loader and the sampler:
6 steps instead of 20, so a draft costs about a
fifth of the GPU time. Same timeline, same prompt, same output.
• Stay in
4-8 steps. 6-8 looks best; past 8 it stops helping and starts over-sharpening.
• Leave LoRA strength at
1.0 and the scheduler on
simple.
• Preview weights:
audio and fast motion are the weak spots. For a final take, use
minimaxh3-director.
---
Lay out shots on the
Director node's timeline; it compiles them into the single
structured prompt MiniMax H3 reads, and snaps the clip to a length H3 accepts
(
length % 17 == 5 at 24 fps).
•
prompt and
report outputs show exactly what was built and what the linter thinks.
• Drop an image on a shot with
Add Image; it becomes
<Picture 1> automatically.
• Models: the H3 bundle (ref2va unet, Qwen3-VL text encoder, video + audio VAEs).
The three prompt buttons
| Button | Track | Makes |
|---|
| Add Video Prompt | MAIN | what happens on screen |
| Add Sound Prompt | AUDIO | what is heard -- H3 generates it, no file |
| Add Camera Prompt | CAMERA | how the camera moves |
Add Image / Add Audio / Add Video attach a real file instead, and the prose is given
a
<Picture n> /
<Audio n> /
<Video n> token pointing at it. With a block selected that
carries no file yet, the file lands on that block; otherwise it gets a block of its own.
Presets, copying, and the keyboard
Clear empties the timeline: every block, the global prompt and the music. The
preset dropdown replaces it with a worked example -- a talking avatar, a
three-shot scene, a two-hander, a reference shot. Each one compiles and runs as it stands,
so it is a starting point rather than a form. One undo step puts your own work back.
| Key | Does |
|---|
Delete | remove the selected blocks |
S | split them at the playhead |
Cmd/Ctrl+C · Cmd/Ctrl+V | copy the selection, paste it at the playhead keeping its spacing |
Cmd/Ctrl+D | duplicate in place at the playhead |
Cmd/Ctrl+Z | undo |
The playhead
The red line is where every Add button puts its block -- click the empty part of a track
to move it, or drag the scrubber. If it is standing inside a block there is no room, so
the new one goes on the end instead.
• Dragging a block or its edge
snaps to the playhead and to the edge of every other
block, on any track, within a few pixels -- so a cue can start exactly where a shot does.
•
S cuts the selected blocks in two at the playhead. The second half keeps the prose
and drops any attached file, so the same picture is never in the prompt twice.
• Zooming with
+ /
- recentres the view on it.
The clip settings, left to right
| Field | What it does |
|---|
duration | Length of the whole piece, in frames. Type it here; the arrows step 17 at a time so you stay on H3's lattice. Zero or empty means the clip follows its content. |
= ... s | The same length in seconds. Read-only -- see below. |
frame rate | Always 24. H3 has no other rate, so this is shown, never chosen. |
width / height | Output resolution, in multiples of 32. Mirrors of the node's own widgets. |
resize | How reference images are fitted. match scales them to the output size; max keeps them larger, which holds a face or a logo together better and costs more time. |
renders N frames | What will actually be generated, after rounding up to 17n+5. If it differs from what you typed, the rounding is shown. |
The three tabs
The panel under the toolbar shows one of three things, and remembers which one across a
reload:
TIMELINE (the tracks and the selected block's fields),
CAST (one card per
person, with the count on the tab) and
GLOBAL (the two clip-wide prompt boxes).
The segment panel, under the timeline
Select a block first -- with nothing selected there is nothing to edit and the fields are
not on screen. Everything here edits that block.
| Field | What it does |
|---|
SEGMENT PROMPT | What happens in this block. On MAIN it becomes the shot's sentence; on AUDIO the sound; on CAMERA a note added to the move. |
start / end / length | The block's span in frames. Editing end moves the right edge and leaves the start alone -- the same edit as dragging the right grip. |
line / faces / how / language | MAIN blocks only. One row per spoken line, + line for another -- see below. |
camera | CAMERA blocks only. The move, from the vocabulary below. |
describes / used as / keep file | Blocks carrying a file only. See the next section. |
detach media | Removes the file, keeps the block and its prose. |
The
GLOBAL tab holds the two fields set once for the whole piece:
| Field | What it does |
|---|
GLOBAL PROMPT | Style and scene constants for the whole clip. Compiles into the opening of Shot 1. |
GLOBAL MUSIC | Score only the audience hears, as instrumentation, tempo and dynamics -- not mood words. Empty compiles to non_diegetic_music: N/A. |
Why seconds are read-only. Every greyed
= ... s box is a reading, not an input. A
second is 24 frames wide, so a block typed as
1.08 came back as 26 frames and was shown
as
1.08 again -- the number actually set was never on screen. The clip is written in
seconds and cut in frames, and only one of those can be the field you edit.
Dialogue
H3 makes the voice and the picture in one pass, and the guide's form for it is exact:
The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>
The editor writes it for you, and splits it the way the work actually splits: **who the
people are
is written once in the CAST tab, and a block only says who talks and what
they say**.
CAST is one card per person:
| Field | What it does |
|---|
name them | A short name, yours, so the faces on a dialogue row are readable. |
from | Which file on the timeline this person is drawn from. That binding is what makes a face and a voice one person; the card then shows the <Subject n> badge the prompt will use. |
keep them | How much of the person survives, compiled as subject_retention. Not the block's keep file: the photo may be fully_preserved while the face taken out of it is an attribute_transfer onto somebody else. |
| what they look like | For a card with a file. Becomes their line in subject_definitions. |
| how they sound | Age, gender, pitch, timbre, accent, on screen or off. H3 fixes the voice from this, so an empty one is a voice nobody chose and the linter says so. |
+ character adds a card;
they speak switches dialogue off for the whole clip --
every row and every
<d> at once, cards kept. Describing the same speaker two different
ways in two shots used to be possible; to the model that reads as two people wearing one
label.
On the block:| Field | What it does |
|---|
line | The words themselves, sent verbatim -- never translated, punctuation kept. |
| faces | Who says it: click a face from the cast. Two lit on one row is the guide's (S1,S2) -- the same words spoken by both at the same instant. |
how | How it is performed. Becomes the verb: says, whispers, shouts, answers -- free text, used as written. |
language | Names the language of the words; it does not translate them. |
+ line adds another row, so one block can hold a conversation: a line each, spoken in
turn, compiled as one
<d> apiece.
× removes a row.
Three readings, kept apart on purpose: a
chorus is one row with two faces, a
conversation is two rows, and an
argument -- overlapping speech with no agreed
words -- is neither. Write that one in the segment prompt and put the sound in an AUDIO
cue; there is nothing for H3 to quote.
Attached files:
used as,
describes and
keep file
used as -- what the file is
for. It decides the task type the summary opens with, and
the guide wants every relationship named:
| Used as | Task type it produces |
|---|
reference | reference generation -- guidance for a character, scene, style or camera move |
first frame / keyframe / last frame | keyframe completion -- the image is a concrete frame of the target video, and retention_analysis says which |
continue from | video continuation |
edit | video editing |
An attached audio adds
audio reuse or
audio reference depending on its
keep. Several
at once combine:
[keyframe completion + video continuation + audio reuse].
A segment holding a real file gets two extra fields, and they are what switch the prompt
into H3's full-reference format -- six sections instead of three.
describes -- what the file is, and what must stay the same. One noun phrase, written
for someone who cannot see it. It becomes that file's line in
subject_definitions:
<Picture 1> is the raccoon, grey and black with a striped tail and a masked face.
Leave it empty and the linter warns: an unnamed reference is one H3 has to guess at.
keep file -- how much of the file survives into the video. One per file, always; it
also sits on the block itself, bottom right. A cast card drawn
from this file carries its
own
keep them for the person, which may differ. The sentence it produces lands in
retention_analysis, as
<Picture 1> (appears in [Shot 2]): fully_preserved - the raccoon, ...| Value | Means |
|---|
fully_preserved | copy it -- same subject, same look, unchanged |
partially_preserved | keep the subject, let pose, angle or lighting change |
attribute_transfer | take one trait -- a face, a colour, a texture -- onto something else |
weak_reference | loose inspiration only: style, palette, grade, nothing literal |
These are H3's own words, not ours. It reads them as instructions, so a wrong one is worse
than a vague
describes:
fully_preserved on a style reference asks the model to reproduce
the whole frame.
Subjects live on cast cards. Point a card's
from at a file and the person becomes a
<Subject n> of their own, tracked apart from the picture they came from:
<Subject 1> is the man's face, from <Picture 2>.
That separation is what a face swap needs. The picture stays a
weak_reference -- you do
not want the whole frame back -- while the card's
keep them is
attribute_transfer onto
the person in another shot. The shot then mentions
<Subject 1> rather than
<Picture 2>,
because naming both asks for two different things at once.
Camera modes
Each mode contributes one sentence to the shot it covers. H3 reads prose, not enum values,
so this is the exact wording it receives:
| Mode | Sentence sent to the model |
|---|
static | The camera holds a static shot. |
dolly_in | The camera pushes in with small amplitude at slow speed. |
dolly_out | The camera pulls out with small amplitude at slow speed. |
pan_left | The camera pans left. |
pan_right | The camera pans right. |
tilt_up | The camera tilts up. |
tilt_down | The camera tilts down. |
orbit | The camera moves in an arc around the subject. |
handheld | The camera shakes slightly. |
crash_zoom | The camera zooms in with large amplitude at fast speed. |
A note typed into a camera segment is appended after the sentence, so
dolly_in +
"closing on the apple" becomes *The camera pushes in with small amplitude at slow speed.
closing on the apple* -- so write the note as a continuation, not as a sentence of its own.
Camera work is its own block rather than folded into the shot line, because a move can
straddle a cut -- and merging them would silently pick a side.