minimax_h3_director.md

MiniMax H3 Director

A timeline-based authoring node for ComfyUI's native MiniMax H3 models. It centralizes media management, ordering, trimming, per-reference prompts, and the global prompt into one workflow node, validates H3 constraints before execution, and routes everything to the installed native MiniMax H3 implementation — no duplicate backend logic.

Changelog (latest additions)

Since the last GitHub release (August 2026):

New features

Earlier additions

Bug fixes

Quick overview

Installation and graph setup

Install dependencies and restart ComfyUI:

pip install -r requirements.txt

Ensure your ComfyUI version includes native MiniMax H3 support. Add these two nodes from DaSiWa/MiniMax H3:

  1. MiniMax H3 Director — your timeline, references, and prompt editor.
  2. MiniMax H3 Director Guide — validation and routing to native H3 nodes.

Wire them like this:

┌──────────────────────────────────────┐
│ UNET Loader                          │     diffusion_models/*.safetensors
│ CLIP Loader                          │     text_encoders/qwen3vl_32b_minimax_h3_*.safetensors
│ VAE Loader (visual)                  │     vae/minimax_h3_video_vae_fp16.safetensors
│ VAE Loader (audio)                   │     vae/minimax_h3_audio_vae_fp32.safetensors
│                                      │     (only for REF2VA mode)
└───┬────────────┬──────────┬──────────┘
    │            │          │
    ▼            ▼          ▼
┌──────────────────────────────────────────────┐
│ DaSiWa MiniMax H3 Director                   │
│  - add/edit references                       │
│  - set trims, ordering, prompts              │
│  - emits structured "guide" dict             │
└────┬─────────────────────────────────────────┘
     │ guide
     ▼
┌──────────────────────────────────────────────┐
│ DaSiwa MiniMax H3 Director Guide             │
│  - validates director output                 │
│  - assembles final prompt                    │
│  - CALLS the native ComfyUI H3 nodes:        │
│      • MiniMaxH3ImageToVideo   (FL2VA)       │
│      • MiniMaxH3ReferenceToVideo (REF2VA)    │
│  - you NEVER wire those native nodes yourself│
└────┬─────────────────────────────────────────┘
     │ positive, latent
     ▼
┌──────────────────────────────────────────────┐
│ Standard ComfyUI sampling/decoding chain     │
│  - KSampler                                  │
│  - VAE Decode                                │
│  - Enhanced Video Combine / Image Save etc.  │
└──────────────────────────────────────────────┘

Connections detail:

Important: the Guide node replaces and wraps ComfyUI's native MiniMaxH3ImageToVideo and MiniMaxH3ReferenceToVideo nodes. You do not add or wire those native nodes yourself — the Guide calls them internally based on the chosen mode.

The Director has optional model sockets (fl2va_model, ref2va_model) for lazy loading: connect whichever model matches your active mode. The Guide refuses REF2VA without an audio VAE connected.

Modes at a glance

FL2VA (First/Last Frame to Video)

Use for text-only generation, single-image starting frames, or first+last frame interpolation.

Alignment instruction patterns:

For one endpoint (first frame only):

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: ...

For two endpoints (first and last frame):

How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the {duration}.00-second mark of the target video.

integrated_multimodal_description: ...

Replace {duration} with your Director duration setting (e.g. 8.00). Images map left-to-right by slot order: first image = Picture 1, second = Picture 2. FL2VA always uses integrated_multimodal_description, not detailed_description.

REF2VA (Reference to Video)

Use when you want the generated video to borrow identity, appearance, motion, composition, or sound from existing media.

When a video provides audio (via A/V+A or attached soundtrack), that audio becomes <Audio N> in the reference numbering. Embedded audio and separate soundtracks share the same time crop as the host video frames.

Image Inpaint (single image)

Use to produce a single refined/edited frame from exactly one image reference — a still, not a video.

UI walkthrough: controls and buttons

Open the node and read top-to-bottom.

Toolbar row

Timeline area

The main workspace has two horizontal lanes stacked vertically.

Lane selection

Adding media

Three ways:

  1. + buttons: each empty slot shows a "+"; click to open a file picker filtered for that lane type.
  2. Drag-and-drop: drag files from your OS directly onto the desired lane.
  3. Paste: select a lane, focus the node, press Ctrl+V with images/videos/audio on your clipboard.

Each uploaded file is stored in ComfyUI's input/ directory and linked by relative path.

Clip tiles

Each media item renders as a labeled tile inside its lane.

Common elements:

Video-specific elements:

Audio-specific elements:

Images:

Removing/disabling items

Prompt editors

Below the timeline is a unified prompt-builder panel whose layout depends on the active mode. Both editors have resizable text areas with drag-handle bars at the bottom; heights persist in the workflow JSON.

FL2VA / I2VA / L2VA / T2VA builder

Three labeled text areas:

Alignment instruction lines (for I2VA/FL2VA/L2VA) are generated automatically based on mode and duration; you do not type them manually.

REF2VA builder

Six labeled text areas matching the official full-reference format. Section headers (subject_definitions: etc.) are appended automatically by the backend; you write only the content:

Helper buttons above the fields:

Limits and validation

MiniMax H3 enforces hard caps; the Director checks these before sending data downstream:

Violations appear as red status messages inside the node. Fix them before queuing.

How processing flows: upstream and downstream

Understanding the data path makes wiring and debugging easier.

Upstream inputs (what feeds into the Director)

Inside the Director

On queue, the Director executes this sequence:

  1. Reads current mode (FL2VA-family, REF2VA, or Image Inpaint).
  2. Iterates over all enabled timeline items in slot order.
    • For FL2VA: keeps only image items (max 2); discards others temporarily.
    • For REF2VA: processes images, videos, and audio respecting slot limits.
    • For Image Inpaint: accepts exactly one image reference; video/audio items are a hard error.
  3. Loads each asset:
    • Images → resized tensors.
    • Videos → decoded to frame_rate fps frame batches (default 24), cropped according to trim_start/trim_end.
    • Audio → decoded waveforms, cropped identically.
    • For videos in A or V+A mode → embedded audio is extracted using the same crop window.
    • Attached soundtracks → loaded and cropped using the host video's trim range.
  4. Builds a structured guide dictionary containing:
    • Mode flag, dimensions, duration, and frame rate.
    • Ordered lists of images, videos, audios with metadata.
    • Endpoint frames (FL2VA).
    • Reference maps keyed as ref_image_N, ref_video_N, ref_audio_N, ref_video_audio_N.
  5. Reads the mode-specific prompt-builder state:
    • FL2VA/I2VA/L2VA/T2VA: uses integrated_multimodal_description, overall_soundscape, non_diegetic_music fields.
    • REF2VA: uses the six free-text sections (subject_definitions, summary, retention_analysis, detailed_description, overall_soundscape, non_diegetic_music).
    • Alignment lines (I2VA/FL2VA/L2VA) are injected automatically based on mode and duration.

This guide object and builder_state are passed out to the Guide node.

The Guide node

The Guide is a thin adapter between your authored timeline and ComfyUI's native H3 nodes:

  1. Validates the incoming guide:
    • Confirms mode consistency.
    • Checks that required models/CLIP/VAEs are connected.
    • For REF2VA, ensures an audio VAE exists.
  2. Assembles the final prompt via the prompt-builder helper:
    • Reads builder_state from the Director.
    • For FL2VA/I2VA/L2VA/T2VA: injects alignment lines (when applicable), combines integrated_multimodal_description + overall_soundscape + non_diegetic_music into the canonical format.
    • For REF2VA: wraps the six user-written sections with their standard headers (subject_definitions:, summary:, etc.). Legacy v1 structured-builder data is merged automatically if present.
    • Writes the result as resolved_prompt.
  3. Routes to the appropriate native node:
    • FL2VA / I2VA / L2VA / T2VA → calls MiniMaxH3ImageToVideo with endpoint frames and prompt.
    • Image Inpaint → calls MiniMaxH3ImageToVideo with the single image as first frame, last_frame = None, and a fixed 5-frame length.
    • REF2VA → calls MiniMaxH3ReferenceToVideo with all reference maps and prompt.
  4. Emits standard ComfyUI outputs:
    • positive (conditioning)
    • latent (image batch — a one-frame batch for Image Inpaint; extract the result with Get Image from Batch)
    • These feed downstream samplers and decoders exactly like any other H3 workflow.

You never call the native MiniMax H3 nodes directly when using Director+Guide; the Guide abstracts that away.

Model chain: patching & preview

The Director's fl2va_model and ref2va_model inputs are plain MODEL sockets, so any node that outputs MODEL can sit upstream of the Director — a LoRA loader, the MiniMax H3 Cache patcher, Patch Comfy Kitchen Attention, or a KJ ModelPreviewOverrideKJ. This is the same forward chain you already use:

Checkpoint.MODEL → LoRALoader.MODEL → [KJ.model → KJ.MODEL] → Director.fl2va_model
Checkpoint.CLIP  → LoRALoader.clip  → Director.clip

Three rules keep the chain valid:

  1. Forward chain only — never a loop. A patcher's output feeds into the Director's model input; it must never come back out of the Director. The Director is a terminal media node (it emits frame_rate, duration, images, never MODEL), and a wire from the Director back into its own model input would be a graph dependency_cycle, which ComfyUI's validation rejects.
  2. One loader per model, in mode order. The Director picks ref2va_model for REF2VA and fl2va_model otherwise; the active input must be connected (the unconnected twin may stay empty).
  3. Type-safe wires. ComfyUI only lets you connect type-compatible sockets, so LoRA.MODEL → Director.fl2va_model is legal but LoRA.MODEL → Director.clip is not. No name or type resolution happens at runtime — the socket you plugged in arrives as the keyword-argument named for that socket.

Dense prompting guide

MiniMax H3 expects structured natural language rather than keyword piles. Use the official field names and shot/timestamp conventions.

FL2VA prompt structure

Text-only (no endpoint images):

integrated_multimodal_description: [Shot 1] ... [Shot 2] At 00:04.500, ...

overall_soundscape: ...

non_diegetic_music: ...

With endpoint images: prepend the alignment instruction line shown earlier, blank line, then the three fields above.

REF2VA prompt structure

Use the six-section format. Define assets once and reuse labels consistently.

Label rules:

Template:

subject_definitions:
<Subject 1> is the woman in <Picture 1>, with short dark hair and a red coat.
<Picture 1> is the opening-frame anchor for [Shot 1].
<Subject 2> is the walking motion taken from <Video 1>.
<Video 1> provides the camera path and pacing structure.
<Audio 1> is the voice-timbre reference for <Subject 1> (S1).

summary:
[reference generation + audio reference] Use <Subject 1> from <Picture 1>, the motion and pacing of <Video 1>, and the voice character of <Audio 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - identity and clothing remain consistent.
<Picture 1> ([Shot 1] first frame): fully_preserved - opening composition anchor.
<Subject 2> (motion transferred to <Subject 1>): attribute_transfer - walk rhythm is applied to <Subject 1>.
<Video 1> (pacing structure): weak_reference - general timing and camera rhythm are retained.
<Audio 1>: reference - timbre and delivery are followed without copying the signal.

detailed_description: [Shot 1] ... [Shot 2] At 00:04.500, ...
overall_soundscape: ...
non_diegetic_music: ...

Summary task-type prefixes (combine with +):

Retention markers:

Number labels by timeline order: images first (Pictures), then videos, then audio. Keep meanings stable everywhere.

Shots and timestamps

Camera motion vocabulary

Use as natural English inside shots:

Example: The camera pushes in with small amplitude at slow speed toward her hands.

Dialogue and special tokens

Practical workflow

  1. Choose FL2VA for endpoint/text work; REF2VA for multi-reference transfer.
  2. Add only references that contribute specific identity, motion, layout, or sound.
  3. Trim videos/audio to the strongest 2–15s segments; respect totals.
  4. Use the mode-specific prompt builder:
    • FL2VA/I2VA/L2VA/T2VA: fill the three guided fields; alignment lines appear automatically.
    • REF2VA: click Prefill Labels & Summary to scaffold your labels, then edit subject definitions and detailed description. Use Insert [Shot N] for clean shot markers.
  5. Click Preview Prompt to verify the exact output before queuing.
  6. Verify duration, aspect ratio, and motion match your references; queue through the Guide.

Official MiniMax H3 guides (canonical conventions):

下载此文件