comfyui minimax h3
Turn text, image, or reference footage into clips with synchronized sound through the comfyui minimax h3 node set.
AI Video Prompt Generator
10s

Feedback

AI Ad Video Example

Loading...

comfyui minimax h3

Generate 2K 24fps clips with matched stereo audio inside ComfyUI using the comfyui minimax h3 workflow for text, image, or footage references.

All Tools

Discover our comprehensive AI-powered animation toolkit

Benefits of running the comfyui minimax h3 workflow in ComfyUI

By integrating MiniMax’s omni-modal model as open weights, the comfyui minimax h3 workflow lets ComfyUI process text, images, video, and audio together. It produces clips with matching stereo sound — dialogue, SFX, and score — in one step, delivering up to 2K at 24fps for around 15 seconds while keeping every node parameter adjustable.

  • Built-in audio that stays in sync
    A single MP4 contains dialogue, effects, and soundtrack that stay perfectly in sync because the comfyui minimax h3 node produces them alongside the moving image.
  • Full local control without API limits
    With open-weight access, you can dial in resolution, length, and diffusion settings for every comfyui minimax h3 video job, free from API quotas or remote queues.
  • Combine multiple reference types
    Feed the comfyui minimax h3 graph a character, style, motion, camera path, or voice from several reference clips and still get one coherent video generation.

A quick-start guide for the comfyui minimax h3 workflow

Begin producing open-weight clips with embedded stereo sound by following these three steps for the comfyui minimax h3 workflow.

Core capabilities of the comfyui minimax h3 workflow

The comfyui minimax h3 workflow combines three built-in ComfyUI templates, open-weight multimodal generation, synchronized audio, reference-aware controls, and optional Sage Attention acceleration to create a full local video production suite.

Three ready-made ComfyUI graph templates

Your comfyui minimax h3 library includes ready-to-run examples for text-to-video, image-to-video, and reference-to-video, one per generation mode, so you can start immediately.

One context for text, image, video, and audio

Because the comfyui minimax h3 model interprets text, stills, footage, and sound inside the same context, you can blend all of those inputs into a single coherent generation.

Generate from a rich set of references

You can pin a character, visual style, movement, camera movement, or vocal timbre using up to 9 images, 3 videos, and 3 audio clips through the comfyui minimax h3 R2V node.

Sharp text rendering and reliable brand marks

The comfyui minimax h3 model reproduces logos and spelled-out copy with clarity while following natural-language instructions about relationships between references.

Approximately 2x faster generation with SageAttention

Insert the Patch Sage Attention KJ node into the comfyui minimax h3 graph to almost double rendering speed with barely any quality trade-off.

Resolution and duration helper

The comfyui minimax h3 resolution selector calculates width and height from aspect ratio and megapixels, snapping to the model's 32-pixel grid and 17-frame chunk duration at 24fps.

FAQ

Frequently asked questions about the comfyui minimax h3 workflow

Straightforward answers about running the comfyui minimax h3 workflow in ComfyUI, covering output, modes, audio, speedups, and setup.

1

What exactly does the comfyui minimax h3 workflow do?

It’s ComfyUI’s built-in support for MiniMax H3, an open-weight, omni-modal generation model. Through the comfyui minimax h3 workflow, text, stills, video, and audio references are converted into a clip with synchronized stereo sound in one pass.

2

What resolution and frame rate can I expect?

The comfyui minimax h3 workflow can render up to 15-second 2K videos at 24fps. The native canvas has a 768px short edge, a maximum of 768x1344, and resolution rounded to multiples of 32.

3

Which generation modes come with the workflow?

The comfyui minimax h3 template set includes three presets: text-to-video, image-to-video with optional start/end frame settings, and reference-to-video for preserving character, style, motion, camera, or voice.

4

Can the workflow produce audio along with video?

Yes. Voice, sound effects, and music are created together with the picture by the comfyui minimax h3 model, then delivered as one MP4 with perfectly synced stereo audio.

5

What do I need to run it for the first time?

Upgrade ComfyUI to 0.30.0+, go to Template Library > Video, select a comfyui minimax h3 workflow, and follow the prompt to fetch models from the Hugging Face Comfy-Org/MiniMax-H3 repo.

6

Is there a way to make generation faster?

Yes. Install SageAttention and KJNodes, then connect a Patch Sage Attention KJ node between UNETLoader and BasicGuider inside the comfyui minimax h3 graph to accelerate rendering by about 2x.

Kick off your next project with the comfyui minimax h3 workflow

Generate locally with MiniMax H3 in ComfyUI, complete with open weights, native audio, and fine-grained controls. Choose a comfyui minimax h3 preset for text, image, or reference video creation and start right away.