Skip to content

MiniMax H3

MiniMax H3 is an omni-modal generative video model with support for unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.

Note

MiniMax H3 is massive! 33B transformer and 32B text-encoder (qwen3-vl) combine to 134GB of model weights in bfloat16!

Important

See offloading and quantization sections below before attempting to load MiniMax-H3

Important

Video support requires ffmpeg to be installed and available in the PATH

Requirements

Due to size, MiniMax-H3 may not run properly on systems with less than 64GB of RAM
Depending on quantization and offloading methods, MiniMax-H3 can run comfortably on GPUs with as little as 16GB of VRAM

Quantization

Tip

SD.Next provides access to SDNQ pre-quantized weights in uint4 precision which reduces the model size to 47GB

To use pre-quantized weights, simply select them from model selection dropdown

If you choose to use non-quantized weights, you should enable desired quantization

For more information see Quantization Wiki and SDNQ Wiki

Offloading

Balanced offloading

SD.Next default offload method is balanced offloading which is the best starting point for most models. However, MiniMax-H3 is so massive that it may require more aggressive offloading to fit into your GPU memory

with balanced offloading, MiniMax-H3 requires a GPU with at least 24GB of VRAM.

Group offloading

recommended offload method for MiniMax-H3 is group offloading which allows MiniMax-H3 to run on a GPU with as little as 16GB of VRAM.

Tip

group offloading is a more aggressive offload method which does result in some performance loss, but it is the only way to run MiniMax-H3 on GPUs with less than 24GB of VRAM.

for more information see Offloading Wiki

Variants

MiniMax H3 has two separate variants which use different transformers weights: - base: used for text-to-video (t2va), image-to-video (i2va) and first-last-frame-to-video (fl2va) video generation
- ref: used for reference-based (ref2va) video generation

Notes: - both base and ref variants use the same text-encoder weights, but have different transformer weights
- both base and ref variants are available in a pruned variant which is slightly smaller as parts of the transformer model are pruned/removed

depending on which variant you load, you will have access to different workflows

Base workflows

available if base model is loaded - prompting guide - t2va: text-to-video is used when there are no input images provided - i2va: image-to-video is used when a single input image is provided - fl2va: first-last-frame-to-video is used when two input images are provided

Ref workflows

available if ref model is loaded - prompting guide - ref2va: reference-to-video

in ref2va workflow, you can upload any combination of: - 9 reference images - 3 reference videos - 3 reference audio files

total number of reference media must be up to 12 files
lengths of reference videos and audio files must be up to duration of generated video

Image workflows

everything as above except that MiniMax-H3 can also be used as a regular text-to-image/image-to-image model, not limited to Video generation
to use MiniMax-H3 as a t2i/i2i model, simply select it from the networks -> reference and use as any other model

LoRA

SD.Next includes native LoRA support for MiniMax-H3

Warning

LoRA variant must match desired model variant e.g. fl2va LoRA can only be used with base model, and ref2va LoRA can only be used with ref model LoRA will load, but will have severe quality issues if used with the wrong model variant

Warning

Pruned variant may not be compatible with some LoRAs if they were trained to include pruned layers

Tip

Issues with native LoRA support should be reported on GitHub or Discord
Diffusers built-in LoRA support remains available as a fallback in Settings -> LoRA -> LoRA load using Diffusers method

Turbo LoRA

Tip

MiniMax-H3 can be used with Turbo LoRA to reduce steps from base 30-50 down to 4-8 steps only! Link to recommended Turbo LoRA for base and turbo variants

Shift

MiniMax-H3 runs two schedules per request, one for video and one for audio, each with its own exponential sigma shift. The model ships with video 12 and audio 3, which are the defaults of the MiniMax video shift and MiniMax audio shift sliders. Over the API, sampler_shift sets the video shift and audio_shift the audio shift, with -1 keeping the shipped value. The applied values are recorded in the video metadata as Video shift and Audio shift.

Distilled LoRAs run at the shift they were trained with. The settings below are taken from each author's model card or spec table where one exists.

LoRA Steps Video shift Audio shift Strength
none (base model) 30-50 12 3 n/a
larryvrh v4, v1 4-8 12 3 1.0
lightx2v 4step v0.1 (544p) 4 12 3 0.0625
lightx2v 4step v1.1, v1.2 (768p) 4 6 3 1.0
lightx2v 8step v1.0 (768p) 8 6 3 1.0
alibaba-pai PDD 8-step 8, pinned 12, pinned 3, pinned 1.0
TaoMate 3-step 3 12 3 1.0
VDN DMD 8step 8 12 3 1.0
Silveroxides dareties, FastH3 extract, SLA 4 12 3 1.0

Note

PDD LoRAs pin the step count and both shifts to their training values; the override is logged and recorded in the metadata

  • larryvrh suggests strength 0.8 to 0.95 when the output looks over-sharp and 1.05 to 1.2 when it smears
  • The lightx2v 768p files were trained at 1344x768
  • One published workflow (drbaph) runs audio shift 4 to 6 with video 12; no other author documents an audio value other than 3
  • Guidance has no effect: the model is guidance-distilled, so the guidance fields are inert on every LoRA

Due to pending lawsuit, MiniMax-H3 license comes with a disclaimer