MiniMax H3
MiniMax H3 is an omni-modal generative video model with support for unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.
Note
MiniMax H3 is massive! 33B transformer and 32B text-encoder (qwen3-vl) combine to 134GB of model weights in bfloat16!
Important
See offloading and quantization sections below before attempting to load MiniMax-H3
Important
Video support requires ffmpeg to be installed and available in the PATH
Requirements
Due to size, MiniMax-H3 may not run properly on systems with less than 64GB of RAM
Depending on quantization and offloading methods, MiniMax-H3 can run comfortably on GPUs with as little as 16GB of VRAM
Quantization
Tip
SD.Next provides access to SDNQ pre-quantized weights in uint4 precision which reduces the model size to 47GB
To use pre-quantized weights, simply select them from model selection dropdown
If you choose to use non-quantized weights, you should enable desired quantization
For more information see Quantization Wiki and SDNQ Wiki
Offloading
Balanced offloading
SD.Next default offload method is balanced offloading which is the best starting point for most models. However, MiniMax-H3 is so massive that it may require more aggressive offloading to fit into your GPU memory
with balanced offloading, MiniMax-H3 requires a GPU with at least 24GB of VRAM.
Group offloading
recommended offload method for MiniMax-H3 is group offloading which allows MiniMax-H3 to run on a GPU with as little as 16GB of VRAM.
Tip
group offloading is a more aggressive offload method which does result in some performance loss, but it is the only way to run MiniMax-H3 on GPUs with less than 24GB of VRAM.
for more information see Offloading Wiki
Variants
MiniMax H3 has two separate variants which use different transformers weights:
- base: used for text-to-video (t2va), image-to-video (i2va) and first-last-frame-to-video (fl2va) video generation
- ref: used for reference-based (ref2va) video generation
Notes:
- both base and ref variants use the same text-encoder weights, but have different transformer weights
- both base and ref variants are available in a pruned variant which is slightly smaller as parts of the transformer model are pruned/removed
depending on which variant you load, you will have access to different workflows
Base workflows
available if base model is loaded - prompting guide - t2va: text-to-video is used when there are no input images provided - i2va: image-to-video is used when a single input image is provided - fl2va: first-last-frame-to-video is used when two input images are provided
Ref workflows
available if ref model is loaded - prompting guide - ref2va: reference-to-video
in ref2va workflow, you can upload any combination of: - 9 reference images - 3 reference videos - 3 reference audio files
total number of reference media must be up to 12 files
lengths of reference videos and audio files must be up to duration of generated video
Image workflows
everything as above except that MiniMax-H3 can also be used as a regular text-to-image/image-to-image model, not limited to Video generation
to use MiniMax-H3 as a t2i/i2i model, simply select it from the networks -> reference and use as any other model
LoRA
SD.Next includes native LoRA support for MiniMax-H3
Warning
LoRA variant must match desired model variant e.g. fl2va LoRA can only be used with base model, and ref2va LoRA can only be used with ref model LoRA will load, but will have severe quality issues if used with the wrong model variant
Warning
Pruned variant may not be compatible with some LoRAs if they were trained to include pruned layers
Tip
Issues with native LoRA support should be reported on GitHub or Discord
Diffusers built-in LoRA support remains available as a fallback in Settings -> LoRA -> LoRA load using Diffusers method
Turbo LoRA
Tip
MiniMax-H3 can be used with Turbo LoRA to reduce steps from base 30-50 down to 4-8 steps only! Link to recommended Turbo LoRA for base and turbo variants
Shift
MiniMax-H3 runs two schedules per request, one for video and one for audio, each with its own exponential sigma shift. The model ships with video 12 and audio 3, which are the defaults of the MiniMax video shift and MiniMax audio shift sliders. Over the API, sampler_shift sets the video shift and audio_shift the audio shift, with -1 keeping the shipped value. The applied values are recorded in the video metadata as Video shift and Audio shift.
Distilled LoRAs run at the shift they were trained with. The settings below are taken from each author's model card or spec table where one exists.
| LoRA | Steps | Video shift | Audio shift | Strength |
|---|---|---|---|---|
| none (base model) | 30-50 | 12 | 3 | n/a |
| larryvrh v4, v1 | 4-8 | 12 | 3 | 1.0 |
| lightx2v 4step v0.1 (544p) | 4 | 12 | 3 | 0.0625 |
| lightx2v 4step v1.1, v1.2 (768p) | 4 | 6 | 3 | 1.0 |
| lightx2v 8step v1.0 (768p) | 8 | 6 | 3 | 1.0 |
| alibaba-pai PDD 8-step | 8, pinned | 12, pinned | 3, pinned | 1.0 |
| TaoMate 3-step | 3 | 12 | 3 | 1.0 |
| VDN DMD 8step | 8 | 12 | 3 | 1.0 |
| Silveroxides dareties, FastH3 extract, SLA | 4 | 12 | 3 | 1.0 |
Note
PDD LoRAs pin the step count and both shifts to their training values; the override is logged and recorded in the metadata
- larryvrh suggests strength 0.8 to 0.95 when the output looks over-sharp and 1.05 to 1.2 when it smears
- The lightx2v 768p files were trained at 1344x768
- One published workflow (drbaph) runs audio shift 4 to 6 with video 12; no other author documents an audio value other than 3
- Guidance has no effect: the model is guidance-distilled, so the guidance fields are inert on every LoRA
Legal
Due to pending lawsuit, MiniMax-H3 license comes with a disclaimer