Experimental
Practical estimate8–12 GB VRAM
Possible only through low-VRAM workflows, quantized weights, reduced output sizes, and heavy CPU offloading.
MiniMax H3 Hardware Guide · Updated August 8, 2026
MiniMax H3 can run locally, but there is no single official minimum VRAM figure. The practical requirement depends on compact INT8 files, a GGUF build, or the original BF16 checkpoints.
Quick answer: 16 GB is workable with quantization and offloading; 24–32 GB is more practical; 8–12 GB is experimental.
Check your hardware before downloading 42.5 GB or more.
Hardware checker
Choose your hardware and intended workflow. We’ll estimate practical compatibility using current compact ComfyUI files and published deployment results.
This setup is suitable for a compact ComfyUI workflow with CPU offloading. Start with preview resolution and short clips before increasing the workload.
Practical estimate only — not an official MiniMax hardware certification.
Quick answer
There is no single official minimum VRAM requirement. The practical target changes with the weight format, workflow, resolution, and amount of CPU offloading.
For most desktop users, 16 GB VRAM with 64 GB of system RAM is a workable starting point only when using quantized weights and CPU offloading. A 24 GB GPU is more practical, while 32 GB provides better headroom for reference workflows and larger outputs.
Routes using 8–12 GB VRAM are experimental. They may require GGUF weights, lower resolutions, aggressive offloading, and significantly longer generation times.
“Can load the model” and “comfortable to use” are not the same requirement.
Experimental
Practical estimatePossible only through low-VRAM workflows, quantized weights, reduced output sizes, and heavy CPU offloading.
Workable
Practical estimateA practical starting point for compact INT8 or GGUF workflows, preferably with at least 64 GB of system RAM.
Practical
Practical estimateBetter suited to local generation, longer clips, and higher working resolutions while still relying on memory management.
Recommended
Practical estimateProvides more headroom for Ref2VA, larger outputs, and fewer aggressive offloading compromises.
Deployment paths
The hardware requirement changes dramatically depending on the files and runtime you choose.
| Deployment path | Best for | Local GPU | File size | Key trait |
|---|---|---|---|---|
| Browser generator | Quick testing and creation | Not required | 0 GB | Fastest start |
| ComfyUI INT8 | Local creators | 16–32 GB practical | 42.5–63.5 GB | Native nodes and workflows |
| Community GGUF | Low-VRAM experiments | 8–16 GB | Depends on quantization | Slower or quality trade-offs |
| Original BF16 / SGLang | Professional deployment | Multi-GPU | Large | Server and production use |
Published evidence
Only first-party specifications and configurations with published results are listed here. Unpublished fields stay explicitly marked instead of being estimated.
| GPU | VRAM | System RAM | Workflow | Resolution | Clip | Steps | Result / time | Evidence |
|---|---|---|---|---|---|---|---|---|
| Not published | Not published | Not published | Native ComfyUI T2V / I2V / Ref2VA | 768p base | 4–15 sec | Not published | Available in ComfyUI 0.30+ | Official spec ComfyUI H3 documentation |
| 2× RTX 5090 | 2× 32 GB | 377 GiB detected | Lossless BF16 / SGLang | 1344×768 | 5 sec | 50 | 559.67 sec generation | Verified deployment SGLang deployment guide |
Community benchmark timings are omitted until the hardware, workflow, settings, output, source, and test date can all be reviewed.
The current community license excludes the United States, European Union, United Kingdom, and South Korea from its standard open-weight territory. Organizations in those regions can apply for authorization; hosted API availability follows a separate policy.
Model files
Storage size is separate from runtime VRAM: a 21 GB model file does not mean that 21 GB of VRAM is sufficient.
| Component | File | Size | Destination |
|---|---|---|---|
| FL2VA diffusion | minimax_h3_fl2va_pruned_int8_convrot.safetensors | 21 GB | models/diffusion_models/ |
| Text encoder | qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors | 15.7 GB | models/text_encoders/ |
| Video VAE | minimax_h3_video_vae_fp16.safetensors | 5.21 GB | models/vae/ |
| Audio VAE | minimax_h3_audio_vae_fp32.safetensors | 0.61 GB | models/vae/ |
| Ref2VA diffusion | minimax_h3_ref2va_pruned_int8_convrot.safetensors | 21 GB | models/diffusion_models/ |
Storage breakdown
Folder layout
Place files in the corresponding ComfyUI model directories. Keep at least 55–60 GB free for T2V/I2V, or 80 GB when adding Ref2VA, so the runtime has space for caches and generated video.
ComfyUI/models/ ├── diffusion_models/ (21–42 GB) ├── text_encoders/ (15.7 GB) └── vae/ (5.82 GB)
Platforms
NVIDIA CUDA is the most practical current consumer path. System memory becomes increasingly important as model components move out of VRAM.
RAM by workflow
64 GB is a more practical compact-workflow starting point than 32 GB. 96–128 GB gives CPU offloading more headroom.
Best current consumer compatibility. Quantized files and CPU offloading can make lower-VRAM cards possible, with workflow-dependent speed and stability.
SGLang validation covers AMD Instinct MI300X and MI355X. Do not interpret that as confirmed support for consumer Radeon cards.
The current native ComfyUI and published deployment paths should not be presented as a verified Apple Silicon workflow.
Local output
The downloadable H3-Base checkpoint generates 768p output. MiniMax’s documented full 2K workflow adds H3-Context-IR and H3-Regenerate-2K around the base generation process.
Downloadable base checkpoint
Context IR + 2K regeneration
Hosted output specification
ComfyUI
Native templates connect model loading, text conditioning, sampling, VAE decoding, and video export into one graph.
FL2VA INT8
Qwen3-VL
Video + audio
MP4 output
Output examples
These are MiniMax H3 creative examples from the existing product media library. They demonstrate output, not local GPU performance.
A cinematic fashion portrait beside a classic car at night. Warm firelight moves across glossy black leather while the camera makes a restrained push toward the subject.
A lone swordsman crosses a mist-covered mountain pass at dawn. Layered peaks drift through pale clouds while the camera follows in a slow, cinematic tracking shot.
Example generated online with MiniMax H3. This demonstrates model output, not local GPU performance.
Try before setup
You do not need to download 42.5 GB of model files just to test a prompt. Start in your browser, then decide whether a local workflow is worth the setup.
No ComfyUI installation · No local GPU required
FAQ
Practical answers for local ComfyUI workflows, low-VRAM experiments, storage, and hosted generation.
MiniMax has not published one universal minimum VRAM figure. The requirement changes by runtime and weight format. For compact ComfyUI workflows, 16 GB VRAM and 64 GB of system RAM is a workable starting point, while 24–32 GB VRAM provides a more practical local experience.
Experimental low-VRAM workflows may run on 8 GB cards by using quantized weights, lower resolutions, and aggressive CPU offloading. This is not a recommended default and may be very slow.
Sixteen gigabytes can be workable for compact INT8 or GGUF workflows with CPU offloading. Start with shorter clips and preview resolution, and use at least 64 GB of system RAM when possible.
The compact T2V/I2V ComfyUI file set is about 42.5 GB. Adding the separate Ref2VA diffusion model brings the combined model files to about 63.5 GB. Additional free space is needed for caches and generated videos.
Thirty-two gigabytes is an experimental starting point. Sixty-four gigabytes is more practical for quantized workflows, while 96–128 GB provides better headroom for CPU offloading.
A stable Apple Silicon workflow should not currently be presented as officially verified. Published ComfyUI and server deployment paths focus primarily on CUDA and validated datacenter hardware.
The downloadable H3-Base checkpoint produces 768p output. The documented 2K workflow adds H3-Regenerate-2K and hosted orchestration components around the base generation.
The official MiniMax H3 workflow guide currently requires ComfyUI 0.30.0 or later.
Yes. A hosted MiniMax H3 generator runs the model remotely, so you only need a modern browser and an internet connection.
Sources and update policy
MiniMax H3’s local ecosystem is evolving quickly. File sizes, runtime support, quantized builds, and hardware behavior may change. This page separates official specifications from practical deployment estimates.
Last reviewed: August 8, 2026
Skip the local setup
Keep your prompt, generation mode, image, aspect ratio, duration, and MiniMax H3 model selection when you continue to the online generator.