MiniMax H3 Image Studio Review: ComfyUI Images, Reference Editing & Video Workflow
Explore MiniMax H3 Image Studio: ComfyUI setup, model downloads, GPU needs, image editing, NSFW limits and a practical reference-image-to-video workflow.
MiniMax H3 Image Studio makes the most sense if you already use H3 for video and want to prepare character, product or scene references in the same ComfyUI environment. It turns H3’s video-generation machinery into a still-image workflow, with text-to-image generation and reference-based editing.
The trade-off is substantial: this is a community extension, not an official MiniMax online image generator or a small standalone image model. You still load H3 and its text encoder. Its strongest argument is workflow reuse; flawless fine detail, exact lettering and automatic character consistency are not guaranteed.
Updated September 28, 2026 · Community extension v23.0.0
What it is—and who should use it
The project is maintained by astropuzzo. Its normal workflow samples a short sequence, decodes the frames and selects a still. That is different from a newly trained MiniMax image model. The current release adds paired four-step draft workflows and eight-step workflows around 1 MP; single-frame and alternative-decoder paths remain experimental. Project repository · Release history.
A sensible use case is a small set of approved reference assets: one clear character portrait, one full-body view and one environment plate. Another is trying a product colour or costume before generating a longer clip. If you only need an occasional illustration, installing a video-model stack solely for this extension is harder to justify.
Think of the output as a candidate that needs review. Check the face, hands, seams, logos and object count before animation: a small still-image error can become much more conspicuous when it moves.

Hardware, cost and online access
There is no universal minimum-VRAM figure for every workflow. The developer’s v23 tests used an RTX 4090 with 24 GB VRAM and 64 GB system RAM. A September 26 community report used an RTX 4070 with 12 GB VRAM and 48 GB RAM for reference editing. The latter demonstrates one working setup, not a promise that every 12 GB system or 4 MP workflow will run. Developer test conditions · RTX 4070 report.
For an existing H3 installation, reuse compatible weights and the encoder instead of downloading duplicates. Keep ample SSD space for large model files and outputs. Limited VRAM can move pressure to system RAM and storage; a successful load does not guarantee responsive iteration. Start around 1 MP with five frames and batch size one.
| Route | What you pay for | Decision point |
|---|---|---|
| Local ComfyUI | The node package has no subscription or per-image credit fee; hardware, electricity, storage and applicable licensing still cost money. | Best value when a suitable H3 machine is already available. |
| A rented ComfyUI GPU | GPU uptime, persistent storage and possibly transfers; provider prices vary. | Check permission to install this exact custom node and its weights before renting. |
| Comfy Cloud | Its current subscription/usage terms; H3 commercial rights are described as included. | Official H3 availability does not by itself confirm support for this community extension. |
Comfy Cloud is an online route for supported H3 workflows, not a verified one-click Image Studio service. For this exact extension, confirm custom-node support, model availability and licence coverage with the host. The repository supplies API-format workflow JSON for ComfyUI clients; it is not a separately sold MiniMax Image Studio API. Compare total cost per accepted reference image, including failed attempts and setup time.
Model downloads and installation
Update ComfyUI to 0.30.0 or later, then use the project’s current workflow rather than combining settings from old tutorials. ComfyUI H3 guide · Model files · Turbo adapters. Download only the diffusion model needed for your first workflow; generation and reference editing use different variants.
| Component | File or download location | ComfyUI folder |
|---|---|---|
| Generation model | minimax_h3_fl2va_pruned_int8_convrot.safetensors | models/diffusion_models/ |
| Reference-edit model | minimax_h3_ref2va_pruned_int8_convrot.safetensors | models/diffusion_models/ |
| Text/vision encoder | qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors | models/text_encoders/ |
| Video VAE | minimax_h3_video_vae_fp16.safetensors | models/vae/ |
| 8-step generation adapter | minimax_h3_fl2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors | models/loras/ |
| 8-step editing adapter | minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16.safetensors | models/loras/ |
In ComfyUI Manager, search for MiniMax H3 Image Studio. Alternatively, run the commands below from the directory containing your ComfyUI folder, then restart ComfyUI. If you use Portable, adjust the path to its nested ComfyUI directory. The node package does not download all model weights for you.
cd ComfyUI/custom_nodes
git clone https://github.com/astropuzzo/ComfyUI-MiniMax-H3-Image-Studio.git Generate and edit your first image
Use UI JSON for the visible canvas. API JSON under examples/api is for programmatic requests and can look like an empty or invalid canvas when imported as a UI workflow. Download the maintainer’s ready-to-use workflows:
- Generate an image: eight steps — H3_IMAGE_GENERATE.json
- Edit an image: eight steps — H3_IMAGE_EDIT.json
- Generate a quick draft: four steps — H3_IMAGE_DRAFT.json
- Edit a quick draft: four steps — H3_EDIT_DRAFT.json
- Two-reference editing: base workflow — H3_REFERENCE_EDIT.json
Start with one image
- Download H3_IMAGE_GENERATE.json and open it in ComfyUI. Select the FL2VA checkpoint, Qwen encoder, video VAE and matching 768p eight-step FL2VA adapter in their loader nodes.
- Set the aspect ratio for the intended video. Keep the approximately 1 MP preset and recommended five-frame profile for the first run. Leave the supplied sampling recipe intact.
- Write a still-image description: subject, framing, lighting and background. Generate a few candidates by changing the seed, then inspect each result at full size.
- Save the selected PNG and the workflow/seed. Keep the unrefined original even if you later sharpen or edit a copy.
Change a reference without locking it in place
- Open H3_IMAGE_EDIT.json, switch to the REF2VA checkpoint and its matching eight-step adapter, then select your source in Load Image.
- Describe the intended change and what must stay the same. Keep native reference transport initially; semantic-only transport is experimental.
- For two different reference roles, use H3_REFERENCE_EDIT.json and fill both Load Image nodes. Identify the inputs explicitly as <Picture 1> and <Picture 2>. This base workflow is not the same as the Turbo editing workflow.
- Keep source_fidelity near its supplied 0.6 starting value. It adjusts preservation wording, not denoise strength. Inspect both the requested edit and any unintended changes.
For the shipped eight-step 768p recipes, FL2VA uses video/audio shifts 6/3 and REF2VA uses 12/3. Match the checkpoint, adapter and recipe as a set. Changing only “8” to “4” is not the four-step workflow. The normal image-output path needs the video VAE, but not the audio VAE; the full video workflow may need both.

Use the image as a MiniMax H3 video reference
The useful connection is asset reuse, not a special file format: a finished PNG can become an H3 input just like a reference from another image model. There is no published evidence that using Image Studio alone guarantees better identity preservation than a good external reference. See our MiniMax H3 video review for the broader video stack.
- Choose the role first. For an opening composition that should animate, use H3 Image-to-Video/FL2VA. For identity, clothing or style that should inform a different scene, use Reference-to-Video/REF2VA.
- Export a clean single view. Avoid giving a three-view character sheet as an opening frame unless you actually want all three figures in the video. Split a sheet into separate reference views when appropriate.
- Open the corresponding native H3 video template, load the saved PNG and explicitly state the reference’s role. A still-image workflow does not become a video workflow simply by saving its output with a different extension.
- Use a short test clip and one simple action. Match the output framing and add the video/audio components required by that template. For the native 16:9 H3 video starting canvas, the Comfy guide specifies 1344×768; do not assume an experimental 4 MP still preset transfers directly.
- Review identity, motion, unwanted text and cropping. Only then add a longer shot, more references or a more complex camera move. Keep image and video seeds/settings separately for reproducibility.
Open the native H3 image-to-video and reference-to-video guide.
Prompts for reference images and animation
These are adaptable starting prompts, not measured test results. English examples preserve the
Character reference for a later shot
A single full-body reference image of an adult explorer in a teal raincoat and brown boots, standing neutrally against a plain light-gray background. The entire figure is visible, with relaxed hands and unobstructed facial features. Soft even studio lighting, realistic proportions, no text, no collage. A single neutral view is easier to inspect than a crowded character sheet. Generate additional views separately if you need more face detail.
Product edit from one source
Use <Picture 1> as the source product. Change only the ceramic cup colour from white to deep blue. Retain the handle shape, camera angle, wooden tabletop and soft side lighting. Output one product photograph, with no additional objects or lettering. Use an unbranded test object first. Verify shape and material as well as colour; an edit may rebuild reflections or proportions.
Two references with distinct roles
Use <Picture 1> for the adult character identity, hairstyle and clothing. Use <Picture 2> only for the standing body pose. Place the character against a plain gray background. Preserve the face and outfit from <Picture 1>; visibly adopt the pose from <Picture 2>. Output a single image, not a comparison sheet. Do not also request an unchanged pose from Picture 1 when Picture 2 is supposed to control the pose.
Animate the approved opening frame
Animate this opening image into a short continuous shot. The adult explorer slowly turns toward the window and lifts one hand. The camera stays still. Preserve the face, teal coat and room layout. Gentle fabric movement, natural timing, no cuts and no on-screen text. Use this in the video template, not in Image Studio. Start with a simple action so errors are easier to diagnose.
Image quality and real community experience
In TechnoEdge’s August 15 hands-on article, photographer Kazuhisa Nishikawa demonstrated text-to-image generation, clothing colour changes and two-reference editing. He found the capabilities useful but described weaker detail than dedicated image models such as Krea 2 and Z-Image. His older workflow/settings should not be treated as the current v23 default. Read the hands-on article.
The September 26 report by Kamimoto shows anime character sheets and changed scenes from two references on an RTX 4070. The author also notes soft full-body details and imperfect style reproduction. This supports character-preparation experiments, not a general guarantee of exact identity or production-ready model sheets. View the community examples.
The developer’s v23 comparison used fixed prompts and seeds, retaining defects. It reports successful wardrobe and pose changes but imperfect preservation; more frames did not consistently help. A readable four-letter sign is not a typography benchmark. For Japanese or Traditional Chinese labels, packaging text and small lettering, inspect every character and consider adding final typography in a conventional editor. Test report and comparison sheets.

| Developer’s example | Base 20 steps | Turbo 8 steps | Turbo 4 steps |
|---|---|---|---|
| Portrait, 1024×1024, seed 105 | 10.87 s | 7.31 s | 4.41 s |
| Jacket edit, 1440×1440, seed 205 | 25.04 s | 13.88 s | 13.26 s |
These are individual RTX 4090/24 GB observations, not our benchmark or average latency. Loading, cached conditioning, model switching and saving affect timing. The optional Qwen refiner is another generative edit pass: it may improve apparent detail while changing identity or lighting. Compare before/after at 100% rather than treating it as lossless restoration.
Alternatives: Qwen Image Edit and FLUX.2 [klein]
| Choice | Good reason to choose it | Main trade-off |
|---|---|---|
| MiniMax H3 Image Studio | Reuse an existing H3 environment for reference preparation and video production. | Large video-model stack; community workflows and H3 licensing conditions. |
| Qwen-Image-Edit-2511 | Dedicated reference editing, including multi-person and material-change use cases. | A separate model stack; edits still need visual review. |
| FLUX.2 [klein] 4B | A smaller dedicated generation/editing model with local and hosted options. | Different output character and reference behaviour; compare on your own assets. |
Qwen’s model card describes improved consistency and geometric editing; it uses Apache 2.0. Image Studio’s optional Qwen detail workflow is an additional pass, so compare standalone Qwen editing if the still image is your main deliverable. Official Qwen model card.

Black Forest Labs offers a 4B [klein] variant under Apache 2.0, plus hosted API/playground access. Do not apply that licence to every [klein] variant: the 9B model has different terms. These are workflow alternatives, not a controlled quality ranking. Official FLUX.2 [klein] page.
![FLUX.2 [klein] 4B official model page.](https://airmore.ai/wp-content/uploads/2026/09/minimax-h3-image-studio-review-en-klein.webp)
Troubleshooting, licensing and NSFW
| Problem | What to check first |
|---|---|
| Missing/red nodes | Update ComfyUI and the extension, restart, then reload a matching current UI workflow. |
| Output barely changes | Check whether an old FL2VA I2I workflow is keeping the source anchor; try the supplied REF2VA Image Edit workflow. |
| Wrong person or pose | Confirm reference order and assign one clear role to each image. Remove contradictory preservation instructions. |
| Out of memory or excessive delay | Return to about 1 MP, five frames and batch one. Check model variants, RAM use and offloading before adding refiners. |
| Blank canvas after importing JSON | Use examples/ui, not examples/api. |
| Bad detail at a larger size | Higher resolution is not proof of higher fidelity. Compare the same task at the default size and revise the source/prompt. |
The extension’s Unlicense does not replace the model licence. The H3 Community License excludes the EU, UK, South Korea and US, and requires separate authorization above US$20 million in annual commercial product/service revenue. It is therefore misleading to describe H3 as unrestricted open source or universally free for commercial use. Comfy’s licensing FAQ explains the community conditions and a separate paid route; its Professional offer starts at US$5,000/month and is not an Image Studio subscription. Confirm coverage for your use and territory. H3 licence · Commercial licensing.
NSFW: local execution is not an official “uncensored” feature or permission for every type of content. No reliable NSFW capability benchmark is established here. H3’s acceptable-use terms and any host rules still apply; consent, rights and restrictions on harmful content matter independently of whether a prompt technically runs.
Frequently asked questions
Is this MiniMax Image-01 or MiniMax Design?
No. This review covers astropuzzo’s ComfyUI extension for H3. Do not use the pricing, API parameters or account limits of another MiniMax product as if they belonged to Image Studio.
Can I use it without a powerful local GPU?
A compatible rented GPU environment may work, but confirm this extension and its model files are supported first. Native H3 cloud access alone does not prove support for this exact workflow.
Should I always choose four steps?
Use the paired Draft workflow for quick exploration. For a selected reference, compare it with the eight-step workflow; adapter and schedule changes affect quality as well as speed.
Can it make anime and realistic references?
Both appear in published examples. That establishes useful possibilities, not equal reliability across every style, body pose or character. Test a small asset set before committing a production.
Will H3 keep the character identical in the video?
Not automatically. Separate identity from pose/style references, use a clean source and review a short clip. The still’s origin is less useful than whether the reference is clear and fits the shot.
For creators who already run H3, Image Studio is a useful extension of the same production environment. Begin with one reference edit and a short image-to-video test, then keep only assets that survive close inspection. If your priority is fast standalone image editing, simpler hardware requirements or a different licence, compare a dedicated image model before investing in the H3 stack.
Review scope: project documentation and attributed published tests, checked September 28, 2026. No local GPU benchmark or end-to-end paid cloud generation was performed for this article. The cover is an editorial illustration, not an H3 sample.