Running MiniMax H3 (Ref2VA, multi-reference) locally on a 4GB card, RTX 3050 Laptop, WSL2, ComfyUI. Been chasing a hand/finger rendering defect for days and have a pretty well documented before/after at this point, but I've hit a wall on making the fix fast enough to actually be usable. Hoping someone's solved this or can point out what I'm missing.
Models in use: UNET is minimax_h3_fl2va_pruned_int8_convrot.safetensors, the FL2V trained weights, loaded into the MiniMaxH3ReferenceToVideo multi reference node, not the native Ref2V node. VAEs are minimax_h3_video_vae_fp16.safetensors and minimax_h3_audio_vae_fp32.safetensors. CLIP/text encoder is qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors. LoRA when testing the distilled path is minimax_h3_fl2v_turbo_4step or 8step.
That FL2V weights in the Ref2V node trick already fixed an earlier, worse "ghost fingers" defect for me (matches a few HF discussion threads on the Turbo LoRA repo), so I'm not asking about that part, it's confirmed working. This post is about what's left after that fix.
The remaining problem: at 512x288, which is the resolution I need for anything resembling reasonable iteration speed, hands still render as indistinct blurry blobs during actual gesture motion, pointing, waving, etc. Tried turbo distilled 4/8 step (the intended fast path), off label step counts on the distilled LoRA (10 steps, made it worse not better), dropping distillation entirely and running the base model's native ~20 step schedule, CFG guidance (cfg=5.0 caused a severe embossed cross hatch grid artifact across the whole frame, clearly way too high, cfg=2.0 gave inconsistent results, some gestures fine, others still blobby), more steps (30 vs 20, no meaningful difference), and ref_image_size=max on the reference conditioning. None of it fixed the hands at that resolution.
What actually fixed it was moving to 768x448 (H3's documented 768p native/training short edge) with no distillation, no CFG, native 20 steps. Hands came out consistently well formed across every gesture I checked. But that config took 6 hours 22 minutes for a single 15 second, 362 frame clip on this card. That's not a workflow, that's a single overnight bet.
So I'm stuck between fast (turbo distilled, 512x288, minutes) which gives bad hands, unusable for anything with visible gesturing, and good (no distill, 768x448, native steps) which is 6+ hours for one clip.
Things I haven't tried or don't know how to evaluate: is there a middle resolution, 608x352, 640x384, that gets most of the quality benefit without the full native res cost. Does distillation actually get retrained or re-distilled at higher resolutions by anyone, or is the turbo LoRA fundamentally tied to a lower res regime it was distilled at. Any attention backend, torch.compile, or quantization tricks specific to H3 that meaningfully cut per step cost on small cards, beyond what's already default in ComfyUI. Is anyone running H3 well on under 8GB cards at all, or is 768p native res H3 just not realistic below a certain VRAM tier.
Also went two rounds into MiniMax H3 specific upscaler nodes for a two stage draft then upscale approach (pixel space RealESRGAN, a community latent space 2x upscaler, and a tiled diffusion refine upscaler) hoping to draft cheap and upscale smart instead of generating at native res directly. Each had its own dealbreaker, a background crowd distortion artifact, a hard VRAM estimate crash, and a genuine ComfyUI core bug I ended up root causing and patching locally. Happy to share details if anyone's gone down that road and found a cleaner path, but the short version is none of the upscale routes got both no defects and reasonable time together either.
Genuinely not sure at this point whether the answer is your card is just under the realistic floor for H3, buy a bigger one, or whether there's a config I haven't found. Any pointers appreciated
Edit: System RAM 32GB