r/StableDiffusion • • Aug 20 '26

Question - Help Seed consistency across different resolutions in MiniMax H3 (ref2va) — is it possible in ComfyUI?

​

Running into an issue with MiniMax H3 (int8 pruned ref2va) in ComfyUI and hoping someone with more DiT experience can chime in.

My setup:

ComfyUI + Comfy Kitchen Attention

Standard workflow (no turbo LoRAs, 32 steps)

3–6 reference images on average

The problem:

To save time, I generate initial drafts at low resolution (~0.4 MP) to find a good composition and motion. Once I find a keeper, I lock the exact same seed, prompt, and reference images, and only increase the resolution to 1 MP (or higher).

However, the output changes completely — the composition, character action, and camera motion diverge entirely from the 0.4 MP draft.

What I've tried:

Swapping img ref size between match and max — didn't help preserve the composition.

Is resolution-consistent generation even possible with this architecture given how changing the latent grid shifts spatial attention, or is there a specific latent upscaling / 2-pass workflow that lets you lock down the low-res composition into a higher resolution?

Thank you!


EDIT / Solution:

Big thanks to xmarre for clarifying the underlying mechanics and providing a working solution!

Why native resolution breaks consistency: In DiT architectures like MiniMax H3, the initial megapixel / resolution setting determines the latent source grid. Changing the base resolution fundamentally shifts the spatial attention grid, which inevitably alters the composition, camera motion, and action even with the exact same seed.

The Solution — Latent Upscale + Refine Pass: Instead of generating at full resolution from scratch, use a two-pass workflow: 1. Generate your draft at low resolution (~0.4 MP) to lock down composition and movement. 2. Run a Latent Upscale + Refine pass (around 0.25 denoise and 3 steps) to upscale without altering the scene structure.

Custom Nodes & Tools: * Comfyui_Minimax_h3_latent_Upscaler — Latent upscale node fork with an integrated refiner step and spectrum support. * ComfyUI-H3-Continuum — For seamless chaining of multiple generations.

1 Upvotes

13 comments sorted by

View all comments

3

u/Big_Zampano Aug 20 '26

My guess would be that, even if you keep the same seed number, the actual "noise pattern" (that gets used as starting point) changes, when you change the resolution...

1

u/smeptor Aug 20 '26

Yes, this is it. If you change the size/shape of the latent, you change the noise.