MiniMax H3 video output is a grid of square tiles !! what am I missing?
▲ 4 r/FluxAI+2 crossposts

MiniMax H3 video output is a grid of square tiles !! what am I missing?

Hey, hoping someone here has run MiniMax H3 in ComfyUI and can tell me what I'm doing wrong.

Setup: RTX 4090, ComfyUI v0.30, running the H3 video nodes (AIMixer Director + Spectrum + video-tiler + KJ). When I generate, the output isn't a normal video. It's a full-frame grid of small square tiles, each decoded on its own with its own texture. The latent clearly got split into patches and never stitched back together. Looks like a mosaic, not pixelation.

Already ruled out:

  • Using minimax_h3_video_vae_fp16 (the official Comfy-Org one), not the audio VAE.
  • The VAE fp16 loads fine.

My guess is the decode step isn't going through the video-tiler, or it's hitting a generic VAE Decode node that doesn't know how to read H3 latents. H3 latents are tiled spatially, so something has to reassemble them before decode.

Anyone seen this exact tiling pattern? Is it the tiler missing from the chain, the wrong decode node, or a tile_size/overlap setting?

Frame attached so you can see the grid. Thanks.

u/Current-Row-159 — 8 days ago

Been stuck on this for a few weeks and wondering if anyone's dealt with something similar.

I do product photography/retouching work for luxury watches and jewelry, and I built a fairly complex pipeline using a mix of open-weight vision-language models to generate ad-quality campaign images. The idea is simple in theory: take a reference image whose lighting/background/style I like, take my actual product photo, and merge them so the final image has my real product sitting in a scene inspired by that reference.

In practice I ended up with several separate analysis passes (one mid-size VLM handling scene, style, and product cataloging separately) feeding into one merge step handled by a different, smaller multimodal model that also sees the actual images directly. Every time I fix one issue, a new one shows up somewhere else. First the style kept getting ignored entirely, with the output defaulting back to a generic version of the scene. Fixed that. Then lighting effects (bloom, sparkle, flare) started getting copy-pasted in a way that made no physical sense, like a sparkle effect that only makes sense on a pavé diamond setting getting slapped onto plain brushed steel, which instantly reads as fake. Fixed that too. Then the dramatic ambient glow from the background in the reference image, which was honestly like 40% of why that image looked so striking, quietly disappeared once I toned down the on-product sparkle, even though those two things had nothing to do with each other.

I keep tightening the instructions and the output keeps getting technically "more correct" without ever feeling like the genuinely impressive, poster-worthy image I'm actually going for. It's like I'm playing whack-a-mole between "photorealistic and coherent" and "actually has the visual punch of the reference."

Has anyone dealt with this kind of multi-stage analysis-then-merge setup for AI image generation? At what point does splitting analysis into specialized passes start hurting more than it helps, versus leaning harder on one strong multimodal model that sees everything directly and makes the creative calls itself? Or is there a better way to keep both technical product accuracy AND the creative/dramatic energy of the reference without this endless loop of fixing one thing and breaking another?

u/Current-Row-159 — 24 days ago

Been stuck on this for a few weeks and wondering if anyone's dealt with something similar.

I do product photography/retouching work for luxury watches and jewelry, and I built a fairly complex pipeline using a mix of open-weight vision-language models to generate ad-quality campaign images. The idea is simple in theory: take a reference image whose lighting/background/style I like, take my actual product photo, and merge them so the final image has my real product sitting in a scene inspired by that reference.

In practice I ended up with several separate analysis passes (one mid-size VLM handling scene, style, and product cataloging separately) feeding into one merge step handled by a different, smaller multimodal model that also sees the actual images directly. Every time I fix one issue, a new one shows up somewhere else. First the style kept getting ignored entirely, with the output defaulting back to a generic version of the scene. Fixed that. Then lighting effects (bloom, sparkle, flare) started getting copy-pasted in a way that made no physical sense, like a sparkle effect that only makes sense on a pavé diamond setting getting slapped onto plain brushed steel, which instantly reads as fake. Fixed that too. Then the dramatic ambient glow from the background in the reference image, which was honestly like 40% of why that image looked so striking, quietly disappeared once I toned down the on-product sparkle, even though those two things had nothing to do with each other.

I keep tightening the instructions and the output keeps getting technically "more correct" without ever feeling like the genuinely impressive, poster-worthy image I'm actually going for. It's like I'm playing whack-a-mole between "photorealistic and coherent" and "actually has the visual punch of the reference."

Has anyone dealt with this kind of multi-stage analysis-then-merge setup for AI image generation? At what point does splitting analysis into specialized passes start hurting more than it helps, versus leaning harder on one strong multimodal model that sees everything directly and makes the creative calls itself? Or is there a better way to keep both technical product accuracy AND the creative/dramatic energy of the reference without this endless loop of fixing one thing and breaking another?

reddit.com
u/Current-Row-159 — 24 days ago
▲ 2 r/ZImageAI+1 crossposts

Been stuck on this for a few weeks and wondering if anyone's dealt with something similar.

I do product photography/retouching work for luxury watches and jewelry, and I built a fairly complex pipeline using a mix of open-weight vision-language models to generate ad-quality campaign images. The idea is simple in theory: take a reference image whose lighting/background/style I like, take my actual product photo, and merge them so the final image has my real product sitting in a scene inspired by that reference.

In practice I ended up with several separate analysis passes (one mid-size VLM handling scene, style, and product cataloging separately) feeding into one merge step handled by a different, smaller multimodal model that also sees the actual images directly. Every time I fix one issue, a new one shows up somewhere else. First the style kept getting ignored entirely, with the output defaulting back to a generic version of the scene. Fixed that. Then lighting effects (bloom, sparkle, flare) started getting copy-pasted in a way that made no physical sense, like a sparkle effect that only makes sense on a pavé diamond setting getting slapped onto plain brushed steel, which instantly reads as fake. Fixed that too. Then the dramatic ambient glow from the background in the reference image, which was honestly like 40% of why that image looked so striking, quietly disappeared once I toned down the on-product sparkle, even though those two things had nothing to do with each other.

I keep tightening the instructions and the output keeps getting technically "more correct" without ever feeling like the genuinely impressive, poster-worthy image I'm actually going for. It's like I'm playing whack-a-mole between "photorealistic and coherent" and "actually has the visual punch of the reference."

Has anyone dealt with this kind of multi-stage analysis-then-merge setup for AI image generation? At what point does splitting analysis into specialized passes start hurting more than it helps, versus leaning harder on one strong multimodal model that sees everything directly and makes the creative calls itself? Or is there a better way to keep both technical product accuracy AND the creative/dramatic energy of the reference without this endless loop of fixing one thing and breaking another?

u/Current-Row-159 — 24 days ago
▲ 11 r/krea+4 crossposts

Does Krea 2 support real photo editing or just generation from a prompt

Hey everyone, I work as a jewelry retoucher and I'm trying to figure out something about Krea 2. Does anyone know if it actually supports real editing of an existing photo, meaning uploading a shot of a ring or a diamond and having it modify that exact image, or is it purely a text to image model that generates something new every time. I mostly need this for retouching jewelry product shots and swapping out backgrounds while keeping the piece itself completely unchanged. If direct editing like that isn't really what Krea 2 is built for, is there any way to use a reference image or some kind of controlnet setup with it so the shape and details of the jewelry stay locked while only the background or lighting changes. I've seen mentions of style references and moodboards but from what I understand those are more about transferring a look or aesthetic rather than preserving exact product geometry, which is the opposite of what I need. Has anyone actually tried this for product photography or anything where precision matters this much. Would love to hear real experiences before I spend more time testing it myself.

u/Current-Row-159 — 1 month ago
▲ 12 r/FluxAI+3 crossposts

What nobody tells you about retouching shiny stuff (and how AI quietly changed my workflow)

I’ve been retouching jewelry photos for a while and honestly it’s the hardest thing I’ve ever edited. Reflections pick up everything, dust becomes boulders, and keeping gold looking like actual gold across dozens of shots is brutal. I got obsessed with how big brands like Tiffany or Mejuri keep their entire catalog visually cohesive so I started experimenting with AI, not to replace the craft but to speed up the boring parts.

What surprised me most is that once you have a clean consistent dataset of a single stone, training a LoRA on a specific brand's lighting style actually works. You can make a diamond look like it was shot in their studio, same warmth, same shadow depth, same mood. It's wild.

I ended up shooting 100 frames of the same emerald cut diamond at 4K because I needed a perfect base to train from. It made such a difference that I wanted to share it, not to sell anything, but because I wish someone had told me earlier that the quality of your training images matters more than the prompt. If you're stuck fighting inconsistent source material, the AI can't learn the subtleties.

Anyway, just wanted to share what I've been tinkering with. If anyone else here retouches shiny reflective stuff I'd love to know your pain points. This niche is lonely.

u/Current-Row-159 — 3 months ago