No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti
So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.
So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.
And it worked.
For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:
https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v
Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:
“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.
Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”
The final result was generated locally on my RTX 5070 Ti using ComfyUI.