▲ 881 r/MinimaxVideo+1 crossposts

Walter White and the Minimax H3 Official Prompting Guide

This post is half a joke and half a plea and public service announcement.

Some people have been complaining they don't get results as good as other people with Minimax H3 videos, or have the following issues:

  • Dialogue being spoken by the wrong characters
  • Dialogue that is just gibberish or random
  • Random video cuts they didn't ask for
  • Characters talking over each other or too fast
  • Prompts not being followed

These things can all be prevented and avoided and not encountered at all if you follow the official prompting guides. Yes, there are two. Both are on the official Huggingspace page for Minimax H3.

One is the Official Prompting Guide for the Text to Video and Image to Video Model.

The other is the Official Prompting Guide for the Reference Video Model.

There is some overlap, but for the most part, each model has it's own prompting syntax, and in particular, the Reference Video Model for H3 is very picky about you using the right keywords and instructions to get what you want.

"But I get decent results with just a couple of sentences typed in natural language of what I want."

That's great, but you're really just relying on the Qwen 32b vision model guessing what you want. It's like pulling a slot machine lever and hoping you get cherries. Only this slot machine can take a few minutes to nearly an hour to stop spinning, based on your hardware.

The great thing about Minimax H3 is for the first time we can truly direct our own AI videos like a director would on set, with the AI providing the actors, scenery, and props. If you write a properly formatted and detailed prompt for Minimax H3, it looks almost like a shooting script.

Why spend time waiting to hit a jackpot when you can take a few minutes to write a detailed, properly formatted prompt that follows the official guides, and get those bright lights and tokens falling into your lap on the first lever pull?

Okay, quick fire problem solving for people who still won't RTFM:

>Dialogue from the wrong characters?

>Dialogue that is just gibberish or random?

Walter White says, <d>[English in Walter White's voice from Breaking Bad] My product is pure, Jesse! There will be no chili powder in my meth.</d>

Always specify the character speaking, either by name, or using the <Subject 1> system in the official guide. In the Text to Video and Image to Video model, always use the <d>[Language Spoken]</d> tags. This will fix BOTH of those issues.

>Random cuts in the video you didn't ask for?

[Shot 1] A medium close-up of Jesse Pinkman from Breaking Bad, pacing back and forth, agitated. He looks up towards the camera, opens his mouth as if he's about to speak, then seems to change his mind, closing his mouth and shaking his head. [Shot 2] At 00:06:000 the camera cuts to a static camera shot framing Walter White from Breaking Bad, sitting on a cheap white plastic lawn chair, his arms crossed and glaring at Jesse. [Shot 3] At 00:10:500 the camera pans quickly back to Jesse, doing a Push In at slow speed to his face as he stops pacing and narrows his eyes at Walter.

This is how you control not only the camera work, but the PACING of your video. You NEVER include a time code on your first shot. You can omit the time code from ALL shots if you want the model to decide on it's own, based on your prompt, when to cut.

BUT, for ultimate control, you want to use time codes. Look at my example above. I just told the model to have Walter glare at Jesse for 4.5 seconds, because I told the model that camera shot starts at 6 seconds into the video, and the next cut doesn't happen until 10.5 seconds into the video. That lets you control the pacing and timing for jokes, punchlines, acting, everything.

>Characters talking over each other or too fast?

This is an old one that anyone familiar with prompting for video models should know by now - what you are asking for in your prompt and the length of your video in time need to match.

The model will try its best to cram every action and piece of dialogue into your video that you asked for, and if that would naturally take 10 seconds and you've only given it 5 seconds? Well, now everything is crammed together, overlapping, or being cut-off.

My recommendation is to generate just a quick 0.2 MP version of your video first after you type your prompt, generate, and see how the timing is working. Is it too fast? Too slow? Do the actions have enough time to happen? Do you want more breathing room?

This is the time to decide all that and lock in a video length. The low resolution of 0.2 MP is quick to generate on most set-ups (mine for this post's video took 3.5 minutes for a 14 second video) and let you work out any issues in your prompt before going in for the long generation at higher resolution.

>Prompts not being followed?

It's because you didn't read the manual!

--------------------------------------------------------------------------------------------------
Now, with all that said, here is the prompt for the video I made:

integrated_multimodal_description: [Shot 1] Live-action film footage of the American drama series Breaking Bad, professionally color graded with a warm color grade, with slightly desaturated colors for a premium film feel, a continuous camera shot with no cuts, medium close-up POV shot of Walter White, bald with a goatee and glasses, as portrayed by Bryan Cranston. He is standing in the Arizona desert next to a parked RV. He is wearing a white PPE protective suit and yellow rubber dish gloves. He is looking directly at the viewer with barely constrained anger. At 00:01:300 he reaches out towards the camera and points his finger at the POV camera with one hand, the camera shaking slightly from the movement. Walter then says angrily, &lt;d&gt;[English with Walter White's voice] Listen, you want to cook Mini Max H3 videos, you follow the recipe!&lt;/d&gt;. At 00:04:500 Walter raises his other hand revealing he is holding a thin stack of white paper pages in portrait orientation. The front of the paper visible on top of the thin paper stack is blank except for the large black printed text "Minimax H3 Official Prompting Guide". The papers are held in front of the camera on the right side of the screen for a moment in portrait orientation, so the text can be clearly read, while Walter glares at the viewer on the left side of the screen. At 00:07:000 Walter then shakes the papers at the camera, then says angrily, &lt;d&gt;[English with Walter White's voice] Read the fucking manual!&lt;/d&gt;. At 00:10:000 the camera does Pan Right and a Pull Out to show a close-up of Jesse Pinkman from Breaking Bad, with his hands held up by his face with fingers spread, an annoyed look on his face. Then he says in frustration, &lt;d&gt;[English in Jesse Pinkman's voice from Breaking Bad] Alright! Damn, Mr. White! I just want to generate memes.&lt;/d&gt;, overall_soundscape: Ambient sounds of an Arizona outdoor desert during the day, non_diegetic_music: none

For those interested, this video was generated at 1 MP on a 3090, using Sage Attention and the Spectrum Node for H3. The final video of 14 seconds at 1 MP took 40 minutes to generate and then was upscaled using RTX Super Resolution.

The workflow was the default Text to Video Minimax H3 template that comes in the latest update of Comfyui.

Now get out there and go cook some memes, everyone!

u/GrayingGamer — 13 days ago

Starting the Buffy Memes; And How to Prompt for TV Shows in H3

So, I figured I would kick off some Buffy the Vampire Slayer meme generations with Minimax H3, while also giving a lesson in how to prompt for any TV show and character the model knows while also getting the correct character voice, all through just pure text to video prompting.

The prompt for this Buffy video was this:

A television scene from the American television drama series Buffy the Vampire Slayer from in 1997, professional color grading, in the style and aesthetics of the drama series Buffy the Vampire Slayer.

Scene overview: Buffy as played by Sarah Michelle Gellar walking through a cemetary at night, with a low hanging fog and cool blue color grading to emphasize the night. Willow as played by Alyson Hannigan is walking next to her.

Shot 1: Medium close-up tracking shot of the camera following Buffy as played by Sarah Michelle Gellar and Willow as played by Alyson Hannigan walking through a cemetary at night, looking bored. Willow is looks at Buffy with an amused expression, saying in a joking tone of voice &lt;d&gt;[English in Willow's voice from Buffy the Vampire Slayer as played by Alyson Hannigan] You keep this up we're going to start calling you the 'Vampire Layer'.&lt;/d&gt; She makes air quotes with her fingers as she says the 'vampire layer' words.

Shot 2: Hard cut close-up tracking shot of the camera on Buffy's face as played by Sarah Michelle Gellar, looking surprised and offended as she turns her head to look at Willow. She mutters quietly but offended, &lt;d&gt;[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.&lt;/d&gt;

overall_soundscape: Quiet ambience of an outdoor cemetary at night.

non_diegetic_music: none

Notice how I am hammering the details of the show in the prompt first, not just "Buffy", or "A scene from Buffy", or just "Buffy the Vampire Slayer". I'm nailing it down to year, genre, format, and repeating myself.

The same for the characters. Notice how I attach the character names to the show every time and not just in the scene description, but every time they appear. This helps lock down the exact look of the character with no drift.

Next, look at the dialogue. You need to follow the official prompting by putting what characters say in dialogue tags, like so: &lt;d&gt;[English] What they say. &lt;/d&gt; But you can add a LOT more detail about the speaker in those [ ] brackets.

Look how I do them EVERY TIME in my prompt, and ensured I got the exact character voice:

&lt;d&gt;[English in Buffy's voice from Buffy the Vampire Slayer as played by Sarah Michelle Gellar] Damn, Willow.&lt;/d&gt;

It's not just English, it's English from Buffy. Not just any Buffy, but from this television show. And whose voice is Buffy actually speaking with? Her actress's voice, Sarah Michelle Gellar. (Check your spelling on names!)

If you do all this and the movie or television show is in the training data, the model WILL generate you a scene with it. If it doesn't? Well, you're out of luck doing T2V and will need to use the Reference H3 model and supply your own character images, audio clips for voices, etc.

I don't use LLMs to write my prompts. I type them all out myself. I find it just works better that way, though I DO copy and paste all those repeating character names / show name / actor name sentences.

If anyone has any prompting questions, just let me know.

Oh, and all these was with just the default T2V workflow template that comes with Comfyui.

u/GrayingGamer — 14 days ago

Minimax H3: Captain Picard Discusses Your Holodeck Use

This was all done in 5 to 7 second clips, text to video only, in Minimax H3.

Scenes were generated at 0.6 MP, then upscaled with RTX Super Resolution and put together in one video with Davinci Resolve. Each clip took about 5 minutes on a 3090. I have 128GB of system RAM.

I'm using the Spectrum node and SageAttention, so quality isn't as good as it could be, but I was happy enough with the results and saw people were struggling with getting Picard's voice right, so I thought I'd share this as an example of what the model can do, and how to do it consistently. No references were used for his voice, only text prompts.

I got his iconic voice in all these clips by asking for it in the proper format:

Captain Picard from Star Trek:TNG then says, &lt;d&gt;[English with Picard's classic British accent] Number One?&lt;/d&gt;

u/GrayingGamer — 15 days ago

Nothing but Prompts. Ideogram 4 Has Scary Control.

These are posters I made in Ideogram 4, using only prompting and bounding boxes. No image reference, no controlnets, or loras.

I wanted to test how much compositional control is really available with Ideogram 4, so I set out to recreate iconic 1980s horror movie posters from scratch using it.

The first poster is always Ideogram, followed up by the original poster so you can easily compare. They aren't perfect recreations (and I would be suspicious if they were), but I continue to be happily surprised by how accurately Ideogram 4 can let me recreate an image I have in my head using precise object placement, color palettes, text styles, etc.

The Poltergeist TV was an especially cool example - the actual TV model was unknown to Ideogram 4 - but that wasn't a problem, because I built the TV piece by piece with bounding boxes and prompting to recreate a close duplicate.

In a few cases I purposely changed composition - like on the Sleepaway Camp poster recreation, I made the title into one line and dropped the text stinger below it since I wasn't reproducing the cast and production text on the posters. Same on Nightmare on Elm Street where I made the title bigger for the same reason. (Just wanted you all to know some changes like that were on purpose!)

Again, I just want to repeat that no image to image, controlnets, inpainting, or photoshop compositing was used to make these - they are pure generated output from Ideogram 4.

You can get my improved Ideogram 4 workflow here. It includes the prompt and bounding boxes I used to make the Poltergeist poster recreation and really shows off how to make effective use of bounding boxes. It uses INT8 models - if you use FP8, just swap out the model loaders for regular "Load Diffusion Model" nodes and you'll be good.

Hopefully this shows the strengths of Ideogram 4 to translate ideas for images in your own head into reality with precise control and shows Ideogram 4 isn't a "pull the lever and see what you get" image generator, but more of a tool.

u/GrayingGamer — 2 months ago
▲ 938 r/generativeAI+1 crossposts

Ideogram 4.0's Understanding of Characters and IP is Crazy for an Open Model

Like I said in the title, Ideogram 4.0 has the absolute best character and IP knowledge I've seen in an open model without loras.

I hated on Ideogram 4.0 when it first came out because of the initial workflow issues and the safety filter, but now that both of those things have been sorted out, I'm having some of the most fun with a model I've had in years.

These were generated locally in Comfyui at 1.5 megapixels - 1440x1024, specifically.

I am using the INT8 versions of the Ideogram 4.0 models and Kijai's Ideogram 4 Prompt Builder KJ node from his KJ Nodes custom pack. Workflow being used is SilverOxide's which you can find here. EDIT: SilverOxide's workflow got deleted, so I cleaned it up, stripped out some unnecessary stuff put my own workflow up on Pastebin here.

If you don't know, or haven't tried it, Ideogram 4.0 also does very well with inpainting. It makes it easy to generate at lower megapixels and then mask and inpaint areas like faces to clean up and correct detail. I use the Comfyui-Inpaint-CropAndStitch custom node found here, personally, but most of the time Ideogram 4.0 doesn't need it.

If anyone wants prompts for a specific image, just ask in the comments below and I'll provide them there to avoid cluttering the main post with a wall of JSON text.

u/GrayingGamer — 2 months ago