Follow-up to my 6-minute TNG video — I changed the workflow a lot for the second one
A few days ago I posted the 6-minute Star Trek: TNG video I made with MiniMax H3 in ComfyUI. I’ve finished the follow-up now, and I changed the workflow quite a bit after seeing what worked and what didn’t on the first one.
The biggest improvement was consistency. For the first video, most shots were generated more independently, and I deliberately built some of the continuity weirdness into the story. That worked for the premise, but for the second one I wanted it to feel much more like an actual TNG episode, so I became much more rigid about shot composition.
A big part of that was using the H3 reference model differently. Instead of just giving it a single image and hoping for the best, I used reference images and told MiniMax to stick very closely to the composition in those images. In practice that sounds a bit like image-to-video, but it worked quite differently for me.
With the reference model I could use up to six photos and be much more deliberate about how the shot should work. I could decide what the starting shot should be, what the end shot should be, whether I wanted a middle reference, a final-frame reference, etc. That gave me a lot more control over blocking, framing and performance than I was getting from the image-to-video model.
I did test the image-to-video model as well. One of the shots that made it into the finished video is the later one where Data has a slightly longer monologue. You can tell he looks a bit more “off” there. The reference model, by comparison, was giving me Data much more accurately, both in terms of how he looked and in terms of his mannerisms. That ended up being the better approach for this project by a long way.
The video is still built from lots of separate short H3 generations rather than one long generation. I wrote the scenes first, then generated individual shots and multiple takes where needed, and assembled everything in Premiere like a normal edit.
I also changed the audio workflow quite a bit. On the first video, one of the main issues was that the generated ambience and background noise varied too much from clip to clip. This time I spent much more time matching dialogue levels in Premiere, cleaning up individual clips, and adding a continuous Enterprise bridge/interior hum underneath scenes so the cuts felt less obvious.
I also handled the music more deliberately this time. Rather than just dropping in whatever worked at the end, I treated it more like proper scene underscore and generated short incidental cues for specific moments.
So the rough workflow for the second one was:
script and shot planning
→ select or build composition references
→ generate short H3 shots in ComfyUI using the reference model
→ do multiple takes where needed
→ edit in Premiere
→ clean dialogue and level-match clips
→ add continuous ambience/room tone
→ add short music cues
→ final upscale/export
The main thing I learned was that H3 works much better for this kind of project when I treat it less like a one-click video generator and more like a production tool. The closer I got to thinking in terms of individual shots, coverage, performance selection and edit assembly, the better the final result got.
Happy to answer questions about the workflow again.