Anyone remembers the Will Smith Spaghetti Video? - (LTX 2.5)
We have come a long way from that video. So I thought why not try to recreate it again?
We have come a long way from that video. So I thought why not try to recreate it again?
WORKFLOW: https://pastebin.com/raw/xPTwRh9P (H3 not embedded works for any video - mainly for H3 upres with LTX 2.5 INT8 Distilled / SageAttn / SolAttention)
This was my very first gen from H3, time to revisit and upscale it with LTX 2.5 + Spatial 2x. If you notice - when she's far away "walk" it's better! This is really great. Of course, Spatial 2x applies some softness here and there, overall is not bad. Didn't test - SeedVR2 and FlashVSR yet
Perhaps, we can still use LTX 2.5 for now for upscaling (including: voice dub, LTX director, connecting/extending shots), and make gems gens with H3.
Till new beast arrives.... Minimax H3-Regenerate-2K (probably will be superior to anything).
H3 original (super low res): 864 x 480 / 18 Steps / Pruned INT8
Spatial 2x LTX 2.5: 1728 x 960 (Distilled INT8, render time pm: 338 sec - at RTX 4090 / 128gb RAM / SageAttn + SolAttn)
Ps. On phone maybe hard to see - checkout on desktop
Cheers
HI everyone,
Today we got release of LTX 2.5, so lets compare with more advance scenes along side Minimax H3.
3 videos in order - same PROMPT (with <Shot 1>...<Shot 2>)
We can definitely say, LTX is upgraded (more detailed textures vs LTX 2.3...lots of snow, snow slush etc), is it better then H3? Judge yourself ;)
cheers
Here we go PROMPT:
[Shot 1] Handheld selfie-style video shot on a smartphone at 60fps, capturing a bleak, emotional scene under the harsh fluorescent lights of a drafty, concrete train station concourse. A young, strikingly beautiful girl—around eighteen years old—sits slumped against a cold tiled wall in a corner. She holds the smartphone in a trembling right hand, filming herself and her surroundings. She wears a thin, threadbare denim jacket covered in dark grease stains and road salt, a stretched-out oversized sweater torn at the collar, and frayed canvas pants soaked with melting slush at the cuffs. Clustered beside her in the freezing draft is an old German Shepherd with a matted, dirt-streaked coat, shivering against her knee. Heavy snow and howling wind blur through the open station doors behind them. Tears stream down her smudge-streaked cheeks as she speaks into the phone in a quiet, trembling voice: <d>[English with a soft, heartbroken voice] It's so cold tonight... please, nobody even looks at us... we just need a little food...</d>. At 00:05:500 Cut to [Shot 2] Low-angle handheld shot from her perspective on the floor. Commuters in thick, clean winter coats and boots hurry past her in a fast blur, heads turned away, completely ignoring her outstretched, gloveless left hand. Her breath forms dense white clouds in the freezing air. The German Shepherd lets out a soft, whimpering whine and rests its heavy head on her lap. She wipes a tear from her nose with her sleeve, her voice cracking with desperation: <d>[English with choked, sobbing breaths] Please... just a piece of bread for him... anything...</d>. At 00:11:000 Cut to [Shot 3] Upward camera angle as a middle-aged man in a dark wool overcoat and scarf suddenly stops in front of her. His boots come to a halt in the wet slush beside her dog. He crouches down to her eye level, looking at her and the shivering German Shepherd with genuine concern and warmth. He gently reaches into his coat pocket, pulling out a warm paper bag from a bakery and a thermal travel mug, speaking in a gentle, compassionate voice: <d>[English with a warm, caring tone] Hey... hey, don't cry. Here, take this—it's hot soup and fresh bread. Are you okay?</d>. At 00:16:000 Cut to [Shot 4] Close emotional selfie framing as her eyes widen in tearful disbelief. She hugs the warm paper bag to her chest with both hands, tears pouring down her face as she smiles through her sobs, looking up at the man and then into the camera lens: <d>[English with a tearful, weeping whisper] Thank you... oh god, thank you so much... bless you...</d>. The German Shepherd gently licks her cold hand as the man reaches out to pet the dog's head, cutting the video to black on a powerful, dramatic note at 00:20:000, overall_soundscape: Howling blizzard winds outside open station doors, heavy footsteps echoing on wet tile floors, distant train arrival announcements, shivering whimpers from the dog, and her quiet, heartbreaking sobs, non_diegetic_music: Soft, somber cinematic violin pads fading in subtly under the diegetic audio to heighten the emotional drama
I discover this by playing around, so good news is - we can have much more references then 9.
<Picture 1> woman in <Picture 2> luxury bathroom is touching her face showing her silver earrings, camera slow motion up-close on face and torso, she puts on glasses, looks at mobile purple phone puts to her ear and smiles to camera.
Once again H3 it's beyond amazing....
Also by including multiple faces - different expressions H3 learns expressions etc. teeth, looks, in a away - we don't need character LORA.
I guess this is great find for all of us.
Cheers
*UPDATE*
Even faster with - https://www.reddit.com/r/StableDiffusion/comments/1vhlfmw/minimax_h3_firstblockcache_for_comfyui_3033_lower/
You can even combine them both for extreme speeds
Hey all,
Lots of addons been released by hours, one interesting speed-up is - Spectrum (cannot be used with EasyCache): https://github.com/xmarre/ComfyUI-Spectrum-MiniMax-H3
Follow more here:
https://www.reddit.com/r/StableDiffusion/comments/1vf1ze3/spectrum_acceleration_for_minimax_h3_in_comfyui/?utm_source=chatgpt.com
Video above rendered with Spectrum / 20 Steps
cu130+SageAttn + rtx 4090, 128gb RAM, Win11
Spectrum: 201 seconds
Default: 320 seconds
Quality still holds, I was testing all current H3 turbo LORA's - all seems to kill details and/or severely degrade sound. I'm sticking with 20 steps - solid.
Cheers