▲ 42 r/comfyui

No Camera. No Model. Just MiniMax H3 Running Locally on a 5070 Ti

So basically, I saw a workflow on ComfyUI’s official LinkedIn where they used a model image, a product image, and a background image with Google and Kling APIs to generate a one-shot ad using a single camera angle.

So I challenged myself to recreate the idea using only local open-weight/open-source models, but make it more ambitious: multiple shots, multiple cuts, and everything directed through a single prompt.

And it worked.

For this, I used the basic MiniMax H3 Reference-to-Video workflow in ComfyUI:

https://docs.comfy.org/tutorials/video/minimax/minimax-h3#minimax-h3-reference-to-video-r2v

Then I used ChatGPT to help structure the video prompt. I provided the reference images and gave it this direction:

“Write a MiniMax H3 reference-to-video generation prompt to create an ad. Add sound FX and music prompts as well.

Shot 1: Medium close-up. She is about to open the can.
Shot 2: Extreme close-up of the can as she opens it. Can-opening sound FX.
Shot 3: Close-up as she drinks from the can. Gulping soda sound FX.
Shot 4: Close-up as she holds the can forward and smiles.”

The final result was generated locally on my RTX 5070 Ti using ComfyUI.

u/Time-Ad-7720 — 8 hours ago
▲ 78 r/comfyui

Turned my son's drawings into an animated skit

Turned my son's drawings into an animated skit locally using open weight models. This was done in ComfyUI using MiniMax reference to video model #minimaxh3

Here's how I did it: My son came up to me today and asked me if I could make his characters talk (he painted them in Paint 3D app on windows). And he wanted it to be silly. So I threw in some classic dad jokes.

I took screenshots of each character, then added them as reference images in the minimax h3 reference to video workflow, then used chatGPT to write the prompt.

u/Time-Ad-7720 — 3 days ago
▲ 202 r/MinimaxVideo+1 crossposts

Minimax H3 + Krea 2 | LoFi Anime short experiment

Workflow: https://drive.google.com/file/d/1B4kODxXQgJ1QOKRsEIkxHbgYmdruPpTK/view?usp=sharing

I made this experimental short scene using ComfyUI.

I generated the characters and backgrounds using the Krea 2 open-weight model with a Turbo LoRA, then animated them using MiniMax H3 with a reference-to-video workflow. With the Turbo LoRA, each 5-second clip took around 2–3 minutes to generate.

I edited everything in CapCut, added some color grading, bloom, and film grain, and it all came together nicely.

For the lo-fi track, I produced it on the Maschine MK3 using a free sample pack.

u/Hefty_Scallion_3086 — 3 days ago
▲ 1.0k r/PlexPrerollStudio+4 crossposts

Assemble The Multiverse | Minimax H3 R2V is awesome!

Workflow: github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json

Used multiple reference images for each scene.

Prompt For Multi Character:
subject_definitions:

<Subject 1> is [CHARACTER 1] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

<Subject 2> is [CHARACTER 2] from <Picture 2>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

<Subject 3> is [CHARACTER 3] from <Picture 3>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

summary:

[reference generation] A 5-second cinematic multiverse portal arrival. Three characters emerge from a consistent amber-orange portal and take a calm, confident formation.

retention_analysis:

<Subject 1> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.

<Subject 2> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 2>.

<Subject 3> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 3>.

detailed_description:

A 5-second cinematic portal-arrival scene at dusk. One stable medium three-shot, framed from the knees up. No dialogue, no combat, no wide landscape, no camera movement, and no crowd.

Portal continuity: a large circular amber-orange portal stands behind the characters. It has a bright rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.

[Shot 1] <Subject 1> steps through the portal first and takes the centre position with quiet confidence. <Subject 2> emerges on one side, naturally adjusts or lowers any item they are carrying if applicable, then gives a focused glance toward the unseen distance. <Subject 3> walks through last, takes position on the opposite side, and calmly surveys the scene. The three hold a poised, united stance as the portal flickers and golden particles drift around them. Their expressions and body language remain confident and appropriate to their individual character identities.

overall_soundscape:

Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.

non_diegetic_music:

A restrained cinematic rise builds across the shot and resolves on a calm, confident note.

Prompt For single characters:
subject_definitions:

<Subject 1> is [CHARACTER] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.

summary:

[reference generation] A 5-second cinematic multiverse portal arrival. One character walks through a consistent amber-orange portal, then takes a confident action stance with a subtle grin.

retention_analysis:

<Subject 1> (appears in [Shot 1] and [Shot 2]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.

detailed_description:

A 5-second cinematic portal-arrival scene at dusk. No dialogue, no crowd, no wide landscape, and no combat.

Portal continuity: a large circular amber-orange portal stands behind <Subject 1>. It has a bright fiery rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.

[Shot 1] Medium knee-up shot. <Subject 1> walks steadily through the portal toward the camera, then comes to a composed stop. Their costume, silhouette, movement style, and any character-specific accessories remain fully consistent with <Picture 1>. Golden sparks drift around them as the portal flickers behind.

[Shot 2] Close-up of <Subject 1>. They shift into a distinctive, character-appropriate action stance, looking directly ahead with calm confidence. Their expression changes into a subtle smile and restrained grin. Keep the movement natural and controlled, with no exaggerated facial distortion. The portal remains softly visible and out of focus in the background.

overall_soundscape:

Low portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.

non_diegetic_music:

A restrained cinematic rise builds through the entrance and resolves as <Subject 1> holds the final stance.

u/CommunicationFit3862 — 11 days ago
▲ 26 r/aigamedev+1 crossposts

Vibecoded a rhythm game for my own album

Building a rhythm game out of my own album. First level is playable.

Each song becomes a level with its own world. Beats on Maschine MK3, vocals generated in Suno, game is WebGL, mostly written with Claude Code.

Two things took most of the time and neither was the graphics.

First was tempo. Tap tempo put me nearly a full BPM off, which sounds like nothing until the notes are two seconds ahead of the music by the outro. So I measure it off the waveform instead, onset detection plus a least squares fit against the beat grid. Got it to 87.7636 BPM with almost no drift over the whole song.

Second was the editor. Charting in raw JSON is awful, so I built a Studio: live game preview on one side, waveform and timeline on the other. Scrub to any beat, everything snaps to the grid, and you place notes and camera cuts while the song plays. Notes go in four lanes as taps, holds or slides, and the patterns follow the song rather than just the beat, so the chorus and the verse feel different to play. Camera cuts land on section boundaries so the level shifts when the track does.

Still very much in progress. Eleven more tracks to chart. Will post as levels get done.

u/Time-Ad-7720 — 18 days ago
▲ 10 r/VSTi

A wavetable Synthesizer that turns images into Music [FREE+Open source+NO AI Backend]

https://preview.redd.it/ubvl9n04efeh1.jpg?width=800&format=pjpg&auto=webp&s=7d9adae1a83553a4d592f5e6cf05e6cb9d1dd71d

Lumen
An open-source wavetable synthesizer and VST3 with an image-to-tone engine. You can drop an image into it and generate playable wavetables and sounds from the image.

https://github.com/pixelsncodes/lumen

Lumena
A C++ image-to-MIDI engine that analyzes an image and turns its visual properties into a reproducible melody.

https://github.com/pixelsncodes/lumena

FREE Download, no account, no email: https://www.kaziahmed.net/lumen

reddit.com
u/Time-Ad-7720 — 1 month ago
▲ 132 r/VibeCodeCamp+1 crossposts

Share your vibecoded OPEN SOURCE projects, and I will check it out

I keep seeing vibe coding framed as a way to ship another SaaS product as quickly as possible.

Wrap an API, add a dashboard, create three pricing tiers, and start charging a monthly subscription.

There is nothing wrong with making money from software. Developers need to eat, infrastructure costs money, and sustainable projects need funding. But I think we are overlooking a much bigger opportunity.

Vibe coding has made software development accessible to far more people. We could use that accessibility to build open-source tools that people can inspect, modify, learn from, self-host, and continue using without another monthly payment.

With subscriptions creeping into nearly every part of technology, I want to focus more of my own work on software that people can actually own.

u/Conscious-Drawer-364 — 30 days ago
▲ 38 r/JUCE+4 crossposts

I built a wave table Synthesizer that turns images into Music [FREE+Open source+NO AI Backend]

u/Time-Ad-7720 — 30 days ago
▲ 85 r/maschine+2 crossposts

I built a FREE Synth plugin for Maschine that turns images into sounds

u/Time-Ad-7720 — 1 month ago

I built a free, open-source synth that turns any image into a playable instrument (no AI backend or API, just DSP)

u/Time-Ad-7720 — 1 month ago

I built a free, open-source synth that turns any image into a playable instrument (no AI, just DSP)

Been working on a VST3 called Lumen. It's completely free and open source. The core idea: drag any image onto it and it becomes a playable instrument. Everything runs locally. No AI, no cloud, just math.

Download: https://www.kaziahmed.net/lumen
Repo: https://github.com/pixelsncodes/lumen/

How the image engine works:

Scan mode: converts the image to grayscale, samples 64 horizontal rows, and each row's brightness curve becomes one wavetable frame (resampled to 2048 points, DC-corrected, normalized). So morphing the wavetable position literally travels down the photo, and a scanline animates across the image in sync. A smooth gradient gives you evolving pad tones. A brick wall gives you buzzy harmonic combs.

Spectral mode: treats the image as a spectrogram instead. Vertical position maps to frequency (log-scaled, ~30 Hz to 16 kHz), brightness drives the amplitude of sine partials in an additive resynthesis bank. Basically the ANS synthesizer trick from the 1950s.

Chroma mapping: reads the color statistics and sets the patch around the waveform. Mean hue picks the filter type, saturation drives resonance and unison detune, brightness sets envelope attack, edge density (Sobel filter) adds drive and noise, hue variance sets LFO depth. The image designs the whole patch, not just the oscillator.

Beyond the image stuff it's a normal wavetable synth: 2 osc + sub/noise, SVF filter (audio-rate-modulation safe), 3 env / 3 LFO / mod matrix, drive-chorus-delay-reverb, 4 macros. C++20, JUCE 8. Presets store the generated wavetable data itself (not a file path), so projects recall bit-exactly even if the original photo is gone.

Second image is the "lens" panel with a bunch of test images. Every one of those thumbnails sounds completely different. Fourth image is the signal flow if you're curious about the architecture.

Happy to go deep on any of the DSP in the comments. Also open to feature ideas. MPE and a wavetable editor are on the maybe-list for v2.

u/Time-Ad-7720 — 2 months ago