I built a pipeline that turns a voiceover audio file into a fully edited, publish-ready youtube video — zero human editing. This is a raw, unedited screen recording of one run.
Full disclosure: I'm the developer and founder.**
I've been building DVOcean — an engine where you drop in a voiceover audio file (or just a topic) and the pipeline does the rest: fetches contextually relevant footage for every scene, cuts it to the audio, and bakes in captions automatically. The goal was simple but technically brutal: zero human editing.
This screen recording is the entire flow, from dropping in the audio to the raw, untouched output.
I'd love blunt feedback:
1. How does the visual pacing and caption timing look to you?
2. If you run a faceless channel or video podcast, what's missing that would make this a daily tool?
(There's a free 1-minute trial if anyone wants to stress-test it with their own audio — link in comments so I'm not spamming the post.)