I made a tool to go from Script to AI Short Film with zero manual prompting - a complete product walkthrough

Hi everyone - I made a post about a month ago here in this subreddit showcasing a tool I've been building for the last several months for AI Filmmakers who want a seamless way to go from script to AI Short Film through a single UI with zero manual prompting required.

I've had a number of questions since, so I made a new short film, recorded the entire process and put together this video as a walkthrough for how the tool actually works.

As someone who's been making AI Films for over a year, it takes considerable time and technique to bring longer-form narrative projects to life in a way that maintains continuity. So I built the tool I wish I had from the very start.

Happy to answer any questions any of you may have about the tool, how it works, etc.

And if you want to try Kimeric for yourself, you can check out the website here: https://kimeric.ai/

u/BradClarkAI — 11 days ago
▲ 63 r/AI_Short_Films+2 crossposts

Second Earth - My Most Ambitious Sci-Fi AI Short Film Ever

Made with Seedance and Kling.

u/BradClarkAI — 12 days ago

Now that Seedance 2.5 is out, choosing the right model for the job will increasingly become an important part of your workflow.

More of a food for thought post - but something that I've been thinking about as I've seen many of the Seedance 2.5 generations coming across my timeline.

For quite some time - and I've been guilty of this, too - I would just default to using the "best" model available. And as a new model came out, I'd quickly adapt to using it.

For example, I was using Veo for quite some time when they were clearly the category leader. But things have changed over the last 6-8 months. Kling quickly became, in my opinion, the best model as of the start of 2026.

Then Seedance 2.0 was released, dethroning Kling - though at a price that was 4-5x the per-second Kling price.

And now that Seedance 2.5 is released (which, although I can't confirm firsthand, I believe is close to ~2x the price of Seedance 2.0 at a per-second basis (someone fact check me if you have firsthand knowledge)), there is potentially a ~10x cost gap between what you might be able to generate from Kling's top-tier model and what Seedance 2.5 can deliver.

Once it's released in the US, I'll be doing some stress-testing to see how it holds up vs. some of the Seedance 2.0 generations I've done, but more just a food for thought post.

This isn't the last of the "next gen" models that we'll see - so in the same way that you might use a frontier LLM (i.e. Fable) for certain high-leverage tasks, while downgrading execution to Opus, Sonnet or even Haiku, I'd recommend finding an equivalent for your generative content workflow.

If money is no object, then this probably isn't relevant to you - but if you're trying to generate AI Films while being conscious of budget, it's probably worth spending some time deciding:

  1. What is your "frontier" model stack?
  2. What is your "day to day" model?
  3. When is the quality gap big enough to justify frontier usage vs. day to day usage?

I'm continuing to ask myself these questions whenever a new model comes out. May be worth the same for you.

Curious how you're all approaching this in your workflows.

reddit.com
u/BradClarkAI — 19 days ago
▲ 6 r/AIToolsAndTips+1 crossposts

Built a AI Filmmaking tool - would love any website feedback

Hi everyone,

My name is Brad and I'm a solo founder. I released an AI Filmmaking tool about two weeks ago and would love any website feedback from the community.

I've gone through a couple rewrites already of the content on the page. I previously was focusing on more problem-solution value propositions (i.e. AI Video has a continuity problem, keeping track of tons of reference images and such is a huge pain, etc.) but got some previous feedback that it wasn't clear what exactly the product was.

So I did a rewrite focusing primarily on the core value prop immediately below the field (the input surface - rather than other image/video generators where you start with a single prompt, with this tool, you start with the entire script / short film, etc.) before pivoting into various USPs or product features.

I don't want to over-tweak it, but I would love any feedback overall. And beat it up if you have real critiques of it - I'd rather get the hard feedback now so I can address any obvious issues on the website.

Link: https://kimeric.ai/

Thanks in advance.

u/BradClarkAI — 22 days ago
▲ 13 r/aifilmmaking+1 crossposts

Everything I've Learned After 1 Year of AI Filmmaking

I've been making AI Films for roughly a year and have learned a LOT about what to do and, more importantly, what not to do if you're trying to create compelling, long-form projects at scale.

One thing has become increasingly clear:

The hardest part is not generating an impressive image or clip.

The hard part is building a production in which every new generation remembers the decisions the film has already made.

  • Who is this character?
  • What are they wearing right now?
  • Where are they standing?
  • What happened in the previous shot?
  • Which direction are they looking?
  • What object are they holding?
  • What is supposed to change during this clip?

When those answers exist only inside a growing collection of prompts, the production becomes increasingly difficult to control.

So I wrote down the full workflow I now use, from screenplay through final render.

This is not the only valid way to make an AI film, and different projects will require different levels of preparation. But this general order has helped me reduce wasted generations and make longer projects feel more like films instead of collections of loosely related clips.

You can do all of this manually with documents, folders, spreadsheets, image tools, and video tools.

The overall order

My basic production order is:

>Story → Breakdown → Visual canon → Scene construction → Shot plan → Storyboards → Video → Edit → Finish

The key principle is simple:

Make the important decisions while they are still cheap to change.

A script is cheap to change.

A shot list is cheap to change.

A character reference is relatively cheap to change.

A storyboard frame is cheaper to change than a video render.

A finished video clip is one of the most expensive places to discover that the character, wardrobe, composition, or scene was wrong.

1. Write the film before generating it

You do not necessarily need a traditional 100-page screenplay.

But before generating final imagery, I want to know:

  • Who is the story about?
  • What do they want?
  • What prevents them from getting it?
  • What changes?
  • Where does the story turn?
  • How does it end?
  • What should the audience feel?

A common early workflow is to generate a cool shot, generate another cool shot, and gradually search for a story that connects them.

That can work for experimental work, but it becomes difficult when the goal is a coherent narrative.

The image and video models should not be responsible for discovering the story while they are also being asked to visualize it.

The clearer the story is, the easier every downstream decision becomes.

2. Turn the script into a production plan

A screenplay is written to communicate a story.

It is not automatically organized around everything an AI production needs.

For each scene, I extract:

  • Characters present
  • Location
  • Important props
  • Wardrobe and physical condition
  • Time and weather
  • The emotional change
  • The information the audience needs
  • The moments that require individual shots

This is also where I separate a scene from a shot.

For example:

>Hob realizes that the rider approaching the crossroads is not his friend.

That may become several separate visual jobs:

  1. Establish Hob waiting at the crossroads.
  2. Reveal the rider approaching.
  3. Show Hob trying to identify the figure.
  4. Show recognition turning into fear.
  5. Show his hand tightening around the sword.

The script gives me the dramatic moment.

The production plan defines the images the edit will need to communicate it.

3. Build a visual canon

Before generating final shots, I create approved reference material for anything the audience will need to recognize again.

I think of this as the film’s visual canon:

>The approved version of the recurring people, objects, places, and visual rules that belong to the film.

That usually includes:

  • Main and supporting characters
  • Character bodies and silhouettes
  • Recurring wardrobe
  • Important story-state variants
  • Hero props
  • Recurring locations
  • Vehicles or creatures
  • Architecture and world design
  • Color, lighting, and atmospheric rules

The goal is not to prevent the film from changing.

The goal is to make changes intentional.

A character can begin clean, become soaked in the rain, get injured, and later carry a sword and shield. Those are all legitimate versions.

Continuity means that the correct version appears at the correct point in the story.

4. Build characters as systems, not single images

One attractive portrait is rarely enough to make a character production-ready.

The reference stack I have found useful is:

Neutral face

A controlled face image that establishes:

  • Facial structure
  • Apparent age
  • Hair
  • Skin texture
  • Eye, nose, and mouth proportions
  • Overall identity

This should be relatively free of dramatic posing, extreme lighting, and scene-specific distractions.

It is the identity anchor.

Neutral body

A full-body reference that establishes:

  • Build
  • Proportions
  • Height impression
  • Silhouette
  • Posture
  • Physical presence

A face reference may work well for portraits but provide very little guidance when the character appears in a wide shot.

Wardrobe face and body

Once the identity is stable, I establish the character in the wardrobe used by the film.

A full wardrobe-body reference is particularly useful because clothing changes:

  • Silhouette
  • Apparent body shape
  • Material behavior
  • Color distribution
  • The way the character reads at a distance

State variants

Then I create variants for recurring story conditions:

  • Wet
  • Muddy
  • Bloody
  • Injured
  • Burned
  • Exhausted
  • Weather-exposed
  • Damaged clothing

Costume or loadout variants

If the character gains armor, bag, shield, tool, or other major equipment, I treat that as another deliberate reference.

A new loadout can significantly change the character’s silhouette. Leaving it to be rediscovered in every shot creates unnecessary variation.

The result is not one image of a character.

It is a small system of approved character states that can be called upon at different points in the film.

5. Treat important props and locations the same way

Character continuity gets most of the attention, but props and locations can drift just as badly.

If an object matters to the story, I try to establish:

  • Shape
  • Materials
  • Scale
  • Color
  • Condition
  • Ownership
  • How it is held or worn

A sword that is straight in one shot and curved in the next may seem like a small generation issue, but it becomes distracting when the object is narratively important.

The same principle applies to recurring locations.

Instead of prompting “a forest” repeatedly, I define the specific forest crossroads:

  • The wooden signpost
  • The road configuration
  • The tree density
  • The ground materials
  • The weather
  • The lighting direction
  • The nearby landmarks

The location reference does not guarantee identical geometry in every generation.

It gives each shot a shared source rather than asking the model to invent a new forest from scratch.

6. Establish the whole scene before generating isolated shots

Once the characters and location exist, I build a wider scene anchor.

This does not always need to be a final production shot.

Its job is to answer:

  • Who is present?
  • Where is each person positioned?
  • What surrounds them?
  • Which props are visible?
  • What are the relative heights and distances?
  • Which direction is each person facing?
  • What does the scene’s lighting and atmosphere look like?

I think of this as constructing the stage before moving the camera around it.

Starting with isolated close-ups creates a common problem: every frame may look good individually, but the images cannot logically coexist in one physical scene.

The wide anchor gives later coverage something to inherit.

7. Give every shot a job

I try not to generate shot variety merely for visual variety.

Each shot should contribute something to the edit.

A simple dialogue or confrontation might include:

  • Wide: Establish the environment and positions.
  • Two-shot or OTS: Show the relationship.
  • Reverse OTS: Show the other side of the exchange.
  • Medium close-up: Read performance and body language.
  • Close-up: Land the emotional change.
  • Insert: Reveal an important object or action.

That does not mean every scene needs all six.

It means the shot list should come from what the audience needs to see.

One of the easiest ways to waste generations is to create ten attractive versions of essentially the same dramatic information.

Before generating a shot, I ask:

>What can the audience understand from this shot that they could not understand as well from the existing coverage?

If I cannot answer that, the edit may not need it.

8. Structure prompts around production information

I have had better results when I treat prompts as production instructions rather than collections of cinematic adjectives.

The general structure I use is:

  1. Story beat
  2. Subject identity
  3. Current wardrobe and state
  4. Location
  5. Camera
  6. Primary action
  7. Lighting and atmosphere
  8. Continuity requirements

Story beat

Start with what happens and why the shot exists.

>Hob realizes the approaching rider is not who he expected.

This gives everything else a dramatic purpose.

Subject identity

Specify the person being photographed.

>Hob, a weathered laborer in his late 40s with cropped dark hair and a broad, tired face.

When reference images are available, they should carry most of the identity burden. The text should reinforce rather than fight them.

Current state

Describe the exact version of the character at this point in the story.

>Rain-soaked work clothes, mud covering his boots, a rusted sword in his right hand, and a battered shield on his left arm.

Not merely “Hob.”

This version of Hob.

Location

Place the shot in the established environment.

>At the muddy forest crossroads during golden hour, with the old signpost behind him and rain hanging in the air.

Camera

Choose one clear visual approach.

>Medium close-up at eye level, slightly off-center, with shallow depth of field.

Primary action

Give the shot a dominant readable movement.

>He slowly raises the sword as recognition turns into fear.

That is generally easier to control than:

>He turns, walks forward, raises the sword, looks behind him, shouts, and falls to his knees.

Several things can happen in a shot, but one should usually be visually dominant.

Atmosphere

Add cinematic texture after the shot itself is clear.

>Warm backlight cuts through the rain and mist while the foreground remains cool and subdued.

Lighting and mood should support the story moment rather than substitute for one.

Continuity requirements

Finish by naming the things that must not be reinvented.

>Preserve Hob’s established identity, wet wardrobe state, sword, shield, screen direction, and position at the crossroads.

9. Separate what stays fixed from what changes

This may be the single most useful prompting principle I have found.

Every shot contains two categories of information.

Keep fixed

  • Character identity
  • Current wardrobe
  • Current physical state
  • Important props
  • Established location
  • Screen direction
  • Story facts

Change for this shot

  • Expression
  • Action
  • Shot size
  • Camera position
  • Camera movement
  • Emotional emphasis
  • Selective lighting emphasis

The new prompt should describe the delta (i.e. what changes) without unnecessarily reopening every decision the production has already approved.

Every time a prompt fully reinvents the character, costume, location, camera, weather, and action simultaneously, it gives the model more opportunities to drift.

10. Solve the still image before paying to animate it

Before video generation, I want an approved storyboard or start frame.

I check:

  • Is this the correct character?
  • Is this the correct wardrobe and story state?
  • Is the location recognizable?
  • Are the props right?
  • Does the composition communicate the beat?
  • Is the eyeline plausible?
  • Does it match the surrounding shots?
  • Would I approve this image even if it never moved?

Animation rarely rescues a fundamentally incorrect frame.

Video generation is a relatively expensive place to discover that the character has the wrong outfit, the weapon is missing, or the composition never worked.

The still does not have to be perfect.

It needs to prove that the shot is ready for motion.

11. Before rendering video, run a preflight check

This is the checklist I use before spending money or credits on a clip.

Is the story beat clear?

Can I explain what changes between the beginning and end of the shot?

Is this the correct character state?

Identity, body, outfit, condition, damage, and held props should match the exact moment in the story.

Does the location match?

Check recurring architecture, landmarks, weather, time of day, and environmental condition.

Does the frame already work?

The subject, framing, gaze, props, and background should communicate the shot before motion is added.

Will it cut with the neighboring shots?

Check:

  • Screen direction
  • Character position
  • Wardrobe state
  • Prop placement
  • Lighting direction
  • Movement direction
  • Emotional continuity

Is there one primary action?

The model should understand what matters most.

Do I know how the shot begins and ends?

The first frame receives the previous cut.

The last frame prepares the next one.

For example:

Entry state: Sword lowered, looking toward the distant rider.

Exit state: Sword raised, eyes fixed on the approaching threat.

Does every reference have a job?

I try to give references explicit roles:

  • Identity reference
  • Wardrobe/state reference
  • Location reference
  • Scene anchor
  • Previous-shot reference

More references are not automatically better. A smaller, purposeful set can be more useful than an undifferentiated pile.

Does the edit actually need this shot?

This final question has probably saved me the most unnecessary renders.

12. Animate the approved plan

Once the frame and references are correct, the video prompt should focus primarily on motion:

  • Character action
  • Emotional transition
  • Camera movement
  • Environmental movement
  • Entry state
  • Exit state

The video model should not be asked to simultaneously design the character, invent the location, determine the composition, discover the story beat, and choreograph the performance.

The more of that work completed upstream, the narrower and clearer the animation problem becomes.

13. Build the rough cut before polishing everything

It is tempting to perfect each clip before placing it into the edit.

I have found it more useful to assemble a rough cut relatively early.

The rough cut reveals:

  • Which shots are actually usable
  • Which shots are redundant
  • Where pacing drags
  • Whether the emotional progression reads
  • Which clips need to be shortened
  • Where audio can solve a visual gap
  • Which missing shots are truly necessary
  • Which rerenders are worth paying for

A clip that looks impressive alone may be the wrong clip for the sequence.

A less spectacular take may cut better because the gaze, position, movement, and emotion connect correctly.

The film is the sequence, not the collection of individual generations.

14. Polish only what survives the edit

Once the structure works, I move into:

  • Dialogue and voice
  • Sound design
  • Music
  • Timing refinements
  • Color
  • Cleanup
  • Upscaling
  • Final export

This is another form of cost control.

There is little value in upscaling, cleaning, and heavily polishing a shot that will later be removed or reduced to one second.

The order I try to preserve is:

  • Story first.
  • Canon second.
  • Scenes third.
  • Shots fourth.
  • Motion fifth.
  • Edit sixth.
  • Polish last

The core idea

The tools and models will keep changing.

The central production problem is more stable:

How do you stop every new generation from forgetting the film you have already built?

My answer is to create a chain of inheritance:

  • The script defines the moment.
  • The visual canon defines the world.
  • The character references define who appears.
  • The state references define which version appears.
  • The scene anchor defines where everyone is.
  • The shot plan defines what the edit needs.
  • The storyboard defines the frame.
  • The video prompt defines the motion.
  • The edit determines what survives.

Each stage narrows the next problem.

That does not eliminate iteration. It makes the iteration more diagnosable.

When something fails, you can ask:

  • Is the story unclear?
  • Is the reference wrong?
  • Is the scene staging wrong?
  • Is the composition wrong?
  • Is the motion instruction wrong?
  • Or did the model simply fail to execute a good plan?

That is much more useful than repeatedly rerolling and hoping the entire production aligns by chance.

Disclosure: I've built a local production tool called Kimeric around this general workflow, so I obviously think about these problems through a systems lens. There are no links or pitch here, and everything above can be done manually with whatever tools you already use. I’m mainly sharing it because I wish I had a clearer end-to-end framework when I started.

Happy to answer any questions or go deeper here if there are topics or techniques folks want to learn more about.

reddit.com
u/BradClarkAI — 24 days ago
▲ 62 r/aivideo+1 crossposts

The Farm (AI Short Film - Life Lessons From an Intergalactic Alien)

Made with Kling. Voice double generated with ElevenLabs.

u/BradClarkAI — 25 days ago

I Made A Fake Superhero Movie Trailer

Made with Kling. Fun little side project, wanted to test and see if I could put together a superhero-style trailer. Was pretty happy with how it came together!

u/BradClarkAI — 26 days ago

Claude Design is UNREAL for Motion Graphics (had a "this is the future" moment)

Not sure if I'm late to the party, but I've been testing out Claude Design primarily for static images over the last few weeks.

I recently started a company and wanted to produce some motion graphic-style teaser promo. Had thought about trying to use various AI Video Generators to generate something or dust off After Effects, but figured I'd try Claude Design to see if it was even possible.

I was blown away. I had no idea that Claude Design was capable of motion graphics in any capacity, let alone something like what I was able to put together.

I figured I'd walk through my exact workflow in case anyone in this community is interested in trying to create something similar.

I first created a Design System before actually creating the design itself. Candidly, I had Fable create a Design Handbook, logo, etc - essentially the entire brand - about a month ago. Really impressive document in and of itself, but really turbocharges the use-cases for a static handbook when it can be used as an input to create a Design System.

I then ran the general idea of what I was looking for through a couple LLMs (I was testing the same prompt request through both ChatGPT and Claude, but tested with the ChatGPT output first and got this so probably don't need to do the Claude version at this point) and had ChatGPT write out a very detailed prompt that I used as the Claude Design input.

Truncating the prompt slightly, but this was generally the prompt I put into ChatGPT to have it write out the prompt:

Attached is the product manual for Kimeric. This is the sum of how the entire product works. I want to use Claude Design to try and make a motion graphic. Think like a 30-60 second extremely polished motion graphic that explains exactly how it works, optimized for consumption on social media (TikTok, Reels, Shorts, etc.) - meaning, it has VERY fast motion, movement, graphics, effects, transitions, etc. Can you review this product manual to understand how it works and then create a copy+paste brief for exactly what I can put into Claude Design for this?

(For context - I had previously sent a Claude Code request to create the product manual for my product which I attached as the input so that ChatGPT had proper understanding of how exactly the product works (side note: I ran this exercise on Fable Ultracode and it produced like a 300 page document which is pretty wild in and of itself))

I then took the prompt it output and pasted it verbatim into Claude Design. It ran for about 30 minutes and gave me a v1 of the design. I had like 20-or-so images I wanted to include in the animation itself, but there's a cap on the input images for a Claude Design request - though there doesn't appear to be one for any in-chat messages within Claude Design.

So I then followed up with the Claude Design chat and simply pasted in a bunch of images from my local computer that I wanted used in the actual design and asked it to incorporate these images into the design:

https://preview.redd.it/4neau0v37weh1.png?width=999&format=png&auto=webp&s=99fd39654339e1901a437cc4311211eb044388f1

It ran for another 20 minutes or so and got the baseline video generated.

The only real technical stumbling block I had was exporting the video. I don't think Claude Design is really made to export out these extremely detailed motion graphics (or maybe my computer just isn't that great), but Chrome kept crashing due to RAM issues after about 2-3 minutes of attempting to export.

Workaround I found - I downloaded it as an HTML file out of Claude Design which worked perfectly fine, loaded that up in a fresh window (so it had 100% smooth playback) and simply did a screen record via Snipping Tool which got me exactly what I needed.

From there, I then went back to ChatGPT, gave it the runtime of the final product and asked it to write me a voiceover script. Made a few edits to the voiceover script and ran it through Eleven Labs a few times on a few voice models until I found one I liked.

Lastly, exported the voiceover out, moved the entire project into Premiere, sourced music and a TON of SFX from Artlist.io (I quite literally just went to the SFX tab of the Adobe Premiere Artlist plugin, sorted by Popular and downloaded more or less everything in the top 25 sound effects and just placed them everywhere), added captions, did some editing to tighten up the pacing and exported.

End to end took maybe 3 hours. This was something unimaginable just a few months ago, and now it's doable in just an evening.

This was my first "wow this is the future" moment I've had in a while (I had one when Fable first came out, but that + this are really the only two big ones I've had this year), so figured I'd do a quick write-up on the workflow here in this community if helpful.

Happy to answer any specific questions about the process/workflow if helpful!

u/BradClarkAI — 28 days ago

The Titan Hunter (Cinematic AI Short Film)

Made with Kling, was a fun project to work on. Learned a few technical things in the process of making this one which I've applied to newer projects but still had a lot of fun working on this!

u/BradClarkAI — 29 days ago

THE NPC (if you've ever played an MMO, you'll like this)

If you've ever played any MMO (i.e. WoW), you'll probably enjoy this :)

Made with mostly Kling & a bit of Seedance.

u/BradClarkAI — 1 month ago

After a year of AI filmmaking the hard way, I built the tool I wish I had from the start: Go from script to AI short film, all through a single interface - with continuity baked in.

For about a year I've been making AI short films the way most of us do: hand-writing hundreds of prompts, building character reference libraries by hand, babysitting consistency across shots, and cutting it all together in Premiere. The generating was never the hard part. It was trying to make dozens of individual prompts *feel* like a single, cohesive project.

I couldn't find a solution, so I built one. It's called Kimeric, and the Beta went live today. And I think it'll be incredibly valuable for this community, so I wanted to share some details in the event any of you would like to try it out!

**What it is (and isn't):** It's not a model and it doesn't generate anything itself. It's a Windows desktop app that orchestrates the models and LLMs a lot of us already use:

  • Anthropic (script breakdown + prompt authoring)
  • Gemini and GPT-Image (image generation)
  • Kling and Seedance 2.0 (video)
  • Topaz (upscaling)

All using your own per-provider API keys rather than a centralized generation tool. It writes the prompts, sequences and queues the renders, tracks the spend, and holds everything for your approval.

TL;DR - Input a script and work through a series of UI menus to ultimately create a finished AI Short Film.

The biggest differentiator to keep in mind for this tool vs. the majority of other AI Generators on the market is the input surface. Most tools have you input a prompt. With Kimeric, the input surface is the script itself. The prompts are created automatically as derivatives from the script, allowing you to focus on writing and build the project rather than managing a series of prompts.

The Pipeline:

- Paste a screenplay, hit "Roll camera." The breakdown comes back: cast, locations, props, a director's plan, per-scene shot lists with dialogue assigned line by line. You review the plan before anything renders, and can choose a general "style" which dictates some actual prompting techniques under the hood as well as dynamically auto-routing for certain models for specific tasks (ex. OpenAI's model is better at certain animated styles vs. Nano Banana, so this routing auto-applies as the default for certain styles).

https://preview.redd.it/8zg6wqjfruch1.png?width=1427&format=png&auto=webp&s=9c782f3927ac4fa72d82cf82ca325b74d7325757

Before you actually generate anything, you can review the script you input and get an estimate for how much it'll cost roughly to "ingest" the script which is where all of the "brain" of the tool goes to work.

https://preview.redd.it/p8v92s6hruch1.png?width=974&format=png&auto=webp&s=6bef386ed3bc019099648842e7d73ec003434ad8

And once you send the script, a very long sequence of backend computation kicks off, translating the entire project into the format needed to actually create an end-to-end project (this can take a while; I had one script for a 8-10 minute short film take about 45 minutes to ingest. This is just because there's a lot of computation and inference happening). It can handle actual, full-length production scripts. This would likely come with processing that spans several hours, but again, that's simply due to the size of the computation.

https://preview.redd.it/n82pwmhiruch1.png?width=929&format=png&auto=webp&s=1bca621e00bef5b73e4e07c92fab8e15c2f615a7

And then, you get a budget estimate for how much (approximately) the project will cost to generate. At this point, it's primarily a planning "calculator" if you will - just so you can align the project quality to any budget constraints you have. No actual spend dispatches until you generate later - this is purely informational so you can make cost-based decisions at the start (as opposed to a surprise cost later).

https://preview.redd.it/i7bc02gjruch1.png?width=869&format=png&auto=webp&s=01482437d2f85057915177e40924e3a26f8dc4ff

You then view a breakdown of all of the identified "pieces" of the project: characters, locations, scenes, planned shots, etc. - all primarily at a high level to make sure nothing is missed. Typically more of a rubber-stamp phase, but if there happen to be any items missing, this is the step where you can make any high-level revisions. Most of the time, however, you can just continue.

https://preview.redd.it/o4b37e7kruch1.png?width=2557&format=png&auto=webp&s=48e2a82b645ed8b28f4cbdbea0a7fc66364f925a

- Every character locks canon first — a studio face (with an automated AI advisory likeness screening for real-person resemblance) and a costume-neutral turnaround — before any scene renders. Same for locations, worlds, props. It's the reference-library grind, automated.

https://preview.redd.it/9vc4znflruch1.png?width=2541&format=png&auto=webp&s=159f68049046e7aa37ca1e1b390521ca0cf03090

And for characters specifically, it's broken out into two phases: the "Neutral Base" (i.e. who the character is; sans any wardrobe for the project (see above) and then any actual "in costume" variants of that character. This allows for multiple wardrobes for a specific character over the course of any given project, where each wardrobe is either seeded directly from the parent "neutral base" or as a horizontal derivative (i.e. if a character has armor that gets damaged, the "damaged" variant is automatically seeded with the "clean base" variant). All of this happens automatically under the hood.)

https://preview.redd.it/nrlt198mruch1.png?width=2556&format=png&auto=webp&s=b30df95fe0fd8ef8178c5f9e0b2d4550fab0c9d9

- Scenes get actual coverage: an establishing wide, then an OTS pair where the reverse angle generates from the *approved* first angle, then per-character MCU/CUs chained down from there. The 180° rule, held by reference chaining instead of luck.

For example, here's one "OTS_1" image:

https://preview.redd.it/sd29lzxmruch1.png?width=2552&format=png&auto=webp&s=277f0b931338aecea20cb7aacc8674d4aac7cbd0

And here's the companion "OTS_2" image:

https://preview.redd.it/m52g4dlnruch1.png?width=2556&format=png&auto=webp&s=ca45cfc2a4fd513361165f231e6e6d747c4e39b9

Worth emphasizing - These are the *most* critical images in your production pipeline to get right, as many other scenes seed off of these.

So if you're going to use some of the "Regeneration Buffer" you planned for, this is the most critical space to use it. If you have continuity errors or issues in either companion OTS images, these will present in many other aspects of a project, so really take your time with these and make sure that they feel like the same space.

From here, you create a library of Medium Close Up & Close Up images seeded directly from the parent Over-The-Shoulder images.

This creates a rich, continuity-adhering library of image assets to actually use in generative AI video production.

https://preview.redd.it/ad2nznforuch1.png?width=2556&format=png&auto=webp&s=71f687083835ae73d77178332470b43f8190b2a9

Once you've created your core library of production assets, you then transition into building out your storyboard.

This takes the plan created during script ingestion + the assets you created in the previous phase and maps them out chronologically.

Here, images are auto-assigned based on a series of underlying logic. If you have dialogue from certain characters, either the OTS, the MCU or the CU image will be dynamically selected based on a cinematic logic layer and assigned to the character speaking for any given frame.

And for extended dialogue from a specific character, this will be broken up into multiple individual prompts to assist with overall quality

(from our testing, the more text you try to fit into the same prompt, quality and lip sync can degrade - but breaking that same dialogue up into multiple individual prompts can greatly improve the quality)

https://preview.redd.it/gvhmjr7pruch1.png?width=2558&format=png&auto=webp&s=bcd0cde96011c6e4ee21fe07cc2d51207e6654ef

Once you've created and approved the entire Storyboard, a single button allows you to review all text prompts for all videos planned for the totality of your project.

The prompts are already written. You just say "go"

From there, they are auto-dispatched to the models.

Depending on the length of your project, this may take a while - and that's by design. Press "generate" and take a break for a bit.

https://preview.redd.it/eor7yc1qruch1.png?width=2554&format=png&auto=webp&s=ded5898f7d70894a2b68b5b9e165f4d7ee7f4162

Across all menu screens, you'll see a series of control buttons. Here, you can either approve the asset, edit the asset, regenerate or iterate.

  • Approving says the image/video is good to go.
  • Editing allows for subtle adjustments.
  • Regenerate re-runs the prompt (i.e. get a new version to see if the results are better/worse)
  • Iterate is essentially a stronger version of "Edit" - rather than trying to make adjustments to the previous take, it will strongly re-work the actual prompt itself based on your feedback and re-generate a new take.

https://preview.redd.it/begna6wqruch1.png?width=1379&format=png&auto=webp&s=0f348e50be962c33f59bbd3923301cd84ed05b20

You'll also see two columns below any given asset:

Refs - These are the actual reference images used as generative inputs. These are auto-assigned, and you can manually add, remove or replace any of these as you see fit.

Takes - If you regenerate, you can see all of your takes here and hot-swap to other variants. Meaning, you're never locked in to a specific take. If one is close but you want to see if you can fine-tune it, you're free to regenerate a few times, review all takes (this works for images + videos) and ultimately approve whichever one is best for the vision you had in mind for the project.

https://preview.redd.it/k561i1nrruch1.png?width=1035&format=png&auto=webp&s=b3f3cd69f05d2e508fd496756e90485e9d9abbdc

And finally, once you've reviewed and approved all footage, you have another single-press button to dispatch all video to be upscaled if you'd like.

Generation can happen natively at 720p, 1080p or 4k depending on your settings.

4k native footage won't trigger the upscale workflow, but 720 and 1080p native will give you the option to upscale if you'd like.

Similar to all other workflow phases, if you decide to upscale, press it once and let it run for a few hours (upscaling is quite time-consuming; can take 20-30 minutes for a single 15 second clip - though some of these can run concurrently).

https://preview.redd.it/kg9aiwcsruch1.png?width=2555&format=png&auto=webp&s=2b9dc0cce565b4f826890e2bdd5ca831db0d7c9a

Once you're done, the last step is simply to "export" the clips.

All this is doing is taking the final clips you've approved and making duplicate copies on your local computer that are pre-named chronologically.

This makes it significantly easier to edit/compose.

https://preview.redd.it/gti9mf0truch1.png?width=970&format=png&auto=webp&s=2258103233954db1c6c19aacda762254415a673d

This entire project is the culmination of nearly a year of a LOT of testing to understand which techniques do/don't produce good results at scale and then working to systematize them into a tool with an input surface of the script itself.

Again, the Beta is live as of today. Currently Windows-only and US-only, though both of those I'm planning to expand beyond in the coming weeks over the course of the Beta.

The long-term goal is to also support centralized generation rather than supporting only a BYOK model, though there's no immediate timeline to support that model.

If you'd like to check it out, the website is below!

Site: https://kimeric.ai

I'm a solo founder and the filmmaker this was built for, and I'll be in the comments - happy to go as deep as you want on the coverage system, the cost math, or anything else.

Feedback is a gift, so if you try it and find issues, bugs, or have a feature request, I'm still actively building and improving the tool - so feel free to share any thoughts!

reddit.com
u/BradClarkAI — 1 month ago
▲ 20 r/aivideomaking+1 crossposts

After a year of AI filmmaking the hard way, I built the tool I wish I had from the start: Go from script to AI short film, all through a single interface - with continuity baked in.

For about a year I've been making AI short films the way most of us do: hand-writing hundreds of prompts, building character reference libraries by hand, babysitting consistency across shots, and cutting it all together in Premiere. The generating was never the hard part. It was trying to make dozens of individual prompts *feel* like a single, cohesive project.

I couldn't find a solution, so I built one. It's called Kimeric, and the Beta went live today. And I think it'll be incredibly valuable for this community, so I wanted to share some details in the event any of you would like to try it out!

**What it is (and isn't):** It's not a model and it doesn't generate anything itself. It's a Windows desktop app that orchestrates the models and LLMs a lot of us already use:

  • Anthropic (script breakdown + prompt authoring)
  • Gemini and GPT-Image (image generation)
  • Kling and Seedance 2.0 (video)
  • Topaz (upscaling)

All using your own per-provider API keys rather than a centralized generation tool. It writes the prompts, sequences and queues the renders, tracks the spend, and holds everything for your approval.

TL;DR - Input a script and work through a series of UI menus to ultimately create a finished AI Short Film.

The biggest differentiator to keep in mind for this tool vs. the majority of other AI Generators on the market is the input surface. Most tools have you input a prompt. With Kimeric, the input surface is the script itself. The prompts are created automatically as derivatives from the script, allowing you to focus on writing and build the project rather than managing a series of prompts.

The Pipeline:

- Paste a screenplay, hit "Roll camera." The breakdown comes back: cast, locations, props, a director's plan, per-scene shot lists with dialogue assigned line by line. You review the plan before anything renders, and can choose a general "style" which dictates some actual prompting techniques under the hood as well as dynamically auto-routing for certain models for specific tasks (ex. OpenAI's model is better at certain animated styles vs. Nano Banana, so this routing auto-applies as the default for certain styles).

https://preview.redd.it/j6j58d9j6och1.png?width=1427&format=png&auto=webp&s=137f8495814ea5ce9467f152af57e38e99cd4359

Before you actually generate anything, you can review the script you input and get an estimate for how much it'll cost roughly to "ingest" the script which is where all of the "brain" of the tool goes to work.

https://preview.redd.it/qwggbnen6och1.png?width=974&format=png&auto=webp&s=8f98c96b57600b104168c0a01b7d94d4d5767fbb

And once you send the script, a very long sequence of backend computation kicks off, translating the entire project into the format needed to actually create an end-to-end project (this can take a while; I had one script for a 8-10 minute short film take about 45 minutes to ingest. This is just because there's a lot of computation and inference happening). It can handle actual, full-length production scripts. This would likely come with processing that spans several hours, but again, that's simply due to the size of the computation.

https://preview.redd.it/ay3ck9tv6och1.png?width=929&format=png&auto=webp&s=30d833398e70e009ec302e565c29e5f1f6cdb8ee

And then, you get a budget estimate for how much (approximately) the project will cost to generate. At this point, it's primarily a planning "calculator" if you will - just so you can align the project quality to any budget constraints you have. No actual spend dispatches until you generate later - this is purely informational so you can make cost-based decisions at the start (as opposed to a surprise cost later).

https://preview.redd.it/1doantg07och1.png?width=869&format=png&auto=webp&s=3503cd1a9eaf911564aed9c9b13faa98fd9eb5dd

You then view a breakdown of all of the identified "pieces" of the project: characters, locations, scenes, planned shots, etc. - all primarily at a high level to make sure nothing is missed. Typically more of a rubber-stamp phase, but if there happen to be any items missing, this is the step where you can make any high-level revisions. Most of the time, however, you can just continue.

https://preview.redd.it/y78s8dz77och1.png?width=2557&format=png&auto=webp&s=f34748a0f341d8d4bbc18076dde2b79baeba9d77

- Every character locks canon first — a studio face (with an automated AI advisory likeness screening for real-person resemblance) and a costume-neutral turnaround — before any scene renders. Same for locations, worlds, props. It's the reference-library grind, automated.

https://preview.redd.it/kiwi3ab97och1.png?width=2541&format=png&auto=webp&s=3e81dc32f90e838235b61f98e90a1d14d2cf5ae5

And for characters specifically, it's broken out into two phases: the "Neutral Base" (i.e. who the character is; sans any wardrobe for the project (see above) and then any actual "in costume" variants of that character. This allows for multiple wardrobes for a specific character over the course of any given project, where each wardrobe is either seeded directly from the parent "neutral base" or as a horizontal derivative (i.e. if a character has armor that gets damaged, the "damaged" variant is automatically seeded with the "clean base" variant). All of this happens automatically under the hood.)

https://preview.redd.it/rb8me36m7och1.png?width=2556&format=png&auto=webp&s=dcb60cd5f18b5494895b4bd52d61fd164b32291b

- Scenes get actual coverage: an establishing wide, then an OTS pair where the reverse angle generates from the *approved* first angle, then per-character MCU/CUs chained down from there. The 180° rule, held by reference chaining instead of luck.

For example, here's one "OTS_1" image:

https://preview.redd.it/qh8vyjwn7och1.png?width=2552&format=png&auto=webp&s=46344ebd7094665099f0ddafb2a8a009a0e91cdc

And here's the companion "OTS_2" image:

https://preview.redd.it/ymv442bp7och1.png?width=2556&format=png&auto=webp&s=29c89559d0029f52e44a812db4894ae948fd67b0

Worth emphasizing - These are the *most* critical images in your production pipeline to get right, as many other scenes seed off of these.

So if you're going to use some of the "Regeneration Buffer" you planned for, this is the most critical space to use it. If you have continuity errors or issues in either companion OTS images, these will present in many other aspects of a project, so really take your time with these and make sure that they feel like the same space.

From here, you create a library of Medium Close Up & Close Up images seeded directly from the parent Over-The-Shoulder images.

This creates a rich, continuity-adhering library of image assets to actually use in generative AI video production.

https://preview.redd.it/31loz72r7och1.png?width=2556&format=png&auto=webp&s=c10b47bb95afb35bddc7cde0126d8e78284af83a

Once you've created your core library of production assets, you then transition into building out your storyboard.

This takes the plan created during script ingestion + the assets you created in the previous phase and maps them out chronologically.

Here, images are auto-assigned based on a series of underlying logic. If you have dialogue from certain characters, either the OTS, the MCU or the CU image will be dynamically selected based on a cinematic logic layer and assigned to the character speaking for any given frame.

And for extended dialogue from a specific character, this will be broken up into multiple individual prompts to assist with overall quality

(from our testing, the more text you try to fit into the same prompt, quality and lip sync can degrade - but breaking that same dialogue up into multiple individual prompts can greatly improve the quality)

https://preview.redd.it/d3bad6lu7och1.png?width=2558&format=png&auto=webp&s=2d6fb1b7f5e719d9baed247e6699f56a0315e7ff

Once you've created and approved the entire Storyboard, a single button allows you to review all text prompts for all videos planned for the totality of your project.

The prompts are already written. You just say "go"

From there, they are auto-dispatched to the models.

Depending on the length of your project, this may take a while - and that's by design. Press "generate" and take a break for a bit.

https://preview.redd.it/fi6m94fx7och1.png?width=2554&format=png&auto=webp&s=dcaa31f9cc9f03d1b044274f9f50c3f54aab0a5c

Across all menu screens, you'll see a series of control buttons. Here, you can either approve the asset, edit the asset, regenerate or iterate.

  • Approving says the image/video is good to go.
  • Editing allows for subtle adjustments.
  • Regenerate re-runs the prompt (i.e. get a new version to see if the results are better/worse)
  • Iterate is essentially a stronger version of "Edit" - rather than trying to make adjustments to the previous take, it will strongly re-work the actual prompt itself based on your feedback and re-generate a new take.

https://preview.redd.it/eujr9jry7och1.png?width=1379&format=png&auto=webp&s=aaaef1ed9fb9ce83abf7236ce2c7501a6f30ad6b

You'll also see two columns below any given asset:

Refs - These are the actual reference images used as generative inputs. These are auto-assigned, and you can manually add, remove or replace any of these as you see fit.

Takes - If you regenerate, you can see all of your takes here and hot-swap to other variants. Meaning, you're never locked in to a specific take. If one is close but you want to see if you can fine-tune it, you're free to regenerate a few times, review all takes (this works for images + videos) and ultimately approve whichever one is best for the vision you had in mind for the project.

https://preview.redd.it/oqogeo008och1.png?width=1035&format=png&auto=webp&s=faba84dd2227b9158a9b3aaf87d6739baf32e9a3

And finally, once you've reviewed and approved all footage, you have another single-press button to dispatch all video to be upscaled if you'd like.

Generation can happen natively at 720p, 1080p or 4k depending on your settings.

4k native footage won't trigger the upscale workflow, but 720 and 1080p native will give you the option to upscale if you'd like.

Similar to all other workflow phases, if you decide to upscale, press it once and let it run for a few hours (upscaling is quite time-consuming; can take 20-30 minutes for a single 15 second clip - though some of these can run concurrently).

https://preview.redd.it/lq39urf18och1.png?width=2555&format=png&auto=webp&s=4210245ff8c72a1863f244967253a52761c01735

Once you're done, the last step is simply to "export" the clips.

All this is doing is taking the final clips you've approved and making duplicate copies on your local computer that are pre-named chronologically.

This makes it significantly easier to edit/compose.

https://preview.redd.it/ob9swdm28och1.png?width=970&format=png&auto=webp&s=5e66e930d7a29d21724eefbcc3ea3e8318f81249

This entire project is the culmination of nearly a year of a LOT of testing to understand which techniques do/don't produce good results at scale and then working to systematize them into a tool with an input surface of the script itself.

Again, the Beta is live as of today. Currently Windows-only and US-only, though both of those I'm planning to expand beyond in the coming weeks over the course of the Beta.

The long-term goal is to also support centralized generation rather than supporting only a BYOK model, though there's no immediate timeline to support that model.

If you'd like to check it out, the website is below!

Site: https://kimeric.ai

I'm a solo founder and the filmmaker this was built for, and I'll be in the comments - happy to go as deep as you want on the coverage system, the cost math, or anything else.

Feedback is a gift, so if you try it and find issues, bugs, or have a feature request, I'm still actively building and improving the tool - so feel free to share any thoughts!

reddit.com
u/BradClarkAI — 1 month ago