From words to pictures
You have a script and a one-page plan. Now something has to appear on screen. This is the part people expect to be magic, and it is genuinely impressive, as long as you know what you are actually buying.
Heads up: this lesson links to InVideo through an affiliate link, so LearnAI earns a commission if you subscribe. It costs you nothing extra, and it is what keeps these courses free. We tell you where these tools fall short too, including who should not use them.
In plain English
- B-roll:
- Background footage that plays while someone is talking, so the video is not just one static shot.
- Text to video:
- You describe a shot in writing and the tool produces moving pictures from that description.
- Stock footage:
- Professionally filmed clips from a library, licensed for reuse. Not generated, just selected.
- Continuity:
- Things staying consistent between shots. The same jacket, the same mug, the same face, in the same room.
- Render:
- The tool building your finished video file. It takes a few minutes and usually costs part of your allowance.
Two different machines wearing one coat
Tools that turn a script into a video generally do one of two things, and often both.
The first is assembly. The tool reads your script, breaks it into scenes, picks matching clips from a big stock library, adds captions, music and a voiceover, and hands you an edit you can adjust. Nothing is invented. Everything on screen was filmed by a real crew somewhere.
The second is generation. The tool invents frames that never existed. This is the headline-grabbing one, and it is the one with the strange fingers.
InVideo sits mostly in the first camp with generative pieces bolted on, which is exactly why it suits the person this course is for. You paste a script, you get a rough cut with scenes, captions and a voice, and then you fix the bits that are wrong by talking to it. Free plans generally exist with watermarks and limits, paid tiers remove them, and the specifics move around often enough that you should read the pricing page yourself rather than trust any article, including this one.
What AI video is genuinely good at
B-roll and atmosphere. Rain on a window, a city street at dusk, hands typing. Nobody needs it to be a specific window.
Simple explainers. A concept, some captions, calm background footage, a clear voice. This used to cost a freelancer's day rate.
Social clips at volume. Ten variations of the same 30-second idea, so you can post consistently without a shoot every week.
Turning writing you already have into video. A blog post, a FAQ answer, a product description. The words exist. The video is a format change.
What it is still poor at
Your actual brand. It will not reproduce your logo, your exact packaging or your precise brand colours. It will produce something that rhymes with them, which is worse than nothing on a product shot.
Continuity. Ask for the same character across four shots and you will often get four cousins. Clothing changes. Rooms rearrange themselves. Hands do things hands do not do.
Specific real people. If the video needs your founder, your customer, or you, the answer is a camera. A generated person who looks a bit like your founder is an uncanny problem, not a solution.
Text on screen inside generated footage. Shop signs and labels come out as confident nonsense. Add real text as a caption layer instead.
Anything the viewer will check. Demonstrations, instructions and safety steps have to be true and visible. Approximate footage of a real procedure is a liability.
Say what you do not want. Adding "no faces, no on-screen text, no logos" removes most of the failure modes in one line.
Prompting a shot
A shot prompt has five parts: subject, action, setting, light, and camera. Length and restrictions go at the end.
Shot: [subject] [doing what] in [where].
Light: [morning sun through a window / overcast / warm lamplight].
Camera: [slow push in / static wide / handheld follow].
Length: [6] seconds. Style: real footage, not animation.
Do not include: faces, on-screen text, logos, brand names.
Here is my two-column plan (narration on the left, picture on the right).
For each row, write a shot prompt using this structure:
subject, action, setting, light, camera move, length in seconds,
then a do-not-include line.
Keep the light and time of day consistent across every shot.
[paste your plan]
Review this list of shot prompts before I generate anything.
Flag any shot that needs: a specific real person, my exact logo or product,
readable text on screen, or the same character appearing twice.
For each one, suggest a version that avoids the problem.
[paste your shot prompts]
โ Weak prompt
Prompt
A video of our bakery.
Output
A generic cafe interior, a smiling person who is not you, a sign above the counter reading BAKREY, a logo that is almost but not quite a logo.
Every weakness at once: a specific place it cannot know, a person it should not invent, and text it cannot spell. You will regenerate this six times and keep none of them.
โ Good prompt
Prompt
Shot: two hands shaping a round loaf of dough on a floured wooden counter. Light: low morning sun from the left, flour dust visible in the beam. Camera: slow push in, no cuts. Length: 6 seconds. Style: real footage, shallow depth of field. Do not include: faces, on-screen text, logos, brand names.
Output
Six clean seconds of hands and dough that cuts neatly under a line of narration and matches the other shots in the sequence.
It asks only for what the tool is good at: texture, light and movement with nothing identifiable to get wrong. This is usable on the first or second try.
This might not be for you
Some honesty, since the link above pays this site if you sign up.
If your videos need your face, your premises or your team, generated footage is not what you want, and no amount of prompting will get you there. Film it on your phone. Modern phone cameras are extraordinary, and viewers forgive shaky footage far more readily than they forgive an uncanny stranger pretending to be your staff.
If you make one video a quarter, a subscription is a poor deal. Pay for a month when you need it, or use a free tier and accept the watermark.
If you sell a physical product where the exact item matters, jewellery, food, clothing, a phone on a windowsill will beat generated footage every time, because the thing on screen is the actual thing you are selling.
Where these tools earn their keep is volume plus generic visuals: a small team posting weekly, a teacher turning lesson notes into clips, a marketer producing twenty variations for testing. If that is you, the trial costs an evening.
Try InVideo freeCheckpoint
AI video is strong on atmosphere, b-roll and volume, and weak on your logo, your people and anything the viewer will check, so give it the shots it can win and film the rest on your phone.
๐ Quiz
Question 1 of 4Which job is AI video genuinely well suited to?