Key takeaways
- Use a consistent grammar: Subject + Action + Setting for each prompt.
- Define composition and camera movement per shot to maintain visual flow.
- Set lighting and mood for scene consistency across all prompts.
- Apply negative constraints to avoid unwanted elements and ensure focus.
The Prompt Grammar: Subject + Action + Setting
Start every prompt with the core structure: Subject + Action + Setting. This gives the generator a clear anchor. For example, 'A woman in a red coat walking through a rainy city street at night.' The subject is 'woman in red coat,' the action is 'walking,' and the setting is 'rainy city street at night.' This grammar works across generators.
Keep the subject description consistent across all prompts for the same scene. If the subject is 'a woman in a red coat,' do not change it to 'a person in a red jacket' in the next shot. Consistency in wording helps the generator maintain visual identity.
For actions, use simple verbs: walking, sitting, looking, turning. Avoid complex or ambiguous actions like 'pondering' or 'meandering.' The generator interprets literal actions more reliably.
- Subject: noun phrase with key descriptors (e.g., 'a young man with glasses').
- Action: single verb or short phrase (e.g., 'walking slowly').
- Setting: location and time (e.g., 'in a busy market at noon').
Authoritative source: Google Cloud: Video generation prompt guide
Define Composition and Camera Movement Per Shot
After the subject-action-setting, specify composition and camera movement. Use terms like 'close-up,' 'wide shot,' 'over-the-shoulder,' or 'panning left.' For example: 'Close-up of the woman's face, camera slowly zooming in.' This gives the generator a clear framing instruction.
Camera movement should be simple and directional: 'pan right,' 'tilt up,' 'dolly forward.' Avoid compound movements like 'pan and zoom simultaneously' unless you have tested them. Many generators handle one movement better.
Composition terms like 'rule of thirds,' 'symmetrical,' or 'leading lines' can help, but keep them brief. Over-specifying composition can confuse the generator. Stick to one or two terms per shot.
- Use standard shot types: 'wide shot,' 'medium shot,' 'close-up.'
- Camera movement: 'static,' 'pan left,' 'tilt up,' 'dolly in.'
- Add composition cues sparingly: 'subject centered,' 'off-center.'
Authoritative source: Google Cloud: Video generation prompt guide
Set Lighting and Mood for Scene Consistency
Lighting and mood should be consistent across all prompts in a scene. Choose one lighting style and repeat it in every prompt. For example, 'soft golden hour light' or 'harsh overhead fluorescent.' Changing lighting between shots breaks continuity.
Mood words like 'tense,' 'cheerful,' or 'melancholic' can guide the generator, but pair them with lighting. 'Melancholic mood, dim blue light' is clearer than just 'melancholic.'
If the scene has a specific light source (e.g., 'a single lamp on the desk'), include it in the setting part of your grammar. This helps the generator render shadows and highlights consistently.
- Pick one lighting style per scene: 'natural daylight,' 'neon signs,' 'candlelight.'
- Use mood words with lighting: 'ominous, with deep shadows.'
- Repeat the exact lighting phrase in every prompt for the scene.
Maintain Continuity Across Shots with Reference Anchors
Continuity is hard for AI generators because each prompt is processed independently. To improve consistency, use reference anchors: repeat key visual details in every prompt. For example, if the subject wears a 'red coat,' mention it in every prompt, not just the first.
Also anchor the setting. If the scene is 'a rainy city street at night,' include 'rainy' and 'night' in each prompt. Do not assume the generator remembers from a previous prompt.
You can also use a 'master prompt' that sets the scene, then vary only the action and composition per shot. For example: 'A woman in a red coat walking through a rainy city street at night, close-up, camera panning right, soft streetlight.' Then for the next shot: 'A woman in a red coat walking through a rainy city street at night, wide shot, static camera, soft streetlight.' The repeated elements act as anchors.
- Repeat subject descriptors exactly: 'woman in red coat' not 'she.'
- Repeat setting details: 'rainy city street at night' every time.
- Use a master prompt template and swap only action/composition.
Apply Negative Constraints to Avoid Unwanted Elements
Most generators accept negative prompts or constraints. Use them to exclude unwanted objects, styles, or artifacts. For example: 'no umbrellas, no cars, no text.' This keeps the scene clean.
Negative constraints should be specific. 'No people' is clearer than 'no crowd.' 'No blur' helps avoid motion blur if you want a sharp image. Experiment with a few negative terms per prompt; too many can break the generation.
Common negative terms: 'no text,' 'no watermarks,' 'no extra subjects,' 'no distortion.' Adjust based on what your generator produces. If you see recurring artifacts, add them to the negative prompt.
- Use 'no' or 'without' for negatives: 'no umbrellas.'
- Limit negatives to 3-5 terms per prompt.
- Test and iterate: add negatives based on observed outputs.
Example: From Script to Prompt Series for a Short Scene
Suppose you have a simple script: A woman in a red coat walks through a rainy city street at night. She stops at a window. She looks inside. Then she walks on.
Convert each shot into a prompt using the grammar. Shot 1: 'A woman in a red coat walking through a rainy city street at night, wide shot, static camera, soft streetlight, no umbrellas, no cars.' Shot 2: 'A woman in a red coat stopping at a shop window on a rainy city street at night, medium shot, camera panning left, soft streetlight, no umbrellas, no cars.' Shot 3: 'A woman in a red coat looking through a shop window on a rainy city street at night, close-up on her face, camera zooming in slowly, soft streetlight, no umbrellas, no cars.' Shot 4: 'A woman in a red coat walking away from a shop window on a rainy city street at night, wide shot, static camera, soft streetlight, no umbrellas, no cars.'
Notice how the subject, setting, lighting, and negative constraints are repeated. Only the action and composition change. This gives the best chance for visual continuity.
- Write your script as a sequence of simple actions.
- For each action, write a prompt: Subject + Action + Setting + Composition + Lighting + Negatives.
- Repeat the subject, setting, lighting, and negatives exactly across prompts.
- Vary only the action and composition per shot.
Review and Iterate: Test Prompts Before Full Generation
Do not generate all prompts at once. Test each prompt individually. Look for consistency in subject appearance, setting, and lighting. If the woman's coat changes color between shots, adjust the prompt wording.
Keep a log of what works. If 'soft streetlight' gives a warm glow, use it in all prompts. If 'no umbrellas' is ignored, try 'without umbrellas' or add it to the negative field. Iteration is part of the workflow.
YouTube's policy on synthetic content applies if your video could mislead viewers. If you use AI-generated video that looks realistic, you may need to disclose it. Check the guidelines at support.google.com/youtube/answer/14328491. Also, avoid mass-produced template content to stay eligible for monetization, as described at support.google.com/youtube/answer/1311392.
- Generate one shot at a time.
- Compare output to your prompt: does the subject match?
- Tweak prompt wording if details change.
- Once consistent, generate the next shot.
- After all shots, review the sequence for continuity.
Authoritative sources: YouTube: Disclosing altered or synthetic contentYouTube channel monetization policies
Frequently asked questions
How do I keep the subject looking the same across shots?
Repeat the exact same subject description in every prompt. For example, always write 'a woman in a red coat' instead of 'she' or 'the woman.' Consistency in wording helps the generator maintain visual identity.
What if the generator ignores my negative constraints?
Try rewording the negative. Instead of 'no umbrellas,' use 'without umbrellas' or 'umbrellas:0.' Some generators have a separate negative prompt field; use that. Also, reduce the number of negatives to 3-5 per prompt.
Do I need to disclose that my video is AI-generated?
YouTube requires disclosure for realistic altered or synthetic content that could be mistaken for a real person, place, or event. If your video meets that standard, you must check the disclosure box. See YouTube's policy for details.
Sources and further reading
Primary documentation used by the Banned.Video editorial team.