Home »
Trending Technologies MCQs
AI Video Generation MCQs (Multiple-Choice Questions)
Practice AI Video Generation MCQs to understand how artificial intelligence creates, edits, and transforms video content from text prompts, images, audio, and existing footage. These multiple-choice questions explore generative models, diffusion techniques, motion synthesis, prompt engineering, and AI-powered video editing. They are useful for students, developers, content creators, and professionals exploring generative AI applications. The collection progresses from foundational concepts to practical workflows and technical challenges.
How AI Video Generation Works
AI video generation combines generative modeling, computer vision, temporal modeling, and sometimes audio synthesis to produce moving visual content. Unlike static image generation, a video model must also maintain motion continuity, object identity, scene consistency, and plausible transitions across frames.
AI Video Generation Concepts and Techniques
The questions examine text-to-video and image-to-video generation, diffusion models, transformers, latent representations, temporal consistency, frame interpolation, motion control, video upscaling, and multimodal conditioning. They also address model evaluation, computational requirements, responsible use, and common production limitations.
Work through the questions to test your understanding of the technologies behind AI-generated videos and the decisions involved in creating usable results. Each question includes the correct answer and an explanation to clarify the underlying concept.
AI Video Generation MCQs
Explore 50 questions covering model architectures, generation pipelines, creative controls, video quality, and real-world problem-solving. The later questions introduce practical scenarios involving inconsistent frames, camera movement, audio synchronization, and deployment constraints.
1. What is AI video generation?
- A technique for storing videos in compressed formats only
- A process that uses AI models to generate or transform video content
- A method for increasing internet bandwidth
- A system limited to manually editing individual frames
Answer: B) A process that uses AI models to generate or transform video content
Explanation:
AI video generation uses trained models to create or modify sequences of images. Depending on the system, it can generate videos from text, animate images, extend existing footage, or produce video variations from reference inputs.
2. Which task is performed by a text-to-video model?
- Converting a written prompt into a video sequence
- Converting a video into a database schema
- Translating source code into machine instructions
- Removing all pixels from a video file
Answer: A) Converting a written prompt into a video sequence
Explanation:
A text-to-video model interprets a natural-language description and generates visual frames that represent the requested scene. The prompt can describe subjects, actions, environments, lighting, camera movement, and visual style.
3. How does image-to-video generation differ from text-to-video generation?
- Image-to-video can only produce black-and-white output
- Text-to-video always requires a reference photograph
- Image-to-video uses an input image to guide the generated motion and appearance
- Image-to-video cannot generate more than one frame
Answer: C) Image-to-video uses an input image to guide the generated motion and appearance
Explanation:
Image-to-video systems use a reference image as a visual starting point or conditioning signal. They generate additional frames that introduce motion while attempting to preserve important characteristics of the original image.
4. What is the primary purpose of temporal consistency in AI-generated video?
- To make every frame use a different color palette
- To ensure the video has the smallest possible file size
- To eliminate the need for video encoding
- To maintain coherent appearance and motion across successive frames
Answer: D) To maintain coherent appearance and motion across successive frames
Explanation:
Temporal consistency helps objects, characters, lighting, and scene details remain stable as a video progresses. Without it, a generated sequence may contain flickering textures, changing facial features, or objects that appear and disappear unexpectedly.
5. What is a diffusion model commonly used for in generative video systems?
- Routing network packets between servers
- Learning to generate content by progressively removing noise from a representation
- Managing user authentication
- Converting every video into a spreadsheet
Answer: B) Learning to generate content by progressively removing noise from a representation
Explanation:
Diffusion models are trained using a noise-corruption process and learn to reverse that process. During generation, they progressively denoise a representation to produce visual content. Video diffusion systems extend this idea to representations containing temporal information.
6. Why are transformer architectures useful in AI video generation?
- They can model relationships among tokens or representations across space and time
- They eliminate the need for training data
- They guarantee physically accurate motion in every output
- They work only with text documents
Answer: A) They can model relationships among tokens or representations across space and time
Explanation:
Transformers use attention mechanisms to model relationships between elements in a sequence or other structured representation. In video generation, they can help capture interactions between visual regions, frames, text instructions, and other conditioning inputs.
7. What does a latent representation do in a video generation pipeline?
- Stores only the video's filename
- Replaces the need for visual information
- Represents video information in a compressed learned space
- Automatically publishes the generated video online
Answer: C) Represents video information in a compressed learned space
Explanation:
A latent representation encodes relevant visual information into a learned space that can be smaller than the original pixel representation. Generating or processing content in latent space can reduce computational requirements, although the exact savings depend on the model and implementation.
8. What is the role of a video decoder in a latent video generation system?
- It writes the text prompt automatically
- It determines the user's subscription plan
- It removes the video's audio track in every case
- It converts generated latent representations into visual frames
Answer: D) It converts generated latent representations into visual frames
Explanation:
When a model generates video in latent space, a decoder maps that representation back into visible frames. These frames can then be assembled, processed, and encoded into a video file.
9. What is prompt engineering in AI video generation?
- Writing and refining instructions to guide the generated video
- Increasing the physical size of a graphics card
- Repairing damaged storage sectors
- Converting all generated frames into text
Answer: A) Writing and refining instructions to guide the generated video
Explanation:
Prompt engineering involves specifying the desired subject, action, environment, composition, lighting, camera behavior, and style. Clear instructions can improve alignment with the intended result, although prompt quality cannot eliminate every model limitation.
10. What is a negative prompt used for when a video generation tool supports it?
- To reduce the video's playback volume only
- To describe unwanted visual characteristics or elements
- To increase the number of internet users
- To prevent the model from using any training data
Answer: B) To describe unwanted visual characteristics or elements
Explanation:
A negative prompt can identify unwanted characteristics, such as excessive blur, distorted anatomy, unwanted text, or particular visual artifacts. Its effectiveness depends on the model and tool; not all video generation systems support negative prompts.
11. What does frame rate measure in a video?
- The number of audio channels in a recording
- The amount of memory installed in a computer
- The number of frames displayed per second
- The number of words in a prompt
Answer: C) The number of frames displayed per second
Explanation:
Frame rate, commonly expressed as frames per second (FPS), indicates how many frames are displayed each second. It affects perceived motion and playback characteristics. Higher frame rates may require more generated frames or additional processing.
12. Which statement best describes video frame interpolation?
- It converts video captions into speech
- It deletes every alternate frame without replacement
- It changes a video's file extension without modifying its content
- It estimates intermediate frames between existing frames
Answer: D) It estimates intermediate frames between existing frames
Explanation:
Frame interpolation generates estimated frames between existing ones to create smoother motion or increase playback frame rate. Difficult cases include occlusions, rapid movement, and objects whose motion is hard to infer from the available frames.
13. What is the main purpose of motion estimation in video processing?
- To estimate how visual content moves between frames
- To calculate the electricity bill of a rendering server
- To assign a password to every frame
- To translate video titles into multiple languages
Answer: A) To estimate how visual content moves between frames
Explanation:
Motion estimation analyzes changes between frames to infer movement, often through motion vectors, optical flow, or learned features. This information can support interpolation, stabilization, compression, tracking, and video generation.
14. What does camera-motion control allow a user to specify?
- The operating system used by the generation service
- Camera behavior such as panning, zooming, tilting, or tracking
- The exact contents of the model's training dataset
- The number of accounts on a social platform
Answer: B) Camera behavior such as panning, zooming, tilting, or tracking
Explanation:
Camera-motion controls guide how the virtual viewpoint moves through a scene. Depending on the tool, these controls may be expressed through prompts, motion paths, reference videos, parameter settings, or specialized conditioning inputs.
15. What is identity consistency in character-focused video generation?
- Using a different character in every frame
- Ensuring every scene has the same background color
- Maintaining recognizable character features across frames and shots
- Keeping the video filename unchanged
Answer: C) Maintaining recognizable character features across frames and shots
Explanation:
Identity consistency means that a character's recognizable appearance, including facial features, hairstyle, clothing, and proportions, remains sufficiently stable throughout the generated sequence. Reference images, identity conditioning, and careful scene design can help, but results vary by model.
16. Which approach can help a video model preserve a particular character's appearance?
- Randomly changing the character description for every frame
- Removing all visual conditioning information
- Increasing audio volume
- Using consistent reference images or identity-conditioning inputs
Answer: D) Using consistent reference images or identity-conditioning inputs
Explanation:
Reference images and identity-conditioning mechanisms provide visual information that can guide the model toward a consistent appearance. The degree of control depends on the model, the reference quality, and whether the workflow supports identity preservation.
17. What is a major challenge when generating long videos with AI?
- Maintaining consistent characters, objects, actions, and scene details over time
- Making every video contain exactly one frame
- Eliminating the need for storage
- Ensuring that every frame contains identical pixels
Answer: A) Maintaining consistent characters, objects, actions, and scene details over time
Explanation:
Longer sequences require the model to maintain visual and narrative continuity across more frames. Errors may accumulate, causing characters to change appearance, objects to move unexpectedly, or the scene to drift away from the original instructions.
18. How can scene decomposition help with complex video generation?
- It removes the need to understand the requested story
- It divides a complex sequence into manageable shots or segments
- It forces all scenes to use one camera angle
- It converts video into an executable program
Answer: B) It divides a complex sequence into manageable shots or segments
Explanation:
Scene decomposition breaks a complex video into smaller shots or stages, each with a defined action, setting, and camera direction. The segments can be generated and reviewed separately, then assembled while checking continuity between shots.
19. What is multimodal conditioning in a video generation model?
- Using only a video filename as input
- Generating content without any conditioning signal
- Guiding generation with multiple input types, such as text, images, audio, or motion information
- Converting all model outputs into numerical tables
Answer: C) Guiding generation with multiple input types, such as text, images, audio, or motion information
Explanation:
Multimodal conditioning combines different forms of information to guide the generation process. For example, a system might use a text description for the scene, an image for visual appearance, and an audio or motion reference for additional guidance.
20. What is text-video alignment?
- Making every caption the same length
- Synchronizing a video file with a computer clock
- Ensuring all generated frames have identical brightness
- The degree to which generated visual content matches the text prompt
Answer: D) The degree to which generated visual content matches the text prompt
Explanation:
Text-video alignment describes how well the generated content represents the requested subjects, actions, attributes, and relationships. A video may look realistic but still have poor alignment if it omits an important object or depicts the wrong action.
21. What does a video generation model learn during training?
- Patterns and relationships in training examples that help it generate video content
- A fixed list of every possible future video
- Only the names of video editing applications
- The physical location of every viewer
Answer: A) Patterns and relationships in training examples that help it generate video content
Explanation:
Training helps a model learn statistical patterns related to visual appearance, motion, scene structure, and any conditioning signals. These learned patterns support generation, but they do not guarantee factual accuracy, physical realism, or perfect reproduction of the training examples.
22. Why do video generation models often require substantial computing resources?
- Every generated video must be manually drawn by a human
- Video generation processes large spatial and temporal representations through computationally intensive models
- Video models cannot run mathematical operations
- Every frame must be stored as a separate operating system
Answer: B) Video generation processes large spatial and temporal representations through computationally intensive models
Explanation:
Video models often process many frames and high-dimensional visual representations. Memory use and processing cost can grow with resolution, duration, frame count, and model complexity. Latent representations, optimized attention, and lower-resolution drafts can help manage resource demands.
23. What is classifier-free guidance commonly used for in diffusion-based generation?
- Converting videos into subtitles automatically
- Determining the copyright owner of every generated frame
- Adjusting how strongly generation follows conditioning information such as a text prompt
- Guaranteeing that no visual artifacts occur
Answer: C) Adjusting how strongly generation follows conditioning information such as a text prompt
Explanation:
Classifier-free guidance combines conditional and unconditional model predictions to influence adherence to the conditioning input. Higher guidance may improve prompt alignment in some settings, but excessive guidance can reduce diversity or introduce visual artifacts.
24. What is the purpose of a video autoencoder in a latent-generation pipeline?
- To automatically upload videos to social media
- To replace the need for any training process
- To generate a written script from a filename only
- To encode video into a compact representation and decode it back into video
Answer: D) To encode video into a compact representation and decode it back into video
Explanation:
An autoencoder typically contains an encoder that compresses input data into a latent representation and a decoder that reconstructs it. In video generation, this architecture can reduce the cost of modeling visual content, though reconstruction may lose some fine details.
25. How does super-resolution improve an AI-generated video?
- It attempts to reconstruct or generate higher-resolution visual details
- It converts a video into an audio-only recording
- It removes the video's temporal dimension
- It guarantees that every generated object is physically correct
Answer: A) It attempts to reconstruct or generate higher-resolution visual details
Explanation:
Video super-resolution increases spatial resolution and attempts to restore or generate finer visual details. AI-based enhancement can improve perceived sharpness, but it may also invent details or introduce inconsistencies that require visual inspection.
26. Which metric can be used to evaluate perceptual similarity between a generated frame and a reference image?
- DNS lookup time
- Structural Similarity Index Measure (SSIM)
- CPU serial number
- HTTP status code
Answer: B) Structural Similarity Index Measure (SSIM)
Explanation:
SSIM compares image structure, luminance, and contrast to estimate perceptual similarity. It can help evaluate reconstructed or generated frames against references, but it does not fully measure temporal coherence, semantic correctness, or overall video quality.
27. Why is Fréchet Video Distance (FVD) used in video-generation research?
- To calculate the physical weight of a camera
- To measure a video's upload speed directly
- To compare distributions of learned video features from generated and reference videos
- To count the words in a video transcript
Answer: C) To compare distributions of learned video features from generated and reference videos
Explanation:
FVD compares statistical distributions of feature representations extracted from generated and reference videos. It is used to assess aspects of generated-video quality and diversity, although results depend on the feature extractor, evaluation setup, and dataset.
28. What does an AI video watermark or provenance marker help communicate?
- That the video is guaranteed to be factually correct
- That the video has no synthetic elements
- That the content cannot be edited or copied
- Information that may help identify generated content or trace its origin
Answer: D) Information that may help identify generated content or trace its origin
Explanation:
Watermarks and provenance mechanisms can help indicate how content was created or modified. Their reliability depends on implementation and preservation; visible marks can be cropped, and some embedded signals may be damaged by transformations.
29. What is lip-sync generation intended to achieve?
- Aligning visible mouth movements with spoken audio
- Increasing a video's file size without changing its appearance
- Replacing every spoken word with background music
- Changing the camera's physical lens
Answer: A) Aligning visible mouth movements with spoken audio
Explanation:
Lip-sync systems generate or adjust mouth movements to correspond to speech timing and phonetic content. Quality depends on audio clarity, face visibility, language, pose, and the model's ability to represent realistic facial motion.
30. What is the role of audio conditioning in audio-driven video generation?
- It prevents the model from generating visual frames
- It provides audio information that can guide timing, speech-related motion, or other visual events
- It converts every video frame into a network address
- It automatically proves that the speaker consented
Answer: B) It provides audio information that can guide timing, speech-related motion, or other visual events
Explanation:
Audio conditioning can help align generated visuals with speech, music, or sound events. Depending on the system, audio may guide mouth movement, body motion, scene timing, or the generation of synchronized audiovisual content.
31. What is a common limitation of AI-generated human hands?
- Hands cannot appear in video frames
- Hands always remain perfectly still
- Finger counts, joint positions, and object interactions may appear anatomically incorrect
- Hands can only be generated in grayscale
Answer: C) Finger counts, joint positions, and object interactions may appear anatomically incorrect
Explanation:
Hands involve complex geometry, articulation, occlusion, and contact with objects. A model may generate implausible finger positions or inconsistent interactions, particularly during fast movement or when hands occupy only a small portion of the training examples.
32. Why can text and logos appear distorted in AI-generated videos?
- All video codecs automatically scramble letters
- Video models are prohibited from representing symbols
- Text rendering is always independent of the generation model
- The model may not reliably preserve exact character shapes and spatial relationships across frames
Answer: D) The model may not reliably preserve exact character shapes and spatial relationships across frames
Explanation:
Generative models may treat text as visual patterns rather than consistently rendering exact characters. This can lead to misspellings, malformed logos, or letters that change between frames. Adding accurate text during post-production is often more reliable.
33. What is a seed value used for in a video generation workflow?
- To initialize random-number generation for a run
- To specify the video's legal license automatically
- To determine the physical location of a GPU
- To convert audio into a video codec
Answer: A) To initialize random-number generation for a run
Explanation:
A seed can help reproduce or compare generation runs when the model, prompt, settings, and execution environment remain compatible. A matching seed does not guarantee identical results across different implementations, versions, or hardware configurations.
34. What does inpainting mean in AI-assisted video editing?
- Increasing the video's playback speed only
- Regenerating or filling selected regions of visual content
- Converting the audio track into subtitles
- Changing the file extension without altering its content
Answer: B) Regenerating or filling selected regions of visual content
Explanation:
Inpainting uses surrounding information and conditioning signals to fill or modify a selected image region. In video workflows, the edit must also be coordinated across frames to avoid flicker, inconsistent object boundaries, or unnatural changes over time.
35. How does outpainting differ from inpainting?
- Outpainting only changes the video's sound level
- Inpainting always increases the frame dimensions
- Outpainting extends visual content beyond the existing image boundaries
- Outpainting can only be used for text documents
Answer: C) Outpainting extends visual content beyond the existing image boundaries
Explanation:
Outpainting generates content outside an existing image's boundaries, such as extending a scene to fit a wider aspect ratio. Inpainting instead focuses on filling or modifying selected regions within the existing frame.
36. What is motion transfer in AI video generation?
- Moving a video file from one folder to another
- Changing the resolution of every frame
- Transferring ownership of a video account
- Applying motion from a reference sequence to another subject or character
Answer: D) Applying motion from a reference sequence to another subject or character
Explanation:
Motion transfer attempts to reproduce the movement of a reference subject using a different character or visual representation. The system must interpret pose, timing, body structure, and scene constraints to generate plausible motion.
37. What is a major benefit of using a reference video for motion guidance?
- It can provide concrete information about movement timing and trajectory
- It eliminates all requirements for computational resources
- It guarantees that the generated character will be photorealistic
- It automatically verifies every movement's legal status
Answer: A) It can provide concrete information about movement timing and trajectory
Explanation:
A reference video supplies an example of how an action unfolds over time. A model that supports motion conditioning can use that information to guide pose, rhythm, or movement patterns, although the generated result may differ from the reference.
38. Why is camera-aware scene generation useful for realistic video?
- It forces all objects to remain at the same depth
- It helps coordinate viewpoint changes with scene geometry and object appearance
- It prevents the use of multiple scenes
- It replaces the need for any visual representation
Answer: B) It helps coordinate viewpoint changes with scene geometry and object appearance
Explanation:
Camera-aware methods attempt to account for how the viewpoint affects perspective, visibility, and the appearance of objects. Better coordination can reduce distortions when the virtual camera moves, although complex geometry and occlusions remain challenging.
39. What is the purpose of a storyboard in an AI video production workflow?
- To compress generated frames into a smaller file
- To replace all video-generation prompts with random text
- To plan the sequence of shots, actions, and visual transitions
- To measure the model's GPU temperature
Answer: C) To plan the sequence of shots, actions, and visual transitions
Explanation:
A storyboard outlines the intended sequence of scenes and shots before full production. It helps creators plan framing, pacing, action, and continuity, making it easier to generate individual clips that fit the larger narrative.
40. What is the main purpose of post-production in an AI-generated video project?
- To guarantee that the model has no limitations
- To eliminate the need for reviewing the generated output
- To train every video model from scratch
- To refine, assemble, correct, and finalize generated clips
Answer: D) To refine, assemble, correct, and finalize generated clips
Explanation:
Post-production can include trimming, sequencing, color correction, audio mixing, subtitles, transitions, visual repairs, and format conversion. Human review is often important for correcting artifacts and ensuring that the final video meets its intended purpose.
41. A generated video shows a character wearing a blue jacket in one frame and a red jacket in the next, although the prompt describes no clothing change. Which problem does this illustrate?
- Temporal inconsistency in object or character appearance
- Successful lossless compression
- Correct audio synchronization
- Improved network routing
Answer: A) Temporal inconsistency in object or character appearance
Explanation:
The changing jacket color is an example of appearance drift across frames. Consistent reference conditioning, more explicit descriptions, shorter shots, and selecting or regenerating problematic segments may help reduce the issue.
42. A developer needs a video in which a product rotates smoothly while its shape and color remain stable. Which workflow is most appropriate?
- Use unrelated prompts for every frame and remove all references
- Provide a product reference and use a generation method with suitable motion or camera controls
- Generate only an audio file and rename it as a video
- Randomize the product description throughout the sequence
Answer: B) Provide a product reference and use a generation method with suitable motion or camera controls
Explanation:
A reference image can help establish the product's appearance, while supported motion controls can guide the rotation. For strict product accuracy, a 3D rendering or conventional animation workflow may be more reliable than unconstrained generative video.
43. An AI-generated educational clip contains accurate narration, but the speaker's mouth movements do not match the audio. What should be addressed first?
- Increase the video's resolution without checking synchronization
- Change the video filename
- Use a lip-sync or audio-driven facial animation process and review timing
- Remove all frames containing the speaker
Answer: C) Use a lip-sync or audio-driven facial animation process and review timing
Explanation:
The mismatch is an audiovisual synchronization problem. A compatible lip-sync or audio-driven facial animation tool can adjust mouth movement to match speech. The result should be reviewed for timing, facial artifacts, and alignment with the actual audio track.
44. A marketing team wants a video to show the exact spelling of a product name on its packaging. What is the most dependable approach when the generation model distorts text?
- Ask the model to invent a new spelling for every frame
- Lower the frame rate until the letters become correct
- Remove the product from every shot
- Add or replace the product text during post-production using a verified graphic or compositing layer
Answer: D) Add or replace the product text during post-production using a verified graphic or compositing layer
Explanation:
Conventional text and compositing tools provide precise control over spelling, typography, placement, and brand appearance. Tracking or perspective adjustments may be necessary to make the added text follow a moving product naturally.
45. A team is comparing two AI video models for generating short product demonstrations. Which evaluation strategy is most informative?
- Compare them using the same prompts and suitable settings, then assess prompt adherence, visual quality, temporal consistency, and generation cost
- Choose whichever model produces the largest files
- Evaluate only the first frame of every video
- Select a model solely because its interface has more buttons
Answer: A) Compare them using the same prompts and suitable settings, then assess prompt adherence, visual quality, temporal consistency, and generation cost
Explanation:
A controlled comparison helps distinguish model performance from differences in prompts or settings. Multiple evaluation dimensions are important because a model may produce attractive frames but struggle with temporal consistency, product accuracy, processing time, or cost.
46. A company wants to generate thousands of personalized training clips. Which factor should be evaluated early when designing the production pipeline?
- Whether every clip can use an unrelated visual style
- Generation throughput, cost, quality control, and integration with existing systems
- Whether filenames can contain spaces
- Whether the clips can be generated without any output checks
Answer: B) Generation throughput, cost, quality control, and integration with existing systems
Explanation:
Large-scale production requires predictable processing capacity, manageable costs, consistent outputs, and a reliable workflow for reviewing or correcting clips. API limits, privacy requirements, retries, and storage also affect the feasibility of deployment.
47. A video generation service must process user uploads while protecting confidential source footage. Which practice is most appropriate?
- Make every uploaded video publicly accessible
- Store all user footage permanently without informing users
- Apply access controls, secure data handling, and retention policies appropriate to the service
- Include confidential footage in public demonstrations by default
Answer: C) Apply access controls, secure data handling, and retention policies appropriate to the service
Explanation:
Video uploads can contain confidential business information, personal data, or identifiable people. Appropriate security controls, limited access, clear retention rules, and transparent data-use policies reduce privacy and security risks.
48. Before publishing an AI-generated video depicting a real person making a statement they never made, what is the most important consideration?
- Whether the video has a high frame rate
- Whether the generated clip uses a common file format
- Whether the background music is sufficiently loud
- Whether the depiction is deceptive, authorized, and appropriately disclosed under applicable rules
Answer: D) Whether the depiction is deceptive, authorized, and appropriately disclosed under applicable rules
Explanation:
Realistic synthetic depictions can mislead viewers or harm a person's reputation. Creators should consider consent, impersonation risks, disclosure, and applicable legal or platform requirements, particularly when a clip could be mistaken for authentic footage.
49. A developer wants to reduce the cost of testing multiple prompts before generating a final high-quality video. Which strategy is generally sensible?
- Generate low-cost draft previews with reduced settings, then render the selected version at the required quality
- Render every experiment at the maximum possible resolution and duration
- Disable all evaluation and quality checks
- Increase the duration of every test regardless of the intended output
Answer: A) Generate low-cost draft previews with reduced settings, then render the selected version at the required quality
Explanation:
Lower-resolution or shorter draft generations can help identify prompt problems before expensive final renders. Once the composition and motion are satisfactory, the chosen result can be regenerated or enhanced using production settings, subject to the tool's capabilities.
50. A production team is generating a 30-second video from several AI-created shots. Each shot looks good individually, but the assembled video contains abrupt changes in lighting, character appearance, and screen direction. Which solution best addresses the underlying problem?
- Increase the audio bitrate and leave all shots unchanged
- Establish a consistent visual reference and shot plan, regenerate mismatched segments, and review transitions and continuity during editing
- Randomize the camera direction in every shot to create variety
- Convert the entire video to a different file extension
Answer: B) Establish a consistent visual reference and shot plan, regenerate mismatched segments, and review transitions and continuity during editing
Explanation:
The issue is continuity across separately generated clips. A shared visual reference, consistent character and lighting descriptions, planned screen direction, and controlled shot transitions can reduce mismatches. Reviewing the assembled sequence is essential because individually convincing shots do not automatically form a coherent video.