Why AI-generated videos are still a long way from automatically being film
9n16 Redaktion

The leap from clip to scene is not an incremental step
Image quality has made great strides, that is undisputed. OpenAI describes Sora 2 as a “big leap forward in controllability” that can follow “intricate instructions spanning multiple shots” “while accurately persisting world state” (OpenAI, “Sora 2 is here”). In its technical report, Google DeepMind even attributes to its model a “bias for generating cinematic footage, with frequent camera cuts and dramatic camera angles” (Veo 3 Technical Report). Anyone seeing such clips for the first time can easily get the impression that filmmaking has largely been solved.
But it is precisely this impression that misleads. A film is not a sequence of beautiful individual images, but a chain of controlled decisions that must fit together across dozens of shots, scenes, and production phases. A character must be the same character in the next shot, the gaze must point in the right direction so that a shot/reverse-shot works, and a prop must not disappear between a wide shot and a close-up. These requirements sound trivial, but they are the actual core of the craft – and this is exactly where generative models hit a higher threshold than mere image quality.
One can distinguish five stages that are often confused: the impressive demo clip, the controllable film scene, the repeatable production process, the complete film, and the long-form narrated series. Each imposes stricter demands on consistency and controllability than the previous one; a model that masters the first brilliantly says almost nothing about the fifth.
Continuity is the real test, not resolution
The progress in consistency is real and deserves recognition. Runway promotes Gen-4 as generating “consistent characters, locations and objects across scenes” – and indeed “with just a single reference image of your characters” (Runway, “Introducing Runway Gen-4”). Reference and character conditioning have changed the field: you can define a character and recall it from different perspectives. The Verge sums up the announcement soberly: Gen-4 is meant to offer “enhanced ‘continuity and control’”, while AI videos “often face challenges in sustaining coherent storytelling” (The Verge).
Research, however, shows how deep the problem runs. A recent survey on controllability notes that “text prompts alone are often insufficient to express complex, multi-modal, and fine-grained user requirements” – which is why additional signals such as camera motion, depth maps, and body pose are being explored in the first place (Ma et al., “Controllable Video Generation: A Survey”). A second survey on spatiotemporal consistency openly names the typical error classes: “subject swapping”, “object teleportation”, “unnatural jumping”. And for long formats: “Existing generation models typically struggle to effectively capture such long-range dependencies, lacking the dynamic modeling and spatiotemporal memory capabilities required for complex and extended-duration relationships” (Survey: Spatiotemporal Consistency in Video Generation).

The demands become concrete as soon as you look at the details. Costume and prop continuity requires an object to remain exactly the same across takes – yet the consistency survey lists “subject swapping” and “object teleportation” as typical errors (Survey: Spatiotemporal Consistency in Video Generation). For body motion and anatomy, Sora 2 is, according to the provider, “more physically accurate” and lets a missed shot rebound off the backboard instead of teleporting, but concedes it is “far from perfect” (OpenAI, “Sora 2 is here”); research here demands that the “motion trajectory … conforms to physical laws” (Survey: Spatiotemporal Consistency in Video Generation). For dialogue scenes, Sora 2 and Movie Gen deliver, according to their providers, native, synchronized audio; for Sora 2, OpenAI explicitly cites “synchronized dialogue and sound effects”, which markedly improves lip sync (OpenAI; Meta, “Movie Gen”). Whether the emphasis, timing, and emotional coloring of a line can be precisely directed, however, is another matter – the model reports document synchronization, not the fine control of the vocal performance. The limit shows most clearly in emotional nuance and controllable acting decisions: in “Here”, the real actors carried “every single one of the character performance moments”, and the attempt with doubles failed because “it wasn't the soul of the actor that was there” (befores & afters, “‘Here’: A test with Tom Hanks”).
On top of this comes the dimension of film language, which is rarely considered. When the Veo 3 report attributes to the model a “bias for generating cinematic footage, with frequent camera cuts and dramatic camera angles” (Veo 3 Technical Report), that is precisely the point: frequent dramatic cuts or camera angles are not yet precise direction. Shot/reverse-shot, match-on-action, the axis of action, clean eyelines, and the rhythm of a sequence live from controllable coverage and repeatable shots, not from attractive random cuts. A generator that chooses beautiful angles therefore does not replace direction that deliberately and reproducibly sustains the same angle across several shots.
This is exactly what separates the technical categories from one another. Text-to-video generates from a description, image-to-video animates a starting image, video-to-video transforms existing material, world models carry states over time, shot extension lengthens clips – and classic VFX and compositing workflows remain the framework in which generated elements are assembled in a controlled and legally secure way. Adobe's Generative Extend, for example, lengthens a shot “up to two seconds for video and ten seconds for audio” and always labels generated images “so you know exactly where the original footage ends and created frames begin” (Adobe, “Generative Extend”). That is a precise editing tool – and at the same time a reminder that we are talking about seconds here, not scenes. Meta's Movie Gen, too, documents for its largest model “a generated video of 16 seconds at 16 frames-per-second” in “1080p HD” with “synchronized audio” (Meta, “Movie Gen”). Impressive – but a feature film consists of thousands of such building blocks, all of which must fit together.
A real example shows: AI shifts the work, it does not replace it
Anyone who wants to believe that AI takes the work off the set should study the production of “Here” – explicitly as a hybrid example of machine learning and classic VFX for performance editing, not as a film generated via text-to-video. For Robert Zemeckis' film, Metaphysic and VFX supervisor Kevin Baillie were responsible for “53-character minutes of full-face replacement” across four lead actors, some in shots “up to four minutes” long (befores & afters, “‘Here’: A test with Tom Hanks”). Classic CGI was not an option: “It was obvious that there was no way we could use traditional CGI methods to do this”, said Baillie – too expensive, too slow, with the risk of the uncanny valley.
The decisive point is how much specialized handwork the AI required. For each actor, several neural networks were trained on curated datasets – Tom Hanks “at 18, at 30, at 45” – with engineers having to actively combat “identity leak” so that “Tom's brother or cousin” did not appear. Eyelines, too, were a perennial issue, because the networks tended to look “magnetically” into the camera; edge cases such as a kiss had to be “backstopped” with classic tricks. Baillie's conclusion is clear: “All these limitations of the AI tools, they need traditional visual effects teams who know how to do this stuff to backstop them and help them succeed” (befores & afters).
Above all, however, the performance remained human. Wired and Variety document the core: a real-time system on two monitors allowed the actors to immediately review and adjust their rejuvenated performance – a tool for creatives, not a replacement for them (Wired, “The $50 Million Movie ‘Here’”; Variety). “Here” proves that AI is arriving in production – and how much iterative, highly specialized effort remains necessary for a result to hold up.

Counterarguments, limits, and risks
One can rightly object that the boundaries shifted in 2025/2026 – and that is true. For the series “El Eternauta”, Netflix used generative AI in an original production for the first time; according to co-CEO Ted Sarandos, the VFX sequence of a building collapse was created “10 times faster than it could have been completed with traditional VFX tools and workflows” (The Guardian). This figure is a corporate claim and refers to a single sequence. Telling, too, is how it came about: Netflix's Eyeline unit, by its own account, did not use in-house model technology but “commercially available video diffusion models under enterprise licenses” – embedded in a classic VFX pipeline (TheWrap).
The festival picture also deceives easily. Runway's AI Film Festival grew from “merely 300 submissions” in 2024 to “6,000 entries” in 2025, and Jacob Adler's “Total Pixel Space” won the Grand Prix (The Hollywood Reporter; Runway AIFF Screening Room). But the works shown were “varying in quality”, often “dream-like, experimental”, shaped by “constraints on sound and the portrayal of real people” (The Hollywood Reporter).
That leaves the hard limits. Controllability is not the same as control: even a model advertised as “more controllable” does not follow instructions perfectly, and fine directions can drift – OpenAI itself concedes for Sora 2 that it is “far from perfect” and makes “plenty of mistakes” (OpenAI, “Sora 2 is here”). The Veo 3 report notes “small hallucinations that mark videos as clearly fake” (Veo 3 Technical Report). Editability, version control, and reproducibility – vital for any production – are barely available with pure generators; “Here” only solved them through purpose-built Nuke tools and “neural animators”. And legal usability varies widely: Adobe advertises being “only trained on content where Adobe has permission … never on Adobe users' content” (Adobe), while Runway's predecessor model was criticized for training on “scraped YouTube videos and pirated films” (The Verge). The fact that OpenAI, according to its own page, reports that “as of April 26, 2026, the Sora product is no longer available” concerns the consumer product (the Sora app) and not necessarily the underlying model or its API availability; it nonetheless recalls how little production security a consumer product offers (OpenAI).
Conclusion: Spectacle is one skill, control is another
The honest assessment is neither euphoria nor rejection. Reference images, camera control, native soundtracks, and shot-extension tools have turned flickering curiosities into serious production building blocks. They save time, open up budgets, and expand what can be told – “El Eternauta” and “Here” show that concretely. But they do not shift the actual threshold: a good film is not made where an image is beautiful, but where a thousand decisions across shots, scenes, and months fit together and remain repeatable. That precisely – character consistency, eyelines, continuity, controllable acting, clean version states – is the higher threshold that generative models cannot yet cross on their own. The interesting question for the coming years is therefore not whether AI makes beautiful videos. It has long been able to do that. The question is whether it learns to narrate in a controlled way. Until then: an impressive clip is a promise, not a film.
Key Takeaways
- A single AI clip proves image quality, but not production control; between demo and series lie five clearly distinguishable stages of maturity.
- The real test is continuity – character, space, gaze direction, props, matching – and it is exactly here that research documents persistent error classes such as “subject swapping” and “object teleportation”.
- Progress is real: reference/character conditioning, native audio, and shot extension have noticeably shifted the boundaries in 2025/2026.
- “Here” proves that AI shifts the work, it does not replace it: 53 minutes of full-face replacement required specialized data curation, eyeline corrections, and classic VFX as a “backstop”.
- Editability, version control, reproducibility, and legal usability remain decisive, often unsolved hurdles for studio use.
Sources and Further Reading
- OpenAI – „Sora 2 is here“ (30.09.2025)
- Runway – „Introducing Runway Gen-4“ (31.03.2025)
- Google DeepMind – Veo 3 Technical Report (Mai 2025)
- Meta AI – „Movie Gen: A Cast of Media Foundation Models“ (16.10.2024)
- Adobe – „Generative Extend“ (Premiere Pro, abgerufen 11.07.2026)
- Ma et al. – „Controllable Video Generation: A Survey“ (arXiv:2507.16869)
- „A Survey: Spatiotemporal Consistency in Video Generation“ (arXiv:2502.17863)
- The Verge – Runway Gen-4 (01.04.2025)
- The Guardian – Netflix nutzt generative KI in „El Eternauta“ (18.07.2025)
- TheWrap – Netflix Eyeline Studios (15.10.2025)
- The Hollywood Reporter – Runway AI Film Festival 2025 (06.06.2025)
- befores & afters – „‚Here‘: A test with Tom Hanks“ (30.12.2024)
- Wired – „The $50 Million Movie ‚Here‘“ (06.11.2024)
- Variety – Robert Zemeckis über die Technik von „Here“ (02.11.2024)
- Runway AI Film Festival – Screening Room (abgerufen 11.07.2026)
