Can artificial intelligence understand cinematic taste?
9n16 Redaktion

What machines really learn about films
A recommendation system knows no films. It knows behavior. Netflix describes its “Foundation Model for Personalized Recommendation,” unveiled in 2025, as a model that works on the principle of autoregressive prediction – “our default approach employs the autoregressive next-token prediction objective, similar to GPT” – and whose primary goal is simply to predict the next title interaction (Netflix Technology Blog). It is trained on the interaction histories of more than 300 million users and hundreds of billions of interactions (Netflix Technology Blog). This is not the formation of taste but probability calculus at scale.
The personalization of title images works similarly: Netflix selects the artwork it displays via so-called contextual bandits, optimizing the probability of a “quality play,” measured as a “take fraction” (Netflix Technology Blog). In doing so the company explicitly personalizes “not just what we recommend but also how we recommend” – and reports a “significant lift in core metrics” (Netflix Technology Blog). YouTube, for its part, has split its recommendations into two stages since 2016, a candidate generation and a ranking model (Google Research).
These systems are undeniably powerful. But they answer a different question than the artistic one. They measure what often works. They do not explain why a particular image is emotionally necessary within a particular story.
From forecast to judgment: numbers without discernment
The leap from recommendation to evaluation is the industry's real ambition. Vendors promise to detect success before shooting even begins. ScriptBook advertises being “87 % financially accurate” and a greenlight accuracy of 80 percent based on the screenplay alone; the underlying patent cites a classification accuracy of up to 86 percent for success and failure (ScriptBook). Largo.ai claims a greenlight accuracy of 82 percent and puts the number of films analyzed at more than 12,000 (Largo.io). Both are vendor claims that did not withstand independent scrutiny in the course of this research.
Serious research is more cautious. A 2024 study published in Scientific Reports achieved an average forecasting accuracy of 83.7 percent for the box office using a neural network augmented with comments; without the comment data, the figures dropped considerably (Scientific Reports). What matters, however, is the object: what is being predicted is commercial success, not artistic quality. The authors themselves name data selection and the COVID effect as open limitations (Scientific Reports).
This is precisely where the blind spot shows. An independent analysis of ScriptBook by Cineuropa notes that the purpose of the analysis is not to “capture the subtleties of a screenplay, but to compare its content with the known quantities of already released films” (Cineuropa). In the documented case of the film Passengers, the box-office prediction was usable, yet the tool could not anticipate the poor critical reception (Cineuropa). The criticism reads: “false objectivity” – objective metrics that can only support a subjective decision but not make it (Cineuropa).

The difference between effect and truthfulness
Can a machine at least measure emotional effect? In part. Affective-computing models such as AttendAffectNet predict a film's evoked emotions from video, audio, and text tracks – but only as the dimensions of valence and arousal, with best correlation values of around 0.655 for arousal and 0.575 for valence (Sensors). This is a correlative approximation of measured reactions, not a judgment about emotional truthfulness. Such research sketches how audience reactions can be modeled as data points; it replaces neither a representative test audience nor an aesthetic judgment, and a productive studio deployment of this very model is not documented in the sources reviewed.
This is exactly where test screenings and audience data come in. They can cluster reactions, make patterns across viewer groups visible, and estimate probabilities. But a test audience remains a sample, not a value judgment, and different viewers react differently to the same scene. Automated trailer and artwork systems, in turn, can optimize the selection and presentation of material – Netflix demonstrates this for personalized artwork selection (Netflix Technology Blog). In none of the sources reviewed for this piece, however, is there evidence that a system understands a director's signature or the artistic necessity of an image; this sentence should be read as a research finding and journalistic assessment, not as a categorical claim about the future.
This allows us to sharpen the central conceptual axis of this piece. Measurable or predictable are: technical quality, statistical probability of success, recognition, popularity and – in a limited and correlative way – emotional effect. Normative, cultural, and context-dependent, by contrast, remain: emotional truthfulness, originality, artistic necessity, and personal taste. A technically clean, familiar, statistically promising film can be emotionally empty. A work that is initially irritating can become culturally formative. No system documented here decides this difference.
The reason is structural, not merely gradual. Models compare the new with the known. This is precisely why research warns against data-driven uniformity: the classic finding by Fleder and Hosanagar shows that recommendation systems can generate “self-reinforcing cycles” in which popular titles are recommended more often, consumed more often, and then recommended even more often – “These cycles reduce diversity” (INFORMS). A later field experiment across 82,290 products and more than 1.1 million users confirms that collaborative filters lower aggregate offering diversity, even if individual diversity can rise at the same time (University of Pennsylvania). A system optimized for the known tends to produce more of the known.
What really happens in practice
The real picture is more nuanced than either extreme – euphoric faith in machines as much as blanket rejection. Perhaps the most robust practical case is Cinelytic: Warner Bros. Pictures International has used the platform since 2020 to evaluate content and talent and to guide release strategies (Business Wire). Remarkable is its self-restraint. Founder Tobias Queisser told Fortune: “We don't believe that an algorithm can run a script and understand whether the story is good or will connect with humans” – and made clear: “We're not in the business of forecasting hits or flops” (Fortune). In 2025 Cinelytic expanded its portfolio by acquiring Jumpcut Media and its script-analysis tool ScriptSense; the customers named include Lionsgate and Warner Bros. (Variety).
At the tool level of post-production, such automations have been commercially available since 2025. Blackmagic Design describes, for DaVinci Resolve 20, that “IntelliScript” automatically builds a timeline “with the best takes of user selected dialogue” based on a text script supplied by the user, that “IntelliCut” can remove silence in dialogue and, via speaker detection, place speakers on separate tracks, and that “Multicam SmartSwitch” on the cut page can “automatically select angles based on who is speaking” (Blackmagic Design). These are assistive suggestions and automations of technical tasks – a rough cut, the trimming of silence, the choice of the speaking camera. They establish neither a camera aesthetic nor a rhythm: not the dramaturgical decision of why a cut in an emotional scene sits a second too early or too late.

An often-cited example of the boundary between analysis and judgment is the AI-assisted trailer IBM created in 2016 for the 20th Century Fox film Morgan. IBM Research describes the project as a human-machine collaboration: multimodal analysis of image, sound, mood, and scene structure identified ten suitable moments, which were then arranged and cut by a professional filmmaker (IBM Research). According to IBM research lead John Smith, the editor used nine of the ten proposed clips, adjusted their boundaries, set transitions, and chose the music – “that was a completely human process” (VICE). The system analyzed and proposed moments; the cinematic selection and the final cut remained human. As a historical experiment, the case illustrates the potential of automated trailer analysis – not an autonomous cinematic taste.
Counterarguments, limits, and risks
The strongest objection to skepticism goes like this: human taste, too, is learned, shaped by experience, culture, and imitation. If a model recognizes patterns that “work,” is it doing anything fundamentally different from an experienced editor? The demonstrable difference lies less in the capacity for pattern recognition than in what can be measured. Systems optimize for observable quantities – clicks, dwell time, revenue. Artistic necessity, cultural imprint, and bold, productive wrong decisions are not defined as a measurable target in any of the sources reviewed here.
From this follow concrete risks. First, the documented danger of homogenization through self-reinforcing popularity cycles (INFORMS). Second, the “false objectivity” of forecasting tools whose numbers suggest certainty where a cultural and aesthetic decision would have to be made (Cineuropa). Third, the confusion of technical quality with truthfulness: an image can be perfectly composed and still inconsequential. At the same time, it would be wrong to deny AI any relevance. As an analytical tool – for risk assessment, for automating technical work, for personalizing distribution – its usefulness is documented. The boundary does not run between human and machine, but between the measurable and judgment.
Conclusion: between correlation and meaning
An AI system can learn which images, stories, or cuts often work. It can cluster test-group reactions, estimate probabilities, and uncover correlations that escape the human eye. It does not follow that it understands why a particular image is emotionally necessary within a particular story. Taste is not mere popularity, and statistical probability is not an aesthetic judgment. The honest state of affairs in 2026 lies between two convenient narratives: the machine replaces neither the director's signature, nor is it irrelevant. It shifts the question. No longer: what is likely to work? But rather: who decides when the data and the intuition diverge – and who bears the creative risk when an irritating work turns out to be the formative one? For now, that decision remains a human judgment.
Key Takeaways
- Recommendation and evaluation systems optimize for measurable quantities such as interactions, dwell time, and revenue; artistic necessity is defined as a target in none of the sources reviewed.
- One study reported 83.7 percent accuracy for the box-office forecast; what is predicted is commercial success, not quality – critical reception remains hard to grasp (the Passengers case).
- Affective computing measures emotional effect only correlatively (valence/arousal) and makes no judgment about emotional truthfulness.
- Data-driven systems carry the documented danger of uniformity through self-reinforcing popularity cycles.
- In practice, in 2025/2026 AI serves above all risk assessment, technical automation, and production acceleration – not the core creative decision.
Sources and Further Reading
- Netflix Technology Blog – „Foundation Model for Personalized Recommendation“ (21.03.2025)
- Netflix Technology Blog – „Artwork Personalization at Netflix“ (07.12.2017)
- Google Research – „Deep Neural Networks for YouTube Recommendations“ (2016)
- Scientific Reports (Nature) – Box-Office-Prognose mit neuronalen Netzen (11.09.2024)
- INFORMS – Fleder & Hosanagar zu Empfehlungssystemen und Vielfalt (14.05.2009)
- University of Pennsylvania – Lee & Hosanagar, Feldexperiment zu Empfehlungssystemen
- Sensors (MDPI) – „AttendAffectNet – Emotion Prediction of Movie Viewers“ (14.12.2021)
- Cineuropa – „GoCritic! Industry: Analyse This“ (ScriptBook, 10.07.2018)
- Business Wire – Cinelytic & Warner Bros. Pictures International (08.01.2020)
- Fortune – „No, A.I. isn't deciding which movies to green-light“ (25.02.2020)
- Variety – Cinelytic übernimmt Jumpcut (11.03.2025)
- Blackmagic Design – „DaVinci Resolve 20“ (04.04.2025)
- IBM Research – „Harnessing A.I. for augmenting creativity“ (23.10.2017)
- VICE – „Watson Made a Movie Trailer“ (2016)
- ScriptBook – Produkt- und Patentseiten (Anbieterbehauptungen)
- Largo.io – Anbieterangaben zur Greenlight-Genauigkeit (28.01.2026)
