When AI meets vertical cinema: stories for the screen people actually hold in their hands
9n16 Redaktion

The portrait format is not cropped cinema
For decades, cinematic storytelling has been conceived for width. The landscape format rewards breadth: two-shots, wide establishing shots, lateral movement. The vertical image works in exactly the opposite way — it rewards height, full-body movement from floor to ceiling, close-ups of faces, and vertical architecture. The jump from 16:9 to 9:16 is therefore not a mere turning of the camera but a compositional grammar all its own (Tone Production).
Anyone who ignores this and crops an existing cinema image after the fact loses more than edges. If you cut a regular 4K landscape frame (3840×2160) down to a true 9:16 section, only about a third of the original remains — a roughly 1215×2160 central strip. The left and right thirds of the image are lost entirely, and any subject not placed at center gets cut off. The trade literature aptly calls this not a crop but an “amputation” (Tone Production). This is precisely why material shot natively in 9:16 enables different decisions: maximum resolution, deliberate composition, and full control over what the viewer sees.
Research also shows that the portrait format is no makeshift solution. A large-scale study in the Journal of Interactive Marketing demonstrates that mobile users process vertical videos more fluently and perceive less effort because they do not have to rotate the device; the effect is especially pronounced among younger users (Mulier, Slabbinck & Vermeir, Sage). That is an empirical finding from the advertising context, not proof for fictional series — but it supports the basic thesis: on the smartphone, the vertical image is the more natural form, not the more compromised one.
A visual grammar for the screen in the hand
As soon as you compose for the portrait format, the craft decisions change. The vertical image draws attention to a single subject; tight and medium close-ups fill the frame meaningfully, while wide shots quickly feel lost, because the subject stands small within a surface whose height it cannot fill. For introducing a location, a short vertical pan or tilt is therefore preferable to a static wide shot (Tone Production).
Spatial depth is created differently: through vertical leading lines such as staircases, doorframes, street lamps, or tall windows running the full height of the image, as well as through foreground elements and deliberately placed negative space above and below the figure. Eyelines, too, follow their own logic: the eyes of a speaking character sit on the upper third line, roughly 33 percent from the top, which at the same time creates room for subtitles (Tone Production). Subtitles here are not an afterthought but part of the composition: in practice the upper region of the frame is kept clear for platform headers, the actual action plays in the middle band, and subtitles sit in the lower third within the safe zone — important because user interfaces and interaction icons cover the edges (Tone Production).
Even camera and lighting have to think along: many cinema cameras, rigs, and accessory systems are set up for horizontal operation; for native 9:16 the camera is often rotated in the rig, the monitor turned 90 degrees, and the gimbal rebalanced. Because wider focal lengths bring the lighting equipment into frame faster, lamps move further back or higher up (Tone Production). None of this is accidental — it is production design for a screen you hold in your hand.

Where AI really helps — and where it should stay silent
The second driver alongside grammar is production economics. This is exactly where AI comes in. Media Partners Asia describes how artificial intelligence is deployed across the entire value chain, globally above all for localization and dubbing, and that its contribution to cost reduction is expected to increase sharply (Deadline, MPA report). At the European provider My Drama, which belongs to the Ukrainian company Holywater, AI is used, by its own account, at every stage of production — from script work to localization, plus editing, music, and the generation of secondary shots (VideoWeek).
Concretely, several workflows can be distinguished. For fast previsualization, storyboards, and alternative framings, AI-assisted reframing tools exist: Adobe's Auto Reframe in Premiere Pro automatically adjusts aspect ratio and framing, creates a duplicate sequence, and offers motion presets from “Slower Motion” to “Faster Motion” — yet it concedes that complex sequences with multiple focal points or fast movement require manual rework on the keyframes (Adobe Help, vendor statement). Google describes a machine-learning method that recomputes landscape material into vertical or square formats by detecting faces, objects, logos, text, and motion and breaking the video into scenes to keep important elements centered (Google, vendor statement). Such tools are useful for versioning and platform variants — but they do not bypass the fundamental decision: automatic reframing only reconstructs what has already been filmed; a natively vertical composition makes decisions about closeness, depth, and movement that an algorithm cannot invent after the fact.
The second major lever is accessibility across language barriers. ElevenLabs advertises for its method that it supports “90+ languages and accents” and conditions on the original performance rather than on a transcript, so that tone, emotion, and delivery are preserved and a voice clone of the original speaker is automatically generated (ElevenLabs, vendor statement). For international exploitation this is relevant: a vertical series can thus be localized for many markets without being re-shot for each language. Questions of consent and rights to voices and training data are central here — they are treated separately in our series.
What remains decisive is the division of roles: AI reduces barriers to production and distribution, but it should not determine the story. In Chinese practice, fully AI-generated microdramas are still rare; more common are hybrid forms in which lead roles are played by real actors and supporting figures are generated — with the sober verdict of a casting manager that the technology is “just not good enough yet” (The Globe and Mail).
From side format to a standalone ecosystem
That vertical fiction has long been more than a social feed is shown by the numbers. According to the industry association cited by The Globe and Mail, the Chinese microdrama market reached around 50.5 billion yuan in 2024, surpassing national box-office revenue for the first time (The Globe and Mail). Media Partners Asia puts Chinese revenue for 2024 at seven billion US dollars, compared with 0.5 billion in 2021 (Deadline, MPA report). According to an Omdia analysis of Sensor Tower data, the ReelShort app grew its revenue from 36 million dollars (2023) to 214 million (2024), and DramaBox from 8 to 217 million (Deadline, Omdia). In the US in 2025, according to Sensor Tower, viewers spent on average 35 minutes per day with ReelShort — more than with the major streaming services in the cited data (The Globe and Mail).
This engagement arises from a dramaturgy all its own: series run across 60 to 80 episodes of one to three minutes each (VideoWeek); individual chapters are short but complete as units, and character development carries across a whole season. According to data provider Enlightent, a title like “My 18-year-old Great-Grandma” reached more than 4.6 billion views across its 90-part first season (The Globe and Mail). A vertical series thus becomes intellectual property that can be developed further — into a film, into a horizontal series, into an audio format. MPA already describes “premium microdramas” with budgets of 400,000 to 600,000 dollars and explicit franchise potential (Deadline, MPA report).

Established houses are repositioning too. FOX Entertainment came in with a stake in Holywater and, according to My Drama studio head Sasha Tkachenko, plans around 200 vertical series, including possible integration of FOX IP (VideoWeek). Netflix, in turn, introduced a vertical, mobile-personalized video feed with “Clips” — but stresses that it is not trying to copy TikTok, instead using the format as a discovery tool for its own titles (The Hollywood Reporter). This distinction is essential: a vertical discovery feed of excerpts is not the same as a series told natively in vertical form.
Within this landscape, 9n16 is to be understood as a platform concept that treats vertical storytelling not as a side format alongside classic moving images but as a standalone entertainment ecosystem with its own grammar, its own episodic structure, and its own IP development. The approach takes seriously what the market figures suggest: that audiences do not experience the vertical screen as a compromise but as their primary one. What matters here is the attitude. It is not about replacing classic cinema or disparaging social-media culture — it is about telling stories for the mobile screen with the same care granted to cinema, and deploying AI where it lowers barriers, not where it takes over authorship.
Counterarguments, limits, and risks
The euphoria has its flip sides. First, reframing is not inherently wrong: for works with dual exploitation, shooting in 4K with a deliberately centered, somewhat looser composition plus post-crop — or a two-camera setup — can be a legitimate compromise (Tone Production). The mistake lies not in the tool but in treating it as a substitute for vertical design from the start. Second, the market figures should be read with caution: revenue figures partly stem from estimates by Omdia, Sensor Tower, or Media Partners Asia; forecasts to 2030 are not actuals, and base years are stated differently (Deadline, MPA report). Third, success depends on regulation and platforms: since February 2025 China has required a distribution license or filing number for microdramas, platforms may neither serve nor promote unlicensed content; in the course of a cleanup, 25,300 series with around 1.4 million episodes were removed (Reuters).
Fourth, European adoption is still thin: according to Ampere, in Spain, the country with the highest engagement in the region, only around seven percent of internet users have ever seen a mini-drama (VideoWeek). And finally there remains the content risk: a co-founder of the UK app Tattle TV explicitly expects “a backlash” to exploiting classic film works in portrait format (VideoWeek). A comparison with Quibi counsels humility: Jeffrey Katzenberg's service launched in early 2020 with ten-minute episodes and was shut down after about half a year (The Globe and Mail) — a format that was precisely not conceived as mobile-native vertical.
Conclusion: the future is decided at the start, not in the crop
The future of vertical storytelling is decided not in the crop but at the start. Stories developed from the outset for closeness, intimacy, and the movement of the portrait format meet a mode of viewing that has become the primary one for many people. AI can lower the barriers to production and distribution — fast previsualization, alternative frames, localization, dubbing, versioning — without taking over authorship. The portrait format demands neither condescension nor imitation of the social feed. It demands a visual grammar of its own and the willingness to take it seriously. Those who do gain not just reach but a new cinematic form.
Key Takeaways
- Genuine 9:16 storytelling is composed from the outset for height, closeness, and vertical movement; cropping after the fact from 16:9 keeps only about a third of the image and cuts off off-center subjects.
- The vertical image has a grammar of its own: close-ups, vertical leading lines for depth, eyelines on the upper third line, and safe zones for subtitles.
- AI lowers production barriers — reframing, previsualization, localization, AI dubbing — but should not determine the story; auto-reframe tools reach their limits with complex scenes.
- Vertical fiction is a market in its own right: China's microdrama revenue surpassed box-office earnings in 2024, and houses like FOX and Netflix are positioning themselves differently.
- Vertical storytelling is not automatically social media; the portrait format can be premium, cinematic, and emotional.
Sources and Further Reading
- The Globe and Mail – „China's microdramas go big“ (11.07.2026)
- Deadline – „Micro-Drama Revenues In China Set To Exceed Box Office“ (MPA, 17.09.2025)
- Deadline – „Micro-Drama Genre Booms In Asia“ (Omdia/Sensor Tower, 04.04.2025)
- Reuters – „China to control micro drama distribution“ (05.02.2025)
- VideoWeek – „Microdramas Come to Europe“ (04.02.2026)
- Adobe – „Add Auto Reframe effect to sequences“ (Anbieterangabe)
- ElevenLabs – „Dubbing v2“ (Anbieterangaben)
- Google – „New ways to make vertical video ads on YouTube“ (15.09.2022)
- The Hollywood Reporter – „Netflix Launches ‚Clips' Feed“ (30.04.2026)
- Journal of Interactive Marketing – Mulier, Slabbinck & Vermeir (2021)
- Tone Production – „Shooting Vertical Video: How To Frame 9:16 The Right Way“ (20.06.2026)
- YouTube – „Building YouTube Shorts“ (14.09.2020)
