The long take in the age of AI: a threatened human feat or an art more necessary than ever?
AI is now generating disturbingly realistic videos. This article compares these new capabilities with the long take, this extreme exercise where everything depends on the human element, risk, and the moment, and shows why "real" has never been more valuable.
Summary
Introduction - When the machine imitates apnea
What AI can film (and where it falls short)
Russian Ark vs. AI: the impossible match
The value of risk: three sequence shots that AI cannot reproduce
AI in the service of the sequence shot: concrete uses
FAQ
Conclusion
1. When the machine imitates apnea
December 23, 2001. Saint Petersburg. Tilman Büttner straps on his Steadicam and enters the Winter Palace of the Hermitage. Before him: 33 rooms to traverse, 180 actors, 1,300 extras, a symphony orchestra. Behind him: Alexander Sokurov, who knows there's only one drum kit left. 87 uncut minutes. No safety net.
Twenty-four years later, in May 2026, Google Veo 3.1 generates 60-second videos in 1080p. Kling 3.0 pushes it up to 3 minutes. Camera movements (tracking, panning, dolly shots) are handled with a fluidity that would have seemed impossible two years ago.
So the question arises: if an AI can simulate a sequence shot, what's the point of shooting one for real?
This article is for film buffs, videographers, and technology enthusiasts who are asking themselves this question. We'll look at what AI can actually do, where it falls short, and why the human long take, like those in Russian Ark, Goodfellas, and Children of Men, is becoming more valuable, not less so.
2. What AI can film (and where it falls short)
Spectacular progress on the court
Let's be honest: over short periods, the result is stunning. A lateral tracking shot in a street, a push-in towards a face, a panoramic shot of a landscape lasting 10 to 20 seconds – the best tools produce images that even a professional can mistake for reality.
Think of the opening shot of Gravity (Cuarón, 2013): that slow, uninterrupted camera drift toward the space station, 13 minutes without any apparent cut. This kind of fluid movement in a stable environment can be convincingly simulated by AI. At least for the first 30 seconds.
The 60-second wall
Beyond that point, problems arise. And they are insurmountable.
The first is called drift. Frame after frame, details drift: a face subtly distorts, a background object shifts, the light oscillates for no apparent reason. In the Copacabana shot from Goodfellas, a three-minute sequence where the camera follows Henry Hill through the kitchens, corridors, and dining room, an AI would lose track of the set and extras before the first doorway.
The second problem is human choreography. A long take isn't just a camera gliding through space. It's an actor opening a door while another turns around, a waiter placing a glass at the right moment, an off-screen director giving a signal. Take the hallway scene in Oldboy (Park Chan-wook, 2003): each extra falls with precise timing, the camera moves laterally, the main character adjusts his actions in real time. AI can't coordinate these interactions.
Third limitation: resolution. Most generators are limited to 1080p. For theatrical projection (DCI 4K), this is insufficient.
Pro tip : Don't confuse "AI-generated long take" with "sequence shot." The former is a visual simulation of continuity. The latter implies real narrative and spatial continuity of the actors, their interactions, and the physical space they traverse. The difference is visible. And above all, it's palpable.
3. Russian Ark vs. AI: the impossible match
To measure the gap, let's go back to Russian Ark (Sokurov, 2002). The most extreme film ever shot in a single take and the best test for AI.
One single shot. 87 minutes. Filmed in one take inside the Hermitage Museum. 33 rooms traversed. 180 actors, 1,300 extras, a full orchestra. More than 20 assistant directors coordinated the different areas. The crew had 36 hours to light and dress each room, and a 4-hour shooting window; the winter light of Saint Petersburg waited for no one.
Four attempts. The first failed at the five-minute mark. Two more were unsuccessful. Only one battery remained. The final take produced the film.
Ask Veo 3.1 to generate that. Not even a simplified version, just 5 minutes of a character walking through three museum rooms, encountering groups of extras. You'll get facial drift from the second room onward, inconsistent lighting between rooms, and extras appearing and disappearing like ghosts.
But the real problem isn't technical. What makes Russian Ark so powerful is that it almost didn't happen. The risk. The urgency. The single battery. The 2,000 people who knew there wouldn't be a second chance. This tension is palpable in every second of the film. The AI knows no fear of failure.
4. The value of risk: three sequence shots that AI cannot reproduce
If AI can simulate perfection, does the value of the long take increase or disappear? Counterintuitive answer: it increases. While anyone can generate a smooth and pretty fake continuous shot, what distinguishes a real one is precisely what AI cannot reproduce: controlled imperfection, calculated risk, physical performance.
Three examples.
Goodfellas - The Copacabana (Scorsese, 1990)
3 minutes. Henry Hill crosses the kitchens, the service corridors, enters the dining room through a back door, sits down at a table being set up in front of him as he walks. Michael Ballhaus's camera never cuts. It hesitates, it adjusts, it breathes with the actors. It's this vibration, this living, imperfect camera, that follows the rhythm of a man who feels untouchable, whom AI would render too smooth. Too clean. And therefore less visceral.
Key takeaway: a successful sequence shot is not a perfect shot. It's a shot that resonates.
Children of Men – The Car Ambush (Cuarón, 2006)
4 minutes inside a car. The camera pans 360 degrees between the passengers as a motorcycle rams them, a motorcyclist shoots at the windshield, and blood splatters the lens. Emmanuel Lubezki passed the camera from hand to hand in a confined space, on a real road, with real stuntmen. The blood on the lens wasn't planned. Cuarón kept it. No AI could produce this fortunate accident, nor the actors' palpable terror.
Key takeaway: the long take captures what cannot be scripted. AI only generates what is predictable.
Victoria (Schipper, 2015)
138 minutes. An entire film shot in a single take. Filmed at night in the streets of Berlin, with a large element of improvisation. The actors didn't know exactly how each scene would unfold. Cinematographer Sturla Brandth Grøvlen carried his camera for over two hours without a break, adjusting his movements to the actors' unpredictable performances. The charm lies in the fact that everything can change in an instant , and the viewer feels it.
Key takeaway: the energy of a long take comes from the danger it contains. No danger, no tension.
The long take is a tightrope walker's art. Remove the void beneath his feet, and all that remains is a guy walking on a beam.
Pro tip : If you're shooting a single take, keep the micro-imperfections. A slight Steadicam shake, an extra glancing at the camera, a reframing mid-shot—these "flaws" are your signature of authenticity in the face of the synthetic perfection of AI.
5. AI in the service of the sequence shot: concrete uses
The real issue isn't "AI versus the long take." It's: how can AI serve the long take?
Preview before shooting
This is the most immediate use case. Planning a 4-minute single take in a restaurant? Generate a video mockup with Runway, the tool that best handles complex camera movements (tracking shots, pans, dolly zooms). In 30 seconds, you'll see if your camera trajectory works, if the timing between two actions is feasible, and if an obstacle is blocking the way. No need to rent the location or mobilize the crew. Sam Mendes could have tested the camera trajectories in 1917 this way; every transition between the continuous "fake shots" could have been simulated before filming.
Saving a shot in post-production
You've shot a 6-minute single take and there's a jarring detail in the background at minute 4? AI can correct that detail, remove a visible microphone, stabilize a shake, digitally extend a set to the edges of the frame without cutting the shot. This is exactly the kind of intervention AI already performs in post-production: invisible, surgical, and focused on maintaining continuity.
Training the videographers of tomorrow
A film student can experiment with ten different camera trajectories using AI before moving on to real-life filming. It's an accessible training ground. Director Nathalie Marchak sums it up well: AI bridges the gap between the image we have in mind and what we manage to show. It's like a pencil we can use even without knowing how to draw. But the film itself is still shot with humans, in a real space.
Pro tip : The World AI Film Festival (Cannes, 2026) showed that award-winning films are those that use AI as a preparation or finishing tool, not as a substitute for directing. The trend is clear: AI enhances the creative process, it doesn't replace it.
FAQ
How can you distinguish a real sequence shot from one generated by AI?
Look for complex interactions between characters (an exchanged object, a coordinated gesture), natural variations in lighting when the camera moves from room to room, and minor framing imperfections. The AI produces overly smooth movements and struggles to maintain consistency between faces and backgrounds beyond 30 seconds.
Are Birdman and 1917 "true" long takes?
No, these are fake long takes, extended shots stitched together with invisible cuts to create the illusion of continuity. Ironically, that's exactly what AI does too: assemble segments to simulate continuity. The difference with a Russian Ark or a Victoria? Zero editing, zero safety net, a single real take.
Will AI encourage or discourage filmmakers from shooting in single takes?
Both. Directors who used the fake long take (with invisible cuts) now have a competitor: AI does the same job for less money. But directors who shoot true long takes (a single shot, no cheating) will see their work gain distinction. The real long take is becoming a marker of authenticity.
What is the maximum length of an AI-generated video in 2026?
Approximately 60 seconds in high quality (Google Veo 3.1) and up to 3 minutes with Kling 3.0, but consistency degrades significantly over longer durations. For shots longer than one minute, drift (detail loss) remains the major unresolved issue.
When will we see an entire film shot in a single take using AI?
Not for a long time. The problem isn't just the duration, it's the narrative complexity. Coordinating dozens of characters, managing realistic interactions, maintaining the coherence of a physical space traversed for tens of minutes: none of this is within the reach of current AI, or even the next generation of models.
Conclusion
AI will not kill the long take. It might even make it indispensable.
By making simulation accessible, and soon commonplace, it transforms the real long take into an act of creative resistance. A deliberate choice to risk, to coordinate dozens of people, to capture a unique moment rather than to calculate it.
Tilman Büttner carrying his Steadicam for 96 minutes in the Hermitage. Lubezki passing the camera from hand to hand in a car under attack. Sturla Brandth Grøvlen following improvisational actors for 138 minutes in Berlin. These feats don't become obsolete because a machine can generate 60 seconds of smooth video. They become rarer, more precious, and more necessary.
The long take will remain what it has always been: a shared moment of breathlessness between those filming and those watching. AI, however, does not hold its breath.
On plan-sequences.com, you will find detailed analyses of more than 100 sequence shots.