Previsualization has always been the cheapest insurance policy in production. Before you build a set, book a crew, or animate a single frame, you make a rough version and find out whether the idea works. Games borrowed the practice from film and studios have been refining it since the early 2000s, usually in a 3D package, usually staffed by people whose entire job is making rough versions of things.
The shift over the past two years is not that previz got better. It is that previz got cheap enough for teams that never had a previz department.
This is a working preview process built around Seedance 2.5, followed through one scene from beat sheet to timeline, with the shot blocking, camera, and pacing decisions written out as they actually happen.
What previz is actually for
There is a persistent misunderstanding that previz exists to show people what the final thing will look like. It does not. It exists to find out whether a sequence works before anyone commits resources to it.
A preview pass answers questions like: does this reveal land, or does the audience see it coming three seconds early? Does the camera move, sell the scale of this environment, or flatten it? Is this fight readable, or does the geography confuse people? Does this cutscene run forty seconds when the design document says twenty?
These are blocking, camera, and timing questions, and they are almost impossible to answer from a document. Timing especially. Nobody estimates duration accurately in their head, and a beat that reads as tense on paper reads as sluggish at actual speed roughly half the time.
The traditional cost of finding this out was a preview artist, a week, and a scene assembled from grey box geometry. Worth it for a hero sequence in a large production, and not worth it for the twelve other sequences, which shipped on faith.
The scene we will break down
A mid-sized studio is building a story sequence for a third person game. The scene is called The Stairwell.
Premise. The player character enters an abandoned parking structure at night, hears something two levels below, descends a stairwell, and finds the source is not what the audio suggested. Roughly forty seconds of non interactive cutscene, ending on a reveal.
The questions the team needs answered. Does the descent read as long enough to build tension without becoming boring? Does the reveal land, or does the camera give it away one shot early? Is the stairwell geography legible, so the player understands where they are when control returns?
None of those can be answered from the design document. All three can be answered from a rough moving version, and the rough version has to exist before the environment art team commits.
The shot breakdown
Six beats, written before anything was generated. Each line carries subject, blocking, camera, setting, and duration, because a generative model faithfully produces whatever ambiguity you hand it.
Beat 1. Establish. Wide of the parking structure entrance at night, sodium lighting, empty. Static camera, high angle. 5s.
Beat 2. Enter. A figure walks from frame right toward the stairwell door. The camera tracks slowly right, holding the figure at frame left. 6s.
Beat 3. Threshold. The stairwell door from inside, low light, the figure pushing through. Static camera at floor level, looking up. 5s.
Beat 4. Descent. The figure descending, seen through the open centre of the stairwell from below. The camera tilts up slowly as the figure passes. 8s.
Beat 5. Landing. The figure arrives on a landing and stops. The camera pushes in slowly from behind, over the shoulder. 8s.
Beat 6. Reveal withheld. Reverse angle on what the figure sees, held in shadow, no clear subject. Static camera. 6s.
Thirty eight seconds. Added on a calculator rather than estimated, because shot lists routinely sum to ninety seconds for a forty second deliverable.
Note the blocking discipline. The figure holds frame left through beat 2, arrives frame centre in beat 5, and the reverse in beat 6 respects the resulting eyeline. That is ordinary film grammar and it is entirely the team’s job. Generation does not do it for you.
Generating the beats
One beat per generation, never the sequence. A single beat is quick to assess and cheap to throw out. Prompting a whole scene produces a folder of nice clips that do not cut together.
Lock the shared parameters. The team wrote one block, night, sodium and fluorescent mixed lighting, wet concrete, shallow depth of field, no on screen text, and pasted it unchanged into all six prompts. Only subject, blocking, and camera varied. This does more for sequence coherence than any individual prompt improvement.
Constrain the recurring figure with references. Seedance 2.5 documents up to fifty reference assets per generation, and conditioning on images is far more dependable than hoping the model reads a written description the same way six times. The team used three character reference images and two location plates from their own concept art. Reviewing existing example outputs beforehand saved attempts, because it made clear quickly which categories of shot the tool handles cleanly and which it does not.
Expect the camera to sort into tiers. Static and push landed within two attempts. The tracking move in beat 2 took four. The upward tilt in beat 4 took seven and was eventually restaged as a static low angle with the figure passing through the frame, which read better anyway. Restaging a shot the tool cannot execute is the correct response. A tenth attempt at the same framing is not.
Change one variable per iteration. Beat 5 came back close but too fast, and the push arrived at the landing before the tension had built. The fix was pace only. Not the lighting, not the framing, not the duration.
Cutting it together, which is the whole point
Six clips went onto a timeline at their intended durations with a temp ambience track and no titles.
Three findings emerged immediately, none of which the shot list revealed:
The descent was too short. Beat 4 at eight seconds felt hurried. Watching it, the team extended the descent to twelve, which meant something else had to go.
Beat 1 was doing nothing. The establishing wide was handsome and carried no information the player did not already have from gameplay. Cutting it entirely bought the four seconds beat 4 needed, and the sequence tightened.
The reveal was safe. Beat 6 held its subject in shadow as written, and in the cut it read as coy rather than tense. The note was to bring the reveal forward into beat 5 as a partial silhouette, so the audience gets ahead of the character rather than waiting on the camera.
That last one is a scene structure note, and it arrived on day two of a process that previously took a week. It is the entire argument for doing this.
Then throw the clips away. Previz is disposable by design. Teams that get precious about generated clips start trying to use them in the shipped product, which is where this goes wrong.
The honest limitations
Worth stating directly, because overpromising here damages the case for the useful applications.
It will not match your IP. Your characters have specific designs and your world has specific rules, and a general model has seen neither. Reference conditioning narrows the gap without closing it. For previz that is acceptable, because you are testing blocking and timing rather than approving a character model. For anything a customer facing it is not.
No usable data output. You get a video file. Not camera paths, not animation curves, not anything a 3D pipeline can ingest. A preview artist in a 3D package produces data that feeds the next stage. This does not, which is why it supplements rather than replaces a preview department.
Continuity across many shots is limited. Two or three beats with a consistent subject is achievable. A twelve shot sequence with a stable protagonist is currently a fight, which is why the Stairwell scene is six beats and not sixteen.
Action legibility is the real weak point. Combat and complex physical interaction, the thing games most need previzzed, is where generated video is least reliable. Bodies intersect, weight reads wrong, geography gets confused. Atmospheric, environmental, and movement based beats work far better, and the Stairwell scene was chosen deliberately on that basis.
Rights and provenance need a policy. If generated material informs your production, know where it came from and what your vendor’s terms permit. Studios with publisher agreements and IP warranties should get this reviewed before it becomes a discovery item.
Who this changes things for
For a large studio with an existing preview department, this is supplementary. It speeds early exploration and does not touch the pipeline that produces shootable data.
For a mid sized team with no dedicated previz staff, it is a real capability gain. Sequences that previously went straight from document to production can now get a blocking and timing pass, and timing passes prevent the most expensive category of rework.
For a small or independent team, it is the difference between previzzing and not previzzing at all. That is the largest delta in the industry and it is why most of the interesting experimentation is happening at that end.
The unglamorous conclusion
The value here is not that a model makes video. It is that finding out a sequence does not work now costs an afternoon instead of a week, which means teams check more sequences, and checking more sequences is how scenes get better.
That is the same argument previz has always made. The tooling changed. The logic did not.






