What must be true
State propagation
Track recurring entities, their locations, and story-relevant attributes. Apply planned events to define the intended start and end state of every shot.
Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.
Follow objects, places, and people through a complete story. Five one-minute films, generated with Veo 3.1.
Watch for the layout of a revisited environment.
The same story, rendered by three methods. Watch how objects, identities, and the effects of earlier actions carry across shots.
Track what stays behind. A student returns for a notebook while the backpack should remain in its locker.
Watch for
Source clips have different durations. Playback preserves their original timing at the selected speed; shorter clips hold their final frame. Videos are not stretched to match.
What must be true
Track recurring entities, their locations, and story-relevant attributes. Apply planned events to define the intended start and end state of every shot.
How it should be shown
Build shared character, object, and environment references. Bind each shot to action-relevant views, camera constraints, and explicit motion criteria.
Which evidence carries forward
Reuse, reference, or discard earlier visual evidence according to the next shot’s intended state. Repair local failures while keeping the semantic plan fixed.
Reading the pipeline: follow the intended states into render planning, then inspect how the continuity gate and repair loop select visual evidence.
When the student returns for a forgotten notebook, earlier actions still matter. Explicit states preserve their consequences across shots.
A diagnostic benchmark evaluates complete stories using Veo 3.1 and Wan2.2-TI2V-5B.
A diagnostic benchmark tests narrative realization, cross-shot coherence, and visual consistency across two video generators.
Read the chart from the three suites of 20 stories to five setting categories: store, food, transit, workshop, and control room.
Automatic benchmark scores · Veo 3.1
| Method | Narrative | Coherence | Consistency | Avg ↑ | |||||
|---|---|---|---|---|---|---|---|---|---|
| PEC | ECS | APR | SPS | LAR | GPC | MIR | LCS | ||
| Direct I2V | — | 0.6225 | 0.5389 | 0.4406 | 0.8121 | 0.7342 | 0.3012 | 0.2347 | 0.5263 |
| ViMAX | 0.9052 | 0.9067 | 0.5500 | 0.5787 | 0.8684 | 0.8540 | 0.3850 | 0.2488 | 0.6274 |
| MovieAgent | 0.9694 | 0.8783 | 0.4333 | 0.5579 | 0.7889 | 0.6535 | 0.2508 | 0.2125 | 0.5393 |
| StoryEngine | 0.9798 | 0.9225 | 0.9389 | 0.6681 | 0.9800 | 0.8971 | 0.4072 | 0.5691 | 0.7690 |
All metrics are higher-is-better. Avg is the equal-weight mean of seven video metrics, excluding plan-only PEC. These are automatic benchmark scores, not human preference ratings.
PEC Plan Event Coverage
ECS Event Completion Score
APR Anchor Persistence Rate
SPS State Progression Score
LAR Location Adherence Rate
GPC Geometric Place Consistency
MIR Minimum Identity Retention
LCS Lighting Coherence Score
@misc{wang2026storyenginestategroundedagenticframework,
title={StoryEngine: A State-Grounded Agentic Framework for Video Storytelling},
author={Yingrui Wang and Zeqing Wang and Yeying Jin},
year={2026},
eprint={2609.33627},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.33627},
}