StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

Yingrui Wang1,2,†Zeqing Wang1,2,†Yeying Jin1,2,‡,*

1 Tencent2 National University of Singapore (NUS)

† Equal contribution (co-first authors)‡ Project lead* Corresponding author

Abstract

Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, undermining both narrative coherence and visual consistency. To address these challenges, we propose StoryEngine, a state-grounded agentic framework for video storytelling. StoryEngine establishes a separation between authoritative semantic plans and unreliable visual observations. Specifically, StoryEngine maintains a structured representation of entity placement and story-relevant states, and propagates event-induced changes to define the intended start and end states of each shot. To visually realize these states, StoryEngine constructs canonical references for recurring entities and environments, and compiles state and visual constraints into executable render plans. Meanwhile, to realize these states correctly, a bounded evaluation-guided repair loop further corrects local state inconsistencies. Together, these mechanisms preserve causal story progression and prevent local visual errors from propagating across shots. To comprehensively evaluate long-form storytelling, we construct a benchmark across diverse scenarios and visual styles, with metrics assessing storytelling quality, narrative coherence, and visual consistency. Experimental results demonstrate that StoryEngine consistently outperforms state-of-the-art methods across all evaluation dimensions, validating its effectiveness for coherent and consistent video storytelling.

Demonstrations

Follow objects, places, and people through a complete story. Five one-minute films, generated with Veo 3.1.

Reading the Signals

An instrument-lined corridor and a service room form a connected working environment.

01 / 05

Watch for the layout of a revisited environment.

Veo 3.1 · 60 seconds

Comparisons

The same story, rendered by three methods. Watch how objects, identities, and the effects of earlier actions carry across shots.

Veo 3.1

Back for the Notebook

Track what stays behind. A student returns for a notebook while the backpack should remain in its locker.

Watch for

  • The backpack remains in its locker.
  • The notebook is retrieved.
  • The student’s identity persists across shots.

StoryEngine

Ours

MovieAgent

Baseline

ViMAX

Baseline
0:00 / 0:28

Source clips have different durations. Playback preserves their original timing at the selected speed; shorter clips hold their final frame. Videos are not stretched to match.

Overview Framework

What must be true

State propagation

Track recurring entities, their locations, and story-relevant attributes. Apply planned events to define the intended start and end state of every shot.

How it should be shown

Grounded rendering

Build shared character, object, and environment references. Bind each shot to action-relevant views, camera constraints, and explicit motion criteria.

Which evidence carries forward

Bounded repair

Reuse, reference, or discard earlier visual evidence according to the next shot’s intended state. Repair local failures while keeping the semantic plan fixed.

Reading the pipeline: follow the intended states into render planning, then inspect how the continuity gate and repair loop select visual evidence.

Fig. 1A fixed semantic plan. Grounded visual execution. Local repair that keeps the story intact.

Tracking consequences across shots

When the student returns for a forgotten notebook, earlier actions still matter. Explicit states preserve their consequences across shots.

Fig. 2Follow the backpack and notebook through the sequence. Each shot inherits the intended state of the story.
Fig. 3Qualitative ablations. Dashed boxes mark failures in the generated sequences.

Benchmark

A diagnostic benchmark evaluates complete stories using Veo 3.1 and Wan2.2-TI2V-5B.

Benchmark at a glance

A diagnostic benchmark tests narrative realization, cross-shot coherence, and visual consistency across two video generators.

60
stories
10
shots per story

Read the chart from the three suites of 20 stories to five setting categories: store, food, transit, workshop, and control room.

Fig. 4Benchmark composition

Results by video backbone

Automatic benchmark scores · Veo 3.1

Automatic benchmark scores. All metrics are higher is better.
MethodNarrativeCoherenceConsistencyAvg ↑
PECECSAPRSPSLARGPCMIRLCS
Direct I2V—0.62250.53890.44060.81210.73420.30120.23470.5263
ViMAX0.90520.90670.55000.57870.86840.85400.38500.24880.6274
MovieAgent0.96940.87830.43330.55790.78890.65350.25080.21250.5393
StoryEngine0.97980.92250.93890.66810.98000.89710.40720.56910.7690

All metrics are higher-is-better. Avg is the equal-weight mean of seven video metrics, excluding plan-only PEC. These are automatic benchmark scores, not human preference ratings.

What do the metrics measure?

Narrative

PEC Plan Event Coverage
ECS Event Completion Score

Coherence

APR Anchor Persistence Rate
SPS State Progression Score
LAR Location Adherence Rate

Consistency

GPC Geometric Place Consistency
MIR Minimum Identity Retention
LCS Lighting Coherence Score

BibTeX
@misc{wang2026storyenginestategroundedagenticframework,
      title={StoryEngine: A State-Grounded Agentic Framework for Video Storytelling},
      author={Yingrui Wang and Zeqing Wang and Yeying Jin},
      year={2026},
      eprint={2609.33627},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.33627},
}

StoryEngine / Research figure