Project page

DiVA

Enabling Interactive Digital Life Simulation via Video Models

Open-ended language and spatial clicks become continuous, speech-aligned, stateful video interaction.

Cheng Chen1,2,3 Hao Ouyang3 Qiuyu Wang3 Ka Leong Cheng3 Wen Wang3 Yihao Meng3 Hanlin Wang3 Yixuan Li3 Jiacheng Wei1 Zhenshan Tan4 Yanhong Zeng3 Yujun Shen3 Guosheng Lin1 Fayao Liu2,*
1 Nanyang Technological University 2 A*STAR 3 Ant Group 4 NUIST * Corresponding author
LanguageHuman character
Spatial clickInteractive animal
PersonaKnowledge-driven role
Open-ended interactionLanguage, click, persona, and memory.
Persistent stateAn MLLM manages what the character should say, do, and become.
Dedicated video expertsWaiting, action, and AVC render the next valid state.
Overview

From generated clips to persistent characters.

DiVA treats video as a renderer, not the controller. A multimodal language model explicitly manages interaction state across turns, then dispatches specialized video modules to generate the next response.

Language Spatial click History Persona
Observed context Input + current state

Who the character is, where it is, and what just happened.

State manager MLLM Router

Decide the next reply, behavior, and anchor.

ReplyBehaviorAnchor
Video renderer Waiting · Action · AVC

Specialized modules render the next valid visual state.

State t
State t + 1 persistent across turns
Full interaction

One character. Many turns.

Sound on available

Auto-plays muted on load. Turn on audio to hear the spoken response in the selected example.

Multi-turn simulation

Conversational character

Language-driven response with a sitting-to-standing transition.

Method

A decoupled, stateful pipeline.

The same interaction engine supports different entities and scenes, while dedicated video stages handle waiting motion, action, and state transition.

PersonaDialogueClickHistory
Semantic planner MLLM Router
ReplyBehaviorAnchor
01Waiting Video

Subtle idle motion keeps the character alive while the next turn is being planned.

02Action Video

Speech-aligned expression and motion render the actual response to the user.

03Anchored Video Continuation

AVC transitions back to a clean anchor while preserving motion and temporal continuity.

Long-horizon evaluation

Stable beyond the first turn.

Across repeated interaction rounds, DiVA maintains identity and visual quality more reliably than existing long-horizon baselines.

0.9451Subject consistency ↑
0.0341Quality drift ↓
0.8200Click accuracy ↑
First round
DiVAAnchored
LongCat-VideoContinuation
InfiniteTalkAudio-driven
FramePackLong video
Last round
DiVAAnchored
LongCat-VideoContinuation
InfiniteTalkAudio-driven
FramePackLong video

Direct side-by-side comparison of the first and last interaction rounds.

Anchored Video Continuation

Transitions remember motion.

Compared with frame-to-frame interpolation, AVC uses the preceding video segment to preserve object, pose, and camera continuity.

Pose + framing

Wan FLF2V
DiVA AVC

Object + camera

Wan FLF2V
DiVA AVC

Camera continuity

Wan FLF2V
DiVA AVC
76.7%preferred transition fidelity
71.7%preferred physical plausibility

30-person study · 60 transition videos

Applications

Interact by language or click.

The same stateful framework supports spatial reactions, conversational characters, and knowledge-driven roles.

Spatial input

Click-grounded reaction

A visual marker converts coordinates into grounded references for interaction.

0.82 click accuracy
Language

Conversational character

Speech, expression, and gesture remain aligned across multi-turn interaction.

Persona

Knowledge-driven role

Open-ended responses are rendered with character-specific behavior and timing.

Waiting dynamics

Alive before the reply.

Waiting videos maintain subtle motion before interaction begins, avoiding the frozen-frame effect.

Animal idle motion

Character waiting loop

Another waiting state

Paper

DiVA: Enabling Interactive Digital Life Simulation via Video Models

A stateful video-generation system for long-term multimodal character interaction.

Read paper
@article{chen2026diva,
  title   = {DiVA: Enabling Interactive Digital Life Simulation via Video Models},
  author  = {Chen, Cheng and Ouyang, Hao and Wang, Qiuyu and Cheng, Ka Leong and Wang, Wen and Meng, Yihao and Wang, Hanlin and Li, Yixuan and Wei, Jiacheng and Tan, Zhenshan and Zeng, Yanhong and Shen, Yujun and Lin, Guosheng and Liu, Fayao},
  year    = {2026}
}