DiVA
Enabling Interactive Digital Life Simulation via Video Models
Open-ended language and spatial clicks become continuous, speech-aligned, stateful video interaction.
From generated clips to persistent characters.
DiVA treats video as a renderer, not the controller. A multimodal language model explicitly manages interaction state across turns, then dispatches specialized video modules to generate the next response.
Who the character is, where it is, and what just happened.
Decide the next reply, behavior, and anchor.
Specialized modules render the next valid visual state.
One character. Many turns.
Sound on availableAuto-plays muted on load. Turn on audio to hear the spoken response in the selected example.
A decoupled, stateful pipeline.
The same interaction engine supports different entities and scenes, while dedicated video stages handle waiting motion, action, and state transition.
Subtle idle motion keeps the character alive while the next turn is being planned.
Speech-aligned expression and motion render the actual response to the user.
AVC transitions back to a clean anchor while preserving motion and temporal continuity.
Stable beyond the first turn.
Across repeated interaction rounds, DiVA maintains identity and visual quality more reliably than existing long-horizon baselines.
Direct side-by-side comparison of the first and last interaction rounds.
Transitions remember motion.
Compared with frame-to-frame interpolation, AVC uses the preceding video segment to preserve object, pose, and camera continuity.
Pose + framing
Object + camera
Camera continuity
30-person study · 60 transition videos
Interact by language or click.
The same stateful framework supports spatial reactions, conversational characters, and knowledge-driven roles.
Click-grounded reaction
A visual marker converts coordinates into grounded references for interaction.
0.82 click accuracyConversational character
Speech, expression, and gesture remain aligned across multi-turn interaction.
Knowledge-driven role
Open-ended responses are rendered with character-specific behavior and timing.
Alive before the reply.
Waiting videos maintain subtle motion before interaction begins, avoiding the frozen-frame effect.
Animal idle motion
Character waiting loop
Another waiting state
DiVA: Enabling Interactive Digital Life Simulation via Video Models
A stateful video-generation system for long-term multimodal character interaction.
@article{chen2026diva,
title = {DiVA: Enabling Interactive Digital Life Simulation via Video Models},
author = {Chen, Cheng and Ouyang, Hao and Wang, Qiuyu and Cheng, Ka Leong and Wang, Wen and Meng, Yihao and Wang, Hanlin and Li, Yixuan and Wei, Jiacheng and Tan, Zhenshan and Zeng, Yanhong and Shen, Yujun and Lin, Guosheng and Liu, Fayao},
year = {2026}
}