EventBank-G: Compact Event Memory for Controllable Multi-Shot Video Generation
Abstract
Long-horizon video generation is often assembled from short clips, but independently prompted clips drift in state while full-story prompts blur adjacent events. We repurpose EventBank from a video-understanding representation into EventBank-G, a model-agnostic event memory layer for generation. Given an ordered script, EventBank-G builds compact event tokens that pair the current action with a persistent state ledger for subject, scene, lighting, and camera style. These tokens condition a frozen text-to-video model one shot at a time. We introduce TimeLens-Stories, a protocol that converts temporally ordered TimeLens annotations into multi-shot generation scripts, and evaluate event fidelity, event-order discriminability, state adherence, and visual continuity. Across two remote LTX-Video spot runs covering 50 stories and 150 generated shots per method, EventBank-G improves event fidelity over independent prompting (.2570 vs. .2505) and state adherence (.2458 vs. .2400), while avoiding the wrong-event leakage that appears when full-story or neighboring-event context is injected into every shot.