Paper #60: NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
Abstract
Long-form serialized fiction, in the form of au- dio dramas with 200 to 800 episode arcs, is one of the largest creative media of the 2020s, yet it is precisely the regime where contemporary frontier language models fail. We benchmark 21 models spanning classical, fine-tuned, open- frontier, closed-frontier, and reasoning tiers on a uniform set of structural narrative metrics, and show that all closed-frontier systems saturate at plot-beat F1 in [0.78,0.81] while collapsing by −0.20 F1 at horizon h=200 episodes. We in- troduce NarrativeWorldBench, an open bench- mark of nine narrative-structure metrics evaluated across horizons h ∈{10,20,50,100,200}with cross-cultural localization across four Indic lan- guages. We then introduce N-VSSM, a Narrative Variational State-Space Model that pairs a R256 episode latent with Neural-ODE continuous-time dynamics, three multi-task narrative heads, and a prefix-tuned Llama-3.1-8B decoder; N-VSSM holds plot-beat F1 ≥0.84 across all horizons at 4×lower compute than the closed-frontier band. In a within-subjects writer study (n=12 profes- sional authors, 240 trials), N-VSSM is preferred over Claude Opus 4.5 on long-arc consistency 71% of the time and rated +1.3 Likert points higher on controllability. We release the bench- mark, all 21 model traces, the evaluation harness, and N-VSSM weights to seed open work on long- horizon co-creative narrative authoring