NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
Abstract
Long-form serialized fiction, in the form of audio dramas with 200 to 800 episode arcs, is one of the largest creative media of the 2020s, yet it is precisely the regime where contemporary frontier language models fail. We benchmark 21 models spanning classical, fine-tuned, open-frontier, closed-frontier, and reasoning tiers on a uniform set of structural narrative metrics, and show that all closed-frontier systems saturate at plot-beat F1 in [0.78, 0.81] while collapsing by -0.20 F1 at horizon h=200 episodes. We introduce NarrativeWorldBench, an open benchmark of nine narrative-structure metrics evaluated across horizons h in {10, 20, 50, 100, 200} with cross-cultural localization across four Indic languages (Hindi, Tamil, Telugu, Marathi). We then introduce N-VSSM, a Narrative Variational State-Space Model that pairs an R^256 episode latent with Neural-ODE continuous-time dynamics, three multi-task narrative heads, and a prefix-tuned Llama-3.1-8B decoder; N-VSSM holds plot-beat F1 >= 0.84 across all horizons at 4x lower compute than the closed-frontier band. We additionally train a learned Cultural Transfer Function that lifts cross-language fidelity by +0.20 to +0.23 Likert points, evidence that cultural alignment is a representational rather than a prompt-level problem. In a within-subjects writer study (n=12 professional authors, 240 trials), N-VSSM is preferred over Claude Opus 4.5 on long-arc consistency 71% of the time and rated +1.3 Likert points higher on controllability. We release the benchmark, all 21 model traces, the evaluation harness, and N-VSSM weights to seed open work on long-horizon co-creative narrative authoring.