NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
Abstract
We document two empirical learning-dynamics phenomena in modern frontier language models that jointly produce the failure mode of long-horizon serialized fiction generation. First, on a 21-model audit, structural narrative metrics follow a soft scaling law of the form F1 ≈ 0.48 + 0.12 log10 N (N in billions; R^2 = 0.92 across 18 non-reasoning models) and saturate inside the closed-frontier band of [0.78, 0.81]; reasoning post-training adds only +0.02. Second, every closed-frontier model collapses by -0.18 to -0.21 F1 between narrative horizon h=10 and h=200 episodes. Extrapolating the joint fit, a hypothetical 405B reasoning model reaches F1 0.84 at h=10 but only 0.66 at h=200, evidence that the residual gap is structural rather than capacity-bound. We close it with N-VSSM, a Narrative Variational State-Space Model that pairs an R^256 episode latent with Neural-ODE continuous-time dynamics, three multi-task narrative heads, and a prefix-tuned Llama-3.1-8B decoder. N-VSSM holds plot-beat F1 ≥ 0.84 across all horizons at 4x lower compute, breaks both saturations, and is preferred by 12 professional writers over Claude Opus 4.5 on long-arc consistency 71% of the time. We release NarrativeWorldBench, an open benchmark of nine narrative-structure metrics across h ∈ {10, 20, 50, 100, 200}, all 21 model traces, and N-VSSM weights.