The Moving Target: A Longitudinal Audit of Trustworthiness Drift Across Twelve Generations of Open-Source Chat LLMs
Abstract
Model cards and governance reviews often report trust-benchmark scores without specifying when they were measured or whether they remain valid for later checkpoints in the same release line. We audit four open-source chat-LLM release lines—Yi, Qwen, Mistral, and Gemma—across three successive generations each, using five trust-related benchmarks and three prompt templates per benchmark. Across 180 evaluations and 36,000 item-level decisions, mean absolute adjacent-generation drift is well above an independence-based pooled no-drift reference null, and the result remains stable under strict scoring, leave-one-benchmark, leave-one-release-line, and drop-low-parse perturbations. The main operational implication is that trust scores attached to a named release line should not be carried forward to later checkpoints without re-measurement. We recommend treating such scores as checkpoint-bound, time-stamped artifacts and re-auditing materially new releases, while noting that closed APIs, larger models, canonical benchmark protocols, and fixed month-based re-evaluation rules remain outside the scope of this audit.