Best-of-N TTS Evaluation is Confounded by ASR Family Alignment
Taehyung Yu ⋅ Seongjae Kang
Abstract
Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an ASR verifier. We identify an underexplored evaluation confound: a verifier's apparent quality depends strongly on which ASR family judges it. On LibriSpeech-PC test-clean with F5-TTS, verifier rankings reverse across Whisper, wav2vec~2.0, and HuBERT evaluators, and same-family verifier--evaluator pairs recover 2--3$\times$ more oracle headroom than cross-family pairs despite near-identical representations (linear CKA $0.978$)---a pattern consistent with identity- or lineage-level coupling rather than representational overlap. We propose two \textbf{cross-family rank ensembles} (rank-averaging and conjunctive max-rank) that attain the lowest mean WER across three independent evaluators; at $N{=}10$, WER drops from $2.06\%$ to $1.72\%$ ($-16.5\%$ relative) under the official F5-TTS evaluator, with no measurable degradation under automatic SIM-o/UTMOS metrics. We recommend cross-evaluator triangulation as default reporting practice.
Chat is not available.
Successful Page Load