Prosodic Differences Between Child-Directed and Adult-Directed Speech in Text-to-Speech Generation
Abstract
Child-directed speech (CDS), speech produced when addressing children, and adult-directed speech (ADS), speech produced when addressing adults, differ systematically in prosodic characteristics such as pitch and speaking rates. We investigate whether text-to-speech (TTS) models fine-tuned on CDS and ADS reproduce the register-specific acoustic differences observed in human speech. The results showed that generated speech replicated several register-related prosodic patterns observed in human speech, including higher pitch and slower articulation rates in CDS relative to ADS. However, the magnitude of CDS-ADS differences in pitch measures was attenuated in generated speech, while articulation rate differences tended to be exaggerated. Overall, the findings suggest that fine-tuned TTS models can reproduce listener-conditioned register distinctions, while also revealing some limitations.