Heiga Zen: From Statistical Models to Foundation Capabilities: The Historical Progression of Speech Generation
Heiga Zen
Abstract
Abstract The field of speech generation has undergone massive transformations, evolving from physical model simulations and concatenative systems to advanced foundational models. This talk will trace the historical progression of generative approaches for speech generation.
Bio Dr. Heiga Zen is a Principal Scientist and the Tokyo Site Lead at Google DeepMind. He leads diverse teams working on frontier AI research in areas including spoken language processing, natural language processing, gaming, and large language models. Dr. Zen is a Fellow of the ISCA (2022) and the IEEE (2025).
Speaker
Heiga Zen
Dr. Heiga Zen is a Principal Scientist (Director) at Google DeepMind (GDM) and the GDM Tokyo site lead. He leads diverse teams working on frontier AI research in areas including speech technology, natural language processing, gaming, and large language models. Before the formation of Google DeepMind in 2023, Dr. Zen was a co-founding member of the Google Brain team in Tokyo, which he joined in November 2018 after seven years with Google's Speech team in London. He received his PhD in Computer Science from Nagoya Institute of Technology in 2006. He also worked at the IBM T.J. Watson Research Center in Yorktown Heights and the Toshiba Cambridge Research Lab in Cambridge before joining Google. His research interests encompass speech technology, machine learning, multimodal LLMs, and artificial intelligence. He has authored over 200 papers, accumulating more than 30,000 citations. Dr. Zen is a Fellow of the ISCA (2022) and the IEEE (2025).
Video
Chat is not available.
Successful Page Load