DSpace at KOASAS: Learning to Maximize Speech Quality Directly Using MOS Prediction for Neural Text-to-Speech

DSpace at KOASAS

College of Engineering(공과대학)School of Electrical Engineering(전기및전자공학부)EE-Journal Papers(저널논문)

Learning to Maximize Speech Quality Directly Using MOS Prediction for Neural Text-to-Speech

Cited 1 time in

Cited 0 time in

Hit : 165
Download : 0

Export

Choi, Yeunju / Jung, Youngmoon / Suh, Youngjoo / Kim, Hoi-Rin researcher

Although recent neural text-to-speech (TTS) systems have achieved high-quality speech synthesis, there are cases where a TTS system generates low-quality speech, mainly caused by limited training data or information loss during knowledge distillation. Therefore, we propose a novel method to improve speech quality by training a TTS model under the supervision of perceptual loss, which measures the distance between the maximum possible speech quality score and the predicted one. We first pre-train a mean opinion score (MOS) prediction model and then train a TTS model to maximize the MOS of synthesized speech using the pre-trained MOS prediction model. The proposed method can be applied independently regardless of the TTS model architecture or the cause of speech quality degradation and efficiently without increasing the inference time or model complexity. The evaluation results for the MOS and phone error rate demonstrate that our proposed approach improves previous models in terms of both naturalness and intelligibility.

Publisher: IEEE-INST ELECTRICAL ELECTRONICS ENGINEERS INC

Issue Date: 2022-05

Language: English

Article Type: Article

Citation: IEEE ACCESS, v.10, pp.52621 - 52629

ISSN: 2169-3536

DOI: 10.1109/ACCESS.2022.3175810

URI: http://hdl.handle.net/10203/296934

Appears in Collection: EE-Journal Papers(저널논문)

Files in This Item: There are no files associated with this item.

This item is cited by other documents in WoS

⊙ Detail Information in WoSⓡ	Click to see
⊙ Cited 1 items in WoS	Click to see citing articles in

Display Full Item Record

qr_code

트윗하기

KOASAS

Knowledge Service Development Team, KAIST 291 Daehak-ro, Yuseong-gu, Daejeon 34141, Republic of Korea. T. 82-42-350-4493 Email. koasas@kaist.ac.kr
Copyright © 2016. Korea Advanced Institute of Science and Technology. All Rights Reserved.

KOASAS

KOASAS

Browse

Learning to Maximize Speech Quality Directly Using MOS Prediction for Neural Text-to-Speech

This item is cited by other documents in WoS

KOASAS

Communities & Collections