HyperAI
Command Palette
Search for a command to run...
LibriTTS-R 음질 개선 영어 음성 데이터셋
LibriTTS-R은 LibriTTS 데이터셋의 오디오 품질을 개선한 버전으로, 관련 논문 결과는 LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus이며, 음성 연구를 위해 더 높은 품질의 영어 음성 데이터를 제공하는 것을 목표로 합니다.
이 데이터세트는 약 585시간 분량의 다화자 영어 낭독 음성을 포함하며, 샘플링 레이트는 24kHz이고 7개의 데이터 세트로 나뉩니다.
인용문헌
@ARTICLE{Koizumi2023-hs,
title = "{LibriTTS-R}: A restored multi-speaker text-to-speech corpus",
author = "Koizumi, Yuma and Zen, Heiga and Karita, Shigeki and Ding,
Yifan and Yatabe, Kohei and Morioka, Nobuyuki and Bacchiani,
Michiel and Zhang, Yu and Han, Wei and Bapna, Ankur",
abstract = "This paper introduces a new speech dataset called
``LibriTTS-R'' designed for text-to-speech (TTS) use. It is
derived by applying speech restoration to the LibriTTS
corpus, which consists of 585 hours of speech data at 24 kHz
sampling rate from 2,456 speakers and the corresponding
texts. The constituent samples of LibriTTS-R are identical
to those of LibriTTS, with only the sound quality improved.
Experimental results show that the LibriTTS-R ground-truth
samples showed significantly improved sound quality compared
to those in LibriTTS. In addition, neural end-to-end TTS
trained with LibriTTS-R achieved speech naturalness on par
with that of the ground-truth samples. The corpus is freely
available for download from
\url{http://www.openslr.org/141/}.",
month = may,
year = 2023,
copyright = "http://creativecommons.org/licenses/by-nc-nd/4.0/",
archivePrefix = "arXiv",
primaryClass = "eess.AS",
eprint = "2305.18802"
}
이 데이터셋은 커뮤니티 사용자가 기여한 것이며 교육 및 정보 제공 목적으로만 사용됩니다. 저작권 침해와 관련된 콘텐츠가 있는 경우 [email protected]로 문의하시면 신속하게 검토 및 삭제 처리하겠습니다.