HyperAI
Command Palette
Search for a command to run...
LibriTTS-R 音質改善版英語音声データセット
LibriTTS-R は、LibriTTS コーパスの音声品質を改善したバージョンであり、関連する論文は「LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus」で発表されています。これは、音声研究により高品質な英語話者データを提供することを目的としています。
このデータセットには、約 585 時間のマルチスピーカーによる英語朗読音声が含まれており、サンプリングレートは 24kHz です。また、7 つのデータ分割に分類されています。
引用文献
@ARTICLE{Koizumi2023-hs,
title = "{LibriTTS-R}: A restored multi-speaker text-to-speech corpus",
author = "Koizumi, Yuma and Zen, Heiga and Karita, Shigeki and Ding,
Yifan and Yatabe, Kohei and Morioka, Nobuyuki and Bacchiani,
Michiel and Zhang, Yu and Han, Wei and Bapna, Ankur",
abstract = "This paper introduces a new speech dataset called
``LibriTTS-R'' designed for text-to-speech (TTS) use. It is
derived by applying speech restoration to the LibriTTS
corpus, which consists of 585 hours of speech data at 24 kHz
sampling rate from 2,456 speakers and the corresponding
texts. The constituent samples of LibriTTS-R are identical
to those of LibriTTS, with only the sound quality improved.
Experimental results show that the LibriTTS-R ground-truth
samples showed significantly improved sound quality compared
to those in LibriTTS. In addition, neural end-to-end TTS
trained with LibriTTS-R achieved speech naturalness on par
with that of the ground-truth samples. The corpus is freely
available for download from
\url{http://www.openslr.org/141/}.",
month = may,
year = 2023,
copyright = "http://creativecommons.org/licenses/by-nc-nd/4.0/",
archivePrefix = "arXiv",
primaryClass = "eess.AS",
eprint = "2305.18802"
}
このデータセットはコミュニティユーザーによって提供されており、教育および情報提供のみを目的としています。著作権侵害に関わるコンテンツがある場合は、[email protected]までご連絡ください。速やかに確認し、削除いたします。