Amphion support the following academic datasets (sort alphabetically):
The downloading link and the file structure tree of each dataset is displayed as follows.
Note: When using Docker to run Amphion, mount the dataset to the container is necessary after downloading. Check Mount dataset in Docker container for more details.
AudioCaps is a dataset of around 44K audio-caption pairs, where each audio clip corresponds to a caption with rich semantic information.
Download AudioCaps dataset here. The file structure looks like below:
[AudioCaps dataset path]
┣ AudioCpas
┃ ┣ wav
┃ ┃ ┣ ---1_cCGK4M_0_10000.wav
┃ ┃ ┣ ---lTs1dxhU_30000_40000.wav
┃ ┃ ┣ ...
Download the official CSD dataset here. The file structure looks like below:
[CSD dataset path]
┣ english
┣ korean
┣ utterances
┃ ┣ en001a
┃ ┃ ┣ {UtterenceID}.wav
┃ ┣ en001b
┃ ┣ en002a
┃ ┣ en002b
┃ ┣ ...
┣ README
We support custom dataset for Singing Voice Conversion. Organize your data in the following structure to construct your own dataset:
[Your Custom Dataset Path]
┣ singer1
┃ ┣ song1
┃ ┃ ┣ utterance1.wav
┃ ┃ ┣ utterance2.wav
┃ ┃ ┣ ...
┃ ┣ song2
┃ ┣ ...
┣ singer2
┣ ...
Download the official Hi-Fi TTS dataset here. The file structure looks like below:
[Hi-Fi TTS dataset path]
┣ audio
┃ ┣ 11614_other {Speaker_ID}_{SNR_subset}
┃ ┃ ┣ 10547 {Book_ID}
┃ ┃ ┃ ┣ thousandnights8_04_anonymous_0001.flac
┃ ┃ ┃ ┣ thousandnights8_04_anonymous_0003.flac
┃ ┃ ┃ ┣ thousandnights8_04_anonymous_0004.flac
┃ ┃ ┃ ┣ ...
┃ ┃ ┣ ...
┃ ┣ ...
┣ 92_manifest_clean_dev.json
┣ 92_manifest_clean_test.json
┣ 92_manifest_clean_train.json
┣ ...
┣ {Speaker_ID}_manifest_{SNR_subset}_{dataset_split}.json
┣ ...
┣ books_bandwidth.tsv
┣ LICENSE.txt
┣ readers_books_clean.txt
┣ readers_books_other.txt
┣ README.txt
Download the official KiSing dataset here. The file structure looks like below:
[KiSing dataset path]
┣ clean
┃ ┣ 421
┃ ┣ 422
┃ ┣ ...
Download the official LibriLight dataset here. The file structure looks like below:
[LibriTTS dataset path]
┣ small (Subset)
┃ ┣ 100 {Speaker_ID}
┃ ┃ ┣ sea_fairies_0812_librivox_64kb_mp3 {Chapter_ID}
┃ ┃ ┃ ┣ 01_baum_sea_fairies_64kb.flac
┃ ┃ ┃ ┣ 02_baum_sea_fairies_64kb.flac
┃ ┃ ┃ ┣ 03_baum_sea_fairies_64kb.flac
┃ ┃ ┃ ┣ 22_baum_sea_fairies_64kb.flac
┃ ┃ ┃ ┣ 01_baum_sea_fairies_64kb.json
┃ ┃ ┃ ┣ 02_baum_sea_fairies_64kb.json
┃ ┃ ┃ ┣ 03_baum_sea_fairies_64kb.json
┃ ┃ ┃ ┣ 22_baum_sea_fairies_64kb.json
┃ ┃ ┃ ┣ ...
┃ ┃ ┣ ...
┃ ┣ ...
┣ medium (Subset)
┣ ...
Download the official LibriTTS dataset here. The file structure looks like below:
[LibriTTS dataset path]
┣ BOOKS.txt
┣ CHAPTERS.txt
┣ eval_sentences10.tsv
┣ LICENSE.txt
┣ NOTE.txt
┣ reader_book.tsv
┣ README_librispeech.txt
┣ README_libritts.txt
┣ speakers.tsv
┣ SPEAKERS.txt
┣ dev-clean (Subset)
┃ ┣ 1272{Speaker_ID}
┃ ┃ ┣ 128104 {Chapter_ID}
┃ ┃ ┃ ┣ 1272_128104_000001_000000.normalized.txt
┃ ┃ ┃ ┣ 1272_128104_000001_000000.original.txt
┃ ┃ ┃ ┣ 1272_128104_000001_000000.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┃ ┣ 1272_128104.book.tsv
┃ ┃ ┃ ┣ 1272_128104.trans.tsv
┃ ┃ ┣ ...
┃ ┣ ...
┣ dev-other (Subset)
┃ ┣ 116 (Speaker)
┃ ┃ ┣ 288045 {Chapter_ID}
┃ ┃ ┃ ┣ 116_288045_000003_000000.normalized.txt
┃ ┃ ┃ ┣ 116_288045_000003_000000.original.txt
┃ ┃ ┃ ┣ 116_288045_000003_000000.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┃ ┣ 116_288045.book.tsv
┃ ┃ ┃ ┣ 116_288045.trans.tsv
┃ ┃ ┣ ...
┃ ┣ ...
┃ ┣ ...
┣ test-clean (Subset)
┃ ┣ {Speaker_ID}
┃ ┃ ┣ {Chapter_ID}
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.normalized.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.original.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.book.tsv
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.trans.tsv
┃ ┃ ┣ ...
┃ ┣ ...
┣ test-other
┃ ┣ {Speaker_ID}
┃ ┃ ┣ {Chapter_ID}
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.normalized.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.original.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.book.tsv
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.trans.tsv
┃ ┃ ┣ ...
┃ ┣ ...
┣ train-clean-100
┃ ┣ {Speaker_ID}
┃ ┃ ┣ {Chapter_ID}
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.normalized.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.original.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.book.tsv
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.trans.tsv
┃ ┃ ┣ ...
┃ ┣ ...
┣ train-clean-360
┃ ┣ {Speaker_ID}
┃ ┃ ┣ {Chapter_ID}
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.normalized.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.original.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.book.tsv
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.trans.tsv
┃ ┃ ┣ ...
┃ ┣ ...
┣ train-other-500
┃ ┣ {Speaker_ID}
┃ ┃ ┣ {Chapter_ID}
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.normalized.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.original.txt
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}_{Utterance_ID}.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.book.tsv
┃ ┃ ┃ ┣ {Speaker_ID}_{Chapter_ID}.trans.tsv
┃ ┃ ┣ ...
┃ ┣ ...
Download the official LJSpeech dataset here. The file structure looks like below:
[LJSpeech dataset path]
┣ metadata.csv
┣ wavs
┃ ┣ LJ001-0001.wav
┃ ┣ LJ001-0002.wav
┃ ┣ ...
┣ README
Download the official M4Singer dataset here. The file structure looks like below:
[M4Singer dataset path]
┣ {Singer_1}#{Song_1}
┃ ┣ 0000.mid
┃ ┣ 0000.TextGrid
┃ ┣ 0000.wav
┃ ┣ ...
┣ {Singer_1}#{Song_2}
┣ ...
┣ {Singer_2}#{Song_1}
┣ {Singer_2}#{Song_2}
┣ ...
┗ meta.json
Download the official NUS-48E dataset here. The file structure looks like below:
[NUS-48E dataset path]
┣ {SpeakerID}
┃ ┣ read
┃ ┃ ┣ {SongID}.txt
┃ ┃ ┣ {SongID}.wav
┃ ┃ ┣ ...
┃ ┣ sing
┃ ┃ ┣ {SongID}.txt
┃ ┃ ┣ {SongID}.wav
┃ ┃ ┣ ...
┣ ...
┣ README.txt
Download the official Opencpop dataset here. The file structure looks like below:
[Opencpop dataset path]
┣ midis
┃ ┣ 2001.midi
┃ ┣ 2002.midi
┃ ┣ 2003.midi
┃ ┣ ...
┣ segments
┃ ┣ wavs
┃ ┃ ┣ 2001000001.wav
┃ ┃ ┣ 2001000002.wav
┃ ┃ ┣ 2001000003.wav
┃ ┃ ┣ ...
┃ ┣ test.txt
┃ ┣ train.txt
┃ ┗ transcriptions.txt
┣ textgrids
┃ ┣ 2001.TextGrid
┃ ┣ 2002.TextGrid
┃ ┣ 2003.TextGrid
┃ ┣ ...
┣ wavs
┃ ┣ 2001.wav
┃ ┣ 2002.wav
┃ ┣ 2003.wav
┃ ┣ ...
┣ TERMS_OF_ACCESS
┗ readme.md
Download the official OpenSinger dataset here. The file structure looks like below:
[OpenSinger dataset path]
┣ ManRaw
┃ ┣ {Singer_1}_{Song_1}
┃ ┃ ┣ {Singer_1}_{Song_1}_0.lab
┃ ┃ ┣ {Singer_1}_{Song_1}_0.txt
┃ ┃ ┣ {Singer_1}_{Song_1}_0.wav
┃ ┃ ┣ ...
┃ ┣ {Singer_1}_{Song_2}
┃ ┣ ...
┣ WomanRaw
┣ LICENSE
┗ README.md
Download the official Opera dataset here. The file structure looks like below:
[Opera dataset path]
┣ monophonic
┃ ┣ chinese
┃ ┃ ┣ {Gender}_{SingerID}
┃ ┃ ┃ ┣ {Emotion}_{SongID}.wav
┃ ┃ ┃ ┣ ...
┃ ┃ ┣ ...
┃ ┣ western
┣ polyphonic
┃ ┣ chinese
┃ ┣ western
┣ CrossculturalDataSet.xlsx
Download the official PopBuTFy dataset here. The file structure looks like below:
[PopBuTFy dataset path]
┣ data
┃ ┣ {SingerID}#singing#{SongName}_Amateur
┃ ┃ ┣ {SingerID}#singing#{SongName}_Amateur_{UtteranceID}.mp3
┃ ┃ ┣ ...
┃ ┣ {SingerID}#singing#{SongName}_Professional
┃ ┃ ┣ {SingerID}#singing#{SongName}_Professional_{UtteranceID}.mp3
┃ ┃ ┣ ...
┣ text_labels
┗ TERMS_OF_ACCESS
Download the official PopCS dataset here. The file structure looks like below:
[PopCS dataset path]
┣ popcs
┃ ┣ popcs-{SongName}
┃ ┃ ┣ {UtteranceID}_ph.txt
┃ ┃ ┣ {UtteranceID}_wf0.wav
┃ ┃ ┣ {UtteranceID}.TextGrid
┃ ┃ ┣ {UtteranceID}.txt
┃ ┃ ┣ ...
┃ ┣ ...
┗ TERMS_OF_ACCESS
Download the official PJS dataset here. The file structure looks like below:
[PJS dataset path]
┣ PJS_corpus_ver1.1
┃ ┣ background_noise
┃ ┣ pjs{SongID}
┃ ┃ ┣ pjs{SongID}_song.wav
┃ ┃ ┣ pjs{SongID}_speech.wav
┃ ┃ ┣ pjs{SongID}.lab
┃ ┃ ┣ pjs{SongID}.mid
┃ ┃ ┣ pjs{SongID}.musicxml
┃ ┃ ┣ pjs{SongID}.txt
┃ ┣ ...
Download the official SVCC dataset here. The file structure looks like below:
[SVCC dataset path]
┣ Data
┃ ┣ CDF1
┃ ┃ ┣ 10001.wav
┃ ┃ ┣ 10002.wav
┃ ┃ ┣ ...
┃ ┣ CDM1
┃ ┣ IDF1
┃ ┣ IDM1
┗ README.md
Download the official VCTK dataset here. The file structure looks like below:
[VCTK dataset path]
┣ txt
┃ ┣ {Speaker_1}
┃ ┃ ┣ {Speaker_1}_001.txt
┃ ┃ ┣ {Speaker_1}_002.txt
┃ ┃ ┣ ...
┃ ┣ {Speaker_2}
┃ ┣ ...
┣ wav48_silence_trimmed
┃ ┣ {Speaker_1}
┃ ┃ ┣ {Speaker_1}_001_mic1.flac
┃ ┃ ┣ {Speaker_1}_001_mic2.flac
┃ ┃ ┣ {Speaker_1}_002_mic1.flac
┃ ┃ ┣ ...
┃ ┣ {Speaker_2}
┃ ┣ ...
┣ speaker-info.txt
┗ update.txt