言語資源検索 - SHACHI: Language Resource Metadata Database

言語資源の登録件数: 3330件 2023 件中 771 - 780 件目

C-001362: Mandarin Chinese Desktop Speech Recognition Corpus - Digit String (98 people)
Desktop/Microphone
This corpus comprises 1,500 entries uttered by 98 speakers of different dialects, ages and various educational levels (46 males and 52 females), recorded over 4 channels (Mic 1: SHURE SM58; Mic 2: Labtec Axis-002; Mic 3: KOSS; Mic 4: ATR 60C). The database comprises digit strings. Speech samples are stored as a sequence of 16-bit 44.1kHz WAV for 8.7 hours of speech per channel. The total capacity of the data is 10 Gb.
Each speaker read 50 items. Text files are stored in Unicode format. All data have been proofread manually.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001363: Mandarin Chinese Desktop Speech Recognition Corpus - Monosyllabic (98 people)
Desktop/Microphone
This corpus comprises 1,267 entries uttered by 98 speakers of different dialects, ages and various educational levels (46 males and 52 females), recorded over 3 channels (Mic 1: SHURE SM58; Mic 2: Labtec Axis-002; Mic 3: ATR 60C). The database comprises monosyllables. Speech samples are stored as a sequence of 16-bit 44.1kHz WAV for 88 hours of speech per channel. The total capacity of the data is 78 Gb.
Text files are stored in Unicode format. All data have been proofread manually.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001364: Mandarin Chinese Desktop Speech Recognition Corpus - Person Name, Place Name (70 people)
Desktop/Microphone
This corpus comprises 9,667 entries uttered by 70 speakers of different dialects, ages and various educational levels (38 males and 32 females), recorded through head-mounted noise-canceling microphone. The database comprises 12,596 items. Speech samples are stored as a sequence of 16-bit 22.05kHz WAV for a total of 15 hours of speech. The total capacity of the data is 2.17 Gb.
Each speaker read 60 person names, 20 country names, 10 Chinese city names, 30 street names, 50 company and organization names, 10 geographical names. Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001365: Mandarin Chinese Desktop Speech Recognition Corpus - Person name (849 people)
Desktop/Microphone
This corpus comprises 2,250 entries uttered by 849 speakers of different dialects, ages and various educational levels (420 males and 429 females), recorded over 2 channels (Mic1: SHURE SM58; Mic2: Labtec Axis-002). The database comprises 12,750 person names per channel. Speech samples are stored as a sequence of 16-bit 44.1kHz WAV for 18.3 hours of speech per channel. The total capacity of the data is 11 Gb.
Each speaker read 15 items. Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001366: Mandarin Chinese Desktop Speech Recognition Corpus - Person name, Place Name (10 people)
Desktop/Microphone
This corpus comprises 782 entries uttered by 10 speakers of different dialects, ages and various educational levels (3 males and 7 females), recorded over 4 channels (Mic1: SHURE SM58; Mic2: ANC-700 Head-mounted; Mic3: TELEX M-60; Mic4: ACOUSTIC MAGIC). The database comprises 800 Chinese items per channel: 30 stocks, 10 nation names, 10 Chinese city names, 30 person names. Speech samples are stored as a sequence of 16-bit 22.05kHz WAV for 0.97 hours of speech per channel. The total capacity of the data is 587 Mb.
Each speaker read 120 items. Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001367: Mandarin Chinese Desktop Speech Recognition Corpus - SMS (120 people)
Desktop/Microphone
This corpus comprises 7,142 entries uttered by 120 speakers of different dialects, ages and various educational levels (59 males and 61 females), recorded through head-mounted noise-canceling microphone. The database comprises 16,499 short messages (SMS). Speech samples are stored as a sequence of 16-bit 22.05kHz WAV for 21.7 hours of speech. The total capacity of the data is 3.2 Gb.
Each speaker read 120-150 items. Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001368: Mandarin Chinese Desktop Speech Recognition Corpus - SMS (200 people)
Desktop/Microphone
This corpus comprises 7,276 entries uttered by 200 speakers of different dialects, ages and various educational levels (87 males and 113 females), recorded over 4 channels (Mic1: SHURE SM58; Mic2: ANC-700 Head-mounted; Mic3: TELEX M-60; Mic4: ACOUSTIC MAGIC). The database comprises 23,949 short messages (SMS) per channel. Speech samples are stored as a sequence of 16-bit 22.05kHz WAV for 35.6 hours of speech per channel. The total capacity of the data is 21.1 Gb.
Each speaker read 120 items. Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001369: Mandarin Chinese Desktop Speech Recognition Corpus - Simple Chinese sentences (850 people)
Desktop/Microphone
This corpus comprises 14,011 entries uttered by 850 speakers of different dialects, ages and various educational levels (420 males and 430 females), recorded over 2 channels (Mic1: SHURE SM58; Mic2: Labtec Axis-002). The database comprises 104,750 sentences per channel. Speech samples are stored as a sequence of 16-bit 44.1kHz WAV for 150 hours of speech per channel. The total capacity of the data is 88 Gb.
600 speakers read 120 sentences and 250 speakers read 131 sentences. Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001370: Mandarin Chinese Desktop Speech Recognition Corpus - Spontaneous Speech (50 people)
Desktop/Microphone
This corpus comprises spontaneous speech (elicited) from 50 speakers of different dialects, ages and various educational levels (21 males and 29 females), who uttered 36 different topics in a working environment, recorded through head-mounted noise-cancelling microphone. The database comprises 600 speech files. Speech samples are stored as a sequence of 16-bit 44.1kHz WAV for a total of 8 hours of speech. The total capacity of the data is 2.37 Gb.
Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.
C-001371: Mandarin Chinese Desktop Speech Recognition Corpus - Spontaneous Speech (849 people)
Desktop/Microphone
This corpus comprises spontaneous speech (elicited) from 849 speakers of different dialects, ages and various educational levels (420 males and 429 females), who uttered 40 different topics, recorded over 2 channels (Mic1: SHURE SM58; Mic2: Labtec Axis-002). Speech samples are stored as a sequence of 16-bit 44.1kHz WAV for a total of 208 hours of speech per channel. The total capacity of the data is 122.8 Gb.
Each speaker read 15 items. Text files are stored in Unicode format. All data have been proofread manually.
The transcriptions include non-speech markers (background noise, background speech, speaker sounds) as well as markers for mispronunciation, channel distortions, words left-out and duplicates.
The corpus aims to be applied to the testing and telephone natural speech recognition system.

SHACHI - Language Resource Metadata Database