TY - JOUR
T1 - Automatic methods for detecting sung lyrics error
AU - Tsai, Wei Ho
AU - Kung, Shiang Shiun
N1 - Publisher Copyright:
© 2020 Institute of Information Science. All rights reserved.
PY - 2020/5
Y1 - 2020/5
N2 - A sung lyrics error detection system is proposed to examine if the lyrics sung by a singer are incorrect, thereby providing a clue for singing skill evaluation. In essence, sung lyrics error detection is similar to the problem of speech utterance verification in the speech recognition research community, and therefore the techniques in the latter can be applied to the former. However, our experiment found that a speech utterance verification system is far from capable of handling singing data, mainly because of the significant difference between singing and speech. To tackle this problem, we develop two strategies, respectively, from a signal processing perspective and from a model processing perspective. In the signal processing perspective, we recognize that the vowels are often lengthened during singing, and thus propose vowel shrinking/decimation to adjust the length of a vowel in singing to a normal length in speaking. In the model processing perspective, we combine a duration modeling concept into the acoustic modeling to reduce the differences between singing and speech. Our experiments show that the proposed methods can improve the performance of the sung lyrics error detection noticeably, compared to a baseline system based on speech utterance verification.
AB - A sung lyrics error detection system is proposed to examine if the lyrics sung by a singer are incorrect, thereby providing a clue for singing skill evaluation. In essence, sung lyrics error detection is similar to the problem of speech utterance verification in the speech recognition research community, and therefore the techniques in the latter can be applied to the former. However, our experiment found that a speech utterance verification system is far from capable of handling singing data, mainly because of the significant difference between singing and speech. To tackle this problem, we develop two strategies, respectively, from a signal processing perspective and from a model processing perspective. In the signal processing perspective, we recognize that the vowels are often lengthened during singing, and thus propose vowel shrinking/decimation to adjust the length of a vowel in singing to a normal length in speaking. In the model processing perspective, we combine a duration modeling concept into the acoustic modeling to reduce the differences between singing and speech. Our experiments show that the proposed methods can improve the performance of the sung lyrics error detection noticeably, compared to a baseline system based on speech utterance verification.
KW - Duration modeling
KW - Singing
KW - Speech
KW - Sung lyrics
KW - Utterance verification
UR - https://www.scopus.com/pages/publications/85093906744
U2 - 10.6688/JISE.202005_36(3).0005
DO - 10.6688/JISE.202005_36(3).0005
M3 - ???researchoutput.researchoutputtypes.contributiontojournal.article???
AN - SCOPUS:85093906744
SN - 1016-2364
VL - 36
SP - 547
EP - 559
JO - Journal of Information Science and Engineering
JF - Journal of Information Science and Engineering
IS - 3
ER -