{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,11]],"date-time":"2026-04-11T05:02:15Z","timestamp":1775883735587,"version":"3.50.1"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2009,5,1]],"date-time":"2009-05-01T00:00:00Z","timestamp":1241136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2009,5]]},"abstract":"<jats:p>Over the past decade, there has been explosive growth in the availability of multimedia data, particularly image, video, and music. Because of this, content-based music retrieval has attracted attention from the multimedia database and information retrieval communities. Content-based music retrieval requires us to be able to automatically identify particular characteristics of music data. One such characteristic, useful in a range of applications, is the identification of the singer in a musical piece. Unfortunately, existing approaches to this problem suffer from either low accuracy or poor scalability. In this article, we propose a novel scheme, called<jats:italic>Hybrid Singer Identifier<\/jats:italic>(HSI), for efficient automated singer recognition. HSI uses multiple low-level features extracted from both vocal and nonvocal music segments to enhance the identification process; it achieves this via a hybrid architecture that builds profiles of individual singer characteristics based on statistical mixture models. An extensive experimental study on a large music database demonstrates the superiority of our method over state-of-the-art approaches in terms of effectiveness, efficiency, scalability, and robustness.<\/jats:p>","DOI":"10.1145\/1508850.1508856","type":"journal-article","created":{"date-parts":[[2009,5,19]],"date-time":"2009-05-19T16:47:42Z","timestamp":1242751662000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":33,"title":["A novel framework for efficient automated singer identification in large music databases"],"prefix":"10.1145","volume":"27","author":[{"given":"Jialie","family":"Shen","sequence":"first","affiliation":[{"name":"Singapore Management University, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"John","family":"Shepherd","sequence":"additional","affiliation":[{"name":"The University of New South Wales, Sidney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bin","family":"Cui","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kian-Lee","family":"Tan","sequence":"additional","affiliation":[{"name":"National University of Singapore, Kent Ridge, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,5,19]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2003.822637"},{"key":"e_1_2_1_2_1","unstructured":"Becchetti C. Ricotti L. and Ricotti L. 1999. Speech Recognition. John Wiley New York NY. Becchetti C. Ricotti L. and Ricotti L. 1999. Speech Recognition. John Wiley New York NY."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the AES 22nd International Conference on Virtual, Synthetic and Entertainment Audio. 119--122","author":"Berenzweig A.","unstructured":"Berenzweig , A. , Ellis , D. P. W. , and Lawrence , S . 2002. Using voice segments to improve artist classification of music . In Proceedings of the AES 22nd International Conference on Virtual, Synthetic and Entertainment Audio. 119--122 . Berenzweig, A., Ellis, D. P. W., and Lawrence, S. 2002. Using voice segments to improve artist classification of music. In Proceedings of the AES 22nd International Conference on Virtual, Synthetic and Entertainment Audio. 119--122."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1162\/014892604323112257"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. 119--122","author":"Berenzweig A. L.","unstructured":"Berenzweig , A. L. and Ellis , D. P. W. 2001. Locating singing voice segments within music signals . In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. 119--122 . Berenzweig, A. L. and Ellis, D. P. W. 2001. Locating singing voice segments within music signals. In Proceedings of the IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. 119--122."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/100216.100224"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2002.1023800"},{"key":"e_1_2_1_8_1","volume-title":"-J","author":"Chang C.-C.","year":"2001","unstructured":"Chang , C.-C. and Lin , C . -J . 2001 . LIBSVM : A library for support vector machines. http:\/\/www.csie.ntu.edu.tw\/~cjlin\/libsvm. Chang, C.-C. and Lin, C.-J. 2001. LIBSVM: A library for support vector machines. http:\/\/www.csie.ntu.edu.tw\/~cjlin\/libsvm."},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 13rd Annual Conference on Computational Learning Theory (COLT'00)","author":"Collins M.","unstructured":"Collins , M. , Schapire , R. E. , and Singer , Y . 2000. Logistic regression, Adaboost and Bregman distances . In Proceedings of the 13rd Annual Conference on Computational Learning Theory (COLT'00) . 158--169. Collins, M., Schapire, R. E., and Singer, Y. 2000. Logistic regression, Adaboost and Bregman distances. In Proceedings of the 13rd Annual Conference on Computational Learning Theory (COLT'00). 158--169."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1065385.1065479"},{"key":"e_1_2_1_11_1","first-page":"12","article-title":"The Music Information Retrieval Evaluation Exchange (MIREX)","volume":"12","author":"Downie J. S.","year":"2006","unstructured":"Downie , J. S. 2006 . The Music Information Retrieval Evaluation Exchange (MIREX) . D-Lib Mag. 12 , 12 (Dec.) Downie, J. S. 2006. The Music Information Retrieval Evaluation Exchange (MIREX). D-Lib Mag. 12, 12 (Dec.)","journal-title":"D-Lib Mag."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 6th International Conference on Music Information Retrieval (ISMIR). 320--323","author":"Downie J. S.","unstructured":"Downie , J. S. , West , K. , Ehmann , A. , and Vincent , E . 2005b. The 2005 Music Information Retrieval Evaluation Exchange (MIREX 2005) preliminary overview . In Proceedings of the 6th International Conference on Music Information Retrieval (ISMIR). 320--323 . Downie, J. S., West, K., Ehmann, A., and Vincent, E. 2005b. The 2005 Music Information Retrieval Evaluation Exchange (MIREX 2005) preliminary overview. In Proceedings of the 6th International Conference on Music Information Retrieval (ISMIR). 320--323."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/948383.948386"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcss.1997.1504"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1006\/cviu.2001.0946"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Hastie T. Tibshirani R. and Friedman J. 2001. The Elements of Statistical Learning: Data Mining Inference and Prediction. Springer Verlag Berlin Germany. Hastie T. Tibshirani R. and Friedman J. 2001. The Elements of Statistical Learning: Data Mining Inference and Prediction. Springer Verlag Berlin Germany.","DOI":"10.1007\/978-0-387-21606-5"},{"key":"e_1_2_1_17_1","volume-title":"The Fifth International Conference on Music Information Retrieval. http:\/\/ismir2004","author":"ISMIR.","year":"2004","unstructured":"ISMIR. 2004 . The Fifth International Conference on Music Information Retrieval. http:\/\/ismir2004 .ismir.net\/index.html. ISMIR. 2004. The Fifth International Conference on Music Information Retrieval. http:\/\/ismir2004.ismir.net\/index.html."},{"key":"e_1_2_1_18_1","volume-title":"Why the logistic function? a tutorial discussion on probabilities and neural networks. Tech. rep. 9503","author":"Jordan M. I.","unstructured":"Jordan , M. I. 1995. Why the logistic function? a tutorial discussion on probabilities and neural networks. Tech. rep. 9503 . MIT , Cambridge, MA . Jordan, M. I. 1995. Why the logistic function? a tutorial discussion on probabilities and neural networks. Tech. rep. 9503. MIT, Cambridge, MA."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 3rd International Conference Music on Information Retrieval (ISMIR). 164--169","author":"Kim Y. E.","unstructured":"Kim , Y. E. and Whitman , B . 2002. Singer identification in popular music recordings using voice coding features . In Proceedings of the 3rd International Conference Music on Information Retrieval (ISMIR). 164--169 . Kim, Y. E. and Whitman, B. 2002. Singer identification in popular music recordings using voice coding features. In Proceedings of the 3rd International Conference Music on Information Retrieval (ISMIR). 164--169."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 7th International Conference Music Information Retrieval (ISMIR'06)","author":"Kim Y. E.","unstructured":"Kim , Y. E. , Williamson , D. , and Pilli , S . 2006. Towards quantifying the album effect in artist identification . In Proceedings of the 7th International Conference Music Information Retrieval (ISMIR'06) . 393--394. Kim, Y. E., Williamson, D., and Pilli, S. 2006. Towards quantifying the album effect in artist identification. In Proceedings of the 7th International Conference Music Information Retrieval (ISMIR'06). 393--394."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/381641.381658"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Lebanon G. and Lafferty J. 2001. Boosting and maximum likelihood for exponential model and Bregman distances. In Advances in Neural Information Processing Systems 14 (Proceedings of NIPS). 110--121. Lebanon G. and Lafferty J. 2001. Boosting and maximum likelihood for exponential model and Bregman distances. In Advances in Neural Information Processing Systems 14 (Proceedings of NIPS). 110--121.","DOI":"10.7551\/mitpress\/1120.003.0062"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1027527.1027612"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/860435.860487"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/584792.584864"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 7th International Conference on Digital Audio Effects (DAFx). 222--227","author":"Livshin A.","unstructured":"Livshin , A. and Rodet , X . 2004. Musical instrument identification in continuous recordings . In Proceedings of the 7th International Conference on Digital Audio Effects (DAFx). 222--227 . Livshin, A. and Rodet, X. 2004. Musical instrument identification in continuous recordings. In Proceedings of the 7th International Conference on Digital Audio Effects (DAFx). 222--227."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-002-0065-0"},{"key":"e_1_2_1_28_1","unstructured":"MIREX. 2005. Artist identification contest track. http:\/\/www.music-ir.org\/evaluation\/mirex-results\/audio-artist\/index.html. MIREX. 2005. Artist identification contest track. http:\/\/www.music-ir.org\/evaluation\/mirex-results\/audio-artist\/index.html."},{"key":"e_1_2_1_29_1","unstructured":"MIREX. 2007. Artist identification contest track. http:\/\/www.music-ir.org\/mirex2007\/index.php\/AudioArtistIdentificationResults. MIREX. 2007. Artist identification contest track. http:\/\/www.music-ir.org\/mirex2007\/index.php\/AudioArtistIdentificationResults."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/641205.641207"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1145287.1145309"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1281485.1281490"},{"key":"e_1_2_1_33_1","unstructured":"Rabiner L. and Juang B. 1993. Fundamentals of Speech Recognition. Prentice-Hall Englewood Cliffs NJ. Rabiner L. and Juang B. 1993. Fundamentals of Speech Recognition. Prentice-Hall Englewood Cliffs NJ."},{"key":"e_1_2_1_34_1","unstructured":"Rabiner L. and Schafer R. 1978. Digital Processing of Speech Signals. Prentice-Hall Englewood Cliffs NJ. Rabiner L. and Schafer R. 1978. Digital Processing of Speech Signals. Prentice-Hall Englewood Cliffs NJ."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/0005-1098(78)90005-5"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2006.79"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/89.876309"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2005.854091"},{"key":"e_1_2_1_39_1","volume-title":"Proceedings of the 4th international Conference on Music Information Retrieval (ISMIR). 167--173","author":"Tsai W. H.","unstructured":"Tsai , W. H. , Wang , H. M. , Rodgers , D. , Cheng , S. S. , and Yu , H. M . 2003. Blind clustering of popular music recordings based on singer voice characteristics . In Proceedings of the 4th international Conference on Music Information Retrieval (ISMIR). 167--173 . Tsai, W. H., Wang, H. M., Rodgers, D., Cheng, S. S., and Yu, H. M. 2003. Blind clustering of popular music recordings based on singer voice characteristics. In Proceedings of the 4th international Conference on Music Information Retrieval (ISMIR). 167--173."},{"key":"e_1_2_1_40_1","volume-title":"Statistical Learning Theory","author":"Vapnik V.","unstructured":"Vapnik , V. 1998. Statistical Learning Theory . John Wiley & amp; Sons. New York, NY. Vapnik, V. 1998. Statistical Learning Theory. John Wiley &amp; Sons. New York, NY."},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the IEEE Workshop on Neural Networks for Signal Processing. 559--568","author":"Whitman B.","unstructured":"Whitman , B. , Flake , G. , and Lawrence , S . 2001. Artist detection in music with Minnowmatch . In Proceedings of the IEEE Workshop on Neural Networks for Signal Processing. 559--568 . Whitman, B., Flake, G., and Lawrence, S. 2001. Artist detection in music with Minnowmatch. In Proceedings of the IEEE Workshop on Neural Networks for Signal Processing. 559--568."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSA.2004.840939"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.5555\/1170745.1171614"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1508850.1508856","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1508850.1508856","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:29:42Z","timestamp":1750253382000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1508850.1508856"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,5]]},"references-count":43,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2009,5]]}},"alternative-id":["10.1145\/1508850.1508856"],"URL":"https:\/\/doi.org\/10.1145\/1508850.1508856","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,5]]},"assertion":[{"value":"2006-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-05-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}