{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T20:31:25Z","timestamp":1761597085950,"version":"3.41.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2008,5,1]],"date-time":"2008-05-01T00:00:00Z","timestamp":1209600000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2008,5]]},"abstract":"<jats:p>Sports video has attracted a global viewership. Research effort in this area has been focused on semantic event detection in sports video to facilitate accessing and browsing. Most of the event detection methods in sports video are based on visual features. However, being a significant component of sports video, audio may also play an important role in semantic event detection. In this paper, we have borrowed the concept of the \u201ckeyword\u201d from the text mining domain to define a set of specific audio sounds. These specific audio sounds refer to a set of game-specific sounds with strong relationships to the actions of players, referees, commentators, and audience, which are the reference points for interesting sports events. Unlike low-level features, audio keywords can be considered as a mid-level representation, able to facilitate high-level analysis from the semantic concept point of view. Audio keywords are created from low-level audio features with learning by support vector machines. With the help of video shots, the created audio keywords can be used to detect semantic events in sports video by Hidden Markov Model (HMM) learning. Experiments on creating audio keywords and, subsequently, event detection based on audio keywords have been very encouraging. Based on the experimental results, we believe that the audio keyword is an effective representation that is able to achieve satisfying results for event detection in sports video. Application in three sports types demonstrates the practicality of the proposed method.<\/jats:p>","DOI":"10.1145\/1352012.1352015","type":"journal-article","created":{"date-parts":[[2008,5,15]],"date-time":"2008-05-15T18:28:05Z","timestamp":1210876085000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":43,"title":["Audio keywords generation for sports video analysis"],"prefix":"10.1145","volume":"4","author":[{"given":"Min","family":"Xu","sequence":"first","affiliation":[{"name":"University of Newcastle, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Changsheng","family":"Xu","sequence":"additional","affiliation":[{"name":"Institute for Infocom Research, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lingyu","family":"Duan","sequence":"additional","affiliation":[{"name":"Institute for Infocom Research, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jesse S.","family":"Jin","sequence":"additional","affiliation":[{"name":"University of Newcastle, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Suhuai","family":"Luo","sequence":"additional","affiliation":[{"name":"University of Newcastle, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2008,5,16]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the 13th International Conference on Pattern Recognition.","volume":"3","author":"Ardizzo E.","unstructured":"Ardizzo , E. , Cascia , M. L. , Gesu , V. D. , and Valenti , C . 1996. Content-based indexing of image and video databases by global and shape features . In Proceedings of the 13th International Conference on Pattern Recognition. Vol. 3 . 140--144. Ardizzo, E., Cascia, M. L., Gesu, V. D., and Valenti, C. 1996. Content-based indexing of image and video databases by global and shape features. In Proceedings of the 13th International Conference on Pattern Recognition. Vol. 3. 140--144."},{"volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo. 825--828","author":"Assfalg J.","key":"e_1_2_1_2_1","unstructured":"Assfalg , J. , Bertini , M. , Bimbo , A. D. , Nunziati , W. , and Pala , P . 2002. Soccer highlights detection and recognition using HMMs . In Proceedings of the IEEE International Conference on Multimedia and Expo. 825--828 . Assfalg, J., Bertini, M., Bimbo, A. D., Nunziati, W., and Pala, P. 2002. Soccer highlights detection and recognition using HMMs. In Proceedings of the IEEE International Conference on Multimedia and Expo. 825--828."},{"volume-title":"Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops.","author":"Baillie M.","key":"e_1_2_1_3_1","unstructured":"Baillie , M. and Jose , J. M . 2004. An audio-based sports video segmentation and event detection algorithm . In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops. Baillie, M. and Jose, J. M. 2004. An audio-based sports video segmentation and event detection algorithm. In Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/846220.1515084"},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Deller J. R. Hansen J. H. and Proakis J. G. 1999. Discrete-Time Processing of Speech Signals. Wiley-IEEE Computer Society.   Deller J. R. Hansen J. H. and Proakis J. G. 1999. Discrete-Time Processing of Speech Signals. Wiley-IEEE Computer Society.","DOI":"10.1109\/9780470544402"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2005.858395"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the 15th International Conference on Pattern Recognition.","volume":"1","author":"Gong Y.-H.","unstructured":"Gong , Y.-H. and Liu , X . 2000. Video shot segmentation and classification . In Proceedings of the 15th International Conference on Pattern Recognition. Vol. 1 . 860--863. Gong, Y.-H. and Liu, X. 2000. Video shot segmentation and classification. In Proceedings of the 15th International Conference on Pattern Recognition. Vol. 1. 860--863."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/76.988656"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2004.826751"},{"key":"e_1_2_1_10_1","first-page":"1066","article-title":"Multi-modal semantic analysis and annotation for basketball video","volume":"7","author":"Liu S.","year":"2005","unstructured":"Liu , S. , Xu , M. , Yi , H. , Chia , L. , and Rajan , D. 2005 . Multi-modal semantic analysis and annotation for basketball video . IEEE Trans. Multimedia 7 , 6, 1066 -- 1083 . Liu, S., Xu, M., Yi, H., Chia, L., and Rajan, D. 2005. Multi-modal semantic analysis and annotation for basketball video. IEEE Trans. Multimedia 7, 6, 1066--1083.","journal-title":"IEEE Trans. Multimedia"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 16th International Conference on Pattern Recognition.","volume":"2","author":"Miyauchi S.","unstructured":"Miyauchi , S. , Hirano , A. , Babaguchi , N. , and Kitahashi , T . 2002. Collaborative multimedia analysis for detecting semantical events from broadcasted sports video . In Proceedings of the 16th International Conference on Pattern Recognition. Vol. 2 . 1009--1012. Miyauchi, S., Hirano, A., Babaguchi, N., and Kitahashi, T. 2002. Collaborative multimedia analysis for detecting semantical events from broadcasted sports video. In Proceedings of the 16th International Conference on Pattern Recognition. Vol. 2. 1009--1012."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/500141.500181"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2001.941253"},{"key":"e_1_2_1_14_1","unstructured":"Rabiner L. R. and Juang B. H. 1993. Fundamentals of Speech Recognition. Prentice-Hall.   Rabiner L. R. and Juang B. H. 1993. Fundamentals of Speech Recognition. Prentice-Hall."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/354384.354443"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of IEEE International Conference on Multimedia and Expo (ICME).","volume":"2","author":"Sadlier D.","unstructured":"Sadlier , D. , Marlow , S. , O'Connor , N. , and Murphy , N . 2002. Mpeg audio bitstream processing towards the automatic generation of sports programme summaries . In Proceedings of IEEE International Conference on Multimedia and Expo (ICME). Vol. 2 . 77--80. Sadlier, D., Marlow, S., O'Connor, N., and Murphy, N. 2002. Mpeg audio bitstream processing towards the automatic generation of sports programme summaries. In Proceedings of IEEE International Conference on Multimedia and Expo (ICME). Vol. 2. 77--80."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2005.854237"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of IEEE Workshop on Content-based Access of Image and Video Libraries. 15--19","author":"Sebe N.","year":"2000","unstructured":"Sebe , N. , Tian , Q. , Loupias , E. , Lew , M. S. , and S. Huang , T. 2000 . Colour indexing using wavelet-based salient points . In Proceedings of IEEE Workshop on Content-based Access of Image and Video Libraries. 15--19 . Sebe, N., Tian, Q., Loupias, E., Lew, M. S., and S.Huang, T. 2000. Colour indexing using wavelet-based salient points. In Proceedings of IEEE Workshop on Content-based Access of Image and Video Libraries. 15--19."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of IEEE International Conference on Image Processing (ICIP).","volume":"3","author":"Chang W. C.","unstructured":"S.F. Chang , W. C. and Sundaram , H . 1998. Semantic visual templates: Linking features to semantics . In Proceedings of IEEE International Conference on Image Processing (ICIP). Vol. 3 . 531--535. S.F. Chang, W. C. and Sundaram, H. 1998. Semantic visual templates: Linking features to semantics. In Proceedings of IEEE International Conference on Image Processing (ICIP). Vol. 3. 531--535."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of IEEE International Conference on Multimedia and Expo.","volume":"3","author":"Snoek C. G. M.","unstructured":"Snoek , C. G. M. and Worring , M . 2003. Time interval maximum entropy based event indexing in soccer video . In Proceedings of IEEE International Conference on Multimedia and Expo. Vol. 3 . 481--484. Snoek, C. G. M. and Worring, M. 2003. Time interval maximum entropy based event indexing in soccer video. In Proceedings of IEEE International Conference on Multimedia and Expo. Vol. 3. 481--484."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of IEEE International Conference on Multimedia and Expo (ICME).","volume":"2","author":"Sundaram H.","unstructured":"Sundaram , H. and Chang , S . -F. 2000. Video scene segmentation using video and audio features . In Proceedings of IEEE International Conference on Multimedia and Expo (ICME). Vol. 2 . 1145--1148. Sundaram, H. and Chang, S.-F. 2000. Video scene segmentation using video and audio features. In Proceedings of IEEE International Conference on Multimedia and Expo (ICME). Vol. 2. 1145--1148."},{"volume-title":"Statistical Learning Theory","author":"Vapnik V.","key":"e_1_2_1_22_1","unstructured":"Vapnik , V. 1998. Statistical Learning Theory . Wiley . Vapnik, V. 1998. Statistical Learning Theory. Wiley."},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of IEEE International Conference on Multimedia Computing and Systems.","volume":"2","author":"Wei J.","unstructured":"Wei , J. , Li , Z.-N. , and Gertner , I . 1999. A novel motion-based active video indexing method . In Proceedings of IEEE International Conference on Multimedia Computing and Systems. Vol. 2 . 460--465. Wei, J., Li, Z.-N., and Gertner, I. 1999. A novel motion-based active video indexing method. In Proceedings of IEEE International Conference on Multimedia Computing and Systems. Vol. 2. 460--465."},{"volume-title":"Proceedings of IEEE International Conference on Multimedia and Expo (ICME). 805--808","author":"Wu C.","key":"e_1_2_1_24_1","unstructured":"Wu , C. , Ma , Y. , Zhang , H. , and Zhong , Y . 2002. Event recognition by semantic inference for sports video . In Proceedings of IEEE International Conference on Multimedia and Expo (ICME). 805--808 . Wu, C., Ma, Y., Zhang, H., and Zhong, Y. 2002. Event recognition by semantic inference for sports video. In Proceedings of IEEE International Conference on Multimedia and Expo (ICME). 805--808."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2004.01.005"},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of IEEE International Conference on Image Processing (ICIP).","volume":"1","author":"Xiong Z.","unstructured":"Xiong , Z. , Radhakrishnan , R. , and Divakaran , A . 2003. Generation of sports highlight using motion activities in combination with a common audio feature extraction framework . In Proceedings of IEEE International Conference on Image Processing (ICIP). Vol. 1 . 14--17. Xiong, Z., Radhakrishnan, R., and Divakaran, A. 2003. Generation of sports highlight using motion activities in combination with a common audio feature extraction framework. In Proceedings of IEEE International Conference on Image Processing (ICIP). Vol. 1. 14--17."},{"volume-title":"IEEE International Conference on Acoustics, Speech, and Signal Processing. V--632--V--635","author":"Xiong Z.","key":"e_1_2_1_27_1","unstructured":"Xiong , Z. , Radhakrishnan , R. , Divakaran , A. , and Huang , T . -S. 2003. Audio events detection based highlights extraction from baseball, golf and soccer games in a unified framework . In IEEE International Conference on Acoustics, Speech, and Signal Processing. V--632--V--635 . Xiong, Z., Radhakrishnan, R., Divakaran, A., and Huang, T.-S. 2003. Audio events detection based highlights extraction from baseball, golf and soccer games in a unified framework. In IEEE International Conference on Acoustics, Speech, and Signal Processing. V--632--V--635."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of IEEE Pacific-Rim Conference on Multimedia (PCM).","volume":"3","author":"Xu M.","unstructured":"Xu , M. , Duan , L.-Y. , Xu , C.-S. , Kankanhalli , M. , and Tian , Q . 2003. Event detection in basketball video using multiple modalities . In Proceedings of IEEE Pacific-Rim Conference on Multimedia (PCM). Vol. 3 . 189--192. Xu, M., Duan, L.-Y., Xu, C.-S., Kankanhalli, M., and Tian, Q. 2003. Event detection in basketball video using multiple modalities. In Proceedings of IEEE Pacific-Rim Conference on Multimedia (PCM). Vol. 3. 189--192."},{"volume-title":"Proceedings of IEEE International Conference on Acoustic, Speech, and Signal Processing (ICASSP).","author":"Xu M.","key":"e_1_2_1_29_1","unstructured":"Xu , M. , Duan , L.-Y. , Xu , C.-S. , and Tian , Q . 2003. A fusion scheme of visual and auditory modalities for event detection in sports video . In Proceedings of IEEE International Conference on Acoustic, Speech, and Signal Processing (ICASSP). Xu, M., Duan, L.-Y., Xu, C.-S., and Tian, Q. 2003. A fusion scheme of visual and auditory modalities for event detection in sports video. In Proceedings of IEEE International Conference on Acoustic, Speech, and Signal Processing (ICASSP)."},{"key":"e_1_2_1_30_1","volume-title":"Proceedings of the International Conference on Multimedia and Expo (ICME).","volume":"2","author":"Xu M.","unstructured":"Xu , M. , Maddage , N. C. , Xu , C.-S. , Kankanhalli , M. , and Tian , Q . 2003. Creating audio keywords for event detection in soccer video . In Proceedings of the International Conference on Multimedia and Expo (ICME). Vol. 2 . 281--284. Xu, M., Maddage, N. C., Xu, C.-S., Kankanhalli, M., and Tian, Q. 2003. Creating audio keywords for event detection in soccer video. In Proceedings of the International Conference on Multimedia and Expo (ICME). Vol. 2. 281--284."},{"key":"e_1_2_1_31_1","volume-title":"et al","author":"Young S.","year":"2002","unstructured":"Young , S. et al . 2002 . The HTK Book (for HTK Version 3.1) http:\/\/htk.eng.cam.edu\/. Cambridge University Engineering Department . Young, S. et al. 2002. The HTK Book (for HTK Version 3.1) http:\/\/htk.eng.cam.edu\/. Cambridge University Engineering Department."},{"volume-title":"Proceedings of the Storage and Retrieval for Image and Video Databases (SPIE). 389--398","author":"Zhang H. J.","key":"e_1_2_1_32_1","unstructured":"Zhang , H. J. , Smoliar , S. W. , and Wu , J. H . 1995. Content-based video browsing tools . In Proceedings of the Storage and Retrieval for Image and Video Databases (SPIE). 389--398 . Zhang, H. J., Smoliar, S. W., and Wu, J. H. 1995. Content-based video browsing tools. In Proceedings of the Storage and Retrieval for Image and Video Databases (SPIE). 389--398."},{"key":"e_1_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Zhang T. and Kuo C. C. J. 2001. Content-Based Audio Classification and Retrieval for Audiovisual Data Parsing. Kluwer Academic Publishers.   Zhang T. and Kuo C. C. J. 2001. Content-Based Audio Classification and Retrieval for Audiovisual Data Parsing. Kluwer Academic Publishers.","DOI":"10.1007\/978-1-4757-3339-6"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1352012.1352015","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1352012.1352015","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T13:57:59Z","timestamp":1750255079000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1352012.1352015"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,5]]},"references-count":33,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2008,5]]}},"alternative-id":["10.1145\/1352012.1352015"],"URL":"https:\/\/doi.org\/10.1145\/1352012.1352015","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2008,5]]},"assertion":[{"value":"2006-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2007-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-05-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}