{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,7,22]],"date-time":"2024-07-22T18:39:03Z","timestamp":1721673543909},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2013,9,2]],"date-time":"2013-09-02T00:00:00Z","timestamp":1378080000000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Image Video Proc"],"published-print":{"date-parts":[[2013,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>In large-scale multimedia event detection, complex target events are extracted from a large set of consumer-generated web videos taken in unconstrained environments. We devised a multimedia event detection method based on Gaussian mixture model (GMM) supervectors and support vector machines. A GMM supervector consists of the parameters of a GMM for the distribution of low-level features extracted from a video clip. A GMM is regarded as an extension of the bag-of-words framework to a probabilistic framework, and thus, it can be expected to be robust against the data insufficiency problem. We also propose a camera motion cancelled feature, which is a spatio-temporal feature robust against camera motions found in consumer-generated web videos. By combining these methods with the existing features, we aim to construct a high-performance event detection system. The effectiveness of our method is evaluated using TRECVID MED task benchmark.<\/jats:p>","DOI":"10.1186\/1687-5281-2013-51","type":"journal-article","created":{"date-parts":[[2013,9,2]],"date-time":"2013-09-02T12:14:07Z","timestamp":1378124047000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Event detection in consumer videos using GMM supervectors and SVMs"],"prefix":"10.1186","volume":"2013","author":[{"given":"Yusuke","family":"Kamishima","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nakamasa","family":"Inoue","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Koichi","family":"Shinoda","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2013,9,2]]},"reference":[{"key":"92_CR1","first-page":"825","volume-title":"IEEE International Conference on Multimedia and Expo, 2002","author":"J Assfalg","year":"2002","unstructured":"Assfalg J, Bertini M, Del Bimbo A, Nunziati W, Pala P: Soccer highlights detection and recognition using HMMs. In IEEE International Conference on Multimedia and Expo, 2002. Lausanne: IEEE; 26\u201329 August 2002:825-828."},{"issue":"8","key":"92_CR2","doi-asserted-by":"publisher","first-page":"1073","DOI":"10.1109\/TCSVT.2004.831968","volume":"14","author":"Y Li","year":"2004","unstructured":"Li Y, Narayanan S, Kuo C: Content-based movie analysis and indexing based on AudioVisual cues. Circuits Syst. Video Tech. IEEE Trans. 2004, 14(8):1073-1085. 10.1109\/TCSVT.2004.831968","journal-title":"Circuits Syst. Video Tech. IEEE Trans"},{"issue":"3","key":"92_CR3","doi-asserted-by":"publisher","first-page":"555","DOI":"10.1109\/TPAMI.2007.70825","volume":"30","author":"A Adam","year":"2008","unstructured":"Adam A, Rivlin E, Shimshoni I, Reinitz D: Robust real-time unusual event detection using multiple fixed-location monitors. Pattern Anal. Mach. Intell. IEEE Trans. 2008, 30(3):555-560.","journal-title":"Pattern Anal. Mach. Intell. IEEE Trans"},{"key":"92_CR4","first-page":"321","volume-title":"Proceedings of ACM International Workshop on Multimedia Information Retrieval, 2006","author":"A Smeaton","year":"2006","unstructured":"Smeaton A, Over P, Kraaij W: Evaluation campaigns and TRECVid. In Proceedings of ACM International Workshop on Multimedia Information Retrieval, 2006. Santa Barbara, CA: ACM; 26\u201327 October 2006:321-330."},{"key":"92_CR5","volume-title":"TRECVid multimedia event detection evaluation track","author":"The National Institute of Stantdards and Technology (NIST)","year":"2009","unstructured":"The National Institute of Stantdards and Technology (NIST): TRECVid multimedia event detection evaluation track. 2009. Accessed 15 Jan 2013 http:\/\/www.nist.gov\/itl\/iad\/mig\/med.cfm"},{"key":"92_CR6","volume-title":"Proceedings of TRECVID 2010 workshop","author":"Y Jiang","year":"2010","unstructured":"Jiang Y, Zeng X, Ye G, Bhattacharya S, Ellis D, Shah M, Chang S: Columbia-UCF TRECVID2010 multimedia event detection: combining multiple modalities, contextual concepts, and temporal matching. In Proceedings of TRECVID 2010 workshop. Gaithersburg, MD: NIST; November 2010."},{"key":"92_CR7","volume-title":"Proceedings of TRECVID 2010 workshop","author":"M Hill","year":"2010","unstructured":"Hill M, Hua G, Natsev A, Smith J, Xie L, Huang B, Merler M, Ouyang H, Zhou M: IBM research TRECVID-2010 video copy detection and multimedia event detection system. In Proceedings of TRECVID 2010 workshop. Gaithersburg, MD: NIST; November 2010."},{"key":"92_CR8","volume-title":"Proceedings of TRECVID 2011 workshop","author":"L Bao","year":"2011","unstructured":"Bao L, Yu S, Lan Z, Overwijk A, Jin Q, Langner B, Garbus M, Burger S, Metze F, Hauptmann A: Informedia@ TRECVID 2011. In Proceedings of TRECVID 2011 workshop. Gaithersburg, MD: NIST; December 2011."},{"key":"92_CR9","doi-asserted-by":"publisher","first-page":"1298","DOI":"10.1109\/CVPR.2012.6247814","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012","author":"P Natarajan","year":"2012","unstructured":"Natarajan P, Wu S, Vitaladevuni S, Zhuang X, Tsakalidis S, Park U, Prasad R: Multimodal feature fusion for robust event detection in web videos. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012. Providence, RI: IEEE; 16\u201321 June 2012:1298-1305."},{"key":"92_CR10","first-page":"59","volume-title":"Proceedings of IEEE European Conference on Computer Vision, 2004","author":"G Csurka","year":"2004","unstructured":"Csurka G, Dance CR, Fan L, Willamowski J, Bray C: Visual categorization with bags of keypoints. In Proceedings of IEEE European Conference on Computer Vision, 2004. Prague: IEEE; 11\u201314 May 2004:59-74."},{"key":"92_CR11","first-page":"197","volume-title":"Proceedings of ACM Multimedia MIR Workshop, 2007","author":"J Yang","year":"2007","unstructured":"Yang J, Jiang YG, Hauptmann AG, Ngo CW: Evaluating bag-of-visual-words representations in scene classification. In Proceedings of ACM Multimedia MIR Workshop, 2007. Augsburg: ACM; 24\u201329 September 2007:197-206."},{"key":"92_CR12","first-page":"50","volume-title":"Proceedings of the 2nd ACM International Conference on Multimedia Retrieval","author":"Z Ma","year":"2012","unstructured":"Ma Z, Hauptmann A, Yang Y, Sebe N: Classifier-specific intermediate representation for multimedia tasks. In Proceedings of the 2nd ACM International Conference on Multimedia Retrieval. Hong Kong: ACM; 05\u201308 June 2012:50-50."},{"key":"92_CR13","doi-asserted-by":"publisher","first-page":"1420","DOI":"10.1109\/ICCV.2009.5459295","volume-title":"Proceedings of IEEE International Conference on Computer Vision, 2009","author":"Y Jiang","year":"2009","unstructured":"Jiang Y, Wang J, Chang S, Ngo C: Domain adaptive semantic diffusion for large scale context-based video annotation. In Proceedings of IEEE International Conference on Computer Vision, 2009. Kyoto: IEEE; 27 September\u201304 October2009:1420-1427."},{"issue":"4","key":"92_CR14","doi-asserted-by":"publisher","first-page":"1196","DOI":"10.1109\/TMM.2012.2191395","volume":"14","author":"N Inoue","year":"2012","unstructured":"Inoue N, Shinoda K: A fast and accurate video semantic-indexing system using fast MAP adaptation and GMM supervectors. Multimedia, IEEE Trans. 2012, 14(4):1196-1205.","journal-title":"Multimedia, IEEE Trans"},{"key":"92_CR15","first-page":"1357","volume-title":"Proceedings of ACM Multimedia, 2011","author":"N Inoue","year":"2011","unstructured":"Inoue N, Shinoda K: A fast MAP adaptation technique for GMM-supervector-based video semantic indexing systems. In Proceedings of ACM Multimedia, 2011. Scottsdale, AZ: ACM; 28 November\u201301 December 2011:1357-1360."},{"key":"92_CR16","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2008","author":"Y Liu","year":"2008","unstructured":"Liu Y, Perronnin F: A similarity measure between unordered vector sets with application to image categorization. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2008. Anchorage, AL: IEEE; 24\u201326 June 2008."},{"key":"92_CR17","first-page":"141","volume-title":"Proceedings of IEEE European Conference on Computer Vision, 2010","author":"X Zhou","year":"2010","unstructured":"Zhou X, Yu K, Zhang T, Huang T: Image classification using super-vector coding of local image descriptors. In Proceedings of IEEE European Conference on Computer Vision, 2010. Heraklion: IEEE; 5\u201311 September 2010:141-154."},{"key":"92_CR18","doi-asserted-by":"publisher","first-page":"3089","DOI":"10.1109\/ICIP.2012.6467553","volume-title":"Proceedings of IEEE International Conference on Image Processing, 2012","author":"Y Kamishima","year":"2012","unstructured":"Kamishima Y, Inoue N, Shinoda K, Sato S: Multimedia event detection using GMM supervectors and SVMS. In Proceedings of IEEE International Conference on Image Processing, 2012. Orlando, FL: IEEE; 30 September\u201303 October 2012:3089-3092."},{"issue":"2","key":"92_CR19","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","volume":"60","author":"D Lowe","year":"2004","unstructured":"Lowe D: Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. 2004, 60(2):91-110.","journal-title":"Int. J. Comput. Vis"},{"key":"92_CR20","first-page":"229","volume-title":"Proceedings of ACM Multimedia, 2008","author":"X Zhou","year":"2008","unstructured":"Zhou X, Zhuang X, Yan S, Chang S, Hasegawa-Johnson M, Huang T: SIFT-bag kernel for video event analysis. In Proceedings of ACM Multimedia, 2008. Vancouver: ACM; 27\u201331 October 2008:229-238."},{"key":"92_CR21","first-page":"490","volume-title":"Proceedings of IEEE European Conference on Computer Vision, 2006","author":"E Nowak","year":"2006","unstructured":"Nowak E, Jurie F, Triggs B: Sampling strategies for bag-of-features image classification. In Proceedings of IEEE European Conference on Computer Vision, 2006. Austria: IEEE; 7\u201313 May 2006:490-503."},{"issue":"3","key":"92_CR22","doi-asserted-by":"publisher","first-page":"346","DOI":"10.1016\/j.cviu.2007.09.014","volume":"110","author":"H Bay","year":"2008","unstructured":"Bay H, Ess A, Tuytelaars T, Van Gool L: SURF: speeded-up robust features (SURF). Comput. Vis. Image Underst. 2008, 110(3):346-359. 10.1016\/j.cviu.2007.09.014","journal-title":"Comput. Vis. Image Underst"},{"key":"92_CR23","doi-asserted-by":"publisher","first-page":"3220","DOI":"10.1109\/ICPR.2010.787","volume-title":"Proceedings of International Conference on Pattern Recognition, 2010","author":"N Inoue","year":"2010","unstructured":"Inoue N, Tatsuhiko S, Shinoda K, Furui S: High-level feature extraction using SIFT GMMs and audio models. In Proceedings of International Conference on Pattern Recognition, 2010. Istanbul: IAPR; 23\u201326 August 2010:3220-3223."},{"issue":"2","key":"92_CR24","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1007\/s11263-005-1838-7","volume":"64","author":"I Laptev","year":"2005","unstructured":"Laptev I: On space-time interest points. Int. J. Comput. Vis. 2005, 64(2):107-123.","journal-title":"Int. J. Comput. Vis"},{"key":"92_CR25","volume-title":"Proceedings of British Machine Vision Conference, 2009","author":"H Wang","year":"2009","unstructured":"Wang H, Ullah M, Klaser A, Laptev I, Schmid C: Evaluation of local spatio-temporal features for action recognition. In Proceedings of British Machine Vision Conference, 2009. London: BMVA; 7\u201310 September 2009."},{"key":"92_CR26","volume-title":"Proceedings of British Machine Vision Conference, 2008","author":"A Kl\u00e4ser","year":"2008","unstructured":"Kl\u00e4ser A, Marszalek M, Schmid C: A spatio-temporal descriptor based on 3Dgradients. In Proceedings of British Machine Vision Conference, 2008. Leeds: BMVA; September 2008."},{"key":"92_CR27","volume-title":"CMU-CS-09-161","author":"M Chen","year":"2009","unstructured":"Chen M, Hauptmann A: Mosift: recognizing human actions in surveillance videos. In CMU-CS-09-161. Carnegie Mellon University; 2009."},{"key":"92_CR28","doi-asserted-by":"publisher","first-page":"1419","DOI":"10.1109\/ICCV.2011.6126397","volume-title":"Proceedings of IEEE International Conference on Computer Vision, 2011","author":"S Wu","year":"2011","unstructured":"Wu S, Oreifej O, Shah M: Action recognition in videos acquired by a moving camera using motion decomposition of lagrangian particle trajectories. In Proceedings of IEEE International Conference on Computer Vision, 2011. Barcelona: IEEE; 6\u201313 November 2011:1419-1426."},{"key":"92_CR29","first-page":"3169","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2011","author":"H Wang","year":"2011","unstructured":"Wang H, Kl\u00e4ser A, Schmid C, Liu CL: Action recognition by dense trajectories. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2011. Colorado Springs, CO: IEEE; 20\u201325 June 2011:3169-3176."},{"issue":"3","key":"92_CR30","doi-asserted-by":"publisher","first-page":"426","DOI":"10.1016\/j.cviu.2010.11.002","volume":"115","author":"K Mikolajczyk","year":"2011","unstructured":"Mikolajczyk K, Uemura H: Action recognition with appearance - motion features and fast search trees. Comput. Vis. Image Underst. 2011, 115(3):426-438. 10.1016\/j.cviu.2010.11.002","journal-title":"Comput. Vis. Image Underst"},{"key":"92_CR31","first-page":"494","volume-title":"Proceedings of IEEE European Conference on Computer Vision, 2010","author":"N Ikizler-Cinbis","year":"2010","unstructured":"Ikizler-Cinbis N, Sclaroff S: Object, scene and actions: combining multiple features for human action recognition. In Proceedings of IEEE European Conference on Computer Vision, 2010. Heraklion: IEEE; 5\u201311 September 2010:494-507."},{"key":"92_CR32","volume-title":"BMVC","author":"K Li","year":"2012","unstructured":"Li K, Oh S, Perera A, Fu Y: A videography analysis framework for video retrieval and summarization. In BMVC. BMVA; 2012."},{"key":"92_CR33","volume-title":"Proceedings of TRECVID 2011 workshop","author":"N Inoue","year":"2011","unstructured":"Inoue N, Wada T, Kamishima Y, Shinoda K, Sato S: TokyoTech+Canon at TRECVID 2011. In Proceedings of TRECVID 2011 workshop. Gaithersburg, MD: NIST; December 2011."},{"key":"92_CR34","doi-asserted-by":"publisher","first-page":"63","DOI":"10.1023\/B:VISI.0000027790.02288.f2","volume":"60","author":"K Mikolajczyk","year":"2004","unstructured":"Mikolajczyk K, Schmid C: Scale & affine invariant interest point detectors. Int. J. Comput. Vis 2004, 60: 63-86.","journal-title":"Int. J. Comput. Vis"},{"key":"92_CR35","first-page":"428","volume-title":"Proceedings of IEEE European Conference on Computer Vision, 2006","author":"N Dalal","year":"2006","unstructured":"Dalal N, Triggs B, Schmid C: Human detection using oriented histograms of flow and appearance. In Proceedings of IEEE European Conference on Computer Vision, 2006. Austria: IEEE; 7\u201313 May 2006:428-441."},{"issue":"9","key":"92_CR36","doi-asserted-by":"publisher","first-page":"1582","DOI":"10.1109\/TPAMI.2009.154","volume":"32","author":"KEA van de Sande","year":"2010","unstructured":"van de Sande KEA, Gevers T, Snoek CGM: Evaluating color descriptors for object and scene recognition. Pattern Anal. Mach. Int. IEEE Trans 2010, 32(9):1582-1596.","journal-title":"Pattern Anal. Mach. Int. IEEE Trans"},{"key":"92_CR37","doi-asserted-by":"publisher","first-page":"2169","DOI":"10.1109\/CVPR.2006.68","volume-title":"Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2006","author":"S Lazebnik","year":"2006","unstructured":"Lazebnik S, Schmid C, Ponce J: Beyond bags of features: spatial pyramid matching for recognizing natural scene categories. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2006. New York, NY: IEEE; 17\u201322 June 2006:2169-2178."},{"key":"92_CR38","first-page":"97","volume-title":"Proceedings of IEEE International Conference on Acoustics, Speech, and Signal Processing, 2006","author":"W Campbell","year":"2006","unstructured":"Campbell W, Sturim D, Reynolds D, Solomonoff A: SVM based speaker verification using a GMM supervector kernel and nap variability compensation. In Proceedings of IEEE International Conference on Acoustics, Speech, and Signal Processing, 2006. Graz: IEEE; 14\u201319 May 2006:97-100."},{"issue":"2","key":"92_CR39","doi-asserted-by":"publisher","first-page":"291","DOI":"10.1109\/89.279278","volume":"2","author":"J Gauvain","year":"1994","unstructured":"Gauvain J, Lee CH: Maximum a posteriori estimation for multivariate Gaussian mixture observations of Markov chains. Speech Audio Proc. IEEE Trans 1994, 2(2):291-298. 10.1109\/89.279278","journal-title":"Speech Audio Proc. IEEE Trans"},{"key":"92_CR40","volume-title":"SIFT++","author":"A Vedaldi","year":"2006","unstructured":"Vedaldi A: SIFT++. 2006. Accessed 15 Jan 2013 http:\/\/www.vlfeat.org\/~vedaldi\/code\/siftpp.html"},{"key":"92_CR41","volume-title":"The hkt book","author":"SJ Young","year":"2006","unstructured":"Young SJ, Evermann G, Gales MJF, Kershaw D, Moore G, Odell JJ, Ollason DG, Povey D, Valtchev D, Woodland PC: The hkt book. 2006. Accessed 15 Jan 2013 http:\/\/htk.eng.cam.ac.uk"},{"key":"92_CR42","volume-title":"Libsvm: A library for support vector machines","author":"CC Chang","year":"2001","unstructured":"Chang CC, Lin CJ: Libsvm: A library for support vector machines. 2001. Accessed 15 Jan 2013 http:\/\/www.csie.ntu.edu.tw\/~cjlin\/libsvm\/"},{"key":"92_CR43","first-page":"1150","volume-title":"Proceedings of IEEE International Conference on Computer Vision, 1999","author":"DG Lowe","year":"1999","unstructured":"Lowe DG: Object recognition from local scale-invariant features. In Proceedings of IEEE International Conference on Computer Vision, 1999. Corfu: IEEE; 20\u201325 September 1999:1150-1157."},{"issue":"2","key":"92_CR44","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1023\/A:1026543900054","volume":"40","author":"Y Rubner","year":"2000","unstructured":"Rubner Y, Tomasi C, Guibas L: The earth mover\u2019s distance as a metric for image retrieval. Int. J. Comput. Vis 2000, 40(2):99-121. 10.1023\/A:1026543900054","journal-title":"Int. J. Comput. Vis"},{"key":"92_CR45","first-page":"151","volume-title":"Multimedia Content Analysis","author":"A Smeaton","year":"2009","unstructured":"Smeaton A, Over P, Kraaij W: High-level feature detection from video in TRECVid: a 5-year retrospective of achievements. In Multimedia Content Analysis. Norwell: Springer US; 2009:151-174."},{"key":"92_CR46","first-page":"143","volume-title":"Proceedings of IEEE European Conference on Computer Vision, 2010","author":"F Perronnin","year":"2010","unstructured":"Perronnin F, S\u00e1nchez J, Mensink T: Improving the fisher kernel for large-scale image classification. In Proceedings of IEEE European Conference on Computer Vision, 2010. Heraklion: IEEE; 5\u201311 September 2010:143-156."}],"container-title":["EURASIP Journal on Image and Video Processing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/1687-5281-2013-51.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/1687-5281-2013-51\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1687-5281-2013-51.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,2]],"date-time":"2021-09-02T01:19:17Z","timestamp":1630545557000},"score":1,"resource":{"primary":{"URL":"https:\/\/jivp-eurasipjournals.springeropen.com\/articles\/10.1186\/1687-5281-2013-51"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,9,2]]},"references-count":46,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2013,12]]}},"alternative-id":["92"],"URL":"https:\/\/doi.org\/10.1186\/1687-5281-2013-51","relation":{},"ISSN":["1687-5281"],"issn-type":[{"value":"1687-5281","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,9,2]]},"assertion":[{"value":"16 January 2013","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 July 2013","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 September 2013","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"51"}}