{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T15:12:30Z","timestamp":1781104350791,"version":"3.54.1"},"reference-count":48,"publisher":"IGI Global Scientific Publishing","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2012,10,1]]},"abstract":"<p>Multimedia data by its very nature contains multimodal information in it. For a successful analysis of multimedia content, all available multimodal information should be utilized. Additionally, since concepts can contain valuable cues about other concepts, concept interaction is a crucial source of multimedia information and helps to increase the fusion performance. The aim of this study is to show that integrating existing modalities along with the concept interactions can yield a better performance in detecting semantic concepts. Therefore, in this paper, the authors present a multimodal fusion approach that integrates semantic information obtained from various modalities along with additional semantic cues. The experiments conducted on TRECVID 2007 and CCV Database datasets validates the superiority of such combination over best single modality and alternative modality combinations. The results show that the proposed fusion approach provides 16.7% relative performance gain on TRECVID dataset and 47.7% relative performance improvement on CCV database over the results of best unimodal approaches.<\/p>","DOI":"10.4018\/jmdem.2012100103","type":"journal-article","created":{"date-parts":[[2013,2,27]],"date-time":"2013-02-27T12:26:23Z","timestamp":1361967983000},"page":"52-74","source":"Crossref","is-referenced-by-count":2,"title":["Multimodal Information Fusion for Semantic Video Analysis"],"prefix":"10.4018","volume":"3","author":[{"given":"Elvan","family":"Gulen","sequence":"first","affiliation":[{"name":"Department of Computer Engineering, Middle East Technical University, Ankara, Turkey"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Turgay","family":"Yilmaz","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Middle East Technical University, Ankara, Turkey and Institute of Industrial Science, University of Tokyo, Tokyo, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adnan","family":"Yazici","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Middle East Technical University, Ankara, Turkey"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"jmdem.2012100103-0","doi-asserted-by":"publisher","DOI":"10.1155\/S1110865703211173"},{"key":"jmdem.2012100103-1","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-010-0182-0"},{"key":"jmdem.2012100103-2","doi-asserted-by":"crossref","unstructured":"Ayache, S., Qu\u00e9not, G., & Gensel, J. (2007). Classifier fusion for SVM-based multimedia semantic indexing. In Proceedings of the 29th European conference on IR research (pp. 494-504). Berlin, Heidelberg, Germany: Springer-Verlag.","DOI":"10.1007\/978-3-540-71496-5_44"},{"key":"jmdem.2012100103-3","doi-asserted-by":"publisher","DOI":"10.1109\/6046.985555"},{"key":"jmdem.2012100103-4","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009715923555"},{"key":"jmdem.2012100103-5","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2011.04.037"},{"key":"jmdem.2012100103-6","doi-asserted-by":"publisher","DOI":"10.1023\/A:1023622605600"},{"key":"jmdem.2012100103-7","doi-asserted-by":"crossref","unstructured":"Chang, C., & Lin, C. (2011). LIBSVM: A library for support vector machines. ACM Transactions on Intelligent Systems and Technology (TIST), 2(3), 27:1-27:27.","DOI":"10.1145\/1961189.1961199"},{"key":"jmdem.2012100103-8","doi-asserted-by":"crossref","unstructured":"Datta, R., Li, J., & Wang, J. Z. (2005). Content-based image retrieval: Approaches and trends of the new age. In Proceedings of the 7th ACM SIGMM International Workshop on Multimedia Information Retrieval (pp. 253-262). New York, NY: ACM Press.","DOI":"10.1145\/1101826.1101866"},{"key":"jmdem.2012100103-9","doi-asserted-by":"publisher","DOI":"10.1007\/s10791-011-9170-z"},{"key":"jmdem.2012100103-10","first-page":"405","article-title":"Particle swarm model selection.","volume":"10","author":"H.Escalante","year":"2009","journal-title":"Journal of Machine Learning Research"},{"key":"jmdem.2012100103-11","doi-asserted-by":"crossref","unstructured":"Escalante, H. J., Montes, M., & Sucar, L. E. (2007). Word co-occurrence and Markov random fields for improving automatic image annotation. In Proceedings of the 18th British Machine Vision Conference (pp. 600-609). BMVA Press.","DOI":"10.5244\/C.21.60"},{"key":"jmdem.2012100103-12","doi-asserted-by":"crossref","unstructured":"Hsu, W., Kennedy, L., Huang, C., Chang, S., Lin, C., & Iyengar, G. (2004). News video story segmentation using fusion of multi-level multi-modal features in trecvid 2003. In Proceedings of IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP'04), 3, pp. iii - 645-8.","DOI":"10.1109\/ICASSP.2004.1326627"},{"key":"jmdem.2012100103-13","doi-asserted-by":"crossref","unstructured":"Iyengar, G., & Nock, H. (2003). Discriminative model fusion for semantic concept detection and annotation in video. In Proceedings of the 11th ACM international conference on Multimedia (pp. 255-258). New York, NY: ACM.","DOI":"10.1145\/957013.957065"},{"key":"jmdem.2012100103-14","unstructured":"Jiang, Y., Wang, J., Chang, S., & Ngo, C. (2009). Domain adaptive semantic diffusion for large scale context-based video annotation. In Proceedings of IEEE 12th International Conference on Computer Vision (pp. 1420-1427). IEEE."},{"key":"jmdem.2012100103-15","unstructured":"Jiang, Y.-G., Yanagawa, A., Chang, S.-F., & Ngo, C.-W. (2008). CU-VIREO374: Fusing Columbia374 and VIREO374 for large scale semantic concept detection. Columbia University ADVENT #223-2008-1."},{"key":"jmdem.2012100103-16","doi-asserted-by":"crossref","unstructured":"Jiang, Y.-G., Ye, G., Chang, S.-F., Ellis, D., & Loui, A. C. (2011). Consumer video understanding: A benchmark database and an evaluation of human and machine performance. In Proceedings of ACM International Conference on Multimedia Retrieval (ICMR), oral session (pp. 29:1--29:8). New York, NY: ACM.","DOI":"10.1145\/1991996.1992025"},{"key":"jmdem.2012100103-17","doi-asserted-by":"crossref","unstructured":"Kira, K., & Rendell, L. (1992). A practical approach to feature selection. In Proceedings of the 9th International Workshop on Machine learning (pp. 249-256). San Francisco, CA: Morgan Kaufmann Publishers Inc.","DOI":"10.1016\/B978-1-55860-247-2.50037-1"},{"key":"jmdem.2012100103-18","unstructured":"Lie, W., & Su, C. (2005). News video classification based on multi-modal information fusion. In Proceedings of IEEE International Conference on Image Processing (ICIP 2005) (Vol. 1, pp. I-1213-16)."},{"key":"jmdem.2012100103-19","doi-asserted-by":"crossref","unstructured":"Lin, L., Ravitz, G., Shyu, M., & Chen, S. (2008). Correlation-based video semantic concept detection using multiple correspondence analysis. In Proceedings of 10. IEEE International Symposium on Multimedia (ISM 2008) (pp. 316-321). Washington, DC: IEEE Computer Society.","DOI":"10.1109\/ISM.2008.111"},{"key":"jmdem.2012100103-20","article-title":"Available online 11 December 2012). Multimodal recognition of visual concepts using histograms of textual concepts and selective weighted late fusion scheme.","author":"N.Liu","journal-title":"Computer Vision and Image Understanding"},{"key":"jmdem.2012100103-21","doi-asserted-by":"crossref","unstructured":"Llorente, A., & R\u00fcger, S. (2009). Using second order statistics to enhance automated image annotation. In Proceedings of the 31th European Conference on IR Research on Advances in Information Retrieval (pp. 570-577). Berlin, Heidelberg, Germany: Springer-Verlag.","DOI":"10.1007\/978-3-642-00958-7_52"},{"key":"jmdem.2012100103-22","author":"L. v.Maaten","year":"2009","journal-title":"Dimensionality reduction: A comparative review (Tech. Rep. TiCC-TR 2009-005)"},{"key":"jmdem.2012100103-23","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-76316-3_1"},{"key":"jmdem.2012100103-24","unstructured":"Mathieu, B., Essid, S., Fillon, T., Prado, J., & Richard, G. (2010). Yaafe, an easy to use and efficient audio feature extraction software. In Proceedings of the 11th International Society for Music Information Retrieval Conference (pp. 441-446). Utrecht, Netherlands: International Society for Music Information Retrieval."},{"key":"jmdem.2012100103-25","doi-asserted-by":"publisher","DOI":"10.1109\/93.713301"},{"key":"jmdem.2012100103-26","unstructured":"Naphade, M., & Huang, T. (2001). Detecting semantic concepts using context and audiovisual features. In Proceedings of IEEE Workshop on Detection and Recognition of Events in Video (pp. 92-98). Los Alamitos, CA: IEEE Computer Society."},{"key":"jmdem.2012100103-27","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2006.63"},{"key":"jmdem.2012100103-28","unstructured":"Natsev, A., Jiang, W., Merler, M., Smith, J., Tesic, J., Xie, L., et al. (2008). IBM Research TRECVID-2008 video retrieval system. In Proceedings of TRECVID 2008."},{"key":"jmdem.2012100103-29","unstructured":"Over, P., Awad, G., Kraaij, W., & Smeaton, A. F. (2007). TRECVID 2007--Overview. In Proceedings of TRECVID 2007."},{"key":"jmdem.2012100103-30","unstructured":"Ping, S., & Xiao-qing, Y. (2009). Goal event detection in soccer videos using multi-clues detection rules. In Proceedings of the International Conference on Management and Service Science (MASS'2009) (pp. 1-4)."},{"key":"jmdem.2012100103-31","first-page":"61","author":"J.Platt","year":"1999","journal-title":"Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Advances in Large Margin Classifiers"},{"key":"jmdem.2012100103-32","doi-asserted-by":"crossref","unstructured":"Qi, G., Hua, X., Rui, Y., Tang, J., Mei, T., & Zhang, H. (2007). Correlative multi-label video annotation. In Proceedings of the 15th international conference on Multimedia (pp. 17-26). New York, NY: ACM.","DOI":"10.1145\/1291233.1291245"},{"key":"jmdem.2012100103-33","doi-asserted-by":"publisher","DOI":"10.1023\/A:1025667309714"},{"key":"jmdem.2012100103-34","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2005.854237"},{"key":"jmdem.2012100103-35","doi-asserted-by":"publisher","DOI":"10.1023\/B:MTAP.0000046380.27575.a5"},{"key":"jmdem.2012100103-36","doi-asserted-by":"crossref","unstructured":"Snoek, C. G., Worring, M., & Smeulders, A. (2005). Early versus late fusion in semantic video analysis. In Proceedings of the 13th Annual ACM international Conference on Multimedia (pp. 399-402). New York, NY: ACM.","DOI":"10.1145\/1101149.1101236"},{"key":"jmdem.2012100103-37","doi-asserted-by":"publisher","DOI":"10.1002\/sam.10028"},{"key":"jmdem.2012100103-38","doi-asserted-by":"crossref","unstructured":"Truong, B., & Dorai, C. (2000). Automatic genre identification for content-based video categorization. In Proceedings of 15th International Conference on Pattern Recognition, (Vol. 4, pp. 230-233).","DOI":"10.1109\/ICPR.2000.902901"},{"key":"jmdem.2012100103-39","doi-asserted-by":"publisher","DOI":"10.1109\/76.915358"},{"key":"jmdem.2012100103-40","doi-asserted-by":"crossref","unstructured":"Weng, M., & Chuang, Y. (2008). Multi-cue fusion for semantic video indexing. In Proceeding of the 16th ACM International Conference on Multimedia (pp. 71-80). New York, NY: ACM.","DOI":"10.1145\/1459359.1459370"},{"key":"jmdem.2012100103-41","doi-asserted-by":"crossref","unstructured":"Wu, Y., Chang, E. Y., Chang, K. C.-C., & Smith, J. R. (2004). Optimal multimodal fusion for multimedia data analysis. In Proceedings of the 12th Annual ACM International Conference on Multimedia (pp. 572-579). New York, NY: ACM.","DOI":"10.1145\/1027527.1027665"},{"key":"jmdem.2012100103-42","doi-asserted-by":"publisher","DOI":"10.1007\/s10044-005-0244-7"},{"key":"jmdem.2012100103-43","unstructured":"Ye, G., Liu, D., Jhuo, I., & Chang, S. (2012). Robust late fusion with rank minimization. IEEE Conference on Computer Vision and Pattern Recognition, (pp. 3021-3028)."},{"key":"jmdem.2012100103-44","doi-asserted-by":"crossref","unstructured":"Yilmaz, T., Gulen, E., Yazici, A., & Kitsuregawa, M. (2012). A RELIEF-based modality weighting approach for multimodal information retrieval. In Proceedings of the 2nd ACM International Conference on Multimedia Retrieval (pp. 54:1--54:8). New York, NY: ACM.","DOI":"10.1145\/2324796.2324858"},{"key":"jmdem.2012100103-45","unstructured":"Yilmaz, T., Yazici, A., & Kitsuregawa, M. (2012). Non-linear weighted averaging for multimodal information fusion by employing analytical network process. In Proceedings of the 21th International Conference on Pattern Recognition (ICPR 2012), Tsukuba, Japan."},{"key":"jmdem.2012100103-46","doi-asserted-by":"crossref","unstructured":"Zhang, D., & Chang, S. (2002). Event detection in baseball video using superimposed caption recognition. In Proceedings of the tenth ACM International Conference on Multimedia (pp. 315-318). New York, NY: ACM.","DOI":"10.1145\/641007.641073"},{"key":"jmdem.2012100103-47","doi-asserted-by":"crossref","unstructured":"Zhu, Q., Yeh, M., & Cheng, K. (2006). Multimodal fusion using learned text concepts for image categorization. In Proceedings of the 14th Annual ACM International Conference on Multimedia (pp. 211-220). New York, NY: ACM.","DOI":"10.1145\/1180639.1180698"}],"container-title":["International Journal of Multimedia Data Engineering and Management"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=75456","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,1]],"date-time":"2022-06-01T17:45:13Z","timestamp":1654105513000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/jmdem.2012100103"}},"subtitle":[""],"short-title":[],"issued":{"date-parts":[[2012,10,1]]},"references-count":48,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2012,10]]}},"URL":"https:\/\/doi.org\/10.4018\/jmdem.2012100103","relation":{},"ISSN":["1947-8534","1947-8542"],"issn-type":[{"value":"1947-8534","type":"print"},{"value":"1947-8542","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,10,1]]}}}