{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:17:43Z","timestamp":1750306663541,"version":"3.41.0"},"reference-count":26,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2016,2,24]],"date-time":"2016-02-24T00:00:00Z","timestamp":1456272000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Basic Research Program of China","doi-asserted-by":"crossref","award":["2012CB316400"],"award-info":[{"award-number":["2012CB316400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000001","name":"US NSF","doi-asserted-by":"crossref","award":["IIS-0812114, CCF-1017828"],"award-info":[{"award-number":["IIS-0812114, CCF-1017828"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2016,2,24]]},"abstract":"<jats:p>Mining knowledge from a multimedia database has received increasing attentions recently since huge repositories are made available by the development of the Internet. In this article, we exploit the relations among different modalities in a multimedia database and present a framework for general multimodal data mining problem where image annotation and image retrieval are considered as the special cases. Specifically, the multimodal data mining problem can be formulated as a structured prediction problem where we learn the mapping from an input to the structured and interdependent output variables. In addition, in order to reduce the demanding computation, we propose a new max margin structure learning approach called Enhanced Max Margin Learning (EMML) framework, which is much more efficient with a much faster convergence rate than the existing max margin learning methods, as verified through empirical evaluations. Furthermore, we apply EMML framework to develop an effective and efficient solution to the multimodal data mining problem that is highly scalable in the sense that the query response time is independent of the database scale. The EMML framework allows an efficient multimodal data mining query in a very large scale multimedia database, and excels many existing multimodal data mining methods in the literature that do not scale up at all. The performance comparison with a state-of-the-art multimodal data mining method is reported for the real-world image databases.<\/jats:p>","DOI":"10.1145\/2742549","type":"journal-article","created":{"date-parts":[[2016,2,26]],"date-time":"2016-02-26T14:29:03Z","timestamp":1456496943000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Multimodal Data Mining in a Multimedia Database Based on Structured Max Margin Learning"],"prefix":"10.1145","volume":"10","author":[{"given":"Zhen","family":"Guo","sequence":"first","affiliation":[{"name":"SUNY Binghamton, NY"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhongfei (Mark)","family":"Zhang","sequence":"additional","affiliation":[{"name":"SUNY Binghamton, NY"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eric P.","family":"Xing","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christos","family":"Faloutsos","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, PA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,2,24]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings ICML","author":"Altun Yasemin","year":"2003","unstructured":"Yasemin Altun , Ioannis Tsochantaridis , and Thomas Hofmann . 2003 . Hidden Markov support vector machines . In Proceedings ICML . Washington, DC. Yasemin Altun, Ioannis Tsochantaridis, and Thomas Hofmann. 2003. Hidden Markov support vector machines. In Proceedings ICML. Washington, DC."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944965"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/860435.860460"},{"volume-title":"Convex Optimization","author":"Boyd Stephen","key":"e_1_2_1_4_1","unstructured":"Stephen Boyd and Lieven Vandenberghe . 2004. Convex Optimization . Cambridge University Press . Stephen Boyd and Lieven Vandenberghe. 2004. Convex Optimization. Cambridge University Press."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143863"},{"key":"e_1_2_1_6_1","volume-title":"VIPS: A vision-based page segmentation algorithm. Microsoft Technical Report MSR-TR-2003-79.","author":"Cai D.","year":"2003","unstructured":"D. Cai , S. Yu , J.-R. Wen , and W.-Y. Ma . 2003 . VIPS: A vision-based page segmentation algorithm. Microsoft Technical Report MSR-TR-2003-79. D. Cai, S. Yu, J.-R. Wen, and W.-Y. Ma. 2003. VIPS: A vision-based page segmentation algorithm. Microsoft Technical Report MSR-TR-2003-79."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2002.808079"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015354"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1180639.1180856"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102373"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/645318.649254"},{"volume-title":"Proceedings International Conference on Computer Vision and Pattern Recognition","author":"Feng S. L.","key":"e_1_2_1_12_1","unstructured":"S. L. Feng , R. Manmatha , and V. Lavrenko . 2004. Multiple Bernoulli relevance models for image and video annotation . In Proceedings International Conference on Computer Vision and Pattern Recognition . Washington, DC. S. L. Feng, R. Manmatha, and V. Lavrenko. 2004. Multiple Bernoulli relevance models for image and video annotation. In Proceedings International Conference on Computer Vision and Pattern Recognition. Washington, DC."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007662407062"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1281192.1281231"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2007.4284697"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings ICML. 282--289","author":"Lafferty John","year":"2001","unstructured":"John Lafferty , Andrew McCallum , and Fernando Pereira . 2001 . Conditional random fields: Probabilistic models for segmenting and labeling sequence data . In Proceedings ICML. 282--289 . John Lafferty, Andrew McCallum, and Fernando Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings ICML. 282--289."},{"key":"e_1_2_1_17_1","volume-title":"Proceedings ICML. 591--598","author":"McCallum Andrew","year":"2000","unstructured":"Andrew McCallum , Dayne Freitag , and Fernando Pereira . 2000 . Maximum Entropy Markov models for information extraction and segmentation . In Proceedings ICML. 591--598 . Andrew McCallum, Dayne Freitag, and Fernando Pereira. 2000. Maximum Entropy Markov models for information extraction and segmentation. In Proceedings ICML. 591--598."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/NNSP.1997.622408"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1014052.1014135"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.895972"},{"volume-title":"Proceedings Neural Information Processing Systems Conference","author":"Taskar B.","key":"e_1_2_1_21_1","unstructured":"B. Taskar , C. Guestrin , and D. Koller . 2003. Max-margin Markov networks . In Proceedings Neural Information Processing Systems Conference . Vancouver, Canada. B. Taskar, C. Guestrin, and D. Koller. 2003. Max-margin Markov networks. In Proceedings Neural Information Processing Systems Conference. Vancouver, Canada."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1102351.1102464"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015341"},{"volume-title":"The Nature of Statistical Learning Theory","author":"Vapnik Vladimir Naumovich","key":"e_1_2_1_24_1","unstructured":"Vladimir Naumovich Vapnik . 1995. The Nature of Statistical Learning Theory . Springer . Vladimir Naumovich Vapnik. 1995. The Nature of Statistical Learning Theory. Springer."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the SIGIR Multimedia Information Retrieval Workshop","author":"Westerveld T.","year":"2003","unstructured":"T. Westerveld and A. de Vries . 2003 . Experimental evaluation of a generative probabilistic image retrieval model on \u2018easy\u2019 data . In Proceedings of the SIGIR Multimedia Information Retrieval Workshop 2003. T. Westerveld and A. de Vries. 2003. Experimental evaluation of a generative probabilistic image retrieval model on \u2018easy\u2019 data. In Proceedings of the SIGIR Multimedia Information Retrieval Workshop 2003."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1101149.1101338"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2742549","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2742549","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:00:33Z","timestamp":1750230033000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2742549"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,2,24]]},"references-count":26,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2016,2,24]]}},"alternative-id":["10.1145\/2742549"],"URL":"https:\/\/doi.org\/10.1145\/2742549","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2016,2,24]]},"assertion":[{"value":"2009-09-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-02-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}