{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,3]],"date-time":"2025-11-03T22:57:55Z","timestamp":1762210675597,"version":"3.41.0"},"reference-count":105,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2012,8,1]],"date-time":"2012-08-01T00:00:00Z","timestamp":1343779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2012,8]]},"abstract":"<jats:p>\n            Along with the explosive growth of multimedia data, automatic multimedia tagging has attracted great interest of various research communities, such as computer vision, multimedia, and information retrieval. However, despite the great progress achieved in the past two decades, automatic tagging technologies still can hardly achieve satisfactory performance on real-world multimedia data that vary widely in genre, quality, and content. Meanwhile, the power of human intelligence has been fully demonstrated in the Web 2.0 era. If well motivated, Internet users are able to tag a large amount of multimedia data. Therefore, a set of new techniques has been developed by combining humans and computers for more accurate and efficient multimedia tagging, such as batch tagging, active tagging, tag recommendation, and tag refinement. These techniques are able to accomplish multimedia tagging by jointly exploring humans and computers in different ways. This article refers to them collectively as\n            <jats:italic>assistive tagging<\/jats:italic>\n            and conducts a comprehensive survey of existing research efforts on this theme. We first introduce the status of automatic tagging and manual tagging and then state why assistive tagging can be a good solution. We categorize existing assistive tagging techniques into three paradigms: (1) tagging with data selection &amp; organization; (2) tag recommendation; and (3) tag processing. We introduce the research efforts on each paradigm and summarize the methodologies. We also provide a discussion on several future trends in this research direction.\n          <\/jats:p>","DOI":"10.1145\/2333112.2333120","type":"journal-article","created":{"date-parts":[[2012,9,11]],"date-time":"2012-09-11T22:21:06Z","timestamp":1347402066000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":182,"title":["Assistive tagging"],"prefix":"10.1145","volume":"44","author":[{"given":"Meng","family":"Wang","sequence":"first","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bingbing","family":"Ni","sequence":"additional","affiliation":[{"name":"Advanced Digital Science Center"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xian-Sheng","family":"Hua","sequence":"additional","affiliation":[{"name":"Microsoft"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tat-Seng","family":"Chua","sequence":"additional","affiliation":[{"name":"National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,9,7]]},"reference":[{"volume-title":"Proceeding of the International Conference on Semantics and Digital Media Technologies.","author":"Abbasi R.","key":"e_1_2_1_1_1","unstructured":"Abbasi , R. , Grzegorzek , M. , and Staab , S . 2009. Tagez: Flickr tag recommendation . In Proceeding of the International Conference on Semantics and Digital Media Technologies. Abbasi, R., Grzegorzek, M., and Staab, S. 2009. Tagez: Flickr tag recommendation. In Proceeding of the International Conference on Semantics and Digital Media Technologies."},{"volume-title":"Proceedings of the TRECVID Workshop.","author":"Amir A.","key":"e_1_2_1_2_1","unstructured":"Amir , A. , Argillander , J. , Campbell , M. , Haubold , A. , Iyengar , G. , Ebadollahi , S. , Kang , F. , Naphade , M. R. , Natsev , A. , Smith , J. R. , Tesic , J. , and Volkmer , T . 2005. IBM research TRECVID-2005 video retrieval system . In Proceedings of the TRECVID Workshop. Amir, A., Argillander, J., Campbell, M., Haubold, A., Iyengar, G., Ebadollahi, S., Kang, F., Naphade, M. R., Natsev, A., Smith, J. R., Tesic, J., and Volkmer, T. 2005. IBM research TRECVID-2005 video retrieval system. In Proceedings of the TRECVID Workshop."},{"volume-title":"Proceeding of the National Conference on Artificial Intelligence (AAAI).","author":"Anderson A.","key":"e_1_2_1_3_1","unstructured":"Anderson , A. , Ranghunathan , K. , and Vogel , A . 2008. Tagez: Flickr tag recommendation . In Proceeding of the National Conference on Artificial Intelligence (AAAI). Anderson, A., Ranghunathan, K., and Vogel, A. 2008. Tagez: Flickr tag recommendation. In Proceeding of the National Conference on Artificial Intelligence (AAAI)."},{"volume-title":"Proceedings of the International Workshop on Content-Based Multimedia Indexing.","author":"Ayache S.","key":"e_1_2_1_4_1","unstructured":"Ayache , S. and Qu\u00e9not , G . 2007. Evaluation of active learning strategies for video indexing . In Proceedings of the International Workshop on Content-Based Multimedia Indexing. Ayache, S. and Qu\u00e9not, G. 2007. Evaluation of active learning strategies for video indexing. In Proceedings of the International Workshop on Content-Based Multimedia Indexing."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1878151.1878155"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1076034.1076139"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/860435.860460"},{"key":"e_1_2_1_8_1","volume-title":"Mcg-webv: A benchmark dataset for Web video analysis. Tech. rep. ICT-MCG-09-001","author":"Cao J.","year":"2009","unstructured":"Cao , J. , Zhang , Y. , Song , Y. , Chen , Z. , Zhang , X. , and Li , J . 2009 . Mcg-webv: A benchmark dataset for Web video analysis. Tech. rep. ICT-MCG-09-001 , Institute of Computing Technology . Cao, J., Zhang, Y., Song, Y., Chen, Z., Zhang, X., and Li, J. 2009. Mcg-webv: A benchmark dataset for Web video analysis. Tech. rep. ICT-MCG-09-001, Institute of Computing Technology."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1459359.1459473"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Chen L.","key":"e_1_2_1_10_1","unstructured":"Chen , L. , Xu , D. , Tsang , I. W. , and Luo , J . 2010. Tag-based Web photo retrieval improved by batch mode re-tagging . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Chen, L., Xu, D., Tsang, I. W., and Luo, J. 2010. Tag-based Web photo retrieval improved by batch mode re-tagging. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1772690.1772813"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1646396.1646452"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1240624.1240684"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Deng J.","key":"e_1_2_1_14_1","unstructured":"Deng , J. , Dong , W. , Socher , R. , Li , L.-J. , Li , K. , and Fei-Fei , L . 2009. Imagenet: A large-scale hierarchical image database . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. 2009. Imagenet: A large-scale hierarchical image database. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"volume-title":"Proceedings of the European Conference on Computer Vision.","author":"Duygulu P.","key":"e_1_2_1_15_1","unstructured":"Duygulu , P. , Barnard , K. , de Freitas , J. , and Forsyth , D . 2002. Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary . In Proceedings of the European Conference on Computer Vision. Duygulu, P., Barnard, K., de Freitas, J., and Forsyth, D. 2002. Object recognition as machine translation: Learning a lexicon for a fixed image vocabulary. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the PASCAL Visual Object Classes Challenge Workshop.","author":"Everingham M.","year":"2010","unstructured":"Everingham , M. 2010 . Overview and results of the classification challenge . In Proceedings of the PASCAL Visual Object Classes Challenge Workshop. Everingham, M. 2010. Overview and results of the classification challenge. In Proceedings of the PASCAL Visual Object Classes Challenge Workshop."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-009-0275-4"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Fan J.","key":"e_1_2_1_18_1","unstructured":"Fan , J. , Shen , Y. , Zhou , N. , and Gao , Y . 2010. Harvesting large-scale weakly-tagged image databases from the Web . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Fan, J., Shen, Y., Zhou, N., and Gao, Y. 2010. Harvesting large-scale weakly-tagged image databases from the Web. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2006.79"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1816041.1816084"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Feng S. L.","key":"e_1_2_1_21_1","unstructured":"Feng , S. L. , Manmatha , R. , and Lavrenko , V . 2004. Multiple Bernoulli relevance models for image and video annotation . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Feng, S. L., Manmatha, R., and Lavrenko, V. 2004. Multiple Bernoulli relevance models for image and video annotation. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"volume-title":"Proceeding of the International Conference on Machine Learning.","author":"Freund Y.","key":"e_1_2_1_22_1","unstructured":"Freund , Y. , Iyer , R. , Schapire , R. E. , and Singer , Y . 1998. An efficient boosting algorithm for combining preferences . In Proceeding of the International Conference on Machine Learning. Freund, Y., Iyer, R., Schapire, R. E., and Singer, Y. 1998. An efficient boosting algorithm for combining preferences. In Proceeding of the International Conference on Machine Learning."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1178677.1178691"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1454008.1454020"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1027527.1027664"},{"key":"e_1_2_1_26_1","unstructured":"Griffin G. Holub A. and Perona P. 2007. Caltech-256 object category dataset. Tech. rep. 7694 California Institute of Technology.  Griffin G. Holub A. and Perona P. 2007. Caltech-256 object category dataset. Tech. rep. 7694 California Institute of Technology."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1180639.1180721"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1282280.1282369"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1600150.1600153"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2008.916364"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1460096.1460104"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1743384.1743475"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/860435.860459"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops.","author":"Jesus R.","key":"e_1_2_1_34_1","unstructured":"Jesus , R. , Goncalves , D. , Abrantes , A. , and Corriea , N . 2008. Playing games as a way to improve automatic image annotation . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops. Jesus, R., Goncalves, D., Abrantes, A., and Corriea, N. 2008. Playing games as a way to improve automatic image annotation. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2009.2036235"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1178677.1178712"},{"volume-title":"Proceedings of the Advances in Neural Information Processing Systems.","author":"Lavrenko V.","key":"e_1_2_1_37_1","unstructured":"Lavrenko , V. , Manmatha , R. , and Jeon , J . 2004. A model for learning the semantics of pictures . In Proceedings of the Advances in Neural Information Processing Systems. Lavrenko, V., Manmatha, R., and Jeon, J. 2004. A model for learning the semantics of pictures. In Proceedings of the Advances in Neural Information Processing Systems."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2009.12.024"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/1126004.1126005"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1991996.1992033"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2007.70847"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2009.2030598"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/1816041.1816044"},{"volume-title":"Proceedings of the International Conference on Multimedia & Expo.","author":"Lin C.","key":"e_1_2_1_44_1","unstructured":"Lin , C. , Tseng , B. , and Smith , J. R . 2003. VideoAnnEx: IBM MPEG-7 annotation tool for multimedia indexing and concept learning . In Proceedings of the International Conference on Multimedia & Expo. Lin, C., Tseng, B., and Smith, J. R. 2003. VideoAnnEx: IBM MPEG-7 annotation tool for multimedia indexing and concept learning. In Proceedings of the International Conference on Multimedia & Expo."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874031"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526757"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-010-0647-3"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2010.2087744"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873958"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/1631272.1631291"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Liu X.","key":"e_1_2_1_51_1","unstructured":"Liu , X. , Yan , S. , Luo , J. , Tang , J. , Huang , Z. , and Jin , H . 2010. Nonparametric label-to-region by search . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Liu, X., Yan, S., Luo, J., Tang, J., Huang, Z., and Jin, H. 2010. Nonparametric label-to-region by search. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/1290082.1290117"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-88690-7_24"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/957013.957070"},{"volume-title":"Proceedings of the International Workshop on Multimedia Intelligent Storage and Retrieval Management.","author":"Mori Y.","key":"e_1_2_1_55_1","unstructured":"Mori , Y. , Takahashi , H. , and Oka , R . 1999. Image-to-word transformation based on dividing and vector quantizing images with words . In Proceedings of the International Workshop on Multimedia Intelligent Storage and Retrieval Management. Mori, Y., Takahashi, H., and Oka, R. 1999. Image-to-word transformation based on dividing and vector quantizing images with words. In Proceedings of the International Workshop on Multimedia Intelligent Storage and Retrieval Management."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2008.69"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/1027527.1027680"},{"key":"e_1_2_1_58_1","first-page":"17","article-title":"What is Web 2.0: Design patterns and business models for the next generation of software","volume":"1","author":"O'Reilly T.","year":"2007","unstructured":"O'Reilly , T. 2007 . What is Web 2.0: Design patterns and business models for the next generation of software . Commun. Strategies 1 , 17 -- 37 . O'Reilly, T. 2007. What is Web 2.0: Design patterns and business models for the next generation of software. Commun. Strategies 1, 17--37.","journal-title":"Commun. Strategies"},{"volume-title":"TRECVID 2010 c goals, tasks, data, evaluation, mechanisms and metrics. In Proceedings of the TRECVID Workshop.","author":"Over P.","key":"e_1_2_1_59_1","unstructured":"Over , P. , Awad , G. , Fiscus , J. , and Michel , M . 2010 . TRECVID 2010 c goals, tasks, data, evaluation, mechanisms and metrics. In Proceedings of the TRECVID Workshop. Over, P., Awad, G., Fiscus, J., and Michel, M. 2010. TRECVID 2010 c goals, tasks, data, evaluation, mechanisms and metrics. In Proceedings of the TRECVID Workshop."},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1291233.1291245"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.5555\/1937055.1937077"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1006\/jvci.1999.0413"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-007-0090-8"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2009.154"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1145\/1743384.1743473"},{"volume-title":"Proceedings of the NIPS Workshop on Cost-Sensitive Learning.","author":"Settles B.","key":"e_1_2_1_66_1","unstructured":"Settles , B. , Craven , M. , and Friedland , L . 2008. Active learning with real annotation costs . In Proceedings of the NIPS Workshop on Cost-Sensitive Learning. Settles, B., Craven, M., and Friedland, L. 2008. Active learning with real annotation costs. In Proceedings of the NIPS Workshop on Cost-Sensitive Learning."},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-009-0394-5"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873956"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/1367497.1367542"},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/34.895972"},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2010.183"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1561\/1500000014"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops.","author":"Sorokin A.","key":"e_1_2_1_73_1","unstructured":"Sorokin , A. and Forsyth , D . 2008. Utility data annotation via Amazon mechanical turk . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops. Sorokin, A. and Forsyth, D. 2008. Utility data annotation via Amazon mechanical turk. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition Workshops."},{"key":"e_1_2_1_74_1","doi-asserted-by":"crossref","unstructured":"Steggink J. and Snoek C. G. 2011. Adding semantics to image region annotations with the name-it-game. Multimedia Syst.  Steggink J. and Snoek C. G. 2011. Adding semantics to image region annotations with the name-it-game. Multimedia Syst.","DOI":"10.1007\/s00530-010-0220-y"},{"key":"e_1_2_1_75_1","unstructured":"Suh B. and Bederson B. B. 2004. Semi-automatic image annotation using event and torso identification. Tech. rep. HCIL-2004-15 Computer Science Department University of Maryland.  Suh B. and Bederson B. B. 2004. Semi-automatic image annotation using event and torso identification. Tech. rep. HCIL-2004-15 Computer Science Department University of Maryland."},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874029"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874139"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/1899412.1899418"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Tian Y.","key":"e_1_2_1_79_1","unstructured":"Tian , Y. , Liu , W. , Xiao , R. , Wen , F. , and Tang , X . 2007. A face annotation framework with partial clustering and interactive labeling . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Tian, Y., Liu, W., Xiao, R., Wen, F., and Tang, X. 2007. A face annotation framework with partial clustering and interactive labeling. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Toderici G.","key":"e_1_2_1_80_1","unstructured":"Toderici , G. , Aradhye , H. , Pasca , M. , Sbaiz , L. , and Yagnik , J . 2010. Finding meaning on youtube: Tag recommendation and category discovery . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Toderici, G., Aradhye, H., Pasca, M., Sbaiz, L., and Yagnik, J. 2010. Finding meaning on youtube: Tag recommendation and category discovery. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2008.128"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/1386352.1386358"},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2010.2091400"},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/1101149.1101341"},{"key":"e_1_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/985692.985733"},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1145\/1124772.1124782"},{"key":"e_1_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1145\/1180639.1180774"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Wang C.","key":"e_1_2_1_88_1","unstructured":"Wang , C. , Jing , F. , Zhang , L. , and Zhang , H . -J. 2007. Content-based image annotation refinement . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Wang, C., Jing, F., Zhang, L., and Zhang, H.-J. 2007. Content-based image annotation refinement. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.1145\/1899412.1899414"},{"key":"e_1_2_1_90_1","doi-asserted-by":"publisher","DOI":"10. 1109\/TCSVT.2009.2017400"},{"key":"e_1_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2009.2012919"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2010.2055045"},{"volume-title":"Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition.","author":"Wang X.","key":"e_1_2_1_93_1","unstructured":"Wang , X. , Zhang , L. , Liu , M. , Li , Y. , and Ma , W. C . 2010. ARISTA c image search to annotation on billions of Web photos . In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition. Wang, X., Zhang, L., Liu, M., Li, Y., and Ma, W. C. 2010. ARISTA c image search to annotation on billions of Web photos. In Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.58"},{"key":"e_1_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.1145\/1816041.1816049"},{"key":"e_1_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/1459359.1459375"},{"key":"e_1_2_1_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873957"},{"key":"e_1_2_1_98_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874129"},{"key":"e_1_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526758"},{"key":"e_1_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1145\/1631272.1631359"},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1145\/1631058.1631067"},{"key":"e_1_2_1_102_1","doi-asserted-by":"publisher","DOI":"10.1109\/MMUL.2009.28"},{"key":"e_1_2_1_103_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874035"},{"volume-title":"Proceedings of the International Workshop on Internet Multimedia Search and Mining.","author":"Yang K.","key":"e_1_2_1_104_1","unstructured":"Yang , K. , Wang , M. , Hua , X. S. , and Zhang , H. J . 2009. Active tagging for image indexing . In Proceedings of the International Workshop on Internet Multimedia Search and Mining. Yang, K., Wang, M., Hua, X. S., and Zhang, H. J. 2009. Active tagging for image indexing. In Proceedings of the International Workshop on Internet Multimedia Search and Mining."},{"key":"e_1_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874028"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2333112.2333120","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2333112.2333120","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T09:21:06Z","timestamp":1750238466000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2333112.2333120"}},"subtitle":["A survey of multimedia tagging with human-computer joint exploration"],"short-title":[],"issued":{"date-parts":[[2012,8]]},"references-count":105,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2012,8]]}},"alternative-id":["10.1145\/2333112.2333120"],"URL":"https:\/\/doi.org\/10.1145\/2333112.2333120","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"type":"print","value":"0360-0300"},{"type":"electronic","value":"1557-7341"}],"subject":[],"published":{"date-parts":[[2012,8]]},"assertion":[{"value":"2011-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-09-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}