{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:11:19Z","timestamp":1750306279495,"version":"3.41.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2016,5,20]],"date-time":"2016-05-20T00:00:00Z","timestamp":1463702400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2016,6,15]]},"abstract":"<jats:p>With the popularity of mobile devices, photo retargeting has become a useful technique that adapts a high-resolution photo onto a low-resolution screen. Conventional approaches are limited in two aspects. The first factor is the de-emphasized role of semantic content that is many times more important than low-level features in photo aesthetics. Second is the importance of image spatial modeling: toward a semantically reasonable retargeted photo, the spatial distribution of objects within an image should be accurately learned. To solve these two problems, we propose a new semantically aware photo retargeting that shrinks a photo according to region semantics. The key technique is a mechanism transferring semantics of noisy image labels (inaccurate labels predicted by a learner like an SVM) into different image regions. In particular, we first project the local aesthetic features (graphlets in this work) onto a semantic space, wherein image labels are selectively encoded according to their noise level. Then, a category-sharing model is proposed to robustly discover the semantics of each image region. The model is motivated by the observation that the semantic distribution of graphlets from images tagged by a common label remains stable in the presence of noisy labels. Thereafter, a spatial pyramid is constructed to hierarchically encode the spatial layout of graphlet semantics. Based on this, a probabilistic model is proposed to enforce the spatial layout of a retargeted photo to be maximally similar to those from the training photos. Experimental results show that (1) noisy image labels predicted by different learners can improve the retargeting performance, according to both qualitative and quantitative analysis, and (2) the category-sharing model stays stable even when 32.36% of image labels are incorrectly predicted.<\/jats:p>","DOI":"10.1145\/2886775","type":"journal-article","created":{"date-parts":[[2016,5,22]],"date-time":"2016-05-22T01:23:59Z","timestamp":1463880239000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["Semantic Photo Retargeting Under Noisy Image Labels"],"prefix":"10.1145","volume":"12","author":[{"given":"Luming","family":"Zhang","sequence":"first","affiliation":[{"name":"Department of CSIE, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuelong","family":"Li","sequence":"additional","affiliation":[{"name":"Xi'an Institute of Optics and Precision Mechanics, Chinese Academy of Sciences, Shaanxi, P. R. China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liqiang","family":"Nie","sequence":"additional","affiliation":[{"name":"School of Computing, National University of Singapore, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yan","family":"Yan","sequence":"additional","affiliation":[{"name":"Department of Information Engineering and Computer Science, University of Trento, Trento, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7410-2590","authenticated-orcid":false,"given":"Roger","family":"Zimmermann","sequence":"additional","affiliation":[{"name":"School of Computing, National University of Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,5,20]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proc. of NIPS, 561--568","author":"Andrews Stuart","year":"2003","unstructured":"Stuart Andrews , Ioannis Tsochantaridis , and Thomas Hofmann . 2003 . Support vector machines for multiple-instance learning . In Proc. of NIPS, 561--568 , 2003. Stuart Andrews, Ioannis Tsochantaridis, and Thomas Hofmann. 2003. Support vector machines for multiple-instance learning. In Proc. of NIPS, 561--568, 2003."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1276377.1276390"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2077434.2077447"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873990"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354875"},{"key":"e_1_2_1_6_1","first-page":"2009","article-title":"Saliency, attention, and visual search: An information theoretic approach","volume":"5","author":"Bruce Neil D. B.","year":"2009","unstructured":"Neil D. B. Bruce and John K. Tsotsos . 2009 . Saliency, attention, and visual search: An information theoretic approach . J. Vision, 9(3), article 5 , 2009 . Neil D. B. Bruce and John K. Tsotsos. 2009. Saliency, attention, and visual search: An information theoretic approach. J. Vision, 9(3), article 5, 2009.","journal-title":"J. Vision, 9(3), article"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873992"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2005.177"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995467"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2009.2021781"},{"key":"e_1_2_1_11_1","volume-title":"Proc. of NIPS, 151--158","author":"Hoffman Judy","year":"2014","unstructured":"Judy Hoffman , Sergio Guadarrama , Eric Tzeng , Ronghang Hu , Jeff Donahue , Ross Girshick , Trevor Darrell , and Kate Saenko . 2014 . LSDA: Large scale detection through adaptation . In Proc. of NIPS, 151--158 , 2014. Judy Hoffman, Sergio Guadarrama, Eric Tzeng, Ronghang Hu, Jeff Donahue, Ross Girshick, Trevor Darrell, and Kate Saenko. 2014. LSDA: Large scale detection through adaptation. In Proc. of NIPS, 151--158, 2014."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4408844"},{"key":"e_1_2_1_13_1","volume-title":"Proc. of ICML, 282--289","author":"Lafferty John D.","year":"2001","unstructured":"John D. Lafferty , Andrew McCallum , and Fernando C. N. Pereira . 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data . In Proc. of ICML, 282--289 , 2001 . John D. Lafferty, Andrew McCallum, and Fernando C. N. Pereira. 2001. Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proc. of ICML, 282--289, 2001."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2012.2228475"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206536"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.270"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1631272.1631384"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354807"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1011139631724"},{"key":"e_1_2_1_20_1","volume-title":"Proc. of ICCV, 151--158","author":"Pritch Yael","year":"2009","unstructured":"Yael Pritch , Eitam Kav-Venaki , and Shmuel Peleg . 2009 . Shift-map image editing . In Proc. of ICCV, 151--158 , 2009. Yael Pritch, Eitam Kav-Venaki, and Shmuel Peleg. 2009. Shift-map image editing. In Proc. of ICCV, 151--158, 2009."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1882261.1866186"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1360612.1360615"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531326.1531329"},{"key":"e_1_2_1_24_1","volume-title":"SIGGRAPH Asia Courses","author":"Shamir Ariel","year":"2012","unstructured":"Ariel Shamir , Alexander Sorkine-Hornung , and Olga Sorkine-Hornung . 2012 . Modern approaches to media retargeting . SIGGRAPH Asia Courses , 2012. Ariel Shamir, Alexander Sorkine-Hornung, and Olga Sorkine-Hornung. 2012. Modern approaches to media retargeting. SIGGRAPH Asia Courses, 2012."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2009.2032939"},{"key":"e_1_2_1_26_1","volume-title":"Similarity of color images. Storage and Retrieval of Image and Video Databases, 381--392","author":"Stricker Markus","year":"1995","unstructured":"Markus Stricker and Markus Orengo . 1995. Similarity of color images. Storage and Retrieval of Image and Video Databases, 381--392 , 1995 . Markus Stricker and Markus Orengo. 1995. Similarity of color images. Storage and Retrieval of Image and Video Databases, 381--392, 1995."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2501643.2501651"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1117\/12.862419"},{"key":"e_1_2_1_29_1","volume-title":"Proc. of CVPR, 1--8","author":"Jakob","year":"2007","unstructured":"Jakob J. Verbeek and Bill Triggs. 2007. Region classification with Markov field aspect models . In Proc. of CVPR, 1--8 , 2007 . Jakob J. Verbeek and Bill Triggs. 2007. Region classification with Markov field aspect models. In Proc. of CVPR, 1--8, 2007."},{"key":"e_1_2_1_30_1","volume-title":"Proc of CVPR, 3249--3256","author":"Vezhnevets Alexander","year":"2010","unstructured":"Alexander Vezhnevets and Joachim M. Buhmann . 2010. Towards weakly supervised semantic segmentation by means of multiple instance and multitask learning . In Proc of CVPR, 3249--3256 , 2010 . Alexander Vezhnevets and Joachim M. Buhmann. 2010. Towards weakly supervised semantic segmentation by means of multiple instance and multitask learning. In Proc of CVPR, 3249--3256, 2010."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126299"},{"key":"e_1_2_1_32_1","volume-title":"Proc. of CVPR, 845--852","author":"Vezhnevets Alexander","year":"2012","unstructured":"Alexander Vezhnevets , Vittorio Ferrari , and Joachim M. Buhmann . 2012. Weakly supervised structured output learning for semantic segmentation . In Proc. of CVPR, 845--852 , 2012 . Alexander Vezhnevets, Vittorio Ferrari, and Joachim M. Buhmann. 2012. Weakly supervised structured output learning for semantic segmentation. In Proc. of CVPR, 845--852, 2012."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354843"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2011.2114354"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/1409060.1409071"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2007.4409010"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-008-0161-3"},{"key":"e_1_2_1_38_1","volume-title":"Proc. of ICPR, 897--900","author":"Xiong Xuejian","year":"2000","unstructured":"Xuejian Xiong and Kap Luk Chan . 2000 . Towards an unsupervised optimal fuzzy clustering algorithm for image database organization . In Proc. of ICPR, 897--900 , 2000. Xuejian Xiong and Kap Luk Chan. 2000. Towards an unsupervised optimal fuzzy clustering algorithm for image database organization. In Proc. of ICPR, 897--900, 2000."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.408"},{"key":"e_1_2_1_40_1","volume-title":"Proc. of EMMCVPR, 169--183","author":"Yanulevskaya Victoria","year":"2007","unstructured":"Victoria Yanulevskaya , Jasper R. R. Uijlings , Elia Bruni , Andreza Sartori , Elisa Zamboni , Francesca Bacci , David Melcher , and Nicu Sebe . 2007 . Introduction to a large scale general purpose ground truth dataset: Methodology, annotation tool, and benchmarks . In Proc. of EMMCVPR, 169--183 , 2007. Victoria Yanulevskaya, Jasper R. R. Uijlings, Elia Bruni, Andreza Sartori, Elisa Zamboni, Francesca Bacci, David Melcher, and Nicu Sebe. 2007. Introduction to a large scale general purpose ground truth dataset: Methodology, annotation tool, and benchmarks. In Proc. of EMMCVPR, 169--183, 2007."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2658981"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2534409"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.249"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2012.2223226"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2659520"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2886775","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2886775","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:38:51Z","timestamp":1750221531000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2886775"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,5,20]]},"references-count":45,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2016,6,15]]}},"alternative-id":["10.1145\/2886775"],"URL":"https:\/\/doi.org\/10.1145\/2886775","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2016,5,20]]},"assertion":[{"value":"2015-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-05-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}