{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,27]],"date-time":"2025-11-27T06:37:15Z","timestamp":1764225435280,"version":"3.41.0"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2018,7,24]],"date-time":"2018-07-24T00:00:00Z","timestamp":1532390400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Zhejiang Provincial Key Science and Technology Project Foundation","award":["2018C01012"],"award-info":[{"award-number":["2018C01012"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61602136, 61622205, 61472110, 61702143 and 61601158"],"award-info":[{"award-number":["61602136, 61622205, 61472110, 61702143 and 61601158"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100000923","name":"Australian Research Council","doi-asserted-by":"crossref","award":["FL-170100117, DP-140102164 and LP-150100671"],"award-info":[{"award-number":["FL-170100117, DP-140102164 and LP-150100671"]}],"id":[{"id":"10.13039\/501100000923","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Zhejiang Provincial Natural Science Foundation of China","award":["LR15F020002"],"award-info":[{"award-number":["LR15F020002"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2018,8,31]]},"abstract":"<jats:p>We present a novel fine-grained image recognition framework using user click data, which can bridge the semantic gap in distinguishing categories that are similar in visual. As query set in click data is usually large-scale and redundant, we first propose a click-feature-based query-merging approach to merge queries with similar semantics and construct a compact click feature. Afterward, we utilize this compact click feature and convolutional neural network (CNN)-based deep visual feature to jointly represent an image. Finally, with the combined feature, we employ the metriclearning-based template-matching scheme for efficient recognition. Considering the heavy noise in the training data, we introduce a reliability variable to characterize the image reliability, and propose a weakly-supervised metric and template leaning with smooth assumption and click prior (WMTLSC) method to jointly learn the distance metric, object templates, and image reliability. Extensive experiments are conducted on a public Clickture-Dog dataset and our newly established Clickture-Bird dataset. It is shown that the click-data-based query merging helps generating a highly compact (the dimension is reduced to 0.9%) and dense click feature for images, which greatly improves the computational efficiency. Also, introducing this click feature into CNN feature further boosts the recognition accuracy. The proposed framework performs much better than previous state-of-the-arts in fine-grained recognition tasks.<\/jats:p>","DOI":"10.1145\/3209666","type":"journal-article","created":{"date-parts":[[2018,7,24]],"date-time":"2018-07-24T15:50:24Z","timestamp":1532447424000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":20,"title":["User-Click-Data-Based Fine-Grained Image Recognition via Weakly Supervised Metric Learning"],"prefix":"10.1145","volume":"14","author":[{"given":"Min","family":"Tan","sequence":"first","affiliation":[{"name":"Hangzhou Dianzi University, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Yu","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhou","family":"Yu","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fei","family":"Gao","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yong","family":"Rui","sequence":"additional","affiliation":[{"name":"Lenovo, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dacheng","family":"Tao","sequence":"additional","affiliation":[{"name":"University of Sydney, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,7,24]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806243"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.259"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2978656"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2007.48"},{"volume-title":"Fine-Grained Image Recognition from Click-Through Logs Using Deep Siamese Network","author":"Feng Wu","key":"e_1_2_1_5_1","unstructured":"Wu Feng and Dong Liu . 2017. Fine-Grained Image Recognition from Click-Through Logs Using Deep Siamese Network . Springer International Publishing . Wu Feng and Dong Liu. 2017. Fine-Grained Image Recognition from Click-Through Logs Using Deep Siamese Network. Springer International Publishing."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2013.2290593"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.313"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2502081.2502283"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2014.2308281"},{"key":"e_1_2_1_10_1","volume-title":"1st Workshop on IEEE Conference on Computer Vision and Pattern Recognition.","author":"Khosla Aditya","year":"2011","unstructured":"Aditya Khosla , Nityananda Jayadevaprakash , Bangpeng Yao , and Li Fei-Fei . 2011 . Novel dataset for fine-grained image categorization . In 1st Workshop on IEEE Conference on Computer Vision and Pattern Recognition. Aditya Khosla, Nityananda Jayadevaprakash, Bangpeng Yao, and Li Fei-Fei. 2011. Novel dataset for fine-grained image categorization. In 1st Workshop on IEEE Conference on Computer Vision and Pattern Recognition."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2018.01.027"},{"key":"e_1_2_1_12_1","volume-title":"IEEE International Conference on Multimedia and Expo. 1--4.","author":"Li Chenghua","year":"2016","unstructured":"Chenghua Li , Qiang Song , Yuhang Wang , Hang Song , Qi Kang , Jian Cheng , and Hanqing Lu . 2016 . Learning to recognition from bing clickture data . In IEEE International Conference on Multimedia and Expo. 1--4. Chenghua Li, Qiang Song, Yuhang Wang, Hang Song, Qi Kang, Jian Cheng, and Hanqing Lu. 2016. Learning to recognition from bing clickture data. In IEEE International Conference on Multimedia and Expo. 1--4."},{"volume-title":"International Conference on Machine Learning. 775--782","author":"McFee Brian","key":"e_1_2_1_13_1","unstructured":"Brian McFee and Gert R. Lanckriet . 2010. Metric learning to rank . In International Conference on Machine Learning. 775--782 . Brian McFee and Gert R. Lanckriet. 2010. Metric learning to rank. In International Conference on Machine Learning. 775--782."},{"key":"e_1_2_1_14_1","unstructured":"L. Meng R. Huang and J. Gu. 2013. A review of semantic similarity measures in WordNet. International Journal of Hybrid Information Technology 6 (2013).  L. Meng R. Huang and J. Gu. 2013. A review of semantic similarity measures in WordNet. International Journal of Hybrid Information Technology 6 (2013)."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298995"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-015-0816-y"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13735-015-0080-5"},{"key":"e_1_2_1_18_1","volume-title":"Very deep convolutional networks for large-scale image recognition. Computer Science","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. Computer Science ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. Computer Science (2014)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2809928"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.04.123"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2013.09.054"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2015.2506182"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36669-7_40"},{"key":"e_1_2_1_24_1","first-page":"1","article-title":"Click data guided query modeling with click propagation and sparse coding","volume":"3","author":"Tan Min","year":"2018","unstructured":"Min Tan , Jun Yu , Qingming Huang , and Weichen Wu . 2018 . Click data guided query modeling with click propagation and sparse coding . Multimedia Tools and Applications 3 (2018), 1 -- 14 . Min Tan, Jun Yu, Qingming Huang, and Weichen Wu. 2018. Click data guided query modeling with click propagation and sparse coding. Multimedia Tools and Applications 3 (2018), 1--14.","journal-title":"Multimedia Tools and Applications"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007669.3007730"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2998574"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.170"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.463"},{"volume-title":"The Caltech-UCSD Birds200-2011 dataset","author":"Wah Catherine","key":"e_1_2_1_29_1","unstructured":"Catherine Wah , Steve Branson , Peter Welinder , Pietro Perona , and Serge Belongie . 2011. The Caltech-UCSD Birds200-2011 dataset . California Institute of Technology . Catherine Wah, Steve Branson, Peter Welinder, Pietro Perona, and Serge Belongie. 2011. The Caltech-UCSD Birds200-2011 dataset. California Institute of Technology."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/1577069.1577078"},{"key":"e_1_2_1_31_1","volume-title":"IEEE International Conference on Multimedia and Expo Workshops. 1--4.","author":"Xie Guotian","year":"2016","unstructured":"Guotian Xie , Kuiyuan Yang , Yalong Bai , Min Shang , Yong Rui , and Jianhuang Lai . 2016 . Improve dog recognition by mining more information from both click-through logs and pre-trained models . In IEEE International Conference on Multimedia and Expo Workshops. 1--4. Guotian Xie, Kuiyuan Yang, Yalong Bai, Min Shang, Yong Rui, and Jianhuang Lai. 2016. Improve dog recognition by mining more information from both click-through logs and pre-trained models. In IEEE International Conference on Multimedia and Expo Workshops. 1--4."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2016.2593653"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2962719"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2700286"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00530-014-0379-8"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2658981"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2013.2284755"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2014.2311377"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2014.2336697"},{"key":"e_1_2_1_40_1","volume-title":"Deep multimodal distance metric learning using click constraints for image ranking","author":"Yu Jun","year":"2016","unstructured":"Jun Yu , Xiaokang Yang , Fei Gao , and Dacheng Tao . 2016. Deep multimodal distance metric learning using click constraints for image ranking . IEEE Transactions on Cybernetics ( 2016 ). Jun Yu, Xiaokang Yang, Fei Gao, and Dacheng Tao. 2016. Deep multimodal distance metric learning using click constraints for image ranking. IEEE Transactions on Cybernetics (2016)."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2637291"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2015.2400821"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.96"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.212"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2017.8019407"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3209666","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3209666","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:39:33Z","timestamp":1750210773000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3209666"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,7,24]]},"references-count":45,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2018,8,31]]}},"alternative-id":["10.1145\/3209666"],"URL":"https:\/\/doi.org\/10.1145\/3209666","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2018,7,24]]},"assertion":[{"value":"2017-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-07-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}