{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,6]],"date-time":"2026-05-06T15:32:07Z","timestamp":1778081527099,"version":"3.51.4"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2015,2,5]],"date-time":"2015-02-05T00:00:00Z","timestamp":1423094400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Basic Research Program of China","doi-asserted-by":"crossref","award":["2012CB316400"],"award-info":[{"award-number":["2012CB316400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61322212, and 61035001"],"award-info":[{"award-number":["61322212, and 61035001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Key Technologies R&D Program of China","award":["2012BAH18B02"],"award-info":[{"award-number":["2012BAH18B02"]}]},{"name":"National Hi-Tech Development Program (863 Program) of China","award":["2014AA015202"],"award-info":[{"award-number":["2014AA015202"]}]},{"name":"Lenovo Outstanding Young Scientists Program"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2015,2,5]]},"abstract":"<jats:p>Over the last several decades, researches on visual object retrieval and recognition have achieved fast and remarkable success. However, while the category-level tasks prevail in the community, the instance-level tasks (especially recognition) have not yet received adequate focuses. Applications such as content-based search engine and robot vision systems have alerted the awareness to bring instance-level tasks into a more realistic and challenging scenario. Motivated by the limited scope of existing instance-level datasets, in this article we propose a new benchmark for INSTance-level visual object REtrieval and REcognition (INSTRE). Compared with existing datasets, INSTRE has the following major properties: (1) balanced data scale, (2) more diverse intraclass instance variations, (3) cluttered and less contextual backgrounds, (4) object localization annotation for each image, (5) well-manipulated double-labelled images for measuring multiple object (within one image) case. We will quantify and visualize the merits of INSTRE data, and extensively compare them against existing datasets. Then on INSTRE, we comprehensively evaluate several popular algorithms to large-scale object retrieval problem with multiple evaluation metrics. Experimental results show that all the methods suffer a performance drop on INSTRE, proving that this field still remains a challenging problem. Finally we integrate these algorithms into a simple yet efficient scheme for recognition and compare it with classification-based methods. Importantly, we introduce the realistic multiobjects recognition problem. All experiments are conducted in both single object case and multiple objects case.<\/jats:p>","DOI":"10.1145\/2700292","type":"journal-article","created":{"date-parts":[[2015,2,10]],"date-time":"2015-02-10T13:19:47Z","timestamp":1423574387000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":62,"title":["INSTRE"],"prefix":"10.1145","volume":"11","author":[{"given":"Shuang","family":"Wang","sequence":"first","affiliation":[{"name":"Institute of Computing Technology, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuqiang","family":"Jiang","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, CAS, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,2,5]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the British Machine Vision Conference.","author":"Alcantarilla P. F."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126265"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1873985"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/11744023_32"},{"key":"e_1_2_1_5_1","first-page":"3","article-title":"Kernel descriptors for visual recognition","volume":"1","author":"Bo Liefeng","year":"2010","journal-title":"Neural Information Processing Systems"},{"key":"e_1_2_1_6_1","unstructured":"L. Bo and C. Sminchisescu. 2009. Efficient match kernel between sets of features for visual recognition. In Neural Information Processing Systems 1730--1731.  L. Bo and C. Sminchisescu. 2009. Efficient match kernel between sets of features for visual recognition. In Neural Information Processing Systems 1730--1731."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2013.2270455"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995601"},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"J. Deng A. C Berg K. Li and F.-F. Li. 2010. What does classifying more than 10 000 image categories tell us&quest; In Proceedings of the European Conference on Computer Vision. Springer 71--84.   J. Deng A. C Berg K. Li and F.-F. Li. 2010. What does classifying more than 10 000 image categories tell us&quest; In Proceedings of the European Conference on Computer Vision. Springer 71--84.","DOI":"10.1007\/978-3-642-15555-0_6"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/358669.358692"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126481"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000042993.50813.60"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2072298.2072035"},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Jawahar C. V."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-88682-2_24"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 1169--1176","author":"J\u00e9gou H."},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3100--3107","author":"Jiang Y."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1631272.1631361"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/1991996.1992016"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1023\/B:VISI.0000029664.99615.94"},{"key":"e_1_2_1_22_1","volume-title":"Tech. Rep. CUCS-005-96.","author":"Nene S. A.","year":"1996"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.264"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.344"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the European Conference on Computer Vision. Springer, 143--156","author":"Perronnin F."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Philbin J."},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.","author":"Philbin J."},{"key":"e_1_2_1_28_1","first-page":"e27","article-title":"Why is real-world visual object recognition hard&quest; PLoS","volume":"4","author":"Pinto N.","year":"2008","journal-title":"Computa. Biol."},{"key":"e_1_2_1_29_1","doi-asserted-by":"crossref","unstructured":"J. Ponce T. L. Berg M. Everingham etal 2006. Dataset Issues in object recognition. In Toward Category-Level Object Recognition Springer 29--48.  J. Ponce T. L. Berg M. Everingham et al. 2006. Dataset Issues in object recognition. In Toward Category-Level Object Recognition Springer 29--48.","DOI":"10.1007\/11957959_2"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1991996.1992021"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-007-0090-8"},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. IEEE, 3013--3020","author":"Shen X."},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision. IEEE, 1470--1477","author":"Sivic J."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995347"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2009.132"},{"key":"e_1_2_1_37_1","unstructured":"A. Vedaldi and B. Fulkerson. 2008. VLFeat: An open and portable library of computer vision algorithms. http:\/\/www.vlfeat.org\/.  A. Vedaldi and B. Fulkerson. 2008. VLFeat: An open and portable library of computer vision algorithms. http:\/\/www.vlfeat.org\/."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461466.2461524"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2010.212"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995368"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2422956.2422960"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874019"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2700292","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2700292","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T05:07:43Z","timestamp":1750223263000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2700292"}},"subtitle":["A New Benchmark for Instance-Level Object Retrieval and Recognition"],"short-title":[],"issued":{"date-parts":[[2015,2,5]]},"references-count":42,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2015,2,5]]}},"alternative-id":["10.1145\/2700292"],"URL":"https:\/\/doi.org\/10.1145\/2700292","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,2,5]]},"assertion":[{"value":"2014-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2014-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-02-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}