{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T21:26:36Z","timestamp":1780435596678,"version":"3.54.1"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2016,11,11]],"date-time":"2016-11-11T00:00:00Z","timestamp":1478822400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2016,11,11]]},"abstract":"<jats:p>\n            We address the problem of autonomously exploring unknown objects in a scene by consecutive depth acquisitions. The goal is to reconstruct the scene while online identifying the objects from among a large collection of 3D shapes. Fine-grained shape identification demands a meticulous series of observations attending to varying views and parts of the object of interest. Inspired by the recent success of attention-based models for 2D recognition, we develop a\n            <jats:italic>3D Attention Model<\/jats:italic>\n            that selects the best views to scan from, as well as the most informative regions in each view to focus on, to achieve efficient object recognition. The region-level attention leads to focus-driven features which are quite robust against object occlusion. The attention model, trained with the 3D shape collection, encodes the temporal dependencies among consecutive views with deep recurrent networks. This facilitates order-aware view planning accounting for robot movement cost. In achieving instance identification, the shape collection is organized into a hierarchy, associated with pre-trained hierarchical classifiers. The effectiveness of our method is demonstrated on an autonomous robot (PR) that explores a scene and identifies the objects to construct a 3D scene model.\n          <\/jats:p>","DOI":"10.1145\/2980179.2980224","type":"journal-article","created":{"date-parts":[[2016,11,11]],"date-time":"2016-11-11T17:02:54Z","timestamp":1478883774000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":32,"title":["3D attention-driven depth acquisition for object identification"],"prefix":"10.1145","volume":"35","author":[{"given":"Kai","family":"Xu","sequence":"first","affiliation":[{"name":"National University of Defense Technology and Shandong University and Shenzhen University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yifei","family":"Shi","sequence":"additional","affiliation":[{"name":"National University of Defense Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lintao","family":"Zheng","sequence":"additional","affiliation":[{"name":"National University of Defense Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"SIAT"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Liu","sequence":"additional","affiliation":[{"name":"National University of Defense Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hui","family":"Huang","sequence":"additional","affiliation":[{"name":"Shenzhen University and SIAT"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Su","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Cohen-Or","sequence":"additional","affiliation":[{"name":"Tel-Aviv University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Baoquan","family":"Chen","sequence":"additional","affiliation":[{"name":"Shandong University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,12,5]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/TRO.2014.2320795"},{"key":"e_1_2_1_2_1","unstructured":"Ba J. Mnih V. and Kavukcuoglu K. 2014. Multiple object recognition with visual attention. arXiv preprint arXiv:1412.7755.  Ba J. Mnih V. and Kavukcuoglu K. 2014. Multiple object recognition with visual attention. arXiv preprint arXiv:1412.7755."},{"key":"e_1_2_1_3_1","unstructured":"Bansal A. Shrivastava A. Doersch C. and Gupta A. 2015. Mid-level elements for object detection. arXiv preprint arXiv:1504.07284.  Bansal A. Shrivastava A. Doersch C. and Gupta A. 2015. Mid-level elements for object detection. arXiv preprint arXiv:1504.07284."},{"key":"e_1_2_1_4_1","volume-title":"Proc. CVPR, IEEE, 1--8.","author":"Bart E.","unstructured":"Bart , E. , Porteous , I. , Perona , P. , and Welling , M . 2008. Unsupervised learning of visual taxonomies . In Proc. CVPR, IEEE, 1--8. Bart, E., Porteous, I., Perona, P., and Welling, M. 2008. Unsupervised learning of visual taxonomies. In Proc. CVPR, IEEE, 1--8."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661239"},{"key":"e_1_2_1_6_1","volume-title":"Proc. CVPR, 5556--5565","author":"Choi S.","unstructured":"Choi , S. , Zhou , Q.-Y. , and Koltun , V . 2015. Robust reconstruction of indoor scenes . In Proc. CVPR, 5556--5565 . Choi, S., Zhou, Q.-Y., and Koltun, V. 2015. Robust reconstruction of indoor scenes. In Proc. CVPR, 5556--5565."},{"key":"e_1_2_1_7_1","unstructured":"Choi S. Zhou Q.-Y. Miller S. and Koltun V. 2016. A large dataset of object scans. arXiv:1602.02481.  Choi S. Zhou Q.-Y. Miller S. and Koltun V. 2016. A large dataset of object scans. arXiv:1602.02481."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1038\/nrn755"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.167"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366154"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126481"},{"key":"e_1_2_1_12_1","volume-title":"Proc. CVPR, 4731--4740","author":"Gupta S.","unstructured":"Gupta , S. , Arbel\u00e1ez , P. , Girshick , R. , and Malik , J . 2015. Aligning 3d models to RGB-D images of cluttered scenes . In Proc. CVPR, 4731--4740 . Gupta, S., Arbel\u00e1ez, P., Girshick, R., and Malik, J. 2015. Aligning 3d models to RGB-D images of cluttered scenes. In Proc. CVPR, 4731--4740."},{"key":"e_1_2_1_13_1","volume-title":"Proc. CVPR.","author":"Haque A.","unstructured":"Haque , A. , Alahi , A. , and Fei-Fei , L . 2016. Recurrent attention models for depth-based person identification . In Proc. CVPR. Haque, A., Alahi, A., and Fei-Fei, L. 2016. Recurrent attention models for depth-based person identification. In Proc. CVPR."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508364"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Huang H. Lischinski D. Hao Z. Gong M. Christie M. and Cohen-Or D. 2016. Trip synopsis: 60km in 60sec. Computer Graphics Forum (Pacific Graphics) to appear.  Huang H. Lischinski D. Hao Z. Gong M. Christie M. and Cohen-Or D. 2016. Trip synopsis: 60km in 60sec. Computer Graphics Forum (Pacific Graphics) to appear.","DOI":"10.1111\/cgf.13008"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818097"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818116"},{"key":"e_1_2_1_19_1","volume-title":"Proc. CVPR, 5546--5555","author":"Krause J.","unstructured":"Krause , J. , Jin , H. , Yang , J. , and Fei-Fei , L . 2015. Fine-grained recognition without part annotations . In Proc. CVPR, 5546--5555 . Krause, J., Jin, H., Yang, J., and Fei-Fei, L. 2015. Fine-grained recognition without part annotations. In Proc. CVPR, 5546--5555."},{"key":"e_1_2_1_20_1","volume-title":"Proc. NIPS, 1097--1105","author":"Krizhevsky A.","unstructured":"Krizhevsky , A. , Sutskever , I. , and Hinton , G. E . 2012. ImageNet classification with deep convolutional neural networks . In Proc. NIPS, 1097--1105 . Krizhevsky, A., Sutskever, I., and Hinton, G. E. 2012. ImageNet classification with deep convolutional neural networks. In Proc. NIPS, 1097--1105."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_1_22_1","volume-title":"Proc. CVPR, IEEE, 3336--3343","author":"Li L.-J.","unstructured":"Li , L.-J. , Wang , C. , Lim , Y. , Blei , D. M. , and Fei-Fei , L . 2010. Building and using a semantivisual image hierarchy . In Proc. CVPR, IEEE, 3336--3343 . Li, L.-J., Wang, C., Lim, Y., Blei, D. M., and Fei-Fei, L. 2010. Building and using a semantivisual image hierarchy. In Proc. CVPR, IEEE, 3336--3343."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12573"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818071"},{"key":"e_1_2_1_25_1","volume-title":"Proc. NIPS, 2204--2212","author":"Mnih V.","year":"2014","unstructured":"Mnih , V. , Heess , N. , Graves , A. , 2014 . Recurrent models of visual attention . In Proc. NIPS, 2204--2212 . Mnih, V., Heess, N., Graves, A., et al. 2014. Recurrent models of visual attention. In Proc. NIPS, 2204--2212."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISMAR.2011.6092378"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508374"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2006.264"},{"key":"e_1_2_1_29_1","unstructured":"ROS 2014. ROS Wiki. http:\/\/wiki.ros.org\/.  ROS 2014. ROS Wiki. http:\/\/wiki.ros.org\/."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.178"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2015.11.003"},{"key":"e_1_2_1_32_1","volume-title":"Proc. CVPR.","author":"Song S.","unstructured":"Song , S. , and Xiao , J . 2016. Deep sliding shapes for amodal 3d object detection in rgb-d images . In Proc. CVPR. Song, S., and Xiao, J. 2016. Deep sliding shapes for amodal 3d object detection in rgb-d images. In Proc. CVPR."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.114"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.308"},{"key":"e_1_2_1_35_1","unstructured":"Su H. Savva M. Yi L. Chang A. X. Song S. Yu F. Li Z. Xiao J. Huang Q. Savarese S. Funkhouser T. Hanrahan P. and Guibas L. J. 2015. ShapeNet: An information-rich 3d model repository. http:\/\/www.shapenet.org\/.  Su H. Savva M. Yi L. Chang A. X. Song S. Yu F. Li Z. Xiao J. Huang Q. Savarese S. Funkhouser T. Hanrahan P. and Guibas L. J. 2015. ShapeNet: An information-rich 3d model repository. http:\/\/www.shapenet.org\/."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0620-5"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2751556"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992696"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2661229.2661242"},{"key":"e_1_2_1_40_1","volume-title":"Proc. CVPR","author":"Wu Z.","unstructured":"Wu , Z. , Song , S. , Khosla , A. , Yu , F. , Zhang , L. , Tang , X. , and Xiao , J . 2015. 3D ShapeNets: A deep representation for volumetric shapes . In Proc. CVPR , 1912--1920. Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. 2015. 3D ShapeNets: A deep representation for volumetric shapes. In Proc. CVPR, 1912--1920."},{"key":"e_1_2_1_41_1","volume-title":"Proc. CVPR, 842--850","author":"Xiao T.","unstructured":"Xiao , T. , Xu , Y. , Yang , K. , Zhang , J. , Peng , Y. , and Zhang , Z . 2015. The application of two-level attention models in deep convolutional neural network for fine-grained image classification . In Proc. CVPR, 842--850 . Xiao, T., Xu, Y., Yang, K., Zhang, J., Peng, Y., and Zhang, Z. 2015. The application of two-level attention models in deep convolutional neural network for fine-grained image classification. In Proc. CVPR, 842--850."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461968"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818075"},{"key":"e_1_2_1_44_1","unstructured":"Xu K. Ba J. Kiros R. Courville A. Salakhutdinov R. Zemel R. and Bengio Y. 2015. Show attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044.  Xu K. Ba J. Kiros R. Courville A. Salakhutdinov R. Zemel R. and Bengio Y. 2015. Show attend and tell: Neural image caption generation with visual attention. arXiv preprint arXiv:1502.03044."},{"key":"e_1_2_1_45_1","volume-title":"Proc. NIPS, 1601--1608","author":"Zelnik-Manor L.","unstructured":"Zelnik-Manor , L. , and Perona , P . 2004. Self-tuning spectral clustering . In Proc. NIPS, 1601--1608 . Zelnik-Manor, L., and Perona, P. 2004. Self-tuning spectral clustering. In Proc. NIPS, 1601--1608."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2768821"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2980179.2980224","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2980179.2980224","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:23:22Z","timestamp":1750220602000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2980179.2980224"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,11,11]]},"references-count":46,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2016,11,11]]}},"alternative-id":["10.1145\/2980179.2980224"],"URL":"https:\/\/doi.org\/10.1145\/2980179.2980224","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,11,11]]},"assertion":[{"value":"2016-12-05","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}