{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T20:32:23Z","timestamp":1769286743196,"version":"3.49.0"},"reference-count":38,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2019,2,21]],"date-time":"2019-02-21T00:00:00Z","timestamp":1550707200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61673136"],"award-info":[{"award-number":["61673136"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Self-Planned Task of State Key Laboratory of Robotics and System (HIT)","award":["NO. SKLRS201715A, NO. SKLRS201609B"],"award-info":[{"award-number":["NO. SKLRS201715A, NO. SKLRS201609B"]}]},{"name":"the Foundation for Innovative Research Groups of the National Natural Science Foundation of China","award":["No.51521003"],"award-info":[{"award-number":["No.51521003"]}]},{"name":"the ST Engineering-NTU Corporate Lab through the NRF corporate lab@ university scheme","award":["SRP4"],"award-info":[{"award-number":["SRP4"]}]},{"name":"the China Scholarship Council","award":["201706120137"],"award-info":[{"award-number":["201706120137"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Environmental perception is a vital feature for service robots when working in an indoor environment for a long time. The general 3D reconstruction is a low-level geometric information description that cannot convey semantics. In contrast, higher level perception similar to humans requires more abstract concepts, such as objects and scenes. Moreover, the 2D object detection based on images always fails to provide the actual position and size of an object, which is quite important for a robot\u2019s operation. In this paper, we focus on the 3D object detection to regress the object\u2019s category, 3D size, and spatial position through a convolutional neural network (CNN). We propose a multi-channel CNN for 3D object detection, which fuses three input channels including RGB, depth, and bird\u2019s eye view (BEV) images. We also propose a method to generate 3D proposals based on 2D ones in the RGB image and semantic prior. Training and test are conducted on the modified NYU V2 dataset and SUN RGB-D dataset in order to verify the effectiveness of the algorithm. We also carry out the actual experiments in a service robot to utilize the proposed 3D object detection method to enhance the environmental perception of the robot.<\/jats:p>","DOI":"10.3390\/s19040893","type":"journal-article","created":{"date-parts":[[2019,2,22]],"date-time":"2019-02-22T03:49:44Z","timestamp":1550807384000},"page":"893","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Multi-Channel Convolutional Neural Network Based 3D Object Detection for Indoor Robot Environmental Perception"],"prefix":"10.3390","volume":"19","author":[{"given":"Li","family":"Wang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruifeng","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hezi","family":"Shi","sequence":"additional","affiliation":[{"name":"EON Reality Pte Ltd, Singapore 138567, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingwen","family":"Sun","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9108-8276","authenticated-orcid":false,"given":"Lijun","family":"Zhao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hock Soon","family":"Seah","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Nanyang Technological University, Singapore 639798, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chee Kwang","family":"Quah","sequence":"additional","affiliation":[{"name":"ST Electronics (Training &amp; Simulation Systems) Pte Ltd, Singapore 567714, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0526-3497","authenticated-orcid":false,"given":"Budianto","family":"Tandianus","sequence":"additional","affiliation":[{"name":"ST Engineering-NTU Corporate Laboratory, School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore 637335, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,2,21]]},"reference":[{"key":"ref_1","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Conference on Neural Information Processing Systems (NIPS 2012), Lake Tahoe, NV, USA."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2014), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV 2015), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? The KITTI vision benchmark suite. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2012), Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The KITTI dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Liu, W., Wu, C., Su, H., and Guibas, L.J. (2018, January 18\u201322). Frustum PointNets for 3D object detection from RGB-D Data. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zhou, Y., and Tuzel, O. (2018, January 18\u201322). VoxelNet: End-to-end learning for point cloud based 3D object detection. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00472"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Choi, W., Lin, Y., and Savarese, S. (2015, January 8\u201310). Data-driven 3D voxel patterns for object category recognition. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298800"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., and Urtasun, R. (2016, January 27\u201330). Monocular 3D object detection for autonomous driving. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.236"},{"key":"ref_14","unstructured":"Chen, X., and Zhu, Y. (2015, January 7\u201312). 3D object proposals for accurate object class detection. Proceedings of the Conference on Neural Information Processing Systems (NIPS 2015), Montreal, QC, Canada."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chen, X., Ma, H., Wan, J., Li, B., and Xia, T. (2017, January 21\u201326). Multi-view 3D object detection network for autonomous driving. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.691"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ku, J., Mozifian, M., Lee, J., Harakeh, A., and Waslander, S.L. (2018, January 1\u20135). Joint 3D proposal generation and object detection from view aggregation. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS 2018), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8594049"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012, January 7\u201313). Indoor segmentation and support inference from RGBD images. Proceedings of the 12th European Conference on Computer Vision (ECCV 2012), Florence, Italy.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Song, S., Lichtenberg, S.P., and Xiao, J. (2015, January 7\u201312). SUN RGB-D: A RGB-D scene understanding benchmark suite. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298655"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1109\/JSEN.2014.2336987","article-title":"Using scale coordination and semantic information for robust 3-D object recognition by a service robot","volume":"15","author":"Zhuang","year":"2015","journal-title":"IEEE Sens. J."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"9370","DOI":"10.1109\/JSEN.2018.2870957","article-title":"Visual object recognition and pose estimation based on a deep semantic segmentation network","volume":"18","author":"Lin","year":"2018","journal-title":"IEEE Sens. J."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object detection with discriminatively trained part-based models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_24","unstructured":"Redmon, J., and Farhadi, A. (arXiv, 2018). YOLOv3: An incremental improvement, arXiv."},{"key":"ref_25","unstructured":"Dai, J., Li, Y., He, K., and Sun, J. (2016, January 5\u201310). R-FCN: Object detection via region-based fully convolutional networks. Proceedings of the Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Engelcke, M., Rao, D., Wang, D., Tong, C., and Posner, I. (June, January 29). Vote3Deep: Fast object detection in 3D point clouds using efficient convolutional neural networks. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA 2017), Singapore, Singapore.","DOI":"10.1109\/ICRA.2017.7989161"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Song, S., and Xiao, J. (2016, January 27\u201330). Deep sliding shapes for amodal 3D object detection in RGB-D images. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.94"},{"key":"ref_28","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). PointNet: Deep learning on point sets for 3D classification and segmentation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA."},{"key":"ref_29","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017, January 4\u20139). PointNet++: Deep hierarchical feature learning on point sets in a metric space. Proceedings of the Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Su, H., Nie\u00dfner, M., Dai, A., Yan, M., and Guibas, L.J. (2016, January 27\u201330). Volumetric and multi-view CNNs for object classification on 3D data. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.609"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"2967","DOI":"10.1109\/TIP.2011.2142006","article-title":"A multi-level mixture-of-experts framework for pedestrian classification","volume":"20","author":"Enzweiler","year":"2011","journal-title":"IEEE Trans. Image Process."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"3980","DOI":"10.1109\/TCYB.2016.2593940","article-title":"On-board object detection: Multicue, multimodal, and multiview random forest of local experts","volume":"47","author":"Gonzalez","year":"2017","journal-title":"IEEE Trans. Cybern."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Lahoud, J., and Ghanem, B. (2017, January 22\u201329). 2D-driven 3D object detection in RGB-D images. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV 2017), Venice, Italy.","DOI":"10.1109\/ICCV.2017.495"},{"key":"ref_34","unstructured":"Zhuo, D., and Latecki, L.J. (2017, January 21\u201326). Amodal detection of 3D objects: Inferring 3D bounding boxes from 2D ones in RGB-Depth images. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Arbelaez, P., Pont-Tuset, J., Barron, J., Marques, F., and Malik, J. (2014, January 23\u201328). Multiscale combinatorial grouping. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2014), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.49"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Ren, Z., and Sudderth, E.B. (2016, January 27\u201330). Three-dimensional object detection and layout prediction using clouds of oriented gradients. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.169"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1147","DOI":"10.1109\/TRO.2015.2463671","article-title":"ORB-SLAM: A versatile and accurate monocular SLAM system","volume":"31","author":"Montiel","year":"2015","journal-title":"IEEE Trans. Rob."},{"key":"ref_38","first-page":"1","article-title":"ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras","volume":"33","author":"Tardos","year":"2017","journal-title":"IEEE Trans. Rob."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/4\/893\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:33:42Z","timestamp":1760186022000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/4\/893"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,2,21]]},"references-count":38,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2019,2]]}},"alternative-id":["s19040893"],"URL":"https:\/\/doi.org\/10.3390\/s19040893","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,2,21]]}}}