{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T03:02:39Z","timestamp":1785466959330,"version":"3.56.0"},"reference-count":31,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2019,9,21]],"date-time":"2019-09-21T00:00:00Z","timestamp":1569024000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key Research and Development Program \u201cIntelligent Robot\u201d Key Special Project","award":["2018YFB1307100"],"award-info":[{"award-number":["2018YFB1307100"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61673136"],"award-info":[{"award-number":["61673136"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Self-Planned Task of State Key Laboratory of Robotics and System (HIT)","award":["NO. SKLRS201906B"],"award-info":[{"award-number":["NO. SKLRS201906B"]}]},{"name":"Self-Planned Task of State Key Laboratory of Robotics and System (HIT)","award":["NO. SKLRS201715A"],"award-info":[{"award-number":["NO. SKLRS201715A"]}]},{"name":"the Foundation for Innovative Research Groups of the National Natural Science Foundation of China","award":["No.51521003"],"award-info":[{"award-number":["No.51521003"]}]},{"name":"the ST Engineering-NTU Corporate Lab through the NRF corporate lab@university scheme","award":["SRP4"],"award-info":[{"award-number":["SRP4"]}]},{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"publisher","award":["201706120137"],"award-info":[{"award-number":["201706120137"]}],"id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>To autonomously move and operate objects in cluttered indoor environments, a service robot requires the ability of 3D scene perception. Though 3D object detection can provide an object-level environmental description to fill this gap, a robot always encounters incomplete object observation, recurring detections of the same object, error in detection, or intersection between objects when conducting detection continuously in a cluttered room. To solve these problems, we propose a two-stage 3D object detection algorithm which is to fuse multiple views of 3D object point clouds in the first stage and to eliminate unreasonable and intersection detections in the second stage. For each view, the robot performs a 2D object semantic segmentation and obtains 3D object point clouds. Then, an unsupervised segmentation method called Locally Convex Connected Patches (LCCP) is utilized to segment the object accurately from the background. Subsequently, the Manhattan Frame estimation is implemented to calculate the main orientation of the object and subsequently, the 3D object bounding box can be obtained. To deal with the detected objects in multiple views, we construct an object database and propose an object fusion criterion to maintain it automatically. Thus, the same object observed in multi-view is fused together and a more accurate bounding box can be calculated. Finally, we propose an object filtering approach based on prior knowledge to remove incorrect and intersecting objects in the object dataset. Experiments are carried out on both SceneNN dataset and a real indoor environment to verify the stability and accuracy of 3D semantic segmentation and bounding box detection of the object with multi-view fusion.<\/jats:p>","DOI":"10.3390\/s19194092","type":"journal-article","created":{"date-parts":[[2019,9,23]],"date-time":"2019-09-23T03:26:32Z","timestamp":1569209192000},"page":"4092","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":29,"title":["Multi-View Fusion-Based 3D Object Detection for Robot Indoor Scene Perception"],"prefix":"10.3390","volume":"19","author":[{"given":"Li","family":"Wang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruifeng","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jingwen","family":"Sun","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xingxing","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Physical Sciences, University of Science and Technology of China, Hefei 230026, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9108-8276","authenticated-orcid":false,"given":"Lijun","family":"Zhao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Robotics and System, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hock Soon","family":"Seah","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Nanyang Technological University, Singapore 639798, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chee Kwang","family":"Quah","sequence":"additional","affiliation":[{"name":"ST Electronics (Training &amp; Simulation Systems) Pte Ltd., Singapore 567714, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0526-3497","authenticated-orcid":false,"given":"Budianto","family":"Tandianus","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Nanyang Technological University, Singapore 639798, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2019,9,21]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelo, D., Erhan, D., Szegedy, C., Reed, S., Fu, C., and Berg, A. (2016, January 8\u201316). SSD: Single Shot MultiBox Detector. Proceedings of the European Conference on Computer Vision (ECCV 2016), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_5","first-page":"1","article-title":"Mask R-CNN","volume":"99","author":"He","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Silberman, N., Hoiem, D., Kohli, P., and Fergus, R. (2012, January 7\u201313). Indoor segmentation and support inference from RGBD images. Proceedings of the 12th European Conference on Computer Vision (ECCV 2012), Florence, Italy.","DOI":"10.1007\/978-3-642-33715-4_54"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Song, S., Lichtenberg, S.P., and Xiao, J. (2015, January 7\u201312). SUN RGB-D: A RGB-D scene understanding benchmark suite. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2015), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298655"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Song, S., and Xiao, J. (2016, January 27\u201330). Deep sliding shapes for amodal 3D object detection in RGB-D images. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.94"},{"key":"ref_9","unstructured":"Qi, C., Su, H., Mo, K., and Guibas, L. (2017, January 21\u201326). PointNet: Deep learning on point sets for 3D classification and segmentation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA."},{"key":"ref_10","unstructured":"Qi, C., Yi, L., Su, H., and Guibas, L. (2017, January 4\u20139). PointNet++: Deep hierarchical feature learning on point sets in a metric space. Proceedings of the Advances in Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Liu, W., Wu, C., Su, H., and Guibas, L.J. (2018, January 18\u201322). Frustum PointNets for 3D object detection from RGB-D Data. Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1109\/TRO.2006.889486","article-title":"Improved techniques for grid mapping with Rao-Blackwellized particle filters","volume":"23","author":"Grisetti","year":"2007","journal-title":"IEEE Trans. Rob."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hess, W., Kohler, D., Rapp, H., and Andor, D. (2016, January 16\u201321). Real-time loop closure in 2D LIDAR SLAM. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2016), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487258"},{"key":"ref_14","first-page":"1","article-title":"ORB-SLAM2: An open-source SLAM system for monocular, stereo, and RGB-D cameras","volume":"33","author":"Tardos","year":"2017","journal-title":"IEEE Trans. Rob."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1697","DOI":"10.1177\/0278364916669237","article-title":"ElasticFusion","volume":"35","author":"Whelan","year":"2016","journal-title":"Int. J. Robot. Res."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Gupta, S., Arbel\u00e1ez, P., and Malik, J. (2013, January 18\u201322). Perceptual organization and recognition of indoor scenes from RGB-D images. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2013), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.79"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Gupta, S., Girshick, R., Arbel\u00e1ez, P., and Malik, J. (2014, January 6\u201312). Learning rich features from RGB-D images for object detection and segmentation. Proceedings of the 12th European Conference on Computer Vision (ECCV 2014), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10584-0_23"},{"key":"ref_18","unstructured":"Ren, X., Bo, L., and Fox, D. (2012, January 16\u201321). RGB-(D) scene labeling: Features and algorithms. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2012), Providence, RI, USA."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Lin, D., Fidler, S., and Urtasun, R. (2013, January 3\u20136). Holistic scene understanding for 3D object detection with RGBD cameras. Proceedings of the 2013 IEEE International Conference on Computer Vision (ICCV 2013), Sydney, NSW, Australia.","DOI":"10.1109\/ICCV.2013.179"},{"key":"ref_20","unstructured":"Zhuo, D., and Latecki, L.J. (2017, January 21\u201326). Amodal detection of 3D objects: Inferring 3D bounding boxes from 2D ones in RGB-Depth images. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2017), Honolulu, HI, USA."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ren, Z., and Sudderth, E.B. (2016, January 27\u201330). Three-dimensional object detection and layout prediction using clouds of oriented gradients. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.169"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lahoud, J., and Ghanem, B. (2017, January 22\u201329). 2D-driven 3D object detection in RGB-D images. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV 2017), Venice, Italy.","DOI":"10.1109\/ICCV.2017.495"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Antonello, M., Wolf, D., Prankl, J., Ghidoni, S., Menegatti, E., and Vincze, M. (2018, January 21\u201325). Multi-view 3D entangled forest for semantic segmentation and mapping. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2018), Brisbane, QLD, Australia.","DOI":"10.1109\/ICRA.2018.8460837"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tateno, K., Tombari, F., and Navab, N. (2016, January 16\u201321). When 2.5D is not enough: Simultaneous reconstruction, segmentation and recognition on dense SLAM. Proceedings of the IEEE International Conference on Robotics and Automation (ICRA 2016), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487378"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"3206","DOI":"10.1109\/ACCESS.2018.2887022","article-title":"Efficient Object-Oriented Semantic Mapping with Object Detector","volume":"7","author":"Nakajima","year":"2019","journal-title":"IEEE Access."},{"key":"ref_26","unstructured":"Prisacariu, V., K\u00e4hler, O., Golodetz, S., Sapienza, M., Cavallari, T., Torr, P., and Murray, D. (2017). InfiniTAM v3: A Framework for Large-Scale 3D Reconstruction with Loop Closure. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"3037","DOI":"10.1109\/LRA.2019.2923960","article-title":"Volumetric instance-aware semantic mapping and 3D object discovery","volume":"4","author":"Grinvald","year":"2019","journal-title":"IEEE Robot. Automat. Lett."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Joo, K., Oh, T., and Kim, I. (2016, January 27\u201330). Globally optimal Manhattan frame estimation in real-time. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.195"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Hua, B., Pham, Q., Nguyen, D., Tran, M., and Yu, L. (2016, January 25\u201328). SceneNN: A Scene Meshes Dataset with aNNotations. Proceedings of the 2016 4th International Conference on 3D Vision (3DV 2016), Stanford, CA, USA.","DOI":"10.1109\/3DV.2016.18"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"3005","DOI":"10.1109\/TVCG.2017.2772238","article-title":"A robust 3D-2D interactive tool for scene segmentation and annotation","volume":"24","author":"Nguyen","year":"2018","journal-title":"IEEE T. Vis. Comput. Gr."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Pham, Q., Hua, B., Nguyen, D., and Yeung, S. (2019, January 7\u201311). Real-time progressive 3D semantic segmentation for indoor scene. Proceedings of the IEEE Winter Conference on Applications of Computer Vision (WACV 2019), Hilton Waikoloa Village, HI, USA.","DOI":"10.1109\/WACV.2019.00121"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/19\/4092\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:22:51Z","timestamp":1760188971000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/19\/4092"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,9,21]]},"references-count":31,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2019,10]]}},"alternative-id":["s19194092"],"URL":"https:\/\/doi.org\/10.3390\/s19194092","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,9,21]]}}}