{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,31]],"date-time":"2026-07-31T17:40:14Z","timestamp":1785519614252,"version":"3.56.0"},"reference-count":40,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2018,9,14]],"date-time":"2018-09-14T00:00:00Z","timestamp":1536883200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In this paper, a novel Pixel-Voxel network is proposed for dense 3D semantic mapping, which can perform dense 3D mapping while simultaneously recognizing and labelling the semantic category each point in the 3D map. In our approach, we fully leverage the advantages of different modalities. That is, the PixelNet can learn the high-level contextual information from 2D RGB images, and the VoxelNet can learn 3D geometrical shapes from the 3D point cloud. Unlike the existing architecture that fuses score maps from different modalities with equal weights, we propose a softmax weighted fusion stack that adaptively learns the varying contributions of PixelNet and VoxelNet and fuses the score maps according to their respective confidence levels. Our approach achieved competitive results on both the SUN RGB-D and NYU V2 benchmarks, while the runtime of the proposed system is boosted to around 13 Hz, enabling near-real-time performance using an i7 eight-cores PC with a single Titan X GPU.<\/jats:p>","DOI":"10.3390\/s18093099","type":"journal-article","created":{"date-parts":[[2018,9,14]],"date-time":"2018-09-14T10:57:59Z","timestamp":1536922679000},"page":"3099","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Dense RGB-D Semantic Mapping with Pixel-Voxel Neural Network"],"prefix":"10.3390","volume":"18","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8502-3233","authenticated-orcid":false,"given":"Cheng","family":"Zhao","sequence":"first","affiliation":[{"name":"Extreme Robotics Lab, University of Birmingham, Birmingham B15 2TT, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0393-8665","authenticated-orcid":false,"given":"Li","family":"Sun","sequence":"additional","affiliation":[{"name":"Lincoln Centre for Autonomous Systems (L-CAS), University of Lincoln, Lincoln LN6 7TS, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0684-1209","authenticated-orcid":false,"given":"Pulak","family":"Purkait","sequence":"additional","affiliation":[{"name":"Cambridge Research Lab, Toshiba Research Europe, Cambridge CB4 0GZ, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2971-7905","authenticated-orcid":false,"given":"Tom","family":"Duckett","sequence":"additional","affiliation":[{"name":"Lincoln Centre for Autonomous Systems (L-CAS), University of Lincoln, Lincoln LN6 7TS, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0890-8836","authenticated-orcid":false,"given":"Rustam","family":"Stolkin","sequence":"additional","affiliation":[{"name":"Extreme Robotics Lab, University of Birmingham, Birmingham B15 2TT, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2018,9,14]]},"reference":[{"key":"ref_1","unstructured":"Purkait, P., Zhao, C., and Zach, C. (arXiv, 2018). SPP-Net: Deep Absolute Pose Regression with Synthetic Views, arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhao, C., Sun, L., Purkait, P., Duckett, T., and Stolkin, R. (arXiv, 2018). Learning monocular visual odometry with dense 3D mapping from dense 3D flow, arXiv.","DOI":"10.1109\/IROS.2018.8594151"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"253","DOI":"10.1016\/j.dcan.2015.09.002","article-title":"Building a grid-semantic map for the navigation of service robots through human-robot interaction","volume":"1","author":"Zhao","year":"2015","journal-title":"Digit. Commun. Netw."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zhao, C., Hu, H., and Gu, D. (2015, January 11\u201312). Building a grid-point cloud-semantic map based on graph for the navigation of intelligent wheelchair. Proceedings of the 2015 IEEE International Conference on Automation and Computing (ICAC), Glasgow, UK.","DOI":"10.1109\/IConAC.2015.7313995"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Sun, L., Yan, Z., Molina, S., Hanheide, M., and Duckett, T. (2018, January 21\u201325). 3DOF pedestrian trajectory prediction learned from long-term autonomous mobile robot deployment data. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8461228"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Valiente, D., Pay\u00e1, L., Jim\u00e9nez, L.M., Sebasti\u00e1n, J.M., and Reinoso, \u00d3. (2018). Visual Information Fusion through Bayesian Inference for Adaptive Probability-Oriented Feature Matching. Sensors, 18.","DOI":"10.3390\/s18072041"},{"key":"ref_7","unstructured":"Sun, L., Zhao, C., Duckett, T., and Stolkin, R. (arXiv, 2017). Weakly-supervised DCNN for RGB-D object recognition in real-world applications which lack large-scale annotated training data, arXiv."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1109\/TRO.2013.2279412","article-title":"3-D mapping with an RGB-D camera","volume":"30","author":"Endres","year":"2014","journal-title":"Trans. Robot."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Newcombe, R.A., Izadi, S., Hilliges, O., Molyneaux, D., Kim, D., Davison, A.J., Kohi, P., Shotton, J., Hodges, S., and Fitzgibbon, A. (2011, January 26\u201329). KinectFusion: Real-time dense surface mapping and tracking. Proceedings of the 2011 10th IEEE International Symposium on Mixed and Augmented Reality, Basel, Switzerland.","DOI":"10.1109\/ISMAR.2011.6092378"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Whelan, T., Leutenegger, S., Salas-Moreno, R., Glocker, B., and Davison, A. (2015, January 13\u201317). ElasticFusion: Dense SLAM without a pose graph. Proceedings of the Robotics: Scienceand Systems, Rome, Italy.","DOI":"10.15607\/RSS.2015.XI.001"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_12","unstructured":"Badrinarayanan, V., Kendall, A., and Cipolla, R. (arXiv, 2015). Segnet: A deep convolutional encoder-decoder architecture for image segmentation, arXiv."},{"key":"ref_13","unstructured":"Chen, L.C., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (arXiv, 2016). Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs, arXiv."},{"key":"ref_14","unstructured":"Hazirbas, C., Ma, L., Domokos, C., and Cremers, D. (2016). Fusenet: Incorporating depth into semantic segmentation via fusion-based cnn architecture. Asian Conference on Computer Vision, Springer."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, Z., Gan, Y., Liang, X., Yu, Y., Cheng, H., and Lin, L. (2016). LSTM-CF: Unifying context modeling and fusion with LSTMS for RGB-D scene labelling. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46475-6_34"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Cheng, Y., Cai, R., Li, Z., Zhao, X., and Huang, K. (2017, January 21\u201326). Locality-Sensitive Deconvolution Networks with Gated Fusion for RGB-D Indoor Semantic Segmentation. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.161"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Lin, D., Chen, G., Cohen-Or, D., Heng, P.A., and Huang, H. (2017, January 22\u201329). Cascaded Feature Network for Semantic Segmentation of RGB-D Images. Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.147"},{"key":"ref_18","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (arXiv, 2016). Pointnet: Deep learning on point sets for 3d classification and segmentation, arXiv."},{"key":"ref_19","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (arXiv, 2017). PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space, arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Salas-Moreno, R.F., Newcombe, R.A., Strasdat, H., Kelly, P.H., and Davison, A.J. (2013, January 1\u20138). Slam++: Simultaneous localisation and mapping at the level of objects. Proceedings of the 2013 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Sydney, Australia.","DOI":"10.1109\/CVPR.2013.178"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Hermans, A., Floros, G., and Leibe, B. (June, January 31). Dense 3D semantic mapping of indoor scenes from RGB-D images. Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA), Hong Kong, China.","DOI":"10.1109\/ICRA.2014.6907236"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"McCormac, J., Handa, A., Davison, A., and Leutenegger, S. (June, January 29). Semanticfusion: Dense 3D semantic mapping with convolutional neural networks. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore.","DOI":"10.1109\/ICRA.2017.7989538"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Xiang, Y., and Fox, D. (arXiv, 2017). DA-RNN: Semantic Mapping with Data Associated Recurrent Neural Networks, arXiv.","DOI":"10.15607\/RSS.2017.XIII.013"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Tateno, K., Tombari, F., and Navab, N. (2016, January 16\u201321). When 2.5 D is not enough: Simultaneous reconstruction, segmentation and recognition on dense SLAM. Proceedings of the 2016 IEEE International Conference on Robotics and Automation (ICRA), Stockholm, Sweden.","DOI":"10.1109\/ICRA.2016.7487378"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Vineet, V., Miksik, O., Lidegaard, M., Nie\u00dfner, M., Golodetz, S., Prisacariu, V.A., K\u00e4hler, O., Murray, D.W., Izadi, S., and P\u00e9rez, P. (2015, January 26\u201330). Incremental dense semantic stereo fusion for large-scale semantic scene reconstruction. Proceedings of the 2015 IEEE International Conference on Robotics and Automation (ICRA), Seattle, WA, USA.","DOI":"10.1109\/ICRA.2015.7138983"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tateno, K., Tombari, F., Laina, I., and Navab, N. (arXiv, 2017). CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction, arXiv.","DOI":"10.1109\/CVPR.2017.695"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhao, C., Sun, L., and Stolkin, R. (2017, January 10\u201312). A fully end-to-end deep learning approach for real-time simultaneous 3D reconstruction and material recognition. Proceedings of the 2017 IEEE International Conference on Advanced Robotics (ICAR), Hong Kong, China.","DOI":"10.1109\/ICAR.2017.8023499"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ma, L., St\u00fcckler, J., Kerl, C., and Cremers, D. (arXiv, 2017). Multi-view deep learning for consistent semantic mapping with RGB-D cameras, arXiv.","DOI":"10.1109\/IROS.2017.8202213"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Mustafa, A., and Hilton, A. (2017, January 21\u201326). Semantically Coherent Co-segmentation and Reconstruction of Dynamic Scenes. Proceedings of the 2017 Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.592"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Noh, H., Hong, S., and Han, B. (2015, January 11\u201318). Learning deconvolution network for semantic segmentation. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Las Condes, Chile.","DOI":"10.1109\/ICCV.2015.178"},{"key":"ref_31","unstructured":"Kr\u00e4henb\u00fchl, P., and Koltun, V. (2011). Efficient inference in fully connected crfs with gaussian edge potentials. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Zheng, S., Jayasumana, S., Romera-Paredes, B., Vineet, V., Su, Z., Du, D., Huang, C., and Torr, P.H. (2015, January 11\u201318). Conditional random fields as recurrent neural networks. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Las Condes, Chile.","DOI":"10.1109\/ICCV.2015.179"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"He, Y., Chiu, W.C., Keuper, M., Fritz, M., and Campus, S.I. (2017, January 21\u201326). STD2P: RGBD Semantic Segmentation using Spatio-Temporal Data-Driven Pooling. Proceedings of the 2017 Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.757"},{"key":"ref_34","unstructured":"Shuai, B., Liu, T., and Wang, G. (arXiv, 2016). Improving Fully Convolution Network for Semantic Segmentation, arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Shuai, B., Zuo, Z., Wang, B., and Wang, G. (July, January 26). Dag-recurrent neural networks for scene labelling. Proceedings of the 2016 Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.394"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1352","DOI":"10.1109\/TPAMI.2017.2708714","article-title":"Exploring context with deep structured models for semantic segmentation","volume":"40","author":"Lin","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Lin, G., Milan, A., Shen, C., and Reid, I. (2017, January 21\u201326). Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. Proceedings of the 2017 Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.549"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Handa, A., Patraucean, V., Badrinarayanan, V., Stent, S., and Cipolla, R. (July, January 26). Understanding real world indoor scenes with synthetic data. Proceedings of the 2016 Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.442"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 11\u201318). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Las Condes, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"3749","DOI":"10.1109\/LRA.2018.2856268","article-title":"Recurrent-OctoMap: Learning State-Based Map Refinement for Long-Term Semantic Mapping With 3-D-Lidar Data","volume":"3","author":"Sun","year":"2018","journal-title":"IEEE Robot. Autom. Lett."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/9\/3099\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:20:36Z","timestamp":1760196036000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/18\/9\/3099"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,9,14]]},"references-count":40,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2018,9]]}},"alternative-id":["s18093099"],"URL":"https:\/\/doi.org\/10.3390\/s18093099","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,9,14]]}}}