{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T16:14:46Z","timestamp":1781108086235,"version":"3.54.1"},"reference-count":55,"publisher":"MDPI AG","issue":"18","license":[{"start":{"date-parts":[[2022,9,7]],"date-time":"2022-09-07T00:00:00Z","timestamp":1662508800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61976227"],"award-info":[{"award-number":["61976227"]}]},{"name":"National Natural Science Foundation of China","award":["62176096"],"award-info":[{"award-number":["62176096"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>3D LiDAR has become an indispensable sensor in autonomous driving vehicles. In LiDAR-based 3D point cloud semantic segmentation, most voxel-based 3D segmentors cannot efficiently capture large amounts of context information, resulting in limited receptive fields and limiting their performance. To address this problem, a sparse voxel-based attention network is introduced for 3D LiDAR point cloud semantic segmentation, termed SVASeg, which captures large amounts of context information between voxels through sparse voxel-based multi-head attention (SMHA). The traditional multi-head attention cannot directly be applied to the non-empty sparse voxels. To this end, a hash table is built according to the incrementation of voxel coordinates to lookup the non-empty neighboring voxels of each sparse voxel. Then, the sparse voxels are grouped into different groups, and each group corresponds to a local region. Afterwards, position embedding, multi-head attention and feature fusion are performed for each group to capture and aggregate the context information. Based on the SMHA module, the SVASeg can directly operate on the non-empty voxels, maintaining a comparable computational overhead to the convolutional method. Extensive experimental results on the SemanticKITTI and nuScenes datasets show the superiority of SVASeg.<\/jats:p>","DOI":"10.3390\/rs14184471","type":"journal-article","created":{"date-parts":[[2022,9,8]],"date-time":"2022-09-08T04:18:32Z","timestamp":1662610712000},"page":"4471","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":50,"title":["SVASeg: Sparse Voxel-Based Attention for 3D LiDAR Point Cloud Semantic Segmentation"],"prefix":"10.3390","volume":"14","author":[{"given":"Lin","family":"Zhao","sequence":"first","affiliation":[{"name":"National Key Laboratory of Science and Technology on Multi-Spectral Information Processing, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Siyuan","family":"Xu","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Science and Technology on Multi-Spectral Information Processing, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3775-8571","authenticated-orcid":false,"given":"Liman","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Biomedical Engineering, South-Central University for Nationalities, Wuhan 430074, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Delie","family":"Ming","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Science and Technology on Multi-Spectral Information Processing, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3284-864X","authenticated-orcid":false,"given":"Wenbing","family":"Tao","sequence":"additional","affiliation":[{"name":"National Key Laboratory of Science and Technology on Multi-Spectral Information Processing, School of Artificial Intelligence and Automation, Huazhong University of Science and Technology, Wuhan 430074, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,9,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Hu, Q., Yang, B., Xie, L., Rosa, S., Guo, Y., Wang, Z., Trigoni, N., and Markham, A. (2020, January 13\u201319). Randla-net: Efficient semantic segmentation of large-scale point clouds. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01112"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Liu, L., Yu, J., Tan, L., Su, W., Zhao, L., and Tao, W. (2021). Semantic Segmentation of 3D Point Cloud Based on Spatial Eight-Quadrant Kernel Convolution. Remote Sens., 13.","DOI":"10.3390\/rs13163140"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Xu, T., Gao, X., Yang, Y., Xu, L., Xu, J., and Wang, Y. (2022). Construction of a Semantic Segmentation Network for the Overhead Catenary System Point Cloud Based on Multi-Scale Feature Fusion. Remote Sens., 14.","DOI":"10.3390\/rs14122768"},{"key":"ref_4","first-page":"12951","article-title":"JSNet: Joint Instance and Semantic Segmentation of 3D Point Clouds","volume":"34","author":"Zhao","year":"2020","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Thomas, H., Qi, C.R., Deschaud, J.E., Marcotegui, B., Goulette, F., and Guibas, L.J. (2019\u20132, January 27). KPConv: Flexible and Deformable Convolution for Point Clouds. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00651"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ballouch, Z., Hajji, R., Poux, F., Kharroubi, A., and Billen, R. (2022). A Prior Level Fusion Approach for the Semantic Segmentation of 3D Point Clouds Using Deep Learning. Remote Sens., 14.","DOI":"10.3390\/rs14143415"},{"key":"ref_7","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017, January 4\u20139). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wu, W., Qi, Z., and Fuxin, L. (2019, January 15\u201320). PointConv: Deep Convolutional Networks on 3D Point Clouds. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00985"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Gao, F., Yan, Y., Lin, H., and Shi, R. (2022). PIIE-DSA-Net for 3D Semantic Segmentation of Urban Indoor and Outdoor Datasets. Remote Sens., 14.","DOI":"10.3390\/rs14153583"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Cortinhal, T., Tzelepis, G., and Aksoy, E.E. (2020, January 5\u20137). SalsaNext: Fast, uncertainty-aware semantic segmentation of LiDAR point clouds. Proceedings of the International Symposium on Visual Computing, San Diego, CA, USA.","DOI":"10.1007\/978-3-030-64559-5_16"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xu, C., Wu, B., Wang, Z., Zhan, W., Vajda, P., Keutzer, K., and Tomizuka, M. (2020, January 23\u201328). Squeezesegv3: Spatially-adaptive convolution for efficient point-cloud segmentation. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58604-1_1"},{"key":"ref_12","unstructured":"Kochanov, D., Nejadasl, F.K., and Booij, O. (2020). KPRNet: Improving projection-based LiDAR semantic segmentation. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhou, Z., David, P., Yue, X., Xi, Z., Gong, B., and Foroosh, H. (2020, January 13\u201319). Polarnet: An improved grid representation for online lidar point clouds semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00962"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Riegler, G., Osman Ulusoy, A., and Geiger, A. (2017, January 21\u201326). Octnet: Learning deep 3d representations at high resolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.701"},{"key":"ref_15","unstructured":"Liu, Z., Tang, H., Lin, Y., and Han, S. (2019). Point-voxel cnn for efficient 3d deep learning. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Graham, B., Engelcke, M., and van der Maaten, L. (2018, January 18\u201323). 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00961"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Tang, H., Liu, Z., Zhao, S., Lin, Y., Lin, J., Wang, H., and Han, S. (2020, January 23\u201328). Searching efficient 3d architectures with sparse point-voxel convolution. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58604-1_41"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhu, X., Zhou, H., Wang, T., Hong, F., Li, W., Ma, Y., Li, H., Yang, R., and Lin, D. (2021). Cylindrical and asymmetrical 3d convolution networks for lidar-based perception. IEEE Trans. Pattern Anal. Mach. Intell.","DOI":"10.1109\/CVPR46437.2021.00981"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Gerdzhev, M., Razani, R., Taghavi, E., and Bingbing, L. (June, January 30). Tornado-net: Multiview total variation semantic segmentation with diamond inception module. Proceedings of the 2021 IEEE International Conference on Robotics and Automation, Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9562041"},{"key":"ref_20","unstructured":"Zhao, L., Zhou, H., Zhu, X., Song, X., Li, H., and Tao, W. (2021). LIF-Seg: LiDAR and Camera Image Fusion for 3D LiDAR Semantic Segmentation. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Choy, C., Gwak, J., and Savarese, S. (2019, January 15\u201320). 4d spatio-temporal convnets: Minkowski convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00319"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021). Swin transformer: Hierarchical vision transformer using shifted windows. arXiv.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_23","unstructured":"Li, Z., Wang, W., Xie, E., Yu, Z., Anandkumar, A., Alvarez, J.M., Lu, T., and Luo, P. (2021). Panoptic SegFormer. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Mao, J., Xue, Y., Niu, M., Bai, H., Feng, J., Liang, X., Xu, H., and Xu, C. (2021, January 11\u201317). Voxel transformer for 3d object detection. Proceedings of the IEEE International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00315"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Fan, L., Pang, Z., Zhang, T., Wang, Y.X., Zhao, H., Wang, F., Wang, N., and Zhang, Z. (2021). Embracing Single Stride 3D Object Detector with Sparse Transformer. arXiv.","DOI":"10.1109\/CVPR52688.2022.00827"},{"key":"ref_26","unstructured":"Behley, J., Garbade, M., Milioto, A., Quenzel, J., Behnke, S., Stachniss, C., and Gall, J. (November, January 27). Semantickitti: A dataset for semantic scene understanding of lidar sequences. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2020, January 13\u201319). nuScenes: A multimodal dataset for autonomous driving. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"ref_28","unstructured":"Cao, H., Lu, Y., Lu, C., Pang, B., Liu, G., and Yuille, A. (2020). Asap-net: Attention and structure aware point cloud sequence segmentation. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Yan, X., Zheng, C., Li, Z., Wang, S., and Cui, S. (2020, January 13\u201319). Pointasnl: Robust point clouds processing using nonlocal neural networks with adaptive sampling. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00563"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"790","DOI":"10.1109\/LRA.2020.2965390","article-title":"Bayesian spatial kernel smoothing for scalable dense semantic mapping","volume":"5","author":"Gan","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Cheng, M., Hui, L., Xie, J., Yang, J., and Kong, H. (January, January 24). Cascaded non-local neural network for point cloud semantic segmentation. Proceedings of the 2020 IEEE International Conference on Intelligent Robots and Systems, Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9341531"},{"key":"ref_32","unstructured":"Fang, Y., Xu, C., Cui, Z., Zong, Y., and Yang, J. (2020). Spatial transformer point convolution. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Geng, X., Ji, S., Lu, M., and Zhao, L. (2021). Multi-scale attentive aggregation for LiDAR point cloud segmentation. Remote Sens., 13.","DOI":"10.3390\/rs13040691"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Milioto, A., Vizzo, I., Behley, J., and Stachniss, C. (2019, January 3\u20138). Rangenet++: Fast and accurate lidar semantic segmentation. Proceedings of the 2019 IEEE International Conference on Intelligent Robots and Systems, Macau, China.","DOI":"10.1109\/IROS40897.2019.8967762"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Duerr, F., Pfaller, M., Weigel, H., and Beyerer, J. (2020, January 25\u201328). LiDAR-based recurrent 3D semantic segmentation with temporal memory alignment. Proceedings of the 2020 International Conference on 3D Vision, Fukuoka, Japan.","DOI":"10.1109\/3DV50981.2020.00088"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Razani, R., Cheng, R., Taghavi, E., and Bingbing, L. (2021\u20135, January 30). Lite-hdseg: Lidar semantic segmentation using lite harmonic dense convolutions. Proceedings of the 2021 IEEE International Conference on Robotics and Automation, Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9561171"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Park, J., Kim, C., and Jo, K. (2022). PCSCNet: Fast 3D Semantic Segmentation of LiDAR Point Cloud for Autonomous Car using Point Convolution and Sparse Convolution Network. arXiv.","DOI":"10.1016\/j.eswa.2022.118815"},{"key":"ref_38","unstructured":"Liong, V.E., Nguyen, T.N.T., Widjaja, S., Sharma, D., and Chong, Z.J. (2020). AMVNet: Assertion-based Multi-View Fusion Network for LiDAR Semantic Segmentation. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wang, Y., Fathi, A., Kundu, A., Ross, D., Pantofaru, C., Funkhouser, T., and Solomon, J. (2020, January 23\u201328). Pillar-based object detection for autonomous driving. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58542-6_2"},{"key":"ref_40","unstructured":"Zhou, Y., Sun, P., Zhang, Y., Anguelov, D., Gao, J., Ouyang, T., Guo, J., Ngiam, J., and Vasudevan, V. (2020, January 16\u201318). End-to-end multi-view fusion for 3d object detection in lidar point clouds. Proceedings of the Conference on Robot Learning, PMLR, Virtual."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhang, F., Fang, J., Wah, B., and Torr, P. (2020, January 23\u201328). Deep fusionnet for point cloud semantic segmentation. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58586-0_38"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Chen, K., Oldja, R., Smolyanskiy, N., Birchfield, S., Popov, A., Wehr, D., Eden, I., and Pehserl, J. (2020). MVLidarNet: Real-Time Multi-Class Scene Understanding for Autonomous Driving Using Multiple Views. arXiv.","DOI":"10.1109\/IROS45743.2020.9341450"},{"key":"ref_43","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhao, H., Jiang, L., Jia, J., Torr, P.H., and Koltun, V. (2021, January 10\u201317). Point transformer. Proceedings of the IEEE International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01595"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Mazur, K., and Lempitsky, V. (2021, January 10\u201317). Cloud transformers: A universal approach to point cloud processing tasks. Proceedings of the IEEE International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01054"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Wang, J., Chakraborty, R., and Stella, X.Y. (2021). Spatial transformer for 3D point clouds. IEEE Trans. Pattern Anal. Mach. Intell.","DOI":"10.1109\/TPAMI.2021.3070341"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1007\/s41095-021-0229-5","article-title":"Pct: Point cloud transformer","volume":"7","author":"Guo","year":"2021","journal-title":"Comput. Vis. Media"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Berman, M., Triki, A.R., and Blaschko, M.B. (2018, January 18\u201323). The lov\u00e1sz-softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00464"},{"key":"ref_49","unstructured":"Shen, Z., Zhang, M., Zhao, H., Yi, S., and Li, H. (2021, January 3\u20138). Efficient attention: Attention with linear complexities. Proceedings of the IEEE Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Zhuang, Z., Li, R., Jia, K., Wang, Q., Li, Y., and Tan, M. (2021, January 10\u201317). Perception-aware Multi-sensor Fusion for 3D LiDAR Semantic Segmentation. Proceedings of the IEEE International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01597"},{"key":"ref_52","unstructured":"Rosu, R.A., Sch\u00fctt, P., Quenzel, J., and Behnke, S. (2019). Latticenet: Fast point cloud segmentation using permutohedral lattices. arXiv."},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"738","DOI":"10.1109\/LRA.2021.3132059","article-title":"Multi-scale interaction for real-time lidar data segmentation on an embedded platform","volume":"7","author":"Li","year":"2021","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"5432","DOI":"10.1109\/LRA.2020.3007440","article-title":"3d-mininet: Learning a 2d representation from point clouds for fast and efficient 3d lidar semantic segmentation","volume":"5","author":"Alonso","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Cheng, R., Razani, R., Taghavi, E., Li, E., and Liu, B. (2021). AF2-S3Net: Attentive Feature Fusion with Adaptive Feature Selection for Sparse Semantic Segmentation Network. arXiv.","DOI":"10.1109\/CVPR46437.2021.01236"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/18\/4471\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:25:23Z","timestamp":1760142323000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/14\/18\/4471"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,7]]},"references-count":55,"journal-issue":{"issue":"18","published-online":{"date-parts":[[2022,9]]}},"alternative-id":["rs14184471"],"URL":"https:\/\/doi.org\/10.3390\/rs14184471","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,7]]}}}