{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,22]],"date-time":"2025-12-22T12:39:47Z","timestamp":1766407187367,"version":"build-2065373602"},"reference-count":45,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2021,2,14]],"date-time":"2021-02-14T00:00:00Z","timestamp":1613260800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Key Research and Development Program of China","award":["2018YFB0505003"],"award-info":[{"award-number":["2018YFB0505003"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Semantic segmentation of LiDAR point clouds has implications in self-driving, robots, and augmented reality, among others. In this paper, we propose a Multi-Scale Attentive Aggregation Network (MSAAN) to achieve the global consistency of point cloud feature representation and super segmentation performance. First, upon a baseline encoder-decoder architecture for point cloud segmentation, namely, RandLA-Net, an attentive skip connection was proposed to replace the commonly used concatenation to balance the encoder and decoder features of the same scales. Second, a channel attentive enhancement module was introduced to the local attention enhancement module to boost the local feature discriminability and aggregate the local channel structure information. Third, we developed a multi-scale feature aggregation method to capture the global structure of a point cloud from both the encoder and the decoder. The experimental results reported that our MSAAN significantly outperformed state-of-the-art methods, i.e., at least 15.3% mIoU improvement for scene-2 of CSPC dataset, 5.2% for scene-5 of CSPC dataset, and 6.6% for Toronto3D dataset.<\/jats:p>","DOI":"10.3390\/rs13040691","type":"journal-article","created":{"date-parts":[[2021,2,14]],"date-time":"2021-02-14T05:54:49Z","timestamp":1613282089000},"page":"691","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":26,"title":["Multi-Scale Attentive Aggregation for LiDAR Point Cloud Segmentation"],"prefix":"10.3390","volume":"13","author":[{"given":"Xiaoxiao","family":"Geng","sequence":"first","affiliation":[{"name":"School of Remote Sensing and Information Engineering, Wuhan University, 129 Luoyu Road, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3088-1481","authenticated-orcid":false,"given":"Shunping","family":"Ji","sequence":"additional","affiliation":[{"name":"School of Remote Sensing and Information Engineering, Wuhan University, 129 Luoyu Road, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Meng","family":"Lu","sequence":"additional","affiliation":[{"name":"Department of Physical Geography, Faculty of Geoscience, Utrecht University, Princetonlaan 8, 3584 CB Utrecht, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lingli","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Remote Sensing and Information Engineering, Wuhan University, 129 Luoyu Road, Wuhan 430079, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,2,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"3749","DOI":"10.3390\/rs5083749","article-title":"SVM-based classification of segmented airborne LiDAR point clouds in urban areas","volume":"5","author":"Zhang","year":"2013","journal-title":"Remote Sens."},{"key":"ref_2","first-page":"207","article-title":"Airborne LiDAR feature selection for urban classification using random forests","volume":"38","author":"Chehata","year":"2009","journal-title":"Geomat. Inf. Sci. Wuhan Univ."},{"key":"ref_3","unstructured":"Zhuang, Y., Liu, Y., He, G., and Wang, W. (October, January 28). Contextual classification of 3D laser points with conditional random fields in urban environments. Proceedings of the IEEE International Conference on Intelligent Robots and Systems, Hamburg, Germany."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Lu, Y., and Rasmussen, C. (2012, January 7\u201312). Simplified Markov random fields for efficient semantic labeling of 3D point clouds. Proceedings of the IEEE International Conference on Intelligent Robots and Systems, Vilamoura, Portugal.","DOI":"10.1109\/IROS.2012.6386039"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"436","DOI":"10.1038\/nature14539","article-title":"Deep learning","volume":"521","author":"LeCun","year":"2015","journal-title":"Nature"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Lawin, F.J., Danelljan, M., Tosteberg, P., Bhat, G., Khan, F.S., and Felsberg, M. (2017, January 22\u201324). Deep projective 3D semantic segmentation. Proceedings of the 17th International Conference on Computer Analysis of Images and Patterns, Ystad, Sweden.","DOI":"10.1007\/978-3-319-64689-3_8"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Boulch, A., Saux, B.L., and Audebert, N. (2017). Unstructured point cloud semantic labeling using deep segmentation networks. Eurographics Workshop on 3D Object Retrieval, The Eurographics Association.","DOI":"10.1016\/j.cag.2017.11.010"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wu, B., Wan, A., Yue, X., and Keutzer, K. (2017). SqueezeSeg: Convolutional Neural Nets with Recurrent CRF for Real-Time Road-Object Segmentation from 3D LiDAR Point Cloud. arXiv.","DOI":"10.1109\/ICRA.2018.8462926"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wu, B., Zhou, X., Zhao, S., Yue, X., and Keutzer, K. (2018). SqueezeSegV2: Improved Model Structure and Unsupervised Domain Adaptation for Road-Object Segmentation from a LiDAR Point Cloud. arXiv.","DOI":"10.1109\/ICRA.2019.8793495"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Milioto, A., Vizzo, I., Behley, J., and Stachniss, C. (2019, January 3\u20138). RangeNet ++: Fast and Accurate LiDAR Semantic Segmentation. Proceedings of the 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China.","DOI":"10.1109\/IROS40897.2019.8967762"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Meng, H.Y., Gao, L., Lai, Y., and Manocha, D. (2018). VV-Net: Voxel Vaenet with Group Convolutions for Point Cloud Segmentation. arXiv.","DOI":"10.1109\/ICCV.2019.00859"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Rethage, D., Wald, J., Sturm, J., Navab, N., and Tombari, F. (2018, January 8\u201314). Fully-convolutional point networks for large-scale point clouds. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01225-0_37"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Graham, B., Engelcke, M., and van der Maaten, L. (2018, January 18\u201322). 3D Semantic Segmentation with Submanifold Sparse Convolutional Networks. Proceedings of the IEEE Computer Vision and Pattern Recognition CVPR, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00961"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Su, H., Jampani, V., Sun, D., Maji, S., Kalogerakis, V., Yang, M.-H., and Kautz, J. (2018, January 19\u201321). SPLATNet: Sparse lattice networks for point cloud processing. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00268"},{"key":"ref_15","unstructured":"Rosu, R.A., Schutt, P., Quenzel, J., and Behnke, S. (2019). Latticenet: Fast Point Cloud Segmentation Using Permutohedral Lattices. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Dai, A., and Nie\u00dfner, M. (2018, January 8\u201314). 3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01249-6_28"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Jaritz, M., Gu, J., and Su, H. (2019, January 1). Multi-view Pointnet for 3D Scene Understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCVW), Seoul, Korea.","DOI":"10.1109\/ICCVW.2019.00494"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Guo, Y., Wang, H., Hu, Q., Liu, H., Liu, L., and Bennamoun, M. (2020). Deep Learning for 3D Point Clouds: A Survey. IEEE Trans. Pattern Anal. and Mach. Intell.","DOI":"10.1109\/TPAMI.2020.3005434"},{"key":"ref_19","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). PointNet: Deep learning on point sets for 3D classification and segmentation. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA."},{"key":"ref_20","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017, January 3\u20139). PointNet++: Deep hierarchical feature learning on point sets in a metric space. Proceedings of the Neural Information Processing Systems (NIPS), Long Beach, CA, USA."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Jiang, M., Wu, Y., Zhao, T., Zhao, Z., and Lu, C. (2018). PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation. arXiv.","DOI":"10.1109\/IGARSS.2019.8900102"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhao, H., Jiang, L., Fu, C.W., and Jia, J. (2019, January 16\u201320). PointWeb: Enhancing local neighborhood features for point cloud processing. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00571"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Hu, Q., Yang, B., Xie, L., Rosa, S., Guo, Y., Wang, Z., Trigoni, N., and Markham, A. (2020, January 16\u201318). RandLA-Net: Efficient Semantic Segmentation of Large-Scale Point Clouds. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01112"},{"key":"ref_24","first-page":"820","article-title":"PointCNN: Convolution on X-Transformed Points","volume":"31","author":"Li","year":"2018","journal-title":"Adv. Neur. Inf."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Thomas, H., Qi, C.R., Deschaud, J.-E., Marcotegui, B., Goulette, F., and Guibas, L.J. (2019). KPConv: Flexible and Deformable Convolution for Point Clouds. arXiv.","DOI":"10.1109\/ICCV.2019.00651"},{"key":"ref_26","unstructured":"Ferrari, V., Hebert, M., Sminchisescu, C., and Weiss, Y. (2018). 3D Recurrent neural networks with context fusion for point cloud semantic segmentation. Computer Vision\u2014ECCV 2018, Springer International Publishing."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Landrieu, L., and Simonovsky, M. (2018, January 18\u201323). Large-scale point cloud semantic segmentation with superpoint graphs. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00479"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wang, L., Huang, Y., Hou, Y., Zhang, S., and Shan, J. (2019, January 16\u201320). Graph Attention Convolution for Point Cloud Semantic Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01054"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Bello, I., Zoph, B., Vaswani, A., Shlens, J., and Le, Q.V. (2019, January 27\u201328). Attention augmented convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00338"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Chen, L., Zhang, H., Xiao, J., Nie, L., Shao, J., Liu, W., and Chua, T.-S. (2017, January 21\u201326). Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.667"},{"key":"ref_31","unstructured":"Li, H., Xiong, P., An, J., and Wang, L. (2018). Pyramid attention network for semantic segmentation. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"71566","DOI":"10.1109\/ACCESS.2018.2880877","article-title":"Exploring new backbone and attention module for semantic segmentation in street scenes","volume":"6","author":"Fan","year":"2018","journal-title":"IEEE Access"},{"key":"ref_33","unstructured":"Wang, X., He, J., and Ma, L. (2019, January 8\u201315). Exploiting Local and Global Structure for Point Cloud Semantic Segmentation with Contextual Point Representations. Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada."},{"key":"ref_34","unstructured":"Jia, M., Li, A., and Wu, Z. (August, January 28). A Global Point-Sift Attention Network for 3d Point Cloud Semantic Segmentation. Proceedings of the International Geoscience and Remote Sensing Symposium, Yokohama, Japan."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"3308","DOI":"10.1080\/01431161.2018.1528024","article-title":"A scale robust convolutional neural network for automatic building extraction from aerial and satellite imagery","volume":"40","author":"Ji","year":"2018","journal-title":"Int. J. Remote Sens."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"2178","DOI":"10.1109\/TGRS.2019.2954461","article-title":"Toward Automatic Building Footprint Delineation from Aerial Images Using CNN and Regularization","volume":"58","author":"Wei","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Pintore, G., Agus, M., and Gobbetti, E. (2020, January 23\u201328). AtlantaNet: Inferring the 3D Indoor Layout from a Single 360 Image Beyond the Manhattan World Assumption. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58598-3_26"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"87695","DOI":"10.1109\/ACCESS.2020.2992612","article-title":"CSPC-Dataset: New LiDAR Point Cloud Dataset and Benchmark for Large-scale Semantic Segmentation","volume":"8","author":"Tong","year":"2020","journal-title":"IEEE Access"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Tan, W., Qin, N., Ma, L., Li, Y., Du, J., Cai, G., Yang, K., and Li, J. (2020). Toronto-3D: A Large-scale Mobile LiDAR Dataset for Semantic Segmentation of Urban Roadways. arXiv.","DOI":"10.1109\/CVPRW50498.2020.00109"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1016\/j.cag.2017.11.010","article-title":"SnapNet: 3D point cloud semantic labeling with 2D deep segmentation networks","volume":"71","author":"Boulch","year":"2017","journal-title":"Comput. Graph."},{"key":"ref_41","unstructured":"Huang, J., and You, S. (2016, January 4\u20138). Point cloud labeling using 3D Convolutional Neural Network. Proceedings of the 2016 23rd International Conference on Pattern Recognition (ICPR), Canc\u00fan, Mexico."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Hackel, T., Savinov, N., Ladicky, L., Wegner, J.D., Schindler, K., and Pollefeys, M. (2017). Semantic3D.net: A new Large-scale Point Cloud Classification Benchmark. arXiv.","DOI":"10.5194\/isprs-annals-IV-1-W1-91-2017"},{"key":"ref_43","first-page":"1","article-title":"Dynamic Graph CNN for Learning on Point Clouds","volume":"38","author":"Wang","year":"2019","journal-title":"Acm Trans. Graphic"},{"key":"ref_44","first-page":"1","article-title":"Multi-scale Point-wise Convolutional Neural Networks for 3D Object Segmentation from LiDAR Point Clouds in Large-scale Environments","volume":"99","author":"Ma","year":"2019","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"3588","DOI":"10.1109\/TGRS.2019.2958517","article-title":"TGNet: Geometric Graph CNN on 3D Point Cloud Segmentation","volume":"58","author":"Li","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/4\/691\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:23:57Z","timestamp":1760160237000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/4\/691"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,14]]},"references-count":45,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2021,2]]}},"alternative-id":["rs13040691"],"URL":"https:\/\/doi.org\/10.3390\/rs13040691","relation":{},"ISSN":["2072-4292"],"issn-type":[{"type":"electronic","value":"2072-4292"}],"subject":[],"published":{"date-parts":[[2021,2,14]]}}}