{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T13:45:23Z","timestamp":1768830323599,"version":"3.49.0"},"reference-count":33,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2019,10,7]],"date-time":"2019-10-07T00:00:00Z","timestamp":1570406400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"The Key Technical Project of Fujian Province","award":["No. 2017H6015."],"award-info":[{"award-number":["No. 2017H6015."]}]},{"DOI":"10.13039\/501100001809","name":"The National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.61702251"],"award-info":[{"award-number":["No.61702251"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Semantic segmentation of 3D point clouds plays a vital role in autonomous driving, 3D maps, and smart cities, etc. Recent work such as PointSIFT shows that spatial structure information can improve the performance of semantic segmentation. Motivated by this phenomenon, we propose Spatial Aggregation Net (SAN) for point cloud semantic segmentation. SAN is based on multi-directional convolution scheme that utilizes the spatial structure information of point cloud. Firstly, Octant-Search is employed to capture the neighboring points around each sampled point. Secondly, we use multi-directional convolution to extract information from different directions of sampled points. Finally, max-pooling is used to aggregate information from different directions. The experimental results conducted on ScanNet database show that the proposed SAN has comparable results with state-of-the-art algorithms such as PointNet, PointNet++, and PointSIFT, etc. In particular, our method has better performance on flat, small objects, and the edge areas that connect objects. Moreover, our model has good trade-off in segmentation accuracy and time complexity.<\/jats:p>","DOI":"10.3390\/s19194329","type":"journal-article","created":{"date-parts":[[2019,10,7]],"date-time":"2019-10-07T10:05:12Z","timestamp":1570442712000},"page":"4329","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Spatial Aggregation Net: Point Cloud Semantic Segmentation Based on Multi-Directional Convolution"],"prefix":"10.3390","volume":"19","author":[{"given":"Guorong","family":"Cai","sequence":"first","affiliation":[{"name":"Computer Engineering College, Jimei University, Xiamen 361021, China"},{"name":"Fujian Collaborative Innovation Center for Big Data Applications in Governments, Fuzhou 350003, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zuning","family":"Jiang","sequence":"additional","affiliation":[{"name":"Computer Engineering College, Jimei University, Xiamen 361021, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zongyue","family":"Wang","sequence":"additional","affiliation":[{"name":"Computer Engineering College, Jimei University, Xiamen 361021, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shangfeng","family":"Huang","sequence":"additional","affiliation":[{"name":"Computer Engineering College, Jimei University, Xiamen 361021, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Chen","sequence":"additional","affiliation":[{"name":"Computer Engineering College, Jimei University, Xiamen 361021, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuyang","family":"Ge","sequence":"additional","affiliation":[{"name":"Computer Engineering College, Jimei University, Xiamen 361021, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yundong","family":"Wu","sequence":"additional","affiliation":[{"name":"Computer Engineering College, Jimei University, Xiamen 361021, China"},{"name":"Fujian Collaborative Innovation Center for Big Data Applications in Governments, Fuzhou 350003, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,10,7]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Su, H., Maji, S., Kalogerakis, E., and Learned-Miller, E. (2015, January 11\u201318). Multi-view convolutional neural networks for 3D shape recognition. Proceedings of the IEEE International Conference on Computer Vision, Las Condes, Chile.","DOI":"10.1109\/ICCV.2015.114"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Maturana, D., and Scherer, S. (October, January 28). Voxnet: A 3D convolutional neural network for real-time object recognition. Proceedings of the 2015 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany.","DOI":"10.1109\/IROS.2015.7353481"},{"key":"ref_3","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3D classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."},{"key":"ref_4","unstructured":"Qi, C., Yi, L., Su, H., and Guibas, L.J. (2017, January 4\u20139). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_5","unstructured":"Li, Y., Bu, R., Sun, M., Wu, W., Di, X., and Chen, B. (2018). Pointcnn: Convolution on x-transformed points. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Jiang, M., Wu, Y., Zhao, T., Zhao, Z., and Lu, C. (2018). Pointsift: A sift-like network module for 3D point cloud semantic segmentation. arXiv.","DOI":"10.1109\/IGARSS.2019.8900102"},{"key":"ref_7","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the 26th Conference on Advances in Neural Information Processing Systems, Lake Tahoe, NV, USA."},{"key":"ref_8","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_9","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Su, H., Niessner, M., Dai, A., Yan, M., and Guibas, L.J. (2016). Volumetric and multi-view cnns for object classification on 3D data. arXiv.","DOI":"10.1109\/CVPR.2016.609"},{"key":"ref_11","unstructured":"Wu, Z., Song, S., Khosla, A., Yu, F., Zhang, L., Tang, X., and Xiao, J. (2015, January 7\u201312). 3D shapenets: A deep representation for volumetric shapes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Garcia-Garcia, A., Orts-Escolano, S., Oprea, S., Villena-Martinez, V., and Garcia-Rodriguez, J. (2017). A review on deep learning techniques applied to semantic segmentation. arXiv.","DOI":"10.1016\/j.asoc.2018.05.018"},{"key":"ref_13","unstructured":"Xie, Y., Tian, J., and Zhu, X.X. (2019). A Review of Point Cloud Semantic Segmentation. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Roveri, R., Rahmann, L., Oztireli, A.C., and Gross, M.H. (2018, January 18\u201323). A network architecture for point cloud classification via automatic depth images generation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00439"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Su, J., Gahelda, M., Wang, R., and Maji, S. (2018, January 8\u201314). A deeper look at 3D shape classifiers. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-11015-4_49"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Milz, S., Simon, M., Fischer, K., and Popperl, M. (2019). Points2Pix: 3D Point-Cloud to Image Translation using conditional Generative Adversarial Networks. arXiv.","DOI":"10.1007\/978-3-030-33676-9_27"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Han, Z., Shang, M., Liu, Y., and Zwicker, M. (2018). View inter-prediction gan: Unsupervised representation learning for 3D shapes by learning global shape memories to support local view predictions. arXiv.","DOI":"10.1609\/aaai.v33i01.33018376"},{"key":"ref_18","unstructured":"You, Y., Lou, Y., Liu, Q., Ma, L., Wang, W., Tai, Y., and Lu, C. (2018). PRIN: Pointwise Rotation-Invariant Network. arXiv."},{"key":"ref_19","unstructured":"Asako, K., Matsushita, Y., and Nishida, Y. (2018, January 18\u201323). Rotationnet: Joint object categorization and pose estimation using multiviews from unsupervised viewpoints. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1016\/j.isprsjprs.2018.01.019","article-title":"Multi-scan segmentation of terrestrial laser scanning data based on normal variation analysis","volume":"143","author":"Che","year":"2018","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Che, E., and Olsen, M.J. (2019). An Efficient Framework for Mobile Lidar Trajectory Reconstruction and Mo-norvana Segmentation. Remote Sens., 11.","DOI":"10.3390\/rs11070836"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1016\/j.isprsjprs.2012.05.001","article-title":"Segmentation of terrestrial laser scanning data using geometry and image information","volume":"76","author":"Barnea","year":"2013","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Song, S., Lichtenberg, S.P., and Xiao, J. (2015, January 7\u201312). Sun rgb-d: A rgb-d scene understanding benchmark suite. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298655"},{"key":"ref_24","unstructured":"Li, Y., Pirk, S., Su, H., Qi, C.R., and Guibas, L.J. (2016, January 5\u201310). Fpnn: Field probing neural networks for 3D data. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Tatarchenko, M., Dosovitskiy, A., and Brox, T. (2017, January 22\u201329). Octree generating networks: Efficient convolutional architectures for high-resolution 3D outputs. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.230"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Su, H., Jampani, V., Sun, D., Maji, S., Kalogerakis, E., Yang, M.H., and Kautz, J. (2018, January 18\u201323). Splatnet: Sparse lattice networks for point cloud processing. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00268"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Liu, X., Han, Z., Liu, Y., and Zwicker, M. (2018). Point2Sequence: Learning the shape representation of 3D point clouds with an attention-based sequence to sequence network. arXiv.","DOI":"10.1609\/aaai.v33i01.33018778"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wu, W., Qi, Z., and Fuxin, L. (2019, January 16\u201320). Pointconv: Deep convolutional networks on 3D point clouds. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00985"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. arXiv.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"1305","DOI":"10.1109\/83.623193","article-title":"The farthest point strategy for progressive image sampling","volume":"6","author":"Eldar","year":"1997","journal-title":"IEEE Trans. Image Process."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T.A., and Niebner, M. (2017, January 21\u201326). Scannet: Richly-annotated 3D reconstructions of indoor scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.261"},{"key":"ref_33","unstructured":"Armeni, I., Sener, O., Zamir, A.R., Jiang, H., Brilakis, I., Fischer, M., and Savarese, S. (July, January 26). 3D semantic parsing of large-scale indoor spaces. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/19\/4329\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T13:28:05Z","timestamp":1760189285000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/19\/19\/4329"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,7]]},"references-count":33,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2019,10]]}},"alternative-id":["s19194329"],"URL":"https:\/\/doi.org\/10.3390\/s19194329","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,10,7]]}}}