{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T05:19:47Z","timestamp":1775625587664,"version":"3.50.1"},"reference-count":40,"publisher":"MDPI AG","issue":"24","license":[{"start":{"date-parts":[[2021,12,14]],"date-time":"2021-12-14T00:00:00Z","timestamp":1639440000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61801359"],"award-info":[{"award-number":["61801359"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>The use of LiDAR point clouds for accurate three-dimensional perception is crucial for realizing high-level autonomous driving systems. Upon considering the drawbacks of the current point cloud object-detection algorithms, this paper proposes HCNet, an algorithm that combines an attention mechanism with adaptive adjustment, starting from feature fusion and overcoming the sparse and uneven distribution of point clouds. Inspired by the basic idea of an attention mechanism, a feature-fusion structure HC module with height attention and channel attention, weighted in parallel, is proposed to perform feature-fusion on multiple pseudo images. The use of several weighting mechanisms enhances the ability of feature-information expression. Additionally, we designed an adaptively adjusted detection head that also overcomes the sparsity of the point cloud from the perspective of original information fusion. It reduces the interference caused by the uneven distribution of the point cloud from the perspective of adaptive adjustment. The results show that our HCNet has better accuracy than other one-stage-network or even two-stage-network RCNNs under some evaluation detection metrics. Additionally, it has a detection rate of 30FPS. Especially for hard samples, the algorithm in this paper has better detection performance than many existing algorithms.<\/jats:p>","DOI":"10.3390\/rs13245071","type":"journal-article","created":{"date-parts":[[2021,12,14]],"date-time":"2021-12-14T22:06:10Z","timestamp":1639519570000},"page":"5071","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":9,"title":["HCNET: A Point Cloud Object Detection Network Based on Height and Channel Attention"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8495-2804","authenticated-orcid":false,"given":"Jing","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Telecommunication Engineering, Xidian University, Xi\u2019an 610100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiajun","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Telecommunication Engineering, Xidian University, Xi\u2019an 610100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Da","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Telecommunication Engineering, Xidian University, Xi\u2019an 610100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunsong","family":"Li","sequence":"additional","affiliation":[{"name":"School of Telecommunication Engineering, Xidian University, Xi\u2019an 610100, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,12,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"113816","DOI":"10.1016\/j.eswa.2020.113816","article-title":"Self-driving cars: A survey","volume":"165","author":"Badue","year":"2020","journal-title":"Expert Syst. Appl."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"58443","DOI":"10.1109\/ACCESS.2020.2983149","article-title":"A survey of autonomous driving: Common practices and emerging technologies","volume":"8","author":"Yurtsever","year":"2020","journal-title":"IEEE Access"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Xiang, Y., Choi, W., Lin, Y., and Savarese, S. (2015, January 7\u201312). Data-driven 3d voxel patterns for object category recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298800"},{"key":"ref_4","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA."},{"key":"ref_5","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. arXiv."},{"key":"ref_6","unstructured":"Yin, Z., and Tuzel, O. (2018, January 18\u201323). VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Yan, Y., Mao, Y., and Li, B. (2018). Second: Sparsely embedded convolutional detection. Sensors, 18.","DOI":"10.3390\/s18103337"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., and Beijbom, O. (2019, January 15\u201320). PointPillars: Fast Encoders for Object Detection from Point Clouds. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01298"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"3212","DOI":"10.1109\/TNNLS.2018.2876865","article-title":"Object detection with deep learning: A review","volume":"30","author":"Zhao","year":"2019","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_10","unstructured":"Li, X., Guivant, J.E., Kwok, N., and Xu, Y. (2019). 3D backbone network for 3D object detection. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Li, B. (2017, January 24\u201328). 3d fully convolutional network for vehicle detection in point cloud. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8205955"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1007\/s13218-010-0059-6","article-title":"Semantic 3d object maps for everyday manipulation in human living environments","volume":"24","author":"Rusu","year":"2010","journal-title":"KI-Kunstl. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, W., Sakurada, K., and Kawaguchi, N. (2016). Incremental and enhanced scanline-based segmentation method for surface reconstruction of sparse LiDAR data. Remote Sens., 8.","DOI":"10.3390\/rs8110967"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Narksri, P., Takeuchi, E., Ninomiya, Y., Morales, Y., Akai, N., and Kawaguchi, N. (2018, January 4\u20137). A slope-robust cascaded ground segmentation in 3D point cloud for autonomous vehicles. Proceedings of the 2018 21st International Conference on Intelligent Transportation Systems (ITSC), Maui, HI, USA.","DOI":"10.1109\/ITSC.2018.8569534"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"598","DOI":"10.20965\/jrm.2018.p0598","article-title":"Tsukuba challenge 2017 dynamic object tracks dataset for pedestrian behavior analysis","volume":"30","author":"Lambert","year":"2018","journal-title":"J. Robot. Mech."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Irshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_18","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Chen, X., Ma, H., Wan, J., Li, B., and Xia, T. (2017, January 21\u201326). Multi-view 3d object detection network for autonomous driving. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.691"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Shi, S., Wang, X., and Li, H.P. (2019, January 16\u201320). 3d object proposal generation and detection from point cloud. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00086"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Liu, W., Wu, C., Su, H., and Guibas, L.J. (2018, January 18\u201323). Frustum PointNets for 3D Object Detection from RGB-D Data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ali, W., Abdelkarim, S., Zidan, M., Zahran, M., and El Sallab, A. (2018, January 8\u201314). YOLO3D: End-to-end real-time 3D Oriented Object Bounding Box Detection from LiDAR Point Cloud. Proceedings of the ECCV 2018: \u201c3D Reconstruction meets Semantics\u201d Workshop, Munich, Germany.","DOI":"10.1007\/978-3-030-11015-4_54"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yang, Z., Sun, Y., Liu, S., and Jia, J. (2020, January 13\u201319). 3dssd: Point-based 3d single stage object detector. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01105"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Li, B., Zhang, T., and Xia, T. (2016). Vehicle detection from 3d lidar using fully convolutional network. arXiv.","DOI":"10.15607\/RSS.2016.XII.042"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Oda, M., Shimizu, N., Roth, H.R., Karasawa, K.I., Kitasaka, T., Misawa, K., Fujiwara, M., Rueckert, D., and Mori, K. (2017). 3D FCN Feature Driven Regression Forest-Based Pancreas Localization and Segmentation. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Springer.","DOI":"10.1007\/978-3-319-67558-9_26"},{"key":"ref_27","unstructured":"Engelcke, M., Rao, D., Wang, D.Z., Tong, C.H., and Posner, I. (June, January 29). Vote3Deep: Fast Object Detection in 3D Point Clouds Using Efficient Convolutional Neural Networks. Proceedings of the International Conference on Robotics and Automation (ICRA), Singapore."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Graham, B., and Maaten, L. (2017). Submanifold Sparse Convolutional Networks. arXiv.","DOI":"10.1109\/CVPR.2018.00961"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.Y., and Berg, A.C. (2016). Ssd: Single shot multibox detector. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3465055","article-title":"An Attentive Survey of Attention Models","volume":"12","author":"Chaudhari","year":"2021","journal-title":"ACM Trans. Intell. Syst. Technol. TIST"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Venugopalan, S., Rohrbach, M., Donahue, J., Mooney, R., Darrell, T., and Saenko, K. (2015, January 7\u201313). Sequence to sequence-video to text. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.515"},{"key":"ref_32","unstructured":"Jaderberg, M., Simonyan, K., and Zisserman, A. (2015). Spatial transformer networks. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., and Tang, X. (2017, January 21\u201326). Residual Attention Network for Image Classification. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.683"},{"key":"ref_35","first-page":"9267","article-title":"3D object detection using scale invariant and feature reweighting networks","volume":"33","author":"Zhao","year":"2019","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_36","first-page":"8778","article-title":"Point2sequence: Learning the shape representation of 3d point clouds with an attention-based sequence to sequence network","volume":"33","author":"Liu","year":"2019","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_37","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Woo, S., Park, J., Lee, J.Y., and Kweon, I.S. (2018, January 8\u201314). Cbam: Convolutional block attention module. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? The kitti vision benchmark suite. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1016\/j.neucom.2019.09.086","article-title":"SARPNET: Shape attention regional proposal network for liDAR-based 3D object detection","volume":"379","author":"Ye","year":"2020","journal-title":"Neurocomputing"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/24\/5071\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:47:32Z","timestamp":1760168852000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/13\/24\/5071"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12,14]]},"references-count":40,"journal-issue":{"issue":"24","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["rs13245071"],"URL":"https:\/\/doi.org\/10.3390\/rs13245071","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,12,14]]}}}