{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,31]],"date-time":"2025-12-31T18:25:39Z","timestamp":1767205539306,"version":"build-2238731810"},"reference-count":30,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2023,10,26]],"date-time":"2023-10-26T00:00:00Z","timestamp":1698278400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Road Meteorological State Sensitive Elements and Sensors","award":["2022YFB3205904"],"award-info":[{"award-number":["2022YFB3205904"]}]}],"content-domain":{"domain":["www.mdpi.com"],"crossmark-restriction":true},"short-container-title":["Sensors"],"abstract":"<jats:p>In the field of autonomous driving, object detection under point clouds is indispensable for environmental perception. In order to achieve the goal of reducing blind spots in perception, many autonomous driving schemes have added low-cost blind-filling LiDAR on the side of the vehicle. Unlike point cloud target detection based on high-performance LiDAR, the blind-filling LiDARs have low vertical angular resolution and are mounted on the side of the vehicle, resulting in easily mixed point clouds of pedestrian targets in close proximity to each other. These characteristics are harmful for target detection. Currently, many research works focus on target detection under high-density LiDAR. These methods cannot effectively deal with the high sparsity of the point clouds, and the recall and detection accuracy of crowded pedestrian targets tend to be low. To overcome these problems, we propose a real-time detection model for crowded pedestrian targets, namely RTCP. To improve computational efficiency, we utilize an attention-based point sampling method to reduce the redundancy of the point clouds, then we obtain new feature tensors by the quantization of the point cloud space and neighborhood fusion in polar coordinates. In order to make it easier for the model to focus on the center position of the target, we propose an object alignment attention module (OAA) for position alignment, and we utilize an additional branch of the targets\u2019 location occupied heatmap to guide the training of the OAA module. These methods improve the model\u2019s robustness against the occlusion of crowded pedestrian targets. Finally, we evaluate the detector on KITTI, JRDB, and our own blind-filling LiDAR dataset, and our algorithm achieved the best trade-off of detection accuracy against runtime efficiency.<\/jats:p>","DOI":"10.3390\/s23218725","type":"journal-article","created":{"date-parts":[[2023,10,26]],"date-time":"2023-10-26T07:22:15Z","timestamp":1698304935000},"page":"8725","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Real-Time 3D Object Detection on Crowded Pedestrians"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-8841-963X","authenticated-orcid":false,"given":"Bin","family":"Lu","sequence":"first","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"University of Chinese Academy of Sciences, Beijing 101408, China"},{"name":"Wuxi IoT Innovation Center Co., Ltd., Wuxi 214028, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qing","family":"Li","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanju","family":"Liang","sequence":"additional","affiliation":[{"name":"Institute of Microelectronics of the Chinese Academy of Sciences, Beijing 100029, China"},{"name":"Wuxi IoT Innovation Center Co., Ltd., Wuxi 214028, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,26]]},"reference":[{"key":"ref_1","unstructured":"Qu, C.R., Yi, L., Su, H., and Guibas, L.J. (2017). PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. NIPS, 5099\u20135108."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_3","first-page":"6748","article-title":"Jrdb: A dataset and benchmark of egocentric robot visual perception of humans in built environments","volume":"45","author":"Patel","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","unstructured":"Maturana, D., and Scherer, S. (October, January 28). Voxnet: A 3d con volutional neural network for real-time object recognition. Proceedings of the Intelligent Robots and Systems (IROS) Conference, Hamburg, Germany."},{"key":"ref_5","unstructured":"Zhou, Y., and Tuzel, O. (June, January 18). VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Riegler, G., Ulusoy, A.O., and Geiger, A. (2017, January 21\u201326). Octnet: Learning deep 3d representations at high resolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.701"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Yan, Y., Mao, Y., and Li, B. (2018). Second: Sparsely embedded convolutional detection. Sensors, 6.","DOI":"10.3390\/s18103337"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., and Beijbom, O. (2019, January 15\u201320). Pointpillars: Fast encoders for object detection from point clouds. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01298"},{"key":"ref_9","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA."},{"key":"ref_10","unstructured":"Zhang, Y., Hu, Q., Xu, G., Ma, Y., Wan, J., and Guo, Y. (June, January 18). Not All Points Are Equal: Learning Highly Efficient Point-Based Detectors for 3D LiDAR Point Clouds. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA."},{"key":"ref_11","unstructured":"Yang, Z., Sun, Y., Liu, S., and Jia, J. (June, January 13). 3DSSD: Point-Based 3D Single Stage Object Detector. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA."},{"key":"ref_12","unstructured":"Graham, B., Engelcke, M., and van der Maaten, L. (June, January 18). 3D Semantic Segmentation With Submanifold Sparse Convolutional Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA."},{"key":"ref_13","unstructured":"Shi, S., Wang, X., and Li, H. (June, January 15). Pointrcnn: 3d object proposal generation and detection from point cloud. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA."},{"key":"ref_14","unstructured":"Chen, Y., Liu, S., Shen, X., and Jia, J. (November, January 27). Fast point r-cnn. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_15","first-page":"1201","article-title":"Voxel R-CNN: Towards High Performance Voxel-based 3D Object Detection","volume":"35","author":"Deng","year":"2021","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_16","unstructured":"Simony, M., Milzy, S., Amendey, K., and Gross, H.-M. (September, January 8). Complex-yolo: An euler-region-proposal for real-time 3d object detection on point clouds. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Wu, B., Wan, A., Yue, X., and Keutzer, K. (May, January 21). Squeezeseg: Convolutional neural nets with recurrent crf for real-time road object segmentation from 3d lidar point cloud. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia.","DOI":"10.1109\/ICRA.2018.8462926"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Li, B. (September, January 24). 3d fully convolutional network for vehicle detection in point cloud. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, Canada.","DOI":"10.1109\/IROS.2017.8205955"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yang, B., Luo, W., and Urtasun, R. (2018, January 18\u201322). PIXOR: Real-Time 3D Object Detection From Point Clouds. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00798"},{"key":"ref_20","first-page":"3555","article-title":"CIA-SSD: Confident IoU-Aware Single-Stage Object Detector From Point Cloud","volume":"35","author":"Zheng","year":"2021","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_21","first-page":"11677","article-title":"TANet: Robust 3D Object Detection from Point Clouds with Triple Attention","volume":"34","author":"Liu","year":"2020","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1159","DOI":"10.1109\/LRA.2022.3233234","article-title":"Accurate and Real-Time 3D Pedestrian Detection Using an Efficient Attentive Pillar Network","volume":"8","author":"Le","year":"2023","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_23","first-page":"221","article-title":"SASA: Semantics-Augmented Set Abstraction for Point-Based 3D Object Detection","volume":"36","author":"Chen","year":"2022","journal-title":"Proc. Aaai Conf. Artif. Intell."},{"key":"ref_24","unstructured":"Bhat, M., and Han, S. (2021). Porikli, Fast Polar Attentive 3D Object Detection on LiDAR Point Clouds. Comput. Sci."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Zhou, Z., David, P., Yue, X., Xi, Z., and Gong, B. (2020, January 13\u201319). Hassan Foroosh, PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00962"},{"key":"ref_26","unstructured":"Qi, C.R., Litany, O., He, K., and Guibas, L.J. (November, January 27). Deep Hough Voting for 3D Object Detection in Point Clouds. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Paigwar, A., Erkent, O., Wolf, C., and Laugier, C. (2019, January 15\u201320). Attentional PointNet for 3D-Object Detection in Point Clouds. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00169"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Wang, L., Huang, Y., Hou, Y., Zhang, S., and Shan, J. (2019, January 16\u201320). Graph attention convolution for point cloud semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2019, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01054"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhang, W., and Xiao, C. (2019, January 15\u201320). PCAN: 3d attention map learning using contextual information for point cloud based retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01272"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"}],"updated-by":[{"DOI":"10.3390\/s25072212","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2023,10,26]],"date-time":"2023-10-26T00:00:00Z","timestamp":1698278400000}}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/21\/8725\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,3]],"date-time":"2025-08-03T14:12:50Z","timestamp":1754230370000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/21\/8725"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,26]]},"references-count":30,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2023,11]]}},"alternative-id":["s23218725"],"URL":"https:\/\/doi.org\/10.3390\/s23218725","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,26]]}}}