{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T12:27:29Z","timestamp":1763468849332,"version":"build-2065373602"},"reference-count":46,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2023,5,22]],"date-time":"2023-05-22T00:00:00Z","timestamp":1684713600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Wuhu and Xidian University special fund for industry-university-research cooperation","award":["XWYCXY-012021019"],"award-info":[{"award-number":["XWYCXY-012021019"]}]},{"name":"General project of key R$\\&amp;$D Plan of Shaanxi Province","award":["2022GY-060"],"award-info":[{"award-number":["2022GY-060"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>In the field of unmanned systems, cameras and LiDAR are important sensors that provide complementary information. However, the question of how to effectively fuse data from two different modalities has always been a great challenge. In this paper, inspired by the idea of deep fusion, we propose a one-stage end-to-end network named FusionPillars to fuse multisensor data (namely LiDAR point cloud and camera images). It includes three branches: a point-based branch, a voxel-based branch, and an image-based branch. We design two modules to enhance the voxel-wise features in the pseudo-image: the Set Abstraction Self (SAS) fusion module and the Pseudo View Cross (PVC) fusion module. For the data from a single sensor, by considering the relationship between the point-wise and voxel-wise features, the SAS fusion module self-fuses the point-based branch and the voxel-based branch to enhance the spatial information of the pseudo-image. For the data from two sensors, through the transformation of the images\u2019 view, the PVC fusion module introduces the RGB information as auxiliary information and cross-fuses the pseudo-image and RGB image of different scales to supplement the color information of the pseudo-image. Experimental results revealed that, compared to existing current one-stage fusion networks, FusionPillars yield superior performance, with a considerable improvement in the detection precision for small objects.<\/jats:p>","DOI":"10.3390\/rs15102692","type":"journal-article","created":{"date-parts":[[2023,5,23]],"date-time":"2023-05-23T01:36:48Z","timestamp":1684805808000},"page":"2692","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["FusionPillars: A 3D Object Detection Network with Cross-Fusion and Self-Fusion"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8495-2804","authenticated-orcid":false,"given":"Jing","family":"Zhang","sequence":"first","affiliation":[{"name":"State Key Laboratory of Integrated Service Network, Xidian University, Xi\u2019an 710071, China"},{"name":"School of Telecommunication Engineering, Xidian University, Xi\u2019an 710071, China"},{"name":"Guangzhou Institute of Technology, Xidian University, Guangzhou 510555, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Da","family":"Xu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Integrated Service Network, Xidian University, Xi\u2019an 710071, China"},{"name":"School of Telecommunication Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yunsong","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Integrated Service Network, Xidian University, Xi\u2019an 710071, China"},{"name":"School of Telecommunication Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liping","family":"Zhao","sequence":"additional","affiliation":[{"name":"National Defense Science and Technology Innovation Research Institute, Beijing 100071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rui","family":"Su","sequence":"additional","affiliation":[{"name":"Xi\u2019an Termony Electronic Technology Co., Ltd., Xi\u2019an 710031, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,5,22]]},"reference":[{"key":"ref_1","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3d classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."},{"key":"ref_2","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017). Pointnet++: Deep hierarchical feature learning on point sets in a metric space. Adv. Neural Inf. Process. Syst., 30."},{"key":"ref_3","first-page":"1","article-title":"Dynamic graph cnn for learning on point clouds","volume":"38","author":"Wang","year":"2019","journal-title":"Acm Trans. Graph. (Tog)"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wang, Y., Chao, W.L., Garg, D., Hariharan, B., Campbell, M., and Weinberger, K.Q. (2019, January 15\u201320). Pseudo-lidar from visual depth estimation: Bridging the gap in 3d object detection for autonomous driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00864"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Engelcke, M., Rao, D., Wang, D.Z., Tong, C.H., and Posner, I. (June, January 29). Vote3deep: Fast object detection in 3d point clouds using efficient convolutional neural networks. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore.","DOI":"10.1109\/ICRA.2017.7989161"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhou, Y., and Tuzel, O. (2018, January 18\u201322). Voxelnet: End-to-end learning for point cloud based 3d object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00472"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Yan, Y., Mao, Y., and Li, B. (2018). Second: Sparsely embedded convolutional detection. Sensors, 18.","DOI":"10.3390\/s18103337"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., and Beijbom, O. (2019, January 15\u201320). Pointpillars: Fast encoders for object detection from point clouds. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01298"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ku, J., Mozifian, M., Lee, J., Harakeh, A., and Waslander, S.L. (2018, January 1\u20135). Joint 3d proposal generation and object detection from view aggregation. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8594049"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Yang, B., Luo, W., and Urtasun, R. (2018, January 18\u201323). Pixor: Real-time 3d object detection from point clouds. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00798"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Liang, M., Yang, B., Chen, Y., Hu, R., and Urtasun, R. (2019, January 15\u201320). Multi-task multi-sensor fusion for 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00752"},{"key":"ref_12","unstructured":"Liang, Z., Zhang, M., Zhang, Z., Zhao, X., and Pu, S. (2020). Rangercnn: Towards fast and accurate 3d object detection with range image representation. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"4722","DOI":"10.1109\/TCSVT.2021.3100848","article-title":"From multi-view to hollow-3D: Hallucinated hollow-3D R-CNN for 3D object detection","volume":"31","author":"Deng","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"3345","DOI":"10.1109\/TCSVT.2019.2957821","article-title":"Three-dimensional point cloud object detection using scene appearance consistency among multi-view projection directions","volume":"30","author":"Sugimura","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? The KITTI vision benchmark suite. Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Xu, D., Anguelov, D., and Jain, A. (2018, January 18\u201323). Pointfusion: Deep sensor fusion for 3d bounding box estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00033"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Vora, S., Lang, A.H., Helou, B., and Beijbom, O. (2020, January 13\u201319). Pointpainting: Sequential fusion for 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00466"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Xie, L., Xiang, C., Yu, Z., Xu, G., Yang, Z., Cai, D., and He, X. (2020, January 7\u201312). PI-RCNN: An efficient multi-sensor 3D object detector with point-based attentive cont-conv fusion module. Proceedings of the AAAI Conference on Artificial Intelligence, Hilton, NY, USA.","DOI":"10.1609\/aaai.v34i07.6933"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"5411","DOI":"10.1109\/TCSVT.2022.3148257","article-title":"AM\u00b3Net: Adaptive Mutual-Learning-Based Multimodal Data Fusion Network","volume":"32","author":"Wang","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Liu, K., Bao, H., Zheng, Y., and Yang, Y. (2023). PMPF: Point-Cloud Multiple-Pixel Fusion-Based 3D Object Detection for Autonomous Driving. Remote Sens., 15.","DOI":"10.3390\/rs15061580"},{"key":"ref_21","unstructured":"Huang, T., Liu, Z., Chen, X., and Bai, X. (2020). Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23\u201328 August 2020, Springer."},{"key":"ref_22","unstructured":"Yoo, J.H., Kim, Y., Kim, J., and Choi, J.W. (2020). Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23\u201328 August 2020, Springer."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Li, Y., Yu, A.W., Meng, T., Caine, B., Ngiam, J., Peng, D., Shen, J., Lu, Y., Zhou, D., and Le, Q.V. (2022, January 18\u201324). Deepfusion: Lidar-camera deep fusion for multi-modal 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01667"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Xu, X., Dong, S., Xu, T., Ding, L., Wang, J., Jiang, P., Song, L., and Li, J. (2023). FusionRCNN: LiDAR-Camera Fusion for Two-Stage 3D Object Detection. Remote Sens., 15.","DOI":"10.3390\/rs15071839"},{"key":"ref_25","unstructured":"Kim, T., and Ghosh, J. (2016, January 1\u20134). Robust detection of non-motorized road users using deep learning on optical and LIDAR data. Proceedings of the 2016 IEEE 19th International Conference on Intelligent Transportation Systems (ITSC), Rio de Janeiro, Brazil."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Liu, J., Zhang, S., Wang, S., and Metaxas, D.N. (2016). Multispectral deep neural networks for pedestrian detection. arXiv.","DOI":"10.5244\/C.30.73"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Pang, S., Morris, D., and Radha, H. (2020\u201324, January 24). CLOCs: Camera-LiDAR object candidates fusion for 3D object detection. Proceedings of the 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9341791"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Liu, W., Wu, C., Su, H., and Guibas, L.J. (2018, January 18\u201323). Frustum pointnets for 3d object detection from rgb-d data. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Shi, S., Wang, X., and Li, H. (2019, January 15\u201320). Pointrcnn: 3d object proposal generation and detection from point cloud. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00086"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yang, Z., Sun, Y., Liu, S., and Jia, J. (2020, January 13\u201319). 3dssd: Point-based 3d single stage object detector. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01105"},{"key":"ref_31","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. Adv. Neural Inf. Process. Syst., 28."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Deng, J., Shi, S., Li, P., Zhou, W., Zhang, Y., and Li, H. (2021, January 2\u20139). Voxel r-cnn: Towards high performance voxel-based 3d object detection. Proceedings of the AAAI Conference on Artificial Intelligence, Virtually.","DOI":"10.1609\/aaai.v35i2.16207"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Shi, S., Guo, C., Jiang, L., Wang, Z., Shi, J., Wang, X., and Li, H. (2020, January 13\u201319). Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01054"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhang, J., Xu, D., Wang, J., and Li, Y. (2021, January 23\u201325). An Improved Detection Algorithm For Pre-processing Problem Based On PointPillars. Proceedings of the 2021 14th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), Shanghai, China.","DOI":"10.1109\/CISP-BMEI53629.2021.9624329"},{"key":"ref_35","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International Conference on Machine Learning, PMLR, Lille, France."},{"key":"ref_36","unstructured":"Nair, V., and Hinton, G.E. (2010, January 21\u201324). Rectified linear units improve restricted boltzmann machines. Proceedings of the 27th International Conference on Machine Learning (ICML-10), Haifa, Israel."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhang, J., Wang, J., Xu, D., and Li, Y. (2021). HCNET: A Point Cloud Object Detection Network Based on Height and Channel Attention. Remote Sens., 13.","DOI":"10.3390\/rs13245071"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Chen, X., Ma, H., Wan, J., Li, B., and Xia, T. (2017, January 21\u201326). Multi-view 3d object detection network for autonomous driving. Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.691"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Liu, Z., Zhao, X., Huang, T., Hu, R., Zhou, Y., and Bai, X. (2020, January 7\u201312). Tanet: Robust 3d object detection from point clouds with triple attention. Proceedings of the AAAI Conference on Artificial Intelligence, Hilton, NY, USA.","DOI":"10.1609\/aaai.v34i07.6837"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Shi, W., and Rajkumar, R. (2020, January 13\u201319). Point-gnn: Graph neural network for 3d object detection in a point cloud. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00178"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Wang, M., Chen, Q., and Fu, Z. (2022). Lsnet: Learned sampling network for 3d object detection from point clouds. Remote Sens., 14.","DOI":"10.3390\/rs14071539"},{"key":"ref_42","unstructured":"Yang, B., Liang, M., and Urtasun, R. (2018, January 29\u201331). Hdnet: Exploiting hd maps for 3d object detection. Proceedings of the Conference on Robot Learning, PMLR, Z\u00fcrich, Switzerland."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Liang, M., Yang, B., Wang, S., and Urtasun, R. (2018, January 8\u201314). Deep continuous fusion for multi-sensor 3d object detection. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01270-0_39"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Sindagi, V.A., Zhou, Y., and Tuzel, O. (2019, January 20\u201324). Mvx-net: Multimodal voxelnet for 3d object detection. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794195"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Yang, Z., Sun, Y., Liu, S., Shen, X., and Jia, J. (2018). Ipod: Intensive point-based object detector for point cloud. arXiv.","DOI":"10.1109\/ICCV.2019.00204"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Wang, Z., and Jia, K. (2019, January 3\u20138). Frustum convnet: Sliding frustums to aggregate local point-wise features for amodal 3d object detection. Proceedings of the 2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China.","DOI":"10.1109\/IROS40897.2019.8968513"}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/10\/2692\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:40:06Z","timestamp":1760125206000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/10\/2692"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,22]]},"references-count":46,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2023,5]]}},"alternative-id":["rs15102692"],"URL":"https:\/\/doi.org\/10.3390\/rs15102692","relation":{},"ISSN":["2072-4292"],"issn-type":[{"type":"electronic","value":"2072-4292"}],"subject":[],"published":{"date-parts":[[2023,5,22]]}}}